Image processing device, image processing method, image processing system, and program
The image processing apparatus anonymizes facial images by replacing faces and maintaining orientation continuity, addressing privacy concerns and improving model training data quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HONDA MOTOR CO LTD
- Filing Date
- 2022-06-30
- Publication Date
- 2026-04-24
AI Technical Summary
Existing techniques for generating learning data for machine learning models from facial images fail to protect privacy while maintaining feature information, leading to ineffective data generation.
An image processing apparatus and method that anonymizes facial images by replacing faces with those of other individuals and ensures consistency in facial orientation and continuity, using threshold-based methods to determine appropriate anonymization techniques.
Generates effective training data for machine learning models while protecting privacy by ensuring facial orientation and continuity, enhancing the accuracy of behavior prediction models.
Smart Images

Figure 0007851199000001 
Figure 0007851199000002 
Figure 0007851199000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, an image processing system, and a program.
Background Art
[0002] In recent years, efforts have been actively made to provide access to a sustainable transportation system that takes into account people in vulnerable positions among traffic participants. Toward this realization, research and development have focused on further improving traffic safety and convenience through research and development related to autonomous driving technology. For example, conventionally, a technique for annotating an individual's face image has been known in order to generate learning data used for learning a machine learning model. For example, Patent Document 1 discloses a technique that enables generating a synthetic face image by referring to a plurality of face images stored in a face image database and performing an annotation operation on the generated synthetic face image.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technique described in Patent Document 1 protects the privacy of these multiple people by an annotator performing an annotation operation on a synthetic face image synthesized from a plurality of face images. However, in the prior art, due to converting the original image to protect privacy, the feature information of the original image may be missing. As a result, it may not be possible to generate learning data effective for learning a machine learning model while protecting the privacy of the person depicted in the face image.
[0005] This invention has been made in consideration of these circumstances, and one of its objectives is to provide an image processing device, an image processing method, and a program that can generate training data effective for training machine learning models while protecting the privacy of people captured in facial images. Ultimately, this will contribute to the development of sustainable transportation systems. [Means for solving the problem]
[0006] The image processing apparatus, image processing method, image processing system, and program according to this invention employ the following configuration. (1) An image processing apparatus according to one aspect of the present invention comprises an image conversion unit that performs anonymization processing on a plurality of input images captured in a time series, and an image determination unit that determines whether the plurality of input images that have undergone the anonymization processing satisfy predetermined requirements, wherein if the image determination unit determines that the plurality of input images that have undergone the anonymization processing satisfy the predetermined requirements, it performs predetermined processing on the plurality of input images that have undergone the anonymization processing, the anonymization processing includes processing to change the faces of people depicted in the plurality of input images to the faces of other people, and the predetermined requirements include that the faces of people tracked as the same person in the plurality of input images are the same faces in each of the plurality of input images that have undergone the anonymization processing.
[0007] (2) In the embodiment of (1) above, the predetermined processing is the process of saving the plurality of input images that have undergone the anonymization process as images to be used for annotation work.
[0008] (3) In the embodiment of (2) above, the predetermined processing is the process of saving all of the consecutive input images that have undergone the anonymization process as the images to be used for the annotation work.
[0009] (4) In the embodiment of (1) above, the predetermined processing is the process of storing the plurality of anonymized input images as training information for generating a behavior prediction model that predicts the actions of people depicted in the input images.
[0010] (5) In the embodiment of (1) above, the predetermined processing is the process of transmitting the plurality of input images that have undergone the anonymization process to an image server via communication means.
[0011] (6) In any of the embodiments described in (1) to (5) above, the image determination unit extracts facial feature points of the person tracked as the same person from each of the plurality of anonymized input images, and determines that the faces of the people tracked as the same person are the same face if the positional relationship of the extracted feature points matches.
[0012] (7) In any embodiment of (1) to (5) above, if each of the plurality of anonymized input images contains the faces of multiple persons, the image determination unit determines whether the predetermined requirements are met with respect to the faces of persons among the plurality of persons that are facing forward in the direction of travel of the vehicle on which the camera that captured the input image is mounted.
[0013] (8) In any embodiment of (1) to (5) above, if each of the multiple anonymized input images contains the faces of multiple persons, the image determination unit determines whether the predetermined requirements are met for the faces of persons whose faces appear in the input images that meet the predetermined criteria.
[0014] (9) In any embodiment of (1) to (5) above, if the image conversion unit determines that the plurality of input images that have undergone the anonymization process do not satisfy the predetermined requirements, the image conversion unit applies the anonymization process to the plurality of input images again.
[0015] (10): In any embodiment of (1) to (5) above, if the image conversion unit determines that the plurality of anonymized input images do not satisfy the predetermined requirements, the image conversion unit does not perform the predetermined processing on the plurality of anonymized input images.
[0016] (11): An image processing system according to another aspect of the present invention comprises an image conversion unit that performs anonymization processing on a plurality of input images captured in a time series, and an image determination unit that determines whether the plurality of input images that have undergone the anonymization processing satisfy predetermined requirements, wherein if the image determination unit determines that the plurality of input images that have undergone the anonymization processing satisfy the predetermined requirements, it performs predetermined processing on the plurality of input images that have undergone the anonymization processing, the anonymization processing includes processing to change the faces of people depicted in the plurality of input images to the faces of other people, and the predetermined requirements include that the faces of people tracked as the same person in the plurality of input images are the same faces in each of the plurality of input images that have undergone the anonymization processing.
[0017] (12): An image processing method according to another aspect of the present invention, wherein a computer performs anonymization processing on a plurality of input images captured in a time series, determines whether the plurality of input images subjected to the anonymization processing satisfy predetermined requirements, and if it is determined that the plurality of input images subjected to the anonymization processing satisfy the predetermined requirements, performs predetermined processing on the plurality of input images subjected to the anonymization processing, wherein the anonymization processing includes processing to change the faces of people depicted in the plurality of input images to the faces of other people, and the predetermined requirements include that the faces of people tracked as the same person in the plurality of input images are the same faces in each of the plurality of input images subjected to the anonymization processing.
[0018] (13): A program according to another aspect of the present invention causes a computer to perform anonymization processing on a plurality of input images captured in time series, determine whether the plurality of input images subjected to the anonymization processing satisfy a predetermined requirement, and when it is determined that the plurality of input images subjected to the anonymization processing satisfy the predetermined requirement, cause the plurality of input images subjected to the anonymization processing to be subjected to a predetermined processing. The anonymization processing includes processing for changing the face of a person depicted in the plurality of input images to the face of another person. The predetermined requirement includes that the face of a person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images subjected to the anonymization processing.
Advantages of the Invention
[0019] (1)~(13) According to the above, while protecting the privacy of the person depicted in the face image, it is possible to generate learning data effective for learning the machine learning model.
Brief Description of the Drawings
[0020] [Figure 1] It is a diagram showing an overview of the system 1 including the image processing apparatus 100 according to the present embodiment. [Figure 2] It is a diagram showing an example of the functional configuration of the image processing apparatus 100 according to the present embodiment. [Figure 3] It is a diagram showing an example of an in-vehicle image and an out-of-vehicle image acquired from the vehicle M1. [Figure 4] It is a diagram for explaining the processing executed by the image processing unit 130. [Figure 5] It is a diagram for explaining the processing executed by the image conversion unit 140. [Figure 6] It is a diagram showing an example of a time series of in-vehicle images converted by the image conversion unit 140. [Figure 7] It is a diagram for explaining the processing executed by the image determination unit 150. [Figure 8] It is a diagram showing an example of an annotation operation executed by an annotator. [Figure 9] It is a diagram showing an example of driving support using the learned model 180. [Figure 10] It is a diagram showing an example of the processing flow executed by the image conversion unit 140. [Figure 11] It is a diagram showing an example of the processing flow executed by the image determination unit 150.
Embodiments for Carrying Out the Invention
[0021] Hereinafter, embodiments of the image processing apparatus, image processing method, image processing system, and program of the present invention will be described with reference to the drawings.
[0022] [Overview] FIG. 1 is a diagram showing an overview of the system 1 including the image processing apparatus 100 according to the present embodiment. As shown in FIG. 1, the system 1 includes at least one or more vehicles M1 and M2, an image processing apparatus 100, and a terminal device 200. For convenience of explanation, vehicles M1 and M2 are illustrated as different vehicles, but these vehicles may be the same.
[0023] The vehicle M1 is, for example, a four-wheel drive vehicle such as a hybrid vehicle or an electric vehicle, and includes at least a camera that images the inside of the vehicle M1 and a camera that images the outside of the vehicle M1. While the vehicle M1 is running, the in-vehicle image and the out-of-vehicle image captured by these cameras are transmitted to the image processing apparatus 100 via a network NW such as a cellular network, a Wi-Fi network, or the Internet.
[0024] When the image processing apparatus 100 receives imaging image data including an in-vehicle image and an out-of-vehicle image from the vehicle M1, it is a server apparatus that performs the following-described image conversion on the received imaging image data. This image conversion is a process for protecting the privacy of the people shown in the in-vehicle image and the out-of-vehicle image. The image processing apparatus 100 transmits the obtained converted image data to the terminal device 200 via the network NW.
[0025] The terminal device 200 is a terminal device such as a desktop computer or a smartphone. When the user of the terminal device 200 obtains converted image data from the image processing device 100, they perform the annotation process described later on the obtained converted image data. Once the annotation process is complete, the user of the terminal device 200 sends the annotated image data, which has been converted image data with annotations added, to the image processing device 100.
[0026] When the image processing device 100 receives annotated image data from the terminal device 200, it uses the received annotated image data as training data and generates a trained model, described later, using an arbitrary machine learning model. This trained model is, for example, a behavior prediction model that, in response to an external image input, outputs the predicted behavior (trajectory) of a person captured in the external image, or, in response to internal and external images input, takes into account the driver's gaze captured in the internal image and prompts attention to pedestrians captured in the external image.
[0027] The image data used as training data in this case may be annotated image data in which annotations have been added to the transformed image data, or it may be annotated image data obtained by re-transforming the transformed image data into captured image data while retaining the annotations (i.e., annotated image data in which annotations have been added to captured image data). By using annotated image data in which annotations have been added to captured image data as training data, it is possible to use more realistic training data in which the influence of image transformation has been removed.
[0028] The image processing device 100 generates a trained model and then distributes the generated trained model to the vehicle M2 via the network NW. Similar to vehicle M1, vehicle M2 is a four-wheel drive vehicle such as a hybrid or electric vehicle. While driving, vehicle M2 inputs at least one of the interior and exterior images captured by the camera into the trained model to obtain behavior prediction data of people present around vehicle M2. The driver of vehicle M2 can refer to the obtained behavior prediction data and use it to guide the driving of vehicle M2. The following describes the details of each process.
[0029] [Functional Configuration of Image Processing Devices] Figure 2 shows an example of the functional configuration of the image processing apparatus 100 according to this embodiment. The image processing apparatus 100 includes, for example, a communication unit 110, a transmission / reception control unit 120, an image processing unit 130, an image conversion unit 140, an image determination unit 150, a trained model generation unit 160, and a storage unit 170. These components are realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or GPU (Graphics Processing Unit), or by the cooperation of software and hardware. The program may be stored in advance on a storage device such as an HDD (Hard Disk Drive) or flash memory (a storage device equipped with a non-transient storage medium), or it may be stored on a removable storage medium such as a DVD or CD-ROM (a non-transient storage medium) and installed when the storage medium is inserted into a drive device. The storage unit 170 is, for example, an HDD, flash memory, RAM (Random Access Memory), etc. The storage unit 170 stores, for example, captured image data 172, converted image data 174, annotation image data 176, annotated image data 178, and a trained model 180. For the sake of explanation, the image processing device 100 is equipped with a trained model generation unit 160 and a storage unit 170 that stores the trained model 180, but the function of generating the trained model and the generated trained model may be held by a server device different from the image processing device 100.
[0030] The communication unit 110 is an interface that communicates with the vehicle M's communication device 10 via the network NW. For example, the communication unit 110 includes a NIC (Network Interface Card) and an antenna for wireless communication.
[0031] The transmission / reception control unit 120 uses the communication unit 110 to transmit and receive data between vehicles M1 and M2 and the terminal device 200. More specifically, the transmission / reception control unit 120 first acquires multiple in-vehicle and out-of-vehicle images captured in time series by a camera mounted on vehicle M1. In this case, the time series refers to images captured at predetermined intervals (for example, every second) during one driving cycle from the start to the stop of vehicle M1.
[0032] Figure 3 shows an example of interior and exterior images acquired from vehicle M1. The left side of Figure 3 represents the interior image acquired from vehicle M1, and the right side of Figure 3 represents the exterior image acquired from vehicle M1. As shown in the left side of Figure 3, the interior image is captured with the camera positioned to capture at least the driver's face area of vehicle M1, and as shown in the right side of Figure 3, the exterior image is captured with the camera positioned to capture at least the area in front of the vehicle M1 in the direction of travel. The transmission / reception control unit 120 associates the interior and exterior images acquired from vehicle M1 with an image ID and stores them in the storage unit 170 as captured image data 172.
[0033] Figure 4 is a diagram illustrating the processing performed by the image processing unit 130. The image processing unit 130 performs image processing on the captured image data 172 and obtains information such as image attributes, face attributes, and orientation for each image contained in the captured image data 172. More specifically, when an image is input, the image processing unit 130 uses a trained model that outputs a classification result indicating whether the image is an interior or exterior image of a vehicle to obtain image attributes indicating whether each image contained in the captured image data 172 is an interior or exterior image of a vehicle.
[0034] Furthermore, when an image is input, the image processing unit 130 uses a trained model that outputs the face region, the size of the face (area of the face region), and the distance from the image capture position to the face for all faces included in the image to acquire the face attributes of each image included in the captured image data 172. In Figure 3, as an example, the face region FA1 of person P1 is acquired from the in-car image, and the face region FA2 of person P2, the face region FA3 of person P3, and the face region FA4 of person P4 are acquired from the out-of-car image. For convenience, the face regions FA1, FA2, FA3, and FA4 are acquired as rectangular regions, but the present invention is not limited to such a configuration, and for example, a trained model that acquires the face region along the contour of a person's face may be used.
[0035] Furthermore, when an image is input, the image processing unit 130 uses a trained model that outputs, for example, a vector, at least one of the face direction and gaze direction for all faces included in the image to acquire the direction information of the faces captured in each image of the captured image data 172. More specifically, for images of the captured image data 172 that have the attributes of an in-car image, the image processing unit 130 uses a trained model that outputs the face direction and gaze direction for all faces included in the image to acquire the direction information. On the other hand, for images of the captured image data 172 that have the attributes of an out-of-car image, the image processing unit 130 uses a trained model that outputs the face direction for all faces included in the image to acquire the direction information. This is because, generally, faces captured in in-car images tend to be closer to the shooting position and larger in size than those captured in out-of-car images, making it possible to extract the gaze direction. In Figure 3, as an example, the face direction FD1 and gaze direction ED1 of person P1 are obtained from the in-car image, while the face direction FD2 of person P2, the face direction FD3 of person P3, and the face direction FD4 of person P4 are obtained from the out-of-car image.
[0036] The image processing unit 130 acquires image attributes, face attributes, and orientation information for each image of the captured image data 172, and records these image attributes, face attributes, and orientation information associated with that image. In the above example, the image processing unit 130 acquires image attributes, face attributes, and orientation information using a trained model, but the present invention is not limited to such a configuration, and the image processing unit 130 may acquire these image attributes, face attributes, and orientation information using any known method.
[0037] The image conversion unit 140 performs a process on the captured image data 172 processed by the image processing unit 130, replacing the faces of the people in each image with the faces of other people, without changing the direction information of the people in each image, using any software that has such a function implemented. Figure 5 is a diagram illustrating the process performed by the image conversion unit 140. As shown in Figure 5, the image conversion unit 140 replaces the faces of people P1, P2, and P3 shown in Figure 4 with the faces of other people without changing the gaze direction ED1 and face directions FD1, FD2, and FD3. On the other hand, the face of person P4 is covered by a mosaic MS as a result of mosaic processing performed by the image conversion unit 140.
[0038] In other words, the image conversion unit 140 determines, based on the face attributes of each face captured in each image of the captured image data 172, whether to replace the face with the face of another person or to apply mosaic processing. More specifically, for each face captured in each image of the captured image data 172, the image conversion unit 140 determines whether the size of the face is greater than or equal to a first threshold Th1. If it is determined that the size of the face is greater than or equal to the first threshold Th1, it decides to replace the face with the face of another person. On the other hand, if it is determined that the size of the face is less than the first threshold Th1, the image conversion unit 140 decides to apply mosaic processing to the face. Replacing the face of a person captured in an image with the face of another person or applying mosaic processing is an example of "anonymization processing".
[0039] Furthermore, the image conversion unit 140 determines whether the distance to each face in each image of the captured image data 172 is less than or equal to the second threshold Th2. If it is determined that the distance to the face is less than or equal to the second threshold Th2, it decides to replace the face with the face of another person. On the other hand, if it is determined that the distance to the face is greater than the second threshold Th2, the image conversion unit 140 decides to apply mosaic processing to the face. The image conversion unit 140 repeatedly performs these determination processes for each face in the image, and replaces each face with the face of another person or applies mosaic processing according to the determination result. The image conversion unit 140 stores the image data obtained by applying such processing to the captured image data 172 as converted image data 174 in the storage unit 170. This allows for the selection of data useful as training data for generating a behavior prediction model, and also protects the privacy of the people in each image when the annotator, described later, performs annotation work.
[0040] Furthermore, the process of determining whether the size of the face is greater than or equal to the first threshold Th1, and the process of determining whether the distance to the face is less than or equal to the second threshold Th2, only need to be performed at least one of these two processes. If both processes are performed, the image conversion unit 140 may decide to replace the face with the face of another person if the size of the face is greater than or equal to the first threshold Th1 and the distance to the face is less than or equal to the second threshold Th2, or it may decide to replace the face with the face of another person if the size of the face is greater than or equal to the first threshold Th1 or the distance to the face is less than or equal to the second threshold Th2.
[0041] Furthermore, the image conversion unit 140 may select faces to be used as training data by applying mosaic processing to faces in each image of the captured image data 172 for which direction information could not be acquired.
[0042] Figure 6 shows an example of a time-series in-vehicle image converted by the image conversion unit 140. As an example, Figure 6 shows an example of converting time-series in-vehicle images at three points in time: t, t+1, and t+2. These time-series in-vehicle images were captured and face-converted of the same person, but as shown in Figure 6, depending on the operation of the face conversion software, the face of the same person may be converted into the faces of multiple different people. It is undesirable to use such converted image data as training data, even if the face of the same person has been converted into the faces of multiple different people, as this will worsen the accuracy of the behavior prediction model. Therefore, the image determination unit 150 determines the continuity of the time-series in-vehicle and out-vehicle images by executing the process described below.
[0043] Figure 7 is a diagram illustrating the processing performed by the image determination unit 150. As shown in Figure 7, the image determination unit 150 first extracts feature points representing the face of a person captured in the converted image. For example, the image determination unit 150 extracts feature points representing the right eye REP, left eye LEP, nose NP, right corner of the mouth RMP, left corner of the mouth LMP, and ear EP from the face of a person captured in the converted image. The image determination unit 150 extracts the feature points of the face of the person tracked as the same person from each of the time-series converted images and matches these feature points. Whether or not a person has been "tracked as the same person" can be determined, for example, by associating the same person captured in the image before the image is converted.
[0044] In the case of Figure 7, the image determination unit 150 extracts the feature points of the person depicted in the transformed image at time t and the feature points of the person depicted in the transformed image at time t+1. The image determination unit 150 performs a comparison by determining whether these two sets of extracted feature points are approximately identical through translation and rotation.
[0045] If the matching results indicate that the extracted feature points are nearly identical, the image determination unit 150 determines that the face of the person tracked as the same person remains the face of the same person after conversion (i.e., there is continuity in the face). On the other hand, if the matching results indicate that the extracted feature points are not nearly identical, the image determination unit 150 determines that the face of the person tracked as the same person is not the face of the same person after conversion (i.e., there is no continuity in the face). In this case, the image conversion unit 140 performs the conversion process again on the face that was determined to lack continuity. At this time, the image conversion unit 140 may perform the conversion process again only on the face that was determined to lack continuity, or it may perform the conversion process again on the faces of all people captured in the time-series converted images. Alternatively, for example, the image conversion unit 140 may apply mosaic processing to the face that was determined to lack continuity without performing the conversion process again, and exclude it from being used as training data. Furthermore, for example, if the image determination unit 150 determines that the faces of individuals tracked as the same person are not the same person's faces after conversion (i.e., there is no continuity in the faces), the image determination unit 150 may restrict the application of predetermined processing to the time-series converted images, that is, exclude the time-series converted images from being used as training data. This prevents the occurrence of discontinuities caused by unintended behavior of the face conversion software.
[0046] The image determination unit 150 further inputs the converted image back into the trained model described above, which outputs at least one of the face direction and gaze direction, to obtain the face direction FD or gaze direction ED in the converted image. The image determination unit 150 determines whether the face direction FD or gaze direction ED of a person's face in the converted image substantially matches the face direction FD or gaze direction ED of the face in the image captured before conversion. As described above, both the face direction FD and gaze direction ED are obtained for in-car images, and only the face direction FD is obtained for out-of-car images. Therefore, for in-car images, the image determination unit 150 determines whether the face direction FD and gaze direction ED substantially match between the image captured before conversion and the converted image, and for out-of-car images, it determines whether the face direction FD substantially matches between the image captured before conversion and the converted image. More specifically, for example, the image determination unit 150 calculates the angular difference between the vector representing the face direction FD in the captured image before conversion and the vector representing the face direction FD in the converted image, and determines that the face direction FD is approximately the same if the calculated angular difference is within a threshold. The same applies to the gaze direction ED. The continuity of the face or the consistency of the direction information is one example of a "predetermined requirement."
[0047] If the image conversion unit 140 determines that the face direction FD or gaze direction ED does not substantially match between the captured image before conversion and the converted image, it performs the conversion process again on the captured image for the faces for which the face direction FD or gaze direction ED is determined to not substantially match. At this time, the image conversion unit 140 may perform the conversion process again only on the faces for which the substantially match is determined to not match, or it may perform the conversion process again on all faces included in the converted image that include the faces for which the substantially match is determined to not match. Alternatively, for example, the image conversion unit 140 may apply mosaic processing to the faces for which the substantially match is determined to not match without performing the conversion process again, and exclude them from being used as training data. Alternatively, for example, if the image determination unit 150 determines that the faces do not substantially match, it may restrict the application of predetermined processing to the time-series converted images, that is, exclude the time-series converted images from being used as training data. This prevents information degradation caused by unintended operation of the face conversion software.
[0048] Furthermore, if there are multiple faces in the converted image (or if the number of faces in the converted image exceeds a predetermined value), the continuity determination process and the direction information consistency determination process performed by the image determination unit 150 described above may be performed only on faces that are assumed to be of higher importance, rather than on all faces in the converted image. As an example of faces assumed to be of higher importance, the image determination unit 150 may perform these determination processes only on faces in the pre-conversion image whose face size is greater than the first threshold Th1 (third threshold Th3 or higher), or on faces whose face distance is less than the second threshold Th2 (fourth threshold Th4 or lower). Also, for example, the image determination unit 150 may assume that faces of people located in front of the vehicle M1 in the direction of travel, or faces of people whose face direction is facing forward in the direction of travel of the vehicle M1, are of higher importance, and perform these determination processes. Furthermore, for example, if continuity or consistency is denied for a particular face in the converted image, a re-conversion process may be performed on that face and on faces that are considered to be of high importance.
[0049] When the image determination unit 150 confirms the continuity and consistency of the time-series converted images, it stores the converted image data 174, which has been confirmed to have continuity and consistency, in the storage unit 170 as annotation image data 176. At this time, the converted image data 174 may be stored in the storage unit 170 as annotation image data 176 along with information indicating the purpose of use, and information indicating that it is annotation image data for generating a behavior prediction model that predicts the actions of people depicted in the input image. The transmission / reception control unit 120 transmits the annotation image data 176 to the terminal device 200. The annotator, who is a user of the terminal device 200, generates annotated image data by performing annotation work on the annotation images contained in the received annotation image data 176 and transmits it to the image processing device 100. The image processing device 100 stores the received annotated image data in the storage unit 170 as annotated image data 178.
[0050] Furthermore, the determination process regarding the continuity of the converted image and the determination process regarding the consistency of the face direction information, both performed by the image determination unit 150 described above, only need to be performed by at least one of them. If at least one of continuity and consistency is met, the converted image data 174 may be stored in the storage unit 170 as annotation image data 176.
[0051] Furthermore, the image determination unit 150 does not have to store all of the time-series captured images (or their converted images) obtained at predetermined intervals (e.g., every second) in one driving cycle as annotation image data 176 in the storage unit 170 if any images are missing due to camera malfunction or the like.
[0052] Figure 8 shows an example of annotation work performed by an annotator. The left side of Figure 8 shows annotation on a converted image of an interior image, and the right side of Figure 8 shows annotation on a converted image of an exterior image. The annotator adds information to the converted interior image, for example, indicating whether the driver's gaze direction ED1 shown in the converted interior image is appropriate in the situation shown in the converted exterior image at the same time (e.g., 1 if appropriate, 0 if inappropriate). For example, in Figure 8, the converted exterior image shows a pedestrian on the left side in the direction of vehicle travel, while the converted interior image shows the driver looking to the left. In other words, since it is assumed that the driver is paying appropriate attention to the pedestrian, the annotator adds information (i.e., 1) indicating that the driver's gaze direction ED1 is appropriate.
[0053] Furthermore, the annotator specifies, for example, the risk area RA of the person depicted in the converted image of the exterior of the vehicle, excluding the person who has been blurred, where the person is expected to be moving. Because the image conversion unit 140 and the image determination unit 150 process the faces of the people depicted in the original image to be replaced with the faces of other people, the privacy of those people is protected. At the same time, since the direction of the person's face and gaze are maintained even after conversion, the annotator can accurately specify the risk area RA by referring to the direction of the face and gaze of the other person depicted in the converted image. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of the person depicted in the face image.
[0054] When the annotated image data 178 is stored in the storage unit 170, the trained model generation unit 160 uses the annotated image data 178 as training data and generates a trained model using an arbitrary machine learning model. As described above, this trained model is, for example, a behavior prediction model that, in response to an external image input, outputs the predicted behavior (trajectory) of a person shown in the external image, or, in response to internal and external images input, takes into account the driver's gaze shown in the internal image and prompts attention to a pedestrian shown in the external image. The trained model generation unit 160 stores the generated trained model as trained model 180 in the storage unit 170.
[0055] When a trained model 180 is generated, the transmission / reception control unit 120 distributes the generated trained model 180 to the vehicle M2 via the NetArc NW. Upon receiving the trained model 180, the vehicle M2 uses the trained model 180 (more precisely, an application program that utilizes the trained model 180) to provide driving assistance to the driver of the vehicle M2.
[0056] Figure 9 shows an example of driver assistance using a trained model 180. Figure 9 illustrates an example where, while the vehicle M2 is in motion, it inputs in-vehicle and out-of-vehicle images captured by its onboard cameras into the trained model 180. The trained model 180 then considers the driver's gaze as seen in the in-vehicle image and outputs information to the HMI (human-machine interface) prompting the driver to pay attention to pedestrians seen in the out-of-vehicle image, thereby providing driver assistance. As shown in Figure 9, for example, the HMI displays a risk area RA2 corresponding to pedestrian P5 in the out-of-vehicle image, and if the driver's gaze as seen in the in-vehicle image is not directed towards pedestrian P5, it outputs a warning message ("Please be careful of distracted driving") as text or audio information. This enables driver assistance that takes the driver's condition into consideration.
[0057] Next, the processing flow performed by the image processing device 100 will be described with reference to Figures 10 and 11. Figure 10 is a diagram showing an example of the processing flow performed by the image conversion unit 140. The processing shown in Figure 10 is performed, for example, when an in-vehicle image or an out-of-vehicle image is captured by a camera mounted on the vehicle M1 and processed by the image processing unit 130.
[0058] First, the image conversion unit 140 acquires an image captured from the image data 172 that has been processed by the image processing unit 130 (step S100). Next, the image conversion unit 140 selects one face to be captured in the acquired image captured (step S102).
[0059] Next, the image conversion unit 140 determines whether the size of the selected face is greater than or equal to the first threshold Th1 (step S104). If it is determined that the size of the selected face is greater than or equal to the first threshold Th1, the image conversion unit 140 converts the face into the face of another person (step S106). On the other hand, if it is determined that the size of the selected face is less than the first threshold Th1, the image conversion unit 140 then determines whether the distance of the selected face is less than or equal to the second threshold Th2 (step S108).
[0060] If the distance of the selected face is determined to be less than or equal to the second threshold Th2, the image conversion unit 140 proceeds to step S106 and converts the face into the face of another person. On the other hand, if the distance of the selected face is determined to be greater than the second threshold Th2, the image conversion unit 140 applies mosaic processing to the face (step S110). Next, the image conversion unit 140 determines whether or not processing has been performed on all faces captured in the acquired image (step S112).
[0061] If it is determined that processing has been performed on all faces in the acquired image, the image conversion unit 140 acquires the image obtained by processing all faces as the converted image and stores it in the storage unit 170 as converted image data 174 (step S114). On the other hand, if it is determined that processing has not been performed on all faces in the acquired image, the image conversion unit 140 returns to step S102. This completes the processing in this flowchart.
[0062] Figure 11 shows an example of the processing flow performed by the image determination unit 150. The processing shown in Figure 11 is performed, for example, when a time-series transformed image is obtained by applying the above transformation processing to the time-series captured images taken during one driving cycle from the start to the stop of the vehicle M1.
[0063] First, the image determination unit 150 acquires a time-series converted image (step S200). Next, the image determination unit 150 selects the face of the person who was tracked as the same person before the conversion from the acquired time-series converted image (step S202).
[0064] Next, the image determination unit 150 extracts feature points from each of the time-series converted images and compares them to determine whether these faces are the same after conversion (step S204). If it is determined that the faces are the same after conversion, the image determination unit 150 then determines whether the acquired time-series converted images are in-vehicle images (step S206). On the other hand, if it is determined that the faces are not the same, the image determination unit 150 instructs the image conversion unit 140 to convert the faces of the people who were tracked as the same person before conversion in the time-series captured images again (step S208). After that, the image determination unit 150 executes the process in step S204 again on the converted faces.
[0065] In step S206, if the acquired time-series converted image is determined to be an in-vehicle image, the image determination unit 150 determines whether the gaze direction and face direction of these faces match the image before conversion (step S210). On the other hand, if the acquired time-series converted image is determined to be not an in-vehicle image, i.e., an out-of-vehicle image, the image determination unit 150 determines whether the face direction of these faces matches the image before conversion (step S212). If it is determined that they do not match in the processing of step S210 or step S212, the image determination unit 150 proceeds to step S208.
[0066] If a match is determined in step S210 or step S212, the image determination unit 150 determines that these faces have been successfully converted and determines whether processing has been performed on all faces captured in the time-series converted images (step S214). If it is determined that processing has been performed on all faces captured in the time-series converted images, the image determination unit 150 acquires these time-series converted images as annotation images and instructs the transmission / reception control unit 120 to transmit the acquired annotation images to the terminal device 200 (step S216). On the other hand, if it is determined that processing has not been performed on all faces captured in the time-series converted images, the image determination unit 150 returns to step S202. This completes the processing of this flowchart.
[0067] As described above, according to this embodiment, when it is determined that multiple anonymized input images satisfy predetermined requirements, predetermined processing is applied to the multiple anonymized input images, and said anonymization processing includes processing to change the faces of people depicted in the multiple input images to the faces of different people, and said predetermined requirements include that the faces of people tracked as the same person, as depicted in the multiple anonymized input images, are the faces of the same person obtained through the anonymization processing. In other words, in this embodiment, faces that were the same person before anonymization processing are guaranteed to remain the faces of the same person after anonymization processing and are used as training data. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of people depicted in face images.
[0068] Furthermore, according to this embodiment, the predetermined requirement includes the fact that the facial orientation information of a person tracked as the same person in multiple input images matches the facial orientation information of the same person in the multiple input images that have undergone the anonymization process. In other words, in this embodiment, it is guaranteed that the facial orientation information of the same person remains unchanged even after anonymization. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of the person captured in the facial image.
[0069] Furthermore, according to this embodiment, the predetermined requirements are determined according to the image attributes, which are the shooting patterns of the multiple input images. That is, in this embodiment, a predetermined process is executed, which is, for example, a process that saves each of the shooting patterns of the multiple input images as training information for generating a behavior prediction model. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of the person captured in the facial image.
[0070] Furthermore, according to this embodiment, it is determined whether to perform anonymization processing using the first method or using a second method different from the first method, based on the size of the face captured in each of the multiple input images or the distance from the shooting location to the face. In other words, in this embodiment, the method of anonymization processing applied to faces is changed depending on whether or not it is useful for training a machine learning model. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of the person captured in the face image.
[0071] [Differentiation] As described above, in this embodiment, if the image determination unit 150 determines that the faces in the converted image do not meet predetermined requirements, an example has been described in which the converted image is re-converted or mosaic processing is applied. However, if the image determination unit 150 determines that the predetermined requirements are not met, the image determination unit 150 may not perform the predetermined processing on the converted image, that is, it may restrict the application of the predetermined processing (such as not storing the image or sending it to the server).
[0072] Furthermore, in this embodiment, an example was described in which the image processing device 100 is implemented as a server device separate from the vehicle M1. However, as a modification of this embodiment, the image processing device 100, more specifically, a device having at least the functions of an image processing unit 130, an image conversion unit 140, and an image determination unit 150, may be mounted on the vehicle M1 as an in-vehicle device. In that case, the in-vehicle device applies the image processing unit 130 described above to the image captured by the in-vehicle camera, performs anonymization by the image conversion unit 140, and performs determination by the image determination unit 150. After that, the in-vehicle device transmits the anonymized image, whose facial continuity and directional information consistency have been confirmed by the image determination unit 150, to an external image server.
[0073] When the image server receives an anonymized image from vehicle M1, it stores the received anonymized image in its storage unit as annotation image data, and either transmits the annotation image data to the annotator's terminal device 200, or allows the terminal device 200 to access the annotation image data. When the image server receives annotated image data from terminal device 200, it generates a trained model 180 based on the annotated image data and distributes the generated trained model 180 to vehicle M2. In this way, as in the embodiment, it is possible to generate training data that is effective for training a machine learning model while protecting the privacy of the person depicted in the face image. Furthermore, according to this modified example, since the in-vehicle device performs anonymization processing on the image and then transmits the anonymized image to the image server, the privacy of the person depicted in the face image can be protected even more reliably.
[0074] Furthermore, in another embodiment, the in-vehicle device may have only some of the functions of the image processing unit 130, the image conversion unit 140, and the image determination unit 150, while the image server has the remaining functions. For example, the in-vehicle device may have the functions of the image processing unit 130 and the image conversion unit 140, and the image server may have the function of the image determination unit 150, or the in-vehicle device may have the function of the image processing unit 130, and the image server may have the functions of the image conversion unit 140 and the image determination unit 150.
[0075] The embodiments described above can be expressed as follows. A storage medium that stores computer-readable instructions, A processor connected to the storage medium, The processor executes the computer-readable instructions to: Anonymization is performed on multiple input images captured in time series. It is determined whether the plurality of input images that have undergone the anonymization process satisfy predetermined requirements. If it is determined that the plurality of input images that have undergone the anonymization process satisfy the predetermined requirements, the plurality of input images that have undergone the anonymization process are subjected to the predetermined processing, The anonymization process includes a process of changing the faces of people depicted in the multiple input images to the faces of other people. The aforementioned predetermined requirement includes the fact that the face of the person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images that have undergone the anonymization process. An image processing device configured in such a way.
[0076] Although embodiments for carrying out the present invention have been described above using examples, the present invention is not limited in any way to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention. [Explanation of Symbols]
[0077] 100 Image Processing Devices 110 Communications Department 120 Transmit / Receive Control Unit 130 Image Processing Unit 140 Image conversion unit 150 Image determination unit 160 Pre-trained model generation unit 170 Storage section 172 Image data 174 Converted Image Data 176 Image data for annotation 178 Annotated image data 180 pre-trained models
Claims
1. An image conversion unit that performs anonymization processing on multiple input images captured in a time series, The system includes an image determination unit that determines whether the plurality of input images subjected to the anonymization process satisfy predetermined requirements, If the image determination unit determines that the plurality of input images that have undergone the anonymization process satisfy the predetermined requirements, it performs the predetermined processing on the plurality of input images that have undergone the anonymization process. The anonymization process includes a process of changing the faces of people depicted in the multiple input images to the faces of other people. The aforementioned predetermined requirement includes the fact that the face of the person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images that have undergone the anonymization process. Image processing device.
2. The predetermined process is the process of saving the multiple input images that have undergone the anonymization process as images to be used for annotation work. The image processing apparatus according to claim 1.
3. The predetermined process is the process of saving all of the consecutive input images that have undergone the anonymization process as the images to be used for the annotation work. The image processing apparatus according to claim 2.
4. The predetermined process involves saving the multiple input images that have undergone the anonymization process as training information for generating a behavior prediction model that predicts the actions of people depicted in the input images. The image processing apparatus according to claim 1.
5. The predetermined process is the process of transmitting the multiple input images that have undergone the anonymization process to an image server via a communication means. The image processing apparatus according to claim 1.
6. The image determination unit extracts facial feature points of the person being tracked as the same person from each of the multiple anonymized input images, and determines that the faces of the people being tracked as the same person are the same face if the positional relationship of the extracted feature points matches. The image processing apparatus according to any one of claims 1 to 5.
7. The image determination unit, when multiple faces of individuals are present in each of the multiple anonymized input images, determines whether the predetermined requirements are met for the faces of individuals among the multiple individuals that are facing forward in the direction of travel of the vehicle on which the camera that captured the input image is mounted. The image processing apparatus according to any one of claims 1 to 5.
8. The image determination unit, when multiple faces of multiple people are present in each of the multiple input images that have undergone the anonymization process, determines whether the predetermined requirements are met for the faces of the multiple people whose faces in the input images before the anonymization process meet the predetermined criteria. The image processing apparatus according to any one of claims 1 to 5.
9. If the image conversion unit determines, based on the image determination unit, that the plurality of input images that have undergone the anonymization process do not meet the predetermined requirements, the image conversion unit applies the anonymization process to the plurality of input images again. The image processing apparatus according to any one of claims 1 to 5.
10. If the image conversion unit determines, based on the image determination unit, that the plurality of input images that have undergone the anonymization process do not satisfy the predetermined requirements, the image conversion unit will not apply the predetermined processing to the plurality of input images that have undergone the anonymization process. The image processing apparatus according to any one of claims 1 to 5.
11. An image conversion unit that performs anonymization processing on multiple input images captured in a time series, The system includes an image determination unit that determines whether the plurality of input images subjected to the anonymization process satisfy predetermined requirements, If the image determination unit determines that the plurality of input images that have undergone the anonymization process satisfy the predetermined requirements, it performs the predetermined processing on the plurality of input images that have undergone the anonymization process. The anonymization process includes a process of changing the faces of people depicted in the multiple input images to the faces of other people. The aforementioned predetermined requirement includes the fact that the face of the person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images that have undergone the anonymization process. Image processing system.
12. Computers Anonymization is performed on multiple input images captured in time series. It is determined whether the plurality of input images that have undergone the anonymization process satisfy predetermined requirements. If it is determined that the plurality of input images that have undergone the anonymization process satisfy the predetermined requirements, the plurality of input images that have undergone the anonymization process are subjected to the predetermined processing, The anonymization process includes a process of changing the faces of people depicted in the multiple input images to the faces of other people. The aforementioned predetermined requirement includes the fact that the face of the person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images that have undergone the anonymization process. Image processing methods.
13. On the computer, Anonymization processing is performed on multiple input images captured in a time series. The system determines whether the plurality of input images that have undergone the anonymization process satisfy predetermined requirements. If it is determined that the plurality of input images that have undergone the anonymization process satisfy the predetermined requirements, the plurality of input images that have undergone the anonymization process are subjected to the predetermined processing. The anonymization process includes a process of changing the faces of people depicted in the multiple input images to the faces of other people. The aforementioned predetermined requirement includes the fact that the face of the person tracked as the same person in the plurality of input images is the same face in each of the plurality of input images that have undergone the anonymization process. program.
Citation Information
Patent Citations
Method for producing base material for additive for casting of steel
JP1984030450A
Person tracking device, person tracking method and person tracking program
JP2007219603A
Information processing device and program
JP2014085796A
Vehicular pedestrian image acquisition system
JP2016126597A
Image processing system, information processing device, and program
JP2017187850A