Image processing device, image processing method, and program

By extracting and anonymizing the facial area of ​​a person in an image processing device, the problem of false detection in the prior art is solved, and simple and reliable anonymization processing and privacy protection are achieved.

CN120660109APending Publication Date: 2025-09-16HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380094644.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the prior art, it is difficult to easily and reliably detect the facial region of an image to be anonymized, and this is particularly prone to false detection when the texture of the person and the background are similar.

Method used

An image processing device is used to extract the human area from the input image through the extraction unit, and the facial area is extracted based on the position and key point detection method. Anonymization processing is performed in combination with the learned model, including replacement or mosaic processing, to ensure the consistency of facial direction.

Benefits of technology

This makes it easier and more reliable to detect and anonymize facial regions in images, protecting privacy and maintaining the accuracy of action prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660109A_ABST
    Figure CN120660109A_ABST
Patent Text Reader

Abstract

An image processing apparatus includes: an extraction unit that extracts a face region of a person from an input image; and a conversion unit that performs anonymization processing on the extracted face region, the extraction unit extracts a person region representing the person from the input image, and extracts the face region from only the extracted person region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing device, an image processing method, and a program. Background Art

[0002] Conventionally, technologies for identifying a person's facial region based on an image captured by a camera are known. For example, Patent Document 1 describes preventing a wall pattern in front of an image input device from being mistakenly detected as a face, even if the pattern is similar to that of a person's face.

[0003] Prior art literature

[0004] Patent Literature

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2005-242777 Summary of the Invention

[0006] Problems to be solved by the invention

[0007] The technology described in Patent Document 1 reliably detects a person's face by lighting a light that illuminates the person's face when a person-sensing sensor detects their presence. However, conventional technologies sometimes fail to easily and reliably detect the facial area of ​​an image to be anonymized.

[0008] The present invention has been made in consideration of such circumstances, and one object of the present invention is to provide an image processing device, an image processing method, and a program that can more easily and reliably detect a face region in an image to be anonymized.

[0009] Means for solving problems

[0010] The image processing device, image processing method, and program according to the present invention have the following configurations.

[0011] (1): One embodiment of the present invention relates to an image processing device, wherein the image processing device comprises: an extraction unit that extracts a facial region of a person from an input image; and a conversion unit that performs anonymization processing on the extracted facial region, wherein the extraction unit extracts a person region representing the person from the input image and extracts the facial region only from the extracted person region.

[0012] (2): In the above-mentioned aspect (1), in the image processing device, the extraction unit further extracts one or more part regions representing body parts of the person from the person region, and extracts the face region from the person region based on positions of the one or more part regions.

[0013] (3): In the above-mentioned aspect (1), the conversion unit performs the anonymization process when the size of the extracted facial region is greater than or equal to a first threshold value and the distance from the shooting position of the input image to the facial region is less than or equal to a second threshold value.

[0014] (4): In the above-mentioned aspect (1), the anonymization process is a process of making the direction of the face of the person consistent before and after the anonymization process, and changing the face of the person to the face of another person.

[0015] (5): In the scheme of (4) above, the image processing device further includes a learning unit, which obtains an annotated image obtained by adding annotations to the input image that has been anonymized, and the annotation indicates whether the direction of the face of the person driving the vehicle is appropriate. The learning unit uses the annotated image as learning data to generate a learned model, and the learned model is used to prompt the person to call attention to pedestrians outside the vehicle.

[0016] (6): Another embodiment of the present invention relates to an image processing method, wherein the image processing method causes a computer to perform the following processing: extracting a facial region of a person from an input image; and anonymizing the extracted facial region, wherein the extraction refers to extracting a person region representing the person from the input image, and extracting the facial region only from the extracted person region.

[0017] (7): Another embodiment of the present invention relates to a program, wherein the program causes a computer to perform the following processing: extracting a facial region of a person from an input image; and anonymizing the extracted facial region, wherein the extraction refers to extracting a person region representing the person from the input image, and extracting the facial region only from the extracted person region.

[0018] Effects of the Invention

[0019] According to the above-mentioned configurations (1) to (7), the face region of an image to be anonymized can be detected more easily and reliably. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 1 is a diagram showing an overview of a system 1 including an image processing apparatus 100 according to this embodiment.

[0021] Figure 2 1 is a diagram showing an example of the functional configuration of the image processing device 100 according to the present embodiment.

[0022] Figure 3 1 is a diagram showing an example of an in-vehicle image and an out-vehicle image acquired from the vehicle M1.

[0023] Figure 4 It is a diagram for explaining the processing executed by the image processing unit 130 .

[0024] Figure 5 It is a diagram for explaining the processing performed by the image conversion unit 140.

[0025] Figure 6 This is a diagram for explaining failure of processing performed by the conventional image processing unit 130 .

[0026] Figure 7 It is a diagram for explaining the processing executed by the image processing unit 130 according to this embodiment.

[0027] Figure 8 This is a diagram showing an example of annotation work performed by an annotator.

[0028] Figure 9 1 is a diagram showing an example of driving support performed using the learned model 180 .

[0029] Figure 10 1 is a diagram showing an example of the flow of processing executed by the image processing unit 130 and the image conversion unit 140 . DETAILED DESCRIPTION

[0030] Hereinafter, embodiments of an image processing device, an image processing method, an image processing system, and a program according to the present invention will be described with reference to the accompanying drawings.

[0031] [summary]

[0032] Figure 1 1 is a diagram showing an overview of a system 1 including an image processing apparatus 100 according to this embodiment. Figure 1 As shown, the system 1 includes at least one vehicle M1 and vehicle M2, an image processing device 100, and a terminal device 200. For ease of explanation, the vehicle M1 and vehicle M2 are shown as different vehicles, but these vehicles may be the same.

[0033] Vehicle M1 is, for example, a hybrid vehicle, an electric vehicle, or the like, and includes at least one camera for capturing images of the interior and exterior of vehicle M1. Vehicle M1 transmits images of the interior and exterior captured by these cameras while driving to image processing device 100 via a network NW such as a cellular network, Wi-Fi network, or the Internet.

[0034] Image processing device 100 is a server device that, upon receiving captured image data including interior and exterior images from vehicle M1, performs image conversion, described below, on the received captured image data. This image conversion is performed to protect the privacy of people reflected in the interior and exterior images. Image processing device 100 transmits the resulting converted image data to terminal device 200 via network NW.

[0035] The terminal device 200 is a terminal device such as a desktop personal computer or a smartphone. When the user of the terminal device 200 obtains transformed image data from the image processing device 100, they perform an annotation operation (described below) on the obtained transformed image data. When the annotation operation is completed, the user of the terminal device 200 transmits the annotated image data, in which the annotations are applied to the transformed image data, to the image processing device 100.

[0036] When the image processing device 100 receives annotated image data from the terminal device 200, it uses the received annotated image data as learning data and generates a learned model (described later) using an arbitrary machine learning model. This learned model is, for example, a behavior prediction model that, given an input image outside the vehicle, outputs a predicted behavior (trajectory) of a person reflected in the image outside the vehicle, or that, given an input image inside and outside the vehicle, draws attention to a pedestrian reflected in the image outside the vehicle, taking into account the driver's line of sight.

[0037] It should be noted that the image data used as learning data in this case can be annotated image data in which annotations are assigned to transformed image data, or annotated image data in which the transformed image data is converted back to captured image data while the annotations remain unchanged (i.e., annotated image data in which annotations are assigned to captured image data). By using annotated image data in which annotations are assigned to captured image data as learning data, more realistic learning data can be used, in which the effects of image transformation are eliminated.

[0038] When the image processing device 100 generates a learned model, it distributes it to vehicle M2 via the network NW. Similar to vehicle M1, vehicle M2 is a hybrid vehicle, electric vehicle, or other similar vehicle. Vehicle M2 inputs at least one of the vehicle's interior and exterior images captured by a camera while driving into the learned model to obtain predicted behavior data for people around vehicle M2. The driver of vehicle M2 can refer to this predicted behavior data and apply it to driving vehicle M2. The following describes each process in more detail.

[0039] [Functional Structure of Image Processing Device]

[0040] Figure 2This figure shows an example of the functional structure of an image processing device 100 according to this embodiment. The image processing device 100 includes, for example, a communication unit 110, a transmission and reception control unit 120, an image processing unit 130, an image conversion unit 140, an image determination unit 150, a learned model generation unit 160, and a storage unit 170. These components are implemented by executing a program (software) on a hardware processor such as a CPU (Central Processing Unit). Some or all of these components can be implemented by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or through the coordinated implementation of software and hardware. The program can be pre-stored on a storage device (including a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or stored on a removable storage medium (non-transitory storage medium) such as a DVD or CD-ROM, and installed by attaching the storage medium to a drive. The storage unit 170 is, for example, an HDD, flash memory, or random access memory (RAM). The storage unit 170 stores, for example, captured image data 172, transformed image data 174, annotation image data 176, annotated image data 178, and a learned model 180. For ease of explanation, the image processing device 100 includes the learned model generation unit 160 and the storage unit 170 that stores the learned model 180. However, the function of generating the learned model and the generated learned model may be stored in a server device separate from the image processing device 100.

[0041] The communication unit 110 is an interface for communicating with the communication device 10 of the host vehicle M via the network NW. For example, the communication unit 110 includes a NIC (Network Interface Card), an antenna for wireless communication, and the like.

[0042] The transmission / reception control unit 120 uses the communication unit 110 to transmit and receive data with vehicles M1 and M2 and the terminal device 200. More specifically, the transmission / reception control unit 120 first obtains from vehicle M1 a plurality of vehicle interior and exterior images captured in a time series by a camera mounted on vehicle M1. In this case, the time series refers to, for example, images captured at predetermined intervals (e.g., every one second) during a driving cycle from the start to the stop of vehicle M1.

[0043] Figure 3 1 is a diagram showing an example of an in-vehicle image and an out-vehicle image acquired from the vehicle M1. Figure 3 The left part of represents the in-vehicle image obtained from vehicle M1. Figure 3 The right side of represents the vehicle exterior image obtained from vehicle M1. Figure 3 As shown in the left part of , the in-vehicle image is captured by a camera in a manner that captures at least the facial area of ​​the driver of the vehicle M1. Figure 3 As shown in the right portion of FIG, the vehicle exterior image is captured by a camera installed so as to capture at least the area in front of the vehicle M1 in its traveling direction. The transmission and reception control unit 120 associates the interior and exterior images acquired from the vehicle M1 with image IDs and stores them in the storage unit 170 as captured image data 172.

[0044] Figure 4 This diagram illustrates the processing performed by image processing unit 130. Image processing unit 130 performs image processing on captured image data 172, acquiring (extracting) information such as image attributes, facial attributes, and orientation for each image included in captured image data 172. More specifically, image processing unit 130 uses a learned model that, when input, outputs a classification result indicating whether the image is an image inside a vehicle or outside a vehicle, to acquire image attributes indicating whether each image included in captured image data 172 is an image inside a vehicle or outside a vehicle. Image processing unit 130 is an example of the "extraction unit" in the technical claims.

[0045] Furthermore, the image processing unit 130 uses a learned model that outputs the facial area, the size of the face (the area of ​​the facial area), and the distance from the image shooting position to the face for all faces included in the image when an image is input, to obtain facial attributes of each image included in the captured image data 172. Figure 3 In this example, facial area FA1 of person P1 is acquired from an image inside the vehicle, while facial area FA2 of person P2, facial area FA3 of person P3, and facial area FA4 of person P4 are acquired from an image outside the vehicle. While facial areas FA1, FA2, FA3, and FA4 are acquired as rectangular areas, the present invention is not limited to such a configuration. For example, a learned model that acquires facial areas along the contours of a person's face may also be used.

[0046] Furthermore, the image processing unit 130 uses a learned model that outputs at least one of the facial direction and the gaze direction, for example, as a vector for all faces included in the image when the image is input, to obtain direction information of the face reflected in each image included in the captured image data 172. More specifically, the image processing unit 130 uses a learned model that outputs the facial direction and the gaze direction for all faces included in the image when the image is input, to obtain direction information for the image of the captured image data 172 having the attribute of an in-vehicle image. On the other hand, the image processing unit 130 uses a learned model that outputs the facial direction for all faces included in the image when the image is input, to obtain direction information for the image of the captured image data 172 having the attribute of an out-of-vehicle image. This is because, generally speaking, compared to out-of-vehicle images, faces reflected in in-vehicle images are closer to the shooting position, and there is a high tendency for faces to be reflected to a large extent to which the gaze direction can be extracted. Figure 3 As an example, the face direction FD1 and the sight line direction ED1 of the person P1 are obtained from the in-vehicle image, and the face directions FD2, FD3, and FD4 of the person P2, P3, and P4 are obtained from the out-vehicle image.

[0047] When the image processing unit 130 obtains image attributes, facial attributes, and orientation information for each image in the captured image data 172, it associates the image attributes, facial attributes, and orientation information with the image. It should be noted that, in the above example, the image processing unit 130 uses a learned model to obtain image attributes, facial attributes, and orientation information. However, the present invention is not limited to such a configuration, and the image processing unit 130 may also use any known method to obtain the image attributes, facial attributes, and orientation information.

[0048] The image conversion unit 140 uses arbitrary software equipped with such a function to execute processing for replacing the face of a person shown in each image of the captured image data 172 processed by the image processing unit 130 with the face of another person without changing the direction information of the person. Figure 5 1 is a diagram for explaining the processing performed by the image conversion unit 140. Figure 5 As shown, the image conversion unit 140 is Figure 4 The faces of persons P1, P2, and P3 are replaced with faces of other persons without changing their eye direction ED1 and facial directions FD1, FD2, and FD3. On the other hand, the face of person P4 is mosaiced by the image conversion unit 140 and is covered by the mosaic MS.

[0049] Specifically, the image conversion unit 140 determines whether to replace each face included in each image of the captured image data 172 with the face of another person or to perform mosaic processing on the face based on the facial attributes of the face. More specifically, the image conversion unit 140 determines whether the size of each face included in each image of the captured image data 172 is greater than a first threshold value Th1. If the size of the face is determined to be greater than the first threshold value Th1, the image conversion unit 140 determines to replace the face with the face of another person. On the other hand, if the size of the face is determined to be less than the first threshold value Th1, the image conversion unit 140 determines to perform mosaic processing on the face. Replacing the face of a person included in a captured image with the face of another person or performing mosaic processing is an example of "anonymization processing."

[0050] Furthermore, the image conversion unit 140 determines whether the distance to each face reflected in each image of the captured image data 172 is less than or equal to a second threshold value Th2. If the distance to the face is determined to be less than the second threshold value Th2, the image conversion unit 140 determines to replace the face with the face of another person. On the other hand, if the distance to the face is determined to be greater than the second threshold value Th2, the image conversion unit 140 determines to perform mosaic processing on the face. The image conversion unit 140 repeats the above determination process for the number of faces reflected in the image and, depending on the determination result, replaces each face with another person's face or performs mosaic processing on the face. The image conversion unit 140 stores the image data obtained by performing this processing on the captured image data 172 as transformed image data 174 in the storage unit 170. This allows the privacy of the people reflected in each image to be protected when selecting useful data for learning data for generating an action prediction model and when the annotator, described later, performs annotation work.

[0051] It should be noted that at least one of the process of determining whether the face size is greater than or equal to the first threshold value Th1 and the process of determining whether the face distance is less than or equal to the second threshold value Th2 may be performed. When both processes are performed, the image conversion unit 140 may decide to replace the face with the face of another person if the face size is greater than or equal to the first threshold value Th1 and the face distance is less than or equal to the second threshold value Th2, or may decide to replace the face with the face of another person if the face size is greater than or equal to the first threshold value Th1 or the face distance is less than or equal to the second threshold value Th2.

[0052] Furthermore, the image conversion unit 140 may perform mosaic processing on faces for which acquisition of direction information failed among faces included in each image of the captured image data 172 , thereby filtering out faces to be used as learning data.

[0053] Thus, when acquiring transformed image data 174 from captured image data 172, image processing unit 130 uses the learned model as described above to acquire information such as the facial region and facial orientation of the person included in the image. However, other software using the learned model or similar functionality may sometimes output areas that are not the person's face as facial regions due to factors such as wrinkles, stains, or brightness on clothing worn by the person, or cracks or dirt on roads or walls.

[0054] Figure 6 This is a diagram for explaining failure of processing performed by the conventional image processing unit 130 . Figure 6 1 and 2 represent the facial regions FA1 to FA6 obtained by the conventional image processing unit 130 by inputting the in-vehicle image and the out-vehicle image into the learned model. Figure 6 As shown, for example, facial area FA5 may be mistakenly detected due to wrinkles, creases, or brightness of clothing worn by a person in the interior image, while facial area FA6 may be mistakenly detected due to cracks or dirt on the road in the exterior image. Using transformed image data 174 containing such mistakenly detected facial areas FA5 and FA6 as training data directly can degrade the accuracy of the behavior prediction model and is therefore undesirable. Therefore, the image processing unit 130 of this embodiment performs the processing described below to prevent misdetection of facial areas and maintain the accuracy of the behavior prediction model.

[0055] Figure 7 1 is a diagram for explaining the processing performed by the image processing unit 130 involved in this embodiment. First, the image processing unit 130 extracts the human figure area HA representing the entire human figure reflected in these images from the obtained in-vehicle image and the outside-vehicle image by any method. The image processing unit 130 can extract these human figure areas HA using, for example, the following learned model, which is learned by outputting a bounding box that surrounds the entire human figure reflected in the image when the image is input. As a result, for example, Figure 7 As shown, human figure areas HA1 to HA4 are obtained.

[0056] When extracting the person area HA, the image processing unit 130 obtains the above-mentioned facial area FA only from the extracted person area HA, and obtains directional information such as the facial direction and the sight line direction from the obtained facial area FA. More specifically, for example, the image processing unit 130 obtains the facial area FA using the following learned model, which is obtained by learning in a manner that when an image in which the person area HA is designated is input, the facial area FA included in the person area HA is output. Next, the image processing unit 130 obtains the directional information of the face using the following learned model, which is obtained by learning in a manner that when an image in which the facial area FA is designated is input, at least one of the facial direction and the sight line direction of the facial area FA is output as, for example, a vector. By such processing, for example, it is possible to prevent Figure 6 The illustrated face area FA6 (ie, a face area erroneously detected due to, for example, cracks carved on the road, dirt attached thereto, etc.) is erroneously detected.

[0057] The image processing unit 130 may further extract one or more body regions (e.g., shoulders, hands, knees, feet, etc.) representing the body parts of the person from the extracted person region HA using any key point detection method, and obtain the face region FA only from the extracted person region HA based on the positions of the one or more body regions. For example, Figure 7 As shown, the image processing unit 130 may also extract two shoulders from the human area HA1 as key points KP1 and KP2, and obtain a facial area FA that is included in the human area HA and intersects with the baseline RL with the center line of the key points KP1 and KP2 as the reference line RL as a regular facial area. More specifically, for example, the image processing unit 130 may also obtain the facial area FA using the following learned model, which is learned in such a way that when the human area HA and an image in which the key points KP1 and KP2 are specified are input, the facial area FA that is included in the human area HA and intersects with the reference line RL obtained based on the key points KP1 and KP2 is output. By such processing, for example, it is possible to prevent Figure 6 The illustrated face area FA5 (ie, a face area that is erroneously detected due to, for example, wrinkles, folds, brightness, etc. of clothing worn by a person) is erroneously detected.

[0058] The image conversion unit 140 performs processing to replace the face area FA of the captured image data 172 thus acquired with the face of another person, without changing the orientation information of the person reflected in each image. This prevents erroneously detected face areas from being used directly as learning data, thereby deteriorating the accuracy of the behavior prediction model.

[0059] The image determination unit 150 re-inputs the transformed image obtained by the image transformation unit 140 into a learned model that has been trained to output at least one of the facial direction and the gaze direction, thereby obtaining the facial direction FD or the gaze direction ED in the transformed image. The image determination unit 150 determines whether the facial direction FD or the gaze direction ED of the face of a person reflected in the transformed image is substantially consistent with the facial direction FD or the gaze direction ED of the face reflected in the captured image before the transformation. As described above, both the facial direction FD and the gaze direction ED are obtained for the in-vehicle image, and the facial direction FD is obtained for the exterior image. Therefore, for the in-vehicle image, the image determination unit 150 determines whether the facial direction FD and the gaze direction ED are substantially consistent between the captured image before the transformation and the transformed image, and for the exterior image, whether the facial direction FD is substantially consistent between the captured image before the transformation and the transformed image. More specifically, for example, the image determination unit 150 calculates the angular difference between the vector indicating the facial direction FD in the captured image before transformation and the vector indicating the facial direction FD in the transformed image. If the calculated angular difference is within a threshold, the image determination unit 150 determines that the facial directions FD are substantially consistent. The same applies to the gaze direction ED. Satisfaction of facial continuity or consistency of directional information is an example of a "prescribed requirement."

[0060] If it is determined that the facial orientation FD or gaze direction ED between the pre-conversion captured image and the converted image is not substantially consistent, the image conversion unit 140 re-converts the captured image for the face determined to have a substantially inconsistent facial orientation FD or gaze direction ED. In this case, the image conversion unit 140 may re-convert only the face determined to be substantially inconsistent, or it may re-convert all faces included in the converted image, including the face determined to be substantially inconsistent. Alternatively, for example, the image conversion unit 140 may not re-convert the face determined to be substantially inconsistent, but instead perform mosaic processing on the face and exclude it from use as learning data. This prevents information degradation caused by unintended operation of the face conversion software.

[0061] It should be noted that, when there are multiple faces reflected in the transformed image (or when the number of faces reflected in the transformed image is greater than a predetermined value), the transformed image determination process performed by the image determination unit 150 may not be performed on all faces reflected in the transformed image, but rather only on faces assumed to be of higher importance. As an example of faces assumed to be of higher importance, the image determination unit 150 may perform the determination process only on faces whose face size in the pre-transformed captured image is greater than or equal to a third threshold value Th3, which is larger than the first threshold value Th1. Alternatively, the image determination unit 150 may perform the determination process only on faces whose face distance is less than or equal to a fourth threshold value Th4, which is smaller than the second threshold value Th2. Furthermore, for example, the image determination unit 150 may perform the determination process based on the assumption that the faces of persons located ahead of the vehicle M1 in the pre-transformed captured image, or the faces of persons facing ahead of the vehicle M1 in the pre-transformed captured image, are of higher importance.

[0062] When the image determination unit 150 completes its determination regarding the transformed image, it stores the transformed image data 174, which resulted in a positive determination, in the storage unit 170 as annotation image data 176. At this time, the transformed image data 174 may be stored in the storage unit 170 as annotation image data 176, along with information indicating the intended use, for example, information indicating that the transformed image data 174 is annotation image data. This annotation image data is used to generate a behavior prediction model for predicting the behavior of a person reflected in the input image. The transmission and reception control unit 120 transmits the annotation image data 176 to the terminal device 200. The annotator, a user of the terminal device 200, performs an annotation operation on the annotation image included in the received annotation image data 176, thereby generating annotated image data and transmitting the generated annotated image data to the image processing device 100. The image processing device 100 stores the received annotated image data as annotated image data 178 in the storage unit 170.

[0063] Figure 8 This is a diagram showing an example of annotation work performed by an annotator. Figure 8 The left part of represents the annotation of the transformed image to the in-car image. Figure 8 The right side of represents the annotation of the transformed image of the vehicle exterior image. The annotator gives information indicating whether the driver's sight line direction ED1 reflected in the transformed image is appropriate under the situation shown in the transformed image of the vehicle exterior image at the same time point (for example, 1 if appropriate, 0 if not appropriate). For example, in Figure 8In this case, the transformed image of the exterior image shows a pedestrian to the left of the vehicle's direction of travel, while the transformed image of the interior image shows the driver's gaze directed to the left. In other words, the annotator assigns appropriate information (i.e., 1) indicating the driver's gaze direction ED1, assuming the driver is paying appropriate attention to the pedestrian.

[0064] Furthermore, the annotator specifies, for the transformed image of the exterior vehicle image, a risk area RA into which, for example, a person reflected in the transformed image, excluding the person being mosaicked, is predicted to travel. Through processing by the image transformation unit 140 and the image determination unit 150, the face of the person reflected in the original image is transformed into the face of another person, thereby protecting the privacy of that person. Furthermore, the face orientation and gaze direction of the person are maintained after the transformation, allowing the annotator to accurately specify the risk area RA while referencing the face orientation and gaze direction of the other person reflected in the transformed image. This makes it possible to more easily and reliably detect the facial area in the image subject to anonymization.

[0065] Once annotated image data 178 is stored in storage unit 170, learned model generation unit 160 generates a learned model using an arbitrary machine learning model, using annotated image data 178 as learning data. As described above, this learned model can be, for example, a behavior prediction model that, given an input image outside the vehicle, outputs a predicted behavior (trajectory) of a person reflected in the image outside the vehicle, or, given an input image inside the vehicle and outside the vehicle, draws attention to pedestrians reflected in the image outside the vehicle, taking into account the driver's line of sight. The learned model generation unit 160 stores the generated learned model as learned model 180 in storage unit 170.

[0066] Once learned model 180 is generated, transmission / reception control unit 120 distributes generated learned model 180 to vehicle M2 via network NW. Upon receiving learned model 180, vehicle M2 uses learned model 180 (more precisely, an application utilizing learned model 180) to provide driving assistance to the driver of vehicle M2.

[0067] Figure 9 1 is a diagram showing an example of driving support using the learned model 180 . Figure 9 The following example is shown: vehicle M2 inputs the in-vehicle image and the exterior image captured by the onboard camera during driving into the learned model 180. The learned model 180 outputs information to the HMI (human machine interface) to draw attention to pedestrians in the exterior image, taking into account the driver's line of sight in the in-vehicle image, thereby providing driving support. Figure 9 As shown, for example, the HMI displays a risk area RA2 corresponding to a pedestrian P5 reflected in the exterior image. Furthermore, if the driver's gaze in the interior image is not directed toward pedestrian P5, a warning message ("Beware of distracted driving") is output in the form of text or audio. This enables driving support that takes the driver's state into consideration.

[0068] Next, refer to Figure 10 The flow of processing executed by the image processing apparatus 100 will be described. Figure 10 1 is a diagram showing an example of the flow of processing executed by the image processing unit 130 and the image conversion unit 140 . Figure 10 The processing shown is executed, for example, when a camera mounted on the vehicle M1 captures an image inside the vehicle or an image outside the vehicle.

[0069] First, the image processing unit 130 acquires a captured image included in the captured image data 172 (step S100 ). Next, the image processing unit 130 extracts a human figure area HA from the acquired captured image (step S102 ). Next, the image processing unit 130 extracts a facial area FA from only the extracted human figure area HA (step S104 ).

[0070] Next, the image conversion unit 140 determines whether the size of the extracted facial area FA is greater than or equal to a first threshold value Th1 (step S106). If the size of the extracted facial area FA is determined to be greater than or equal to the first threshold value Th1, the image conversion unit 140 converts the selected face into the face of another person (step S108). On the other hand, if the size of the selected face is determined to be less than the first threshold value Th1, the image conversion unit 140 next determines whether the distance between the extracted facial area FA is less than or equal to a second threshold value Th2 (step S110).

[0071] If the distance between the extracted facial region FA and the image conversion unit 140 is determined to be less than the second threshold value Th2, the process proceeds to step S106 to convert the face into that of another person. On the other hand, if the distance between the extracted facial region FA and the image conversion unit 140 is determined to be greater than the second threshold value Th2, the image conversion unit 140 performs mosaic processing on the face (step S112). Next, the image conversion unit 140 determines whether processing has been performed on all faces included in the acquired captured image (step S114).

[0072] If the image conversion unit 140 determines that processing has been performed on all faces included in the acquired captured image, it obtains the image resulting from the processing as a converted image and stores it as converted image data 174 in the storage unit 170 (step S116). On the other hand, if the image conversion unit 140 determines that processing has not been performed on all faces included in the acquired captured image, it returns the process to step S104. The process in this flowchart thus ends.

[0073] According to the embodiment described above, a person's facial region is extracted from an input image, and only the facial region is extracted from the extracted person region, and anonymization processing is performed on the extracted facial region. This makes it possible to more easily and reliably detect the facial region of an image to be anonymized.

[0074] [Variation]

[0075] In this embodiment, an example is described in which the image processing device 100 is installed as a server device separate from the vehicle M1. However, as a variation of this embodiment, the image processing device 100, more specifically, a device having at least the functions of the image processing unit 130, the image conversion unit 140, and the image determination unit 150, may also be installed as an onboard device in the vehicle M1. In this case, the onboard device performs the processing described above by the image processing unit 130 on images captured by the onboard camera, performs anonymization by the image conversion unit 140, and performs determination by the image determination unit 150. The onboard device then transmits the anonymized image, for which facial continuity and orientation information consistency have been verified by the image determination unit 150, to an external image server.

[0076] Upon receiving an anonymized image from vehicle M1, the image server stores the received anonymized image as annotation image data in a storage unit and permits the annotator's terminal device 200 to transmit the annotated image data or the terminal device 200 to access the annotation image data. Upon receiving annotated image data from the terminal device 200, the image server generates a learned model 180 based on the annotated image data and distributes the generated learned model 180 to vehicle M2. This, like the present embodiment, allows for simpler and more reliable detection of facial regions in images subject to anonymization. Furthermore, according to this variation, the in-vehicle device anonymizes the image before transmitting the anonymized image to the image server, further reliably protecting the privacy of individuals reflected in the facial image.

[0077] Furthermore, as another embodiment, the in-vehicle device may only have some of the functions of the image processing unit 130, the image conversion unit 140, and the image determination unit 150, and the image server may have the remaining functions. For example, the in-vehicle device may have the functions of the image processing unit 130 and the image conversion unit 140, and the image server may have the functions of the image determination unit 150. Alternatively, the in-vehicle device may have the functions of the image processing unit 130, and the image server may have the functions of the image conversion unit 140 and the image determination unit 150.

[0078] The above-described embodiment can be expressed as follows.

[0079] An image processing device comprising:

[0080] a storage medium storing computer-readable instructions; and

[0081] a processor connected to the storage medium,

[0082] The processor executes the computer-readable instructions to:

[0083] extracting a face region of a person from an input image; and

[0084] Anonymizing the extracted facial area,

[0085] The extraction means extracting a person region representing the person from the input image, and extracting the face region only from the extracted person region.

[0086] While specific embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention.

[0087] Description of Reference Numerals

[0088] 100 Image processing device

[0089] 110 Ministry of Communications

[0090] 120 Transceiver Control Unit

[0091] 130 Image Processing Department

[0092] 140 Image Conversion Unit

[0093] 150 Image determination unit

[0094] 160 Learning Complete Model Generation

[0095] 170 Storage Department

[0096] 172 captured image data

[0097] 174 Transform Image Data

[0098] 176 Annotation image data

[0099] 178 annotated image data

[0100] 180 Learning completed model.

Claims

1. An image processing device, wherein: The image processing device comprises: an extraction unit that extracts a face region of a person from an input image; and a conversion unit, which performs anonymization processing on the extracted facial area, The extraction unit extracts a person region representing the person from the input image, and extracts the face region only from the extracted person region.

2. The image processing apparatus according to claim 1, wherein: The extraction unit further extracts one or more part regions representing body parts of the person from the person region, and extracts the face region from the person region based on positions of the one or more part regions.

3. The image processing apparatus according to claim 1, wherein: The conversion unit performs the anonymization process when the size of the extracted face region is equal to or larger than a first threshold value and the distance from the shooting position of the input image to the face region is equal to or smaller than a second threshold value.

4. The image processing apparatus according to claim 1, wherein: The anonymization process is a process of aligning the directions of the face of the person before and after the anonymization process and changing the face of the person to the face of another person.

5. The image processing apparatus according to claim 4, wherein: The image processing device further includes a learning unit. The learning unit obtains an annotated image obtained by adding an annotation to the input image subjected to the anonymization process, the annotation indicating whether the direction of the face of the person driving the vehicle is appropriate, The learning unit generates a learned model using the annotated image as learning data, the learned model being used to prompt the person to draw attention to a pedestrian outside the vehicle.

6. An image processing method, wherein: The image processing method enables the computer to perform the following processing: extracting a face region of a person from an input image; and Anonymizing the extracted facial area, The extraction means extracting a person region representing the person from the input image, and extracting the face region only from the extracted person region.

7. A program, wherein The program causes the computer to perform the following processing: extracting a face region of a person from an input image; and Anonymizing the extracted facial area, The extraction means extracting a person region representing the person from the input image, and extracting the face region only from the extracted person region.

Citation Information

Patent Citations

  • Face collation device and passage control device

    JP2005242777A