Image processing device, image processing method, and program
Through image processing devices and methods, image transformation and judgment technology is used to replace or mosaic the faces of people in images, and a reliable anonymized image is generated through a learning model. This solves the problem of difficulty in simply and reliably detecting facial areas in existing technologies and achieves more efficient anonymization processing.
Patent Information
- Application Number
- CN202380094662.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2025-09-16
AI Technical Summary
In the prior art, it is difficult to easily and reliably detect the facial region of an image to be anonymized. In particular, when the texture of the person and the background are similar, misdetection is likely to occur.
An image processing device and method are used to replace or mosaic the face through an image transformation unit, determine the suitability of the person's presence in conjunction with an image determination unit, delete or reprocess the erroneously detected area, and use a learned model to generate and save a reliable anonymized image.
This makes it easier and more reliable to detect and protect the privacy of people in images, improves the accuracy and precision of anonymization processing, and prevents false detections.
Smart Images

Figure CN120660110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing device, an image processing method, and a program. Background Art
[0002] Conventionally, technologies for identifying a person's facial region based on an image captured by a camera are known. For example, Patent Document 1 describes preventing a wall pattern in front of an image input device from being mistakenly detected as a face, even if the pattern is similar to that of a person's face.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2005-242777 Summary of the Invention
[0006] Problems to be solved by the invention
[0007] The technology described in Patent Document 1 reliably detects a person's face by lighting a light that illuminates the person's face when a person-sensing sensor detects their presence. However, conventional technologies sometimes fail to easily and reliably detect the facial area of an image to be anonymized.
[0008] The present invention has been made in consideration of such circumstances, and one object of the present invention is to provide an image processing device, an image processing method, and a program that can more easily and reliably detect a face region in an image to be anonymized.
[0009] Solutions to Problems
[0010] The image processing device, image processing method, and program according to the present invention have the following configurations.
[0011] (1): An image processing device according to one embodiment of the present invention comprises: an extraction unit that extracts a person from a transformed image obtained by performing image transformation processing on an input image; and a processing unit that determines whether the extracted person satisfies requirements related to suitability of the person's presence and performs prescribed processing on the transformed image based on the result of the determination.
[0012] (2): In the embodiment of (1), the extraction unit extracts the face of the person as the person, and the requirement related to the suitability of the person's presence is that parts other than the face of the person are recognized around the face.
[0013] (3): In the above-mentioned solution (1), the requirement related to the suitability of the presence of the person is that the area where the person exists in the transformed image is recognized as an area where pedestrians can pass.
[0014] (4): In the above-mentioned embodiment (1), the requirement related to the suitability of the presence of the person is that the person also exists in the transformed images at the previous and next time points in the time series of the transformed images.
[0015] (5): Based on the above-mentioned solution (1), when the result of the determination as to whether the requirement related to the suitability of the presence of the person is satisfied is negative, the processing unit deletes the transformed image as the prescribed processing or performs the image transformation processing on the input image again.
[0016] (6): In the above-mentioned aspect (1), when the determination result of whether the requirement related to the suitability of the person's presence is satisfied is affirmative, the processing unit stores the transformed image as learning data as the predetermined processing.
[0017] (7): Based on any one of the above-mentioned schemes (1) to (6), the image transformation processing is a processing that makes the direction of the face of the person consistent before and after the image transformation processing, and changes the face of the person to the face of another person.
[0018] (8): Another embodiment of the present invention relates to an image processing method, wherein the image processing method causes a computer to perform the following processing: extracting a person from a transformed image obtained by performing image transformation processing on an input image; and determining whether the extracted person satisfies requirements related to suitability of the person's existence, and performing prescribed processing on the transformed image based on the result of the determination.
[0019] (9): Another embodiment of the present invention relates to a program, wherein the program causes a computer to perform the following processing: extracting a person from a transformed image obtained by performing image transformation processing on an input image; and determining whether the extracted person satisfies requirements related to suitability of the person's existence, and performing prescribed processing on the transformed image based on the result of the determination.
[0020] Effects of the Invention
[0021] According to the above-mentioned configurations (1) to (9), the face region of an image to be anonymized can be detected more easily and reliably. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 11 is a diagram showing an overview of a system 1 including an image processing apparatus 100 according to this embodiment.
[0023] Figure 2 1 is a diagram showing an example of the functional configuration of the image processing device 100 according to the present embodiment.
[0024] Figure 3 1 is a diagram showing an example of an in-vehicle image and an out-vehicle image acquired from the vehicle M1.
[0025] Figure 4 It is a diagram for explaining the processing executed by the image processing unit 130 .
[0026] Figure 5 It is a diagram for explaining the processing performed by the image conversion unit 140.
[0027] Figure 6 This is a diagram for explaining failure of processing performed by the conventional image processing unit 130 .
[0028] Figure 7 It is a diagram for explaining the processing executed by the image processing unit 130 according to this embodiment.
[0029] Figure 8 This is a diagram showing an example of annotation work performed by an annotator.
[0030] Figure 9 1 is a diagram showing an example of driving support performed using the learned model 180 .
[0031] Figure 10 14 is a diagram showing an example of the flow of processing executed by the image conversion unit 140 .
[0032] Figure 11 This is a diagram showing an example of the flow of processing executed by the image determination unit 150 . DETAILED DESCRIPTION
[0033] Hereinafter, embodiments of an image processing device, an image processing method, an image processing system, and a program according to the present invention will be described with reference to the accompanying drawings.
[0034] [summary]
[0035] Figure 1 1 is a diagram showing an overview of a system 1 including an image processing apparatus 100 according to this embodiment. Figure 1 As shown, the system 1 includes at least one vehicle M1 and vehicle M2, an image processing device 100, and a terminal device 200. For ease of explanation, the vehicle M1 and vehicle M2 are shown as different vehicles, but these vehicles may be the same.
[0036] Vehicle M1 is, for example, a hybrid vehicle, an electric vehicle, or the like, and includes at least one camera for capturing images of the interior and exterior of vehicle M1. Vehicle M1 transmits images of the interior and exterior captured by these cameras while driving to image processing device 100 via a network NW such as a cellular network, Wi-Fi network, or the Internet.
[0037] Image processing device 100 is a server device that, upon receiving captured image data including interior and exterior images from vehicle M1, performs image conversion, described below, on the received captured image data. This image conversion is performed to protect the privacy of people reflected in the interior and exterior images. Image processing device 100 transmits the resulting converted image data to terminal device 200 via network NW.
[0038] The terminal device 200 is a terminal device such as a desktop personal computer or a smartphone. When the user of the terminal device 200 obtains transformed image data from the image processing device 100, they perform an annotation operation (described below) on the obtained transformed image data. When the annotation operation is completed, the user of the terminal device 200 transmits the annotated image data, in which the annotations are applied to the transformed image data, to the image processing device 100.
[0039] When the image processing device 100 receives annotated image data from the terminal device 200, it uses the received annotated image data as learning data and generates a learned model (described later) using an arbitrary machine learning model. This learned model, for example, is a behavior prediction model that, given an input image outside the vehicle, outputs a predicted behavior (trajectory) of a person reflected in the image outside the vehicle, or, given an input image inside the vehicle and an image outside the vehicle, urges attention to pedestrians reflected in the image outside the vehicle, taking into account the driver's line of sight.
[0040] It should be noted that the image data used as learning data in this case can be annotated image data in which annotations are assigned to transformed image data, or annotated image data in which the transformed image data is converted back to captured image data while the annotations remain unchanged (i.e., annotated image data in which annotations are assigned to captured image data). By using annotated image data in which annotations are assigned to captured image data as learning data, more realistic learning data can be used, in which the effects of image transformation are eliminated.
[0041] When the image processing device 100 generates a learned model, it distributes it to vehicle M2 via the network NW. Similar to vehicle M1, vehicle M2 is a hybrid vehicle, electric vehicle, or other similar vehicle. Vehicle M2 inputs at least one of the vehicle's interior and exterior images captured by a camera while driving into the learned model to obtain predicted behavior data for people around vehicle M2. The driver of vehicle M2 can refer to this predicted behavior data and apply it to driving vehicle M2. The following describes each process in more detail.
[0042] [Functional Structure of Image Processing Device]
[0043] Figure 2 This figure shows an example of the functional structure of an image processing device 100 according to this embodiment. The image processing device 100 includes, for example, a communication unit 110, a transmission and reception control unit 120, an image processing unit 130, an image conversion unit 140, an image determination unit 150, a learned model generation unit 160, and a storage unit 170. These components are implemented by executing a program (software) on a hardware processor such as a CPU (Central Processing Unit). Some or all of these components can be implemented by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or through the coordinated implementation of software and hardware. The program can be pre-stored on a storage device (including a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or stored on a removable storage medium (non-transitory storage medium) such as a DVD or CD-ROM, and installed by attaching the storage medium to a drive. The storage unit 170 is, for example, an HDD, flash memory, or random access memory (RAM). The storage unit 170 stores, for example, captured image data 172, transformed image data 174, annotation image data 176, annotated image data 178, and a learned model 180. For ease of explanation, the image processing device 100 includes the learned model generation unit 160 and the storage unit 170 that stores the learned model 180. However, the function of generating the learned model and the generated learned model may be stored in a server device separate from the image processing device 100.
[0044] The communication unit 110 is an interface for communicating with the communication device 10 of the host vehicle M via the network NW. For example, the communication unit 110 includes a NIC (Network Interface Card), an antenna for wireless communication, and the like.
[0045] The transmission / reception control unit 120 uses the communication unit 110 to transmit and receive data with vehicles M1 and M2 and the terminal device 200. More specifically, the transmission / reception control unit 120 first obtains from vehicle M1 a plurality of vehicle interior and exterior images captured in a time series by a camera mounted on vehicle M1. In this case, the time series refers to, for example, images captured at predetermined intervals (e.g., every one second) during a driving cycle from the start to the stop of vehicle M1.
[0046] Figure 3 1 is a diagram showing an example of an in-vehicle image and an out-vehicle image acquired from the vehicle M1. Figure 3 The left part of represents the in-vehicle image obtained from vehicle M1. Figure 3 The right side of represents the vehicle exterior image obtained from vehicle M1. Figure 3 As shown in the left part of , the in-vehicle image is captured by a camera in a manner that captures at least the facial area of the driver of the vehicle M1. Figure 3 As shown in the right portion of FIG, the vehicle exterior image is captured by a camera installed so as to capture at least the area in front of the vehicle M1 in its traveling direction. The transmission and reception control unit 120 associates the interior and exterior images acquired from the vehicle M1 with image IDs and stores them in the storage unit 170 as captured image data 172.
[0047] Figure 4 This figure illustrates the processing performed by image processing unit 130. Image processing unit 130 performs image processing on captured image data 172 to obtain (extract) information such as image attributes, facial attributes, and orientation for each image included in captured image data 172. More specifically, image processing unit 130 uses a learned model that, when input, outputs a classification result indicating whether the image is an image inside a vehicle or outside a vehicle, to obtain image attributes indicating whether each image included in captured image data 172 is an image inside a vehicle or outside a vehicle.
[0048] Furthermore, the image processing unit 130 uses a learned model that outputs the facial area, the size of the face (the area of the facial area), and the distance from the image shooting position to the face for all faces included in the image when an image is input, to obtain facial attributes of each image included in the captured image data 172. Figure 3In this example, facial area FA1 of person P1 is acquired from an image inside the vehicle, while facial area FA2 of person P2, facial area FA3 of person P3, and facial area FA4 of person P4 are acquired from an image outside the vehicle. While facial areas FA1, FA2, FA3, and FA4 are acquired as rectangular areas, the present invention is not limited to such a configuration. For example, a learned model that acquires facial areas along the contours of a person's face may also be used.
[0049] Furthermore, the image processing unit 130 uses a learned model that outputs at least one of the facial direction and the gaze direction, for example, as a vector for all faces included in the image when the image is input, to obtain direction information of the face reflected in each image included in the captured image data 172. More specifically, the image processing unit 130 uses a learned model that outputs the facial direction and the gaze direction for all faces included in the image when the image is input, to obtain direction information for the image of the captured image data 172 having the attribute of an in-vehicle image. On the other hand, the image processing unit 130 uses a learned model that outputs the facial direction for all faces included in the image when the image is input, to obtain direction information for the image of the captured image data 172 having the attribute of an out-of-vehicle image. This is because, generally speaking, compared to out-of-vehicle images, faces reflected in in-vehicle images are closer to the shooting position, and there is a high tendency for faces to be reflected to a large extent to which the gaze direction can be extracted. Figure 3 As an example, the face direction FD1 and the sight line direction ED1 of the person P1 are obtained from the in-vehicle image, and the face directions FD2, FD3, and FD4 of the person P2, P3, and P4 are obtained from the out-vehicle image.
[0050] When the image processing unit 130 obtains image attributes, facial attributes, and orientation information for each image in the captured image data 172, it associates the image attributes, facial attributes, and orientation information with the image. It should be noted that, in the above example, the image processing unit 130 uses a learned model to obtain image attributes, facial attributes, and orientation information. However, the present invention is not limited to such a configuration, and the image processing unit 130 may also use any known method to obtain the image attributes, facial attributes, and orientation information.
[0051] The image conversion unit 140 uses arbitrary software equipped with such a function to execute processing for replacing the face of a person shown in each image of the captured image data 172 processed by the image processing unit 130 with the face of another person without changing the direction information of the person. Figure 5 1 is a diagram for explaining the processing performed by the image conversion unit 140. Figure 5 As shown, the image conversion unit 140 is Figure 4The faces of persons P1, P2, and P3 are replaced with faces of other persons without changing their eye direction ED1 and facial directions FD1, FD2, and FD3. On the other hand, the face of person P4 is mosaiced by the image conversion unit 140 and is covered by the mosaic MS.
[0052] Specifically, the image conversion unit 140 determines whether to replace each face included in each image of the captured image data 172 with the face of another person or to perform mosaic processing on the face based on the facial attributes of the face. More specifically, the image conversion unit 140 determines whether the size of each face included in each image of the captured image data 172 is greater than a first threshold value Th1. If the size of the face is determined to be greater than the first threshold value Th1, the image conversion unit 140 determines to replace the face with the face of another person. On the other hand, if the size of the face is determined to be less than the first threshold value Th1, the image conversion unit 140 determines to perform mosaic processing on the face. Replacing the face of a person included in a captured image with the face of another person or performing mosaic processing is an example of "anonymization processing."
[0053] Furthermore, the image conversion unit 140 determines whether the distance to each face reflected in each image of the captured image data 172 is less than or equal to a second threshold value Th2. If the distance to the face is determined to be less than the second threshold value Th2, the image conversion unit 140 determines to replace the face with the face of another person. On the other hand, if the distance to the face is determined to be greater than the second threshold value Th2, the image conversion unit 140 determines to perform mosaic processing on the face. The image conversion unit 140 repeats the above determination process for the number of faces reflected in the image and, depending on the determination result, replaces each face with another person's face or performs mosaic processing on the face. The image conversion unit 140 stores the image data obtained by performing this processing on the captured image data 172 as transformed image data 174 in the storage unit 170. This allows the privacy of the people reflected in each image to be protected when selecting useful data for learning data for generating an action prediction model and when the annotator, described later, performs annotation work.
[0054] It should be noted that at least one of the process of determining whether the face size is greater than or equal to the first threshold value Th1 and the process of determining whether the face distance is less than or equal to the second threshold value Th2 may be performed. When both processes are performed, the image conversion unit 140 may decide to replace the face with the face of another person if the face size is greater than or equal to the first threshold value Th1 and the face distance is less than or equal to the second threshold value Th2, or may decide to replace the face with the face of another person if the face size is greater than or equal to the first threshold value Th1 or the face distance is less than or equal to the second threshold value Th2.
[0055] Furthermore, the image conversion unit 140 may perform mosaic processing on faces for which acquisition of direction information failed among faces included in each image of the captured image data 172 , thereby filtering out faces to be used as learning data.
[0056] Thus, when acquiring transformed image data 174 from captured image data 172, image processing unit 130 uses the learned model as described above to acquire information such as the facial region and facial orientation of the person included in the image. However, other software using the learned model or similar functionality may sometimes output non-persons as people due to factors such as wrinkles, stains, or brightness on clothing worn by the person, cracks or dirt on roads or walls, or dirt, stickers, or advertisements attached to other vehicles.
[0057] Figure 6 This is a diagram for explaining failure of processing performed by the conventional image processing unit 130 and the image conversion unit 140 . Figure 6 The conventional image processing unit 130 obtains persons P1 to P6 by inputting the in-car image and the out-of-car image into the learned model, and the image conversion unit 140 performs anonymization processing on the facial areas FA1 to FA6 of the persons P1 to P6 to obtain the converted images. Figure 6 As shown, for example, the following situation may occur: person P5 is mistakenly captured and transformed due to wrinkles, stains, brightness, etc. on the clothing worn by the person reflected in the vehicle interior image, and person P6 is mistakenly captured and transformed due to cracks, dirt attached to the road, or dirt, stickers, advertisements, etc. attached to other vehicles reflected in the vehicle exterior image. In this way, using the transformed image data 174 including the mistakenly captured persons P5 and P6 as learning data directly will cause the accuracy of the behavior prediction model to deteriorate, which is not preferable. Therefore, as described below, the image determination unit 150 involved in this embodiment determines the transformed image including the mistakenly captured person based on whether each person reflected in the transformed image meets the prescribed requirements related to the suitability of existence, and deletes the transformed image or re-anonymizes the input image corresponding to the transformed image, thereby maintaining the accuracy of the behavior prediction model.
[0058] Figure 7This figure illustrates the processing performed by the image determination unit 150 according to this embodiment. First, the image determination unit 150 extracts the persons P1 to P6 reflected in the vehicle interior and exterior images from the transformed images obtained by the image conversion unit 140. As described above, the persons P1 to P6 are previously identified by the image processing unit 130 as corresponding to the facial areas FA1 to FA6 before the image conversion unit 140 performs the transformation. It should be noted that the mosaic-processed person P4 may be excluded from the processing by the image determination unit 150.
[0059] For example, the image determination unit 150 searches for one or more body parts (e.g., shoulders, hands, knees, feet, etc.) around the facial regions FA1 to FA6 of the extracted persons P1 to P6 using an arbitrary key point detection method. If one or more body parts are detected, the image determination unit 150 determines that the person and the facial region are real humans. Figure 7 In the case of the transformed image of the vehicle interior shown on the left, image determination unit 150 detects key points KP1 and KP2 representing the shoulders of person P1 based on facial area FA1 (in other words, taking into account the relative position from facial area FA1). As a result, image determination unit 150 determines that person P1 is a real human. On the other hand, for person P5, no key points can be extracted based on facial area FA5, and image determination unit 150 determines that person P5 is not a real human. Figure 7 Similarly, in the case of the transformed image of the vehicle exterior image shown on the right, image determination unit 150 can extract key points such as shoulders, hands, knees, and feet for persons P2 through P4 based on facial areas FA2 through FA4. However, for person P6, no key points can be extracted based on facial area FA6, and image determination unit 150 can determine that person P6 is not a real person. Detecting one or more body parts based on facial areas is an example of a "prescribed requirement."
[0060] As another method, the image determination unit 150 may determine whether the person is a real human being based on whether the area where the person is reflected in the transformed image of the vehicle exterior image is a free space (in other words, an area where pedestrians can pass). Figure 7In the case of the transformed image of the vehicle exterior image shown on the right, the image determination unit 150 detects free space FS by inputting the transformed image into a learned model that has been trained to output free space in the image when an image is input. Next, the image determination unit 150 determines whether each of the persons P2 to P4 and P6 exists within free space FS. The image determination unit 150 determines that persons P2 to P4 exist within free space FS, while person P6 exists within the vehicle and therefore does not exist within free space FS. Consequently, the image determination unit 150 determines that persons P2 to P4 are real humans, while person P6 is not real humans. For the transformed image of the vehicle exterior image, the presence of a person within free space is an example of a "prescribed requirement."
[0061] Alternatively, as another method, the image determination unit 150 can determine whether a person is a real person based on whether the same person appears in the transformed images at each point in time, using transformed images of the vehicle interior and exterior images captured in a time series. For example, if a person detected based on the transformed images at a certain point in time also appears in transformed images at a previous or subsequent point in time, the image determination unit 150 can determine that the person is a real person. In this case, the image determination unit 150 can also consider the positional changes (speed) of the person reflected in the captured image and perform this determination only at the point in time when the person is assumed to be included in the captured image. This can prevent temporary misdetection of non-human beings due to wrinkles, stains, brightness, cracks in the road, dirt, etc. in clothing. The presence of a person detected based on the transformed images at a certain point in time in transformed images at previous or subsequent points in time is an example of a "prescribed requirement."
[0062] If a person detected based on the transformed image is determined not to be a real person, the image determination unit 150 deletes the transformed image or re-anonymizes the input image corresponding to the transformed image, thereby obtaining a new transformed image. Deleting the transformed image or re-anonymizing the input image corresponding to the transformed image is an example of "prescribed processing."
[0063] On the other hand, if it is determined that all persons detected from the transformed image are real people, the image determination unit 150 re-inputs the transformed image into a learned model that has been trained to output at least one of the facial direction and the gaze direction, thereby obtaining the facial direction FD or the gaze direction ED in the transformed image. The image determination unit 150 then determines whether the facial direction FD or the gaze direction ED of the person's face in the transformed image is substantially consistent with the facial direction FD or the gaze direction ED of the face in the pre-transformed captured image. As described above, both the facial direction FD and the gaze direction ED are obtained for the in-vehicle image, and the facial direction FD is obtained for the exterior image. Therefore, for the in-vehicle image, the image determination unit 150 determines whether the facial direction FD and the gaze direction ED are substantially consistent between the pre-transformed captured image and the transformed image, and for the exterior image, whether the facial direction FD is substantially consistent between the pre-transformed captured image and the transformed image. More specifically, for example, the image determination unit 150 calculates the angular difference between the vector indicating the facial direction FD in the captured image before transformation and the vector indicating the facial direction FD in the transformed image. If the calculated angular difference is within a threshold, the image determination unit 150 determines that the facial directions FD are substantially consistent. The same applies to the gaze direction ED.
[0064] If it is determined that the facial orientation FD or gaze direction ED between the pre-conversion captured image and the converted image is not substantially consistent, the image conversion unit 140 re-converts the captured image for the face determined to have a substantially inconsistent facial orientation FD or gaze direction ED. In this case, the image conversion unit 140 may re-convert only the face determined to be substantially inconsistent, or it may re-convert all faces included in the converted image, including the face determined to be substantially inconsistent. Alternatively, for example, the image conversion unit 140 may not re-convert the face determined to be substantially inconsistent, but instead perform mosaic processing on the face and exclude it from use as learning data. This prevents information degradation caused by unintended operation of the face conversion software.
[0065] It should be noted that, when there are multiple faces reflected in the transformed image (or when the number of faces reflected in the transformed image is greater than a predetermined value), the transformed image determination process performed by the image determination unit 150 may not be performed on all faces reflected in the transformed image, but rather only on faces assumed to be of higher importance. As an example of faces assumed to be of higher importance, the image determination unit 150 may perform the determination process only on faces whose face size in the pre-transformed captured image is greater than or equal to a third threshold value Th3, which is larger than the first threshold value Th1. Alternatively, the image determination unit 150 may perform the determination process only on faces whose face distance is less than or equal to a fourth threshold value Th4, which is smaller than the second threshold value Th2. Furthermore, for example, the image determination unit 150 may perform the determination process based on the assumption that the faces of persons located ahead of the vehicle M1 in the pre-transformed captured image, or the faces of persons facing ahead of the vehicle M1 in the pre-transformed captured image, are of higher importance.
[0066] When the image determination unit 150 completes its determination regarding the transformed image, it stores the transformed image data 174, which resulted in a positive determination, in the storage unit 170 as annotation image data 176. At this time, the transformed image data 174 may be stored in the storage unit 170 as annotation image data 176, along with information indicating the intended use, for example, information indicating that the transformed image data 174 is annotation image data. This annotation image data is used to generate a behavior prediction model for predicting the behavior of a person reflected in the input image. The transmission and reception control unit 120 transmits the annotation image data 176 to the terminal device 200. The annotator, a user of the terminal device 200, performs an annotation operation on the annotation image included in the received annotation image data 176, thereby generating annotated image data and transmitting the generated annotated image data to the image processing device 100. The image processing device 100 stores the received annotated image data as annotated image data 178 in the storage unit 170.
[0067] Figure 8 This is a diagram showing an example of annotation work performed by an annotator. Figure 8 The left part of represents the annotation of the transformed image to the in-car image. Figure 8 The right side of represents the annotation of the transformed image of the vehicle exterior image. The annotator gives information indicating whether the driver's sight line direction ED1 reflected in the transformed image is appropriate under the situation shown in the transformed image of the vehicle exterior image at the same time point (for example, 1 if appropriate, 0 if not appropriate). For example, in Figure 8In this case, the transformed image of the exterior image shows a pedestrian to the left of the vehicle's direction of travel, while the transformed image of the interior image shows the driver's gaze directed to the left. In other words, the annotator assigns appropriate information (i.e., 1) indicating the driver's gaze direction ED1, assuming the driver is paying appropriate attention to the pedestrian.
[0068] Furthermore, the annotator specifies, for the transformed image of the exterior vehicle image, a risk area RA into which, for example, a person reflected in the transformed image, excluding the person being mosaicked, is predicted to travel. Through processing by the image transformation unit 140 and the image determination unit 150, the face of the person reflected in the original image is transformed into the face of another person, thereby protecting the privacy of that person. Furthermore, the face orientation and gaze direction of the person are maintained after the transformation, allowing the annotator to accurately specify the risk area RA while referencing the face orientation and gaze direction of the other person reflected in the transformed image. This makes it possible to more easily and reliably detect the facial area in the image subject to anonymization.
[0069] Once annotated image data 178 is stored in storage unit 170, learned model generation unit 160 generates a learned model using an arbitrary machine learning model, using annotated image data 178 as learning data. As described above, this learned model can be, for example, a behavior prediction model that, when inputted as an image outside the vehicle, outputs a predicted behavior (trajectory) of a person reflected in the image outside the vehicle, or, when inputted as an image inside the vehicle and outside the vehicle, urges attention to pedestrians reflected in the image outside the vehicle, taking into account the driver's line of sight. The learned model generation unit 160 stores the generated learned model as learned model 180 in storage unit 170.
[0070] Once learned model 180 is generated, transmission / reception control unit 120 distributes generated learned model 180 to vehicle M2 via network NW. Upon receiving learned model 180, vehicle M2 uses learned model 180 (more precisely, an application utilizing learned model 180) to provide driving assistance to the driver of vehicle M2.
[0071] Figure 9 1 is a diagram showing an example of driving support using the learned model 180 . Figure 9 The following example is shown: vehicle M2 inputs the in-vehicle image and the exterior image captured by the onboard camera during driving to the learned model 180. The learned model 180 outputs information to the HMI (human machine interface) urging the driver to pay attention to the pedestrians in the exterior image, taking into account the driver's line of sight in the in-vehicle image, thereby providing driving support. Figure 9 As shown, for example, the HMI displays a risk area RA2 corresponding to a pedestrian P5 reflected in the exterior image. Furthermore, if the driver's gaze in the interior image is not directed toward pedestrian P5, a warning message ("Beware of distracted driving") is output in the form of text or audio. This enables driving support that takes the driver's state into consideration.
[0072] Next, refer to Figure 10 and Figure 11 The flow of processing executed by the image processing apparatus 100 will be described. Figure 10 14 is a diagram showing an example of the flow of processing executed by the image conversion unit 140 . Figure 10 The processing shown is executed, for example, when a camera mounted on the vehicle M1 captures an image inside the vehicle or an image outside the vehicle and the processing by the image processing unit 130 is performed.
[0073] First, the image conversion unit 140 acquires a captured image included in the captured image data 172 processed by the image processing unit 130 (step S100 ). Next, the image conversion unit 140 selects one face reflected in the acquired captured image (step S102 ).
[0074] Next, the image conversion unit 140 determines whether the size of the selected face is greater than or equal to a first threshold value Th1 (step S104). If the size of the selected face is determined to be greater than or equal to the first threshold value Th1, the image conversion unit 140 converts the face into that of another person (step S106). On the other hand, if the size of the selected face is determined to be less than or equal to the first threshold value Th1, the image conversion unit 140 next determines whether the distance of the selected face is less than or equal to a second threshold value Th2 (step S108).
[0075] If the distance to the selected face is determined to be less than the second threshold Th2, the image conversion unit 140 proceeds to step S106 and converts the selected face into the face of another person. On the other hand, if the distance to the selected face is determined to be greater than the second threshold Th2, the image conversion unit 140 performs mosaic processing on the selected face (step S110). Next, the image conversion unit 140 determines whether processing has been performed on all faces included in the acquired captured image (step S112).
[0076] If it is determined that processing has been performed on all faces included in the acquired captured image, the image conversion unit 140 obtains the image resulting from the processing as a converted image and stores it as converted image data 174 in the storage unit 170 (step S114). On the other hand, if it is determined that processing has not been performed on all faces included in the acquired captured image, the image conversion unit 140 returns the process to step S102. The process in this flowchart thus ends.
[0077] Figure 11 This is a diagram showing an example of the flow of processing executed by the image determination unit 150 . Figure 11 The processing shown is executed, for example, at a timing when a time-series converted image is obtained by performing the above-described conversion processing on the time-series captured images captured in one driving cycle from start to stop of the vehicle M1 .
[0078] First, the image determination unit 150 obtains a transformed image (step S200). Next, the image determination unit 150 selects a person from the obtained transformed image (step S202). Next, the image determination unit 150 determines whether the obtained person meets the prescribed requirements related to presence suitability (step S204). If the obtained person is determined to meet the prescribed requirements related to presence suitability, the image determination unit 150 then determines whether the obtained transformed image is an image of the interior of a vehicle (step S206). On the other hand, if the obtained person is determined not to meet the prescribed requirements related to presence suitability, the image determination unit 150 instructs the image conversion unit 140 to re-anonymize the input image corresponding to the transformed image (step S208). The image determination unit 150 then repeats the process of step S202 on the re-transformed image.
[0079] If the converted image is determined to be an in-vehicle image in step S206, the image determination unit 150 determines whether the gaze direction and facial orientation of the faces match those in the pre-conversion image (step S210). On the other hand, if the converted image is determined to be an exterior image, not an in-vehicle image, the image determination unit 150 determines whether the facial orientation of the faces matches those in the pre-conversion image (step S212). If a mismatch is determined in steps S210 or S212, the image determination unit 150 proceeds to step S208.
[0080] If a match is determined in step S210 or step S212, the image determination unit 150 determines that the faces have been properly transformed and determines whether processing has been performed on all faces reflected in the transformed images (step S214). If it is determined that processing has been performed on all faces reflected in the time-series transformed images, the image determination unit 150 obtains these time-series transformed images as annotation images and instructs the transmission and reception control unit 120 to transmit the obtained annotation images to the terminal device 200 (step S216). On the other hand, if it is determined that processing has not been performed on all faces reflected in the time-series transformed images, the image determination unit 150 returns the process to step S202. This concludes the process in this flowchart.
[0081] According to the present embodiment described above, a person is extracted from a transformed image obtained by performing image transformation processing on an input image, a determination is made as to whether the extracted person satisfies the requirements related to the suitability of the person's presence, and a predetermined process is performed on the transformed image based on the determination result. This makes it possible to more easily and reliably detect the facial region of an image to be anonymized.
[0082] The above-described embodiment can be expressed as follows.
[0083] An image processing device comprising:
[0084] a storage medium storing computer-readable instructions; and
[0085] a processor connected to the storage medium,
[0086] The processor executes the computer-readable instructions to:
[0087] extracting a person region representing a person from a transformed image obtained by performing image transformation processing on an input image; and
[0088] It is determined whether the extracted human figure region satisfies requirements related to suitability of the human figure's presence, and predetermined processing is performed on the transformed image based on the result of the determination.
[0089] While specific embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention.
[0090] Description of Reference Numerals
[0091] 100 Image processing device
[0092] 110 Ministry of Communications
[0093] 120 Transceiver Control Unit
[0094] 130 Image Processing Department
[0095] 140 Image Conversion Unit
[0096] 150 Image determination unit
[0097] 160 Learning Complete Model Generation
[0098] 170 Storage Department
[0099] 172 captured image data
[0100] 174 Transform Image Data
[0101] 176 Annotation image data
[0102] 178 annotated image data
[0103] 180 Learning completed model.
Claims
1. An image processing device, wherein: The image processing device comprises: an extraction unit that extracts a person from a transformed image obtained by performing image transformation processing on an input image; and The processing unit determines whether the extracted person satisfies a requirement related to suitability of the person's presence, and performs a predetermined process on the transformed image based on a result of the determination.
2. The image processing apparatus according to claim 1, wherein: The extraction unit extracts the face of the person as the person from the transformed image. The requirement related to the suitability of the presence of the person is that parts other than the face of the person are recognized around the face.
3. The image processing apparatus according to claim 1, wherein: The requirement related to the suitability of the presence of the person is that the region in the transformed image where the person exists is recognized as a region where a pedestrian can pass.
4. The image processing apparatus according to claim 1, wherein: The requirement related to the suitability of the presence of the person is that the person also exists in the transformed images at the previous and next time points in the time series of the transformed images.
5. The image processing apparatus according to claim 1, wherein: The processing unit deletes the converted image or performs the image conversion process again on the input image as the predetermined process when a negative result is determined as to whether the requirement related to suitability of the person's presence is satisfied. The image processing apparatus according to claim 1 , wherein: The processing unit stores the transformed image as learning data as the predetermined processing when a result of determination as to whether or not the requirement related to suitability of the person's presence is satisfied is affirmative.
7. The image processing apparatus according to any one of claims 1 to 6, wherein: The image conversion process is a process of aligning the directions of the face of the person before and after the image conversion process and changing the face of the person to the face of another person.
8. An image processing method, wherein: The image processing method enables the computer to perform the following processing: extracting a person from a transformed image obtained by performing image transformation processing on an input image; and It is determined whether the extracted person satisfies a requirement related to suitability of the person's presence, and a predetermined process is performed on the transformed image based on a result of the determination.
9. A program, wherein The program causes the computer to perform the following processing: extracting a person from a transformed image obtained by performing image transformation processing on an input image; and It is determined whether the extracted person satisfies a requirement related to suitability of the person's presence, and a predetermined process is performed on the transformed image based on a result of the determination.
Citation Information
Patent Citations
Face collation device and passage control device
JP2005242777A