Image storage system and image storage method
The image storage system derives facial orientation from three-dimensional data to predict behavior by masking faces, addressing the challenge of privacy protection without losing predictive information.
Patent Information
- Application Number
- JP2022110667
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-07-08
AI Technical Summary
Existing image processing technologies mask a subject's face to protect privacy, making it difficult to predict their behavior due to the loss of facial orientation information.
An image storage system and method that uses an optical distance measuring device to derive the orientation of a subject's face from three-dimensional information, masks the face, and adds orientation information to the processed image data, associating it with timestamps.
Facial orientation information is retained, enabling accurate prediction of the subject's behavior while maintaining privacy by masking their face.
Smart Images

Figure 0007735948000001 
Figure 0007735948000002 
Figure 0007735948000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for processing and storing captured images of a surrounding environment. [Background technology]
[0002] The drive recorder described in Patent Document 1 includes an image extraction unit that extracts an area in which another vehicle is captured from an image captured by a camera mounted on the vehicle, and an image processing unit that converts the area in which the other vehicle is captured into a vehicle characteristic image that reflects characteristics related to the vehicle type, and superimposes an occupant characteristic image that reflects the characteristics of the vehicle occupant on the vehicle characteristic image. Vehicle occupants are classified into adult males and adult females. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-176235 Summary of the Invention [Problem to be solved by the invention]
[0004] When a captured image includes a subject on the road, it is desirable to mask the subject's face to protect the subject's privacy; however, masking the subject's face makes it difficult to read the subject's intentions, which may indicate the subject's future direction of travel.
[0005] An object of the present invention is to provide a technology that masks the face of a subject while retaining information that is useful for predicting the subject's behavior. [Means for solving the problem]
[0006] In order to solve the above problem, an image storage system according to one aspect of the present invention is By camera Captured image 、 and The distance to the object was measured using an optical distance measuring device that emits and receives light.The device comprises an acquisition unit that acquires three-dimensional information, a derivation unit that detects the face of a subject included in a captured image and derives the orientation of the detected subject's face from the three-dimensional information, an image processing unit that masks the face of the subject included in the captured image and generates a processed image in which information indicating the orientation of the face is added to the data of the captured image, and a storage unit that stores the data of the processed image generated by the image processing unit. The captured image and the 3D information are associated based on timestamps indicating the detection times attached to the captured image and the 3D information, respectively. The derivation unit derives the facial features of the subject based on the 3D information and derives the facial orientation according to the positions of the derived facial features. The image processing unit adds the subject's position information to the processed image data. The added subject's position information is calculated based on the vehicle's position information and the subject's distance and direction from the vehicle as shown in the 3D information.
[0007] Another aspect of the present invention is an image storage method, in which each step is executed by a computer, the image storage method including: 、 and The distance to the object was measured using an optical distance measuring device that emits and receives light. The method includes a step of acquiring three-dimensional information, a step of detecting the face of the subject included in the captured image and deriving the orientation of the detected subject's face from the three-dimensional information, a step of masking the face of the subject included in the captured image and adding information indicating the orientation of the face to the captured image data to generate a processed image, and a step of storing the generated processed image data. The acquired captured image and 3D information are associated with each other based on timestamps indicating the detection times attached to the captured image and the 3D information, respectively. In the deriving step, facial features of the subject are derived based on the 3D information, and the facial orientation is derived according to the positions of the derived facial features. In the processed image generating step, position information of the subject is added to the processed image data. The added position information of the subject is calculated based on the position information of the vehicle and the distance and direction from the vehicle of the subject shown in the 3D information. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a technology that masks the face of a subject while retaining information that is useful for predicting the subject's behavior. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a diagram showing a captured image of the surrounding environment of the vehicle. [Figure 2] 2 is a diagram showing a processed image obtained by processing the captured image shown in FIG. 1. FIG. [Figure 3] FIG. 1 is a diagram illustrating a functional configuration of an image storage system. [Figure 4] FIG. 10 is a diagram illustrating a face image and a process for deriving the face direction. [Figure 5] 1. FIG. 4 is a diagram showing a modified example of a processed image obtained by processing the captured image shown in FIG. [Figure 6] 10 is a flowchart of an image storage process according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Fig. 1 is a diagram showing a captured image 10 of the environment surrounding a vehicle. The captured image 10 is generated by capturing an image of the area in front of the vehicle using a camera mounted on the vehicle. While Fig. 1 shows a captured image 10 of the area in front of the vehicle, the image is not limited to this and may be a captured image of the area behind or to the side of the vehicle.
[0011] Captured image 10 includes subject 12 and subject 14. Subject 12 is walking along the sidewalk facing forward, while subject 14 is walking from the sidewalk toward the road. Face recognition processing has been performed on captured image 10, and a detection frame 16 surrounding the face of subject 12 and a detection frame 18 surrounding the face of subject 14 have been added.
[0012] The captured image 10 is processed and transmitted to a server device, and used for driving assistance of the vehicle. For example, if it is estimated that the subject 14 is about to enter the road, information indicating the presence of the subject 14 is notified to the driver as driving assistance information. The driving assistance information is output by a speaker, a display, or the like. The vehicle may be capable of autonomous driving, and driving assistance information for autonomous driving may be generated based on the captured image 10.
[0013] Fig. 2 is a diagram showing a processed image 11 obtained by processing the captured image 10 shown in Fig. 1. Fig. 2 is an enlarged view focusing on a subject 12 in the captured image 10 shown in Fig. 1. The subject 12 included in the processed image 11 has been masked to create a mask area 20 where the face of the subject 12 was located. This makes it difficult to identify the subject who appears in the captured image 10 when the captured image 10 is used by a system including a vehicle, and makes it possible to protect personal information.
[0014] The subject often faces the direction of travel, and information indicating the facial orientation is useful for collision detection, etc. While the face of the subject 12 is rendered unidentifiable as a mask area 20, facial orientation information 22 indicating the facial orientation of the subject 12 is added to the processed image 11. This allows the vehicle system to determine the facial orientation of the subject 12 even when the subject's face is masked, making it easier to predict the subject's 12's traveling direction based on the facial orientation information 22. By processing the captured image 10 in this manner, it is possible to protect the personal information of the subject 12 appearing in the captured image 10 while making it easier to predict the behavior of the subject 12. While FIGS. 1 and 2 show an example in which the subject 12 is a pedestrian, the subject 12 may also be a cyclist.
[0015] 3 is a diagram showing the functional configuration of the image storage system 1. In terms of hardware, each function of the image storage system 1 can be configured using circuit blocks, memory, and other LSIs, and in terms of software, it is realized by system software, application programs, etc. loaded into memory. Therefore, it will be understood by those skilled in the art that each function of the image storage system 1 can be realized in various ways using only hardware, only software, or a combination of both, and is not limited to any one of these.
[0016] The image storage system 1 includes an image storage device 24, a camera 26, an optical sensor 28, and a server device 30. The image storage device 24 includes an acquisition unit 32, a location information acquisition unit 34, a processing unit 36, a map information holding unit 38, a storage unit 40, and a communication unit 42.
[0017] The camera 26 is mounted on the vehicle, captures images of the environment surrounding the vehicle, and sends the captured images to the image storage device 24. The camera 26 may capture images not only of the area in front of the vehicle, but also of the areas behind and to the sides of the vehicle.
[0018] The optical sensor 28 is, for example, a LiDAR (light detection and ranging) optical distance measuring device. The optical sensor 28 irradiates an object within its measurement range with light such as a laser, receives light reflected from the object in response to the irradiation of the irradiated light, and measures the distance to the object based on a signal corresponding to the received light. The distance to the object is calculated based on the time the irradiated light is emitted and the time the reflected light is received. The direction of the object is determined by the direction in which the irradiated light is emitted. In this way, the optical sensor 28 measures the distance and direction of the object from the vehicle by irradiating the object in the direction of travel of the vehicle. The optical sensor 28 obtains three-dimensional information consisting of a point cloud indicating the distance and direction from the vehicle. The camera 26 and the optical sensor 28 send their detection results to the image storage device 24. The detection results of the camera 26 and the optical sensor 28 are assigned time-synchronized timestamps, allowing the two sets of data to be associated using the timestamps.
[0019] The image storage device 24 receives the captured image 10 from the camera 26, performs masking on the captured image 10 to generate a processed image 11, and transmits the processed image 11 to the server device 30. The information transmitted from the image storage device 24 includes not only the processed image 11 but also vehicle driving information and surrounding environment information detected by an on-board sensor (not shown). The on-board sensor may be, for example, a vehicle speed sensor, a shift position detection sensor, a steering angle sensor, a millimeter-wave radar, or an ultrasonic sensor. The server device 30 analyzes the processed image 11 received from the image storage device 24 and generates driving assistance information.
[0020] The acquisition unit 32 of the image storage device 24 acquires the captured image from the camera 26 and acquires the 3D information from the optical sensor 28. The captured image and the 3D information are each provided with a timestamp indicating the time of detection. Therefore, the 3D information detected at approximately the same time as the captured image can be associated with the captured image.
[0021] The location information acquisition unit 34 acquires vehicle location information using the Global Navigation Satellite System (GNSS). A timestamp is added to the vehicle location information. The vehicle location information is indicated by latitude and longitude. The map information storage unit 38 stores map information indicating addresses and roads in association with latitude and longitude. Therefore, it is possible to derive which road the vehicle is traveling on from the location information acquired by the location information acquisition unit 34. The map information also includes road attribute information such as vehicles, sidewalks, and crosswalks.
[0022] The processing unit 36 performs image processing and has a derivation unit 44 and an image processing unit 46. The derivation unit 44 detects the face of the subject included in the captured image 10 and derives the orientation of the detected face of the subject from the three-dimensional information. The face of the subject is detected by analyzing the captured image 10 using a known face recognition process. The process of detecting the face of the subject from the captured image 10 may be a pattern matching technique or may be performed using a model learned by machine learning. The derivation unit 44 detects the face of the subject and generates detection frames 16 and 18 as shown in FIG. 1.
[0023] FIG. 4 illustrates a face image 50 and is a diagram for explaining a process for deriving a face direction. The derivation unit 44 derives the face direction of the detected subject using three-dimensional information. The derivation unit 44 derives facial features of the subject using three-dimensional information consisting of a point cloud. For example, the derivation unit 44 derives that the most protruding part of the point cloud included in the face is the nose and that the most depressed part is the eye. The derivation unit 44 may use information indicating the positional relationship of the facial features. Once a specific facial feature is derived, the derivation unit 44 derives other facial features based on other point cloud regions indicating concavity and convexity and information indicating the positional relationship of the face. The derivation unit 44 derives a nose 52a, eyes 52b, mouth 52c, and ears 52d included in the face image 50. Note that the derivation unit 44 may derive only the nose 52a. The nose 52a may be located at the center of the point cloud region indicating the convexity.
[0024] The derivation unit 44 sets a reference position 54 of the face image 50, derives the positional relationship between the reference position 54 and facial parts, and derives the face orientation based on the derived positional relationship. The reference position 54 may be set to the center position of the face, or may be set to the average position of the nose 52a on a face facing straight ahead. Simply put, the derivation unit 44 derives the face orientation based on the positional relationship between the nose 52a and the reference position 54. If the nose 52a and the reference position 54 are in the same position, the face is facing straight ahead.
[0025] As shown in FIG. 4 , if the nose 52a is located to the left of the reference position 54 in the horizontal direction, the face is facing rightward. If the nose 52a is located at the same height as the reference position 54 in the vertical direction, the face is facing horizontally. In this way, the derivation unit 44 derives the facial direction in three-dimensional directions based on the positional relationship between the nose 52a and the reference position 54. The derivation unit 44 may derive the facial direction based on the positional relationship between the midpoint of the two eyes 52b and the reference position, or may derive the facial direction based on the positional relationship between the mouth 52c and the reference position. The reference position may be set for each facial feature. The derivation unit 44 may derive the facial direction based on the positional relationship between the positions of multiple facial features and the reference positions set for each feature. In either case, the derivation unit 44 derives the facial features of the subject based on the three-dimensional information and derives the facial direction based on the derived positions of the facial features. The derivation unit 44 holds a map or function that associates the positional relationship between the nose 52a and the reference position 54 with the facial orientation, and derives the facial orientation by inputting the map or function with the positional relationship between the nose 52a and the reference position 54.
[0026] The optical sensor 28 can detect a distance of several tens of meters, and people who are located at a distance that cannot be detected by the optical sensor 28 are not considered as subjects and are excluded from this process.
[0027] Returning to FIG. 3, the image processing unit 46 masks the subject's face included in the captured image 10 and adds information indicating the orientation of the face to the data of the captured image 10, thereby generating the processed image 11 shown in FIG. 2. The masking process for masking the subject's face may be a filling process, a blurring process, a process for converting the face into a computer graphics facial image, or the like. In FIG. 2, a mask area 20 is filled. Furthermore, the mask area 20 may be the area inside the detection frames 16 and 18 set in the facial recognition process shown in FIG. 1. Although the subject's facial image is masked, the subject's body from the shoulders down is left untouched.
[0028] The image processing unit 46 masks the subject's face and adds information indicating the facial orientation, i.e., vector information in three-dimensional directions, to the captured image 10 to generate a processed image 11. The information indicating the facial orientation may be a vector in which a yaw angle, a pitch angle, and a roll angle are determined. The derivation unit 44 and the image processing unit 46 execute a process to generate a processed image for each frame of captured image. In FIG. 2, the information indicating the facial orientation added to the processed image 11 is indicated by an arrow, but it may be a vector associated with the subject 12.
[0029] The image processing unit 46 sends the generated processed image 11 to the memory unit 40, and the memory unit 40 stores the data of the processed image 11 generated by the image processing unit 46. The communication unit 42 transmits the processed image 11 stored in the memory unit 40 to the server device 30. This makes it possible to transmit the processed image 11 that makes it difficult to identify the subject. Furthermore, the server device 30 can accurately predict the subject's future behavior based on the direction of the subject's face. The communication unit 42 may transmit the processed image 11 to the server device 30 by including it in probe information. The probe information includes a vehicle ID, time, vehicle speed, vehicle acceleration, vehicle position information, etc., and is periodically transmitted to the server device 30.
[0030] Fig. 5 is a diagram showing a modified processed image 56 obtained by processing the captured image 10 shown in Fig. 1. The processed image 56 shown in Fig. 5 differs from the processed image 11 shown in Fig. 2 in that eye line information 58 and movement direction information 60 are added.
[0031] The derivation unit 44 analyzes the captured images 10 and derives eye line information 58 based on the position of the iris in the eye. The derivation unit 44 also derives the movement direction of the subject 12 based on the captured images 10 in chronological order. The derivation unit 44 tracks the subjects 12 included in the captured images 10 in chronological order and derives the movement direction of the subject 12. The derivation unit 44 assigns an ID to each subject for tracking. The image processing unit 46 adds eye line information 58 of the subject 12 and movement direction information 60 of the subject 12 to the data of the processed image 56. This makes it possible to determine whether the subject 12 is walking or stationary, thereby enabling accurate prediction of the future behavior of the subject 12.
[0032] The image processing unit 46 may add location information of the subject 12 to the data of the processed image 56. The location information of the subject 12 is calculated based on the vehicle's location information and the distance and direction from the vehicle of the subject shown in the 3D information. The location information of the subject 12 may include not only latitude and longitude, but also road attribute information such as the roadway, sidewalk, and crosswalk. In other words, information indicating where the subject 12 is located on the road is included in the location information of the subject 12. The road attribute information is calculated based on map information. This makes it possible to accurately calculate the possibility of a collision with a vehicle based on the location information of the subject 12. In addition, driving assistance information to be notified to the driver can be generated using geographical names.
[0033] 6 is a flowchart of the image storage process according to the embodiment. The acquisition unit 32 acquires an image of the vehicle's surroundings from the camera 26 (S10), and acquires three-dimensional information of the vehicle's surroundings detected from the optical sensor 28 (S12).
[0034] The derivation unit 44 performs face recognition processing on the captured image to detect an image of the subject's face (S14). The derivation unit 44 determines whether the subject is present, that is, whether a facial image of the subject has been detected (S16). Furthermore, the subjects may be limited to those detected using three-dimensional information, and the derivation unit 44 may determine whether there is a subject whose three-dimensional information corresponding to the facial image has been detected.
[0035] If the subject is not present (N in S16), the image processing unit 46 does not process the captured image, and the communication unit 42 transmits the captured image to the server device 30 (S24). If the subject is present (Y in S16), the derivation unit 44 derives the facial features of the subject based on the three-dimensional information, and derives the facial orientation according to the positions of the derived facial features (S18).
[0036] The image processing unit 46 performs a masking process to mask the facial image in the captured image (S20), and generates a processed image by adding information indicating the orientation of the face (S22). The communication unit 42 transmits the generated processed image to the server device 30 (S24). This process is repeated for each captured image.
[0037] The present disclosure has been described above based on examples. The present disclosure is not limited to the above examples, and various modifications such as design changes may be made based on the knowledge of those skilled in the art.
[0038] For example, in the embodiment, the server device 30 receives the processed image, generates the driving assistance information, and transmits the driving assistance information to the vehicle, but the present invention is not limited to this. For example, the server device 30 may calculate the driving skills and driving tendencies of the driver by analyzing the processed image after the fact.
[0039] Furthermore, although the embodiment shows an embodiment in which the optical sensor 28 is a LiDAR, the present invention is not limited to this embodiment. For example, the optical sensor 28 may be a stereo camera. In either case, the depth of the target person's face can be calculated based on the detection result of the optical sensor 28, and the facial parts can be identified.
[0040] Furthermore, in the embodiment, the processed image is generated on the vehicle side, but the present invention is not limited to this, and the processed image may be generated on the server device 30 side. [Explanation of symbols]
[0041] 1 Image storage system, 10 Captured image, 11 Processed image, 12, 14 Subject, 16, 18 Detection frame, 20 Mask area, 22 Face direction information, 24 Image storage device, 26 Camera, 28 Optical sensor, 30 Server device, 32 Acquisition unit, 34 Location information acquisition unit, 36 Processing unit, 38 Map information storage unit, 40 Memory unit, 42 Communication unit, 44 Derivation unit, 46 Image processing unit.
Claims
1. an acquisition unit that acquires a captured image of the surrounding environment captured by a camera and three-dimensional information obtained by measuring the distance to an object using an optical distance measuring device that emits and receives light; a derivation unit that detects a face of a subject included in a captured image and derives a direction of the detected face of the subject from three-dimensional information; an image processing unit that masks the face of the subject included in the captured image and generates a processed image by adding information indicating the orientation of the face to the captured image data; a storage unit that stores data of the processed image generated by the image processing unit, The captured image and the three-dimensional information are associated with each other based on time stamps indicating the detection times added to the captured image and the three-dimensional information, respectively; the deriving unit derives facial features of the subject based on the three-dimensional information, and derives a facial orientation according to the positions of the derived facial features; The image processing unit adds position information of the subject to data of the processed image, An image storage system characterized in that the added position information of the subject is calculated based on the position information of the vehicle and the distance and direction from the vehicle of the subject shown in the three-dimensional information.
2. The image storage system according to claim 1, characterized in that the derivation unit derives facial features of the subject based on three-dimensional information, and derives the facial orientation based on the positional relationship of the derived facial features from a reference position set on the facial image.
3. the derivation unit derives a moving direction of the subject based on the captured images in chronological order; 3. The image storage system according to claim 1, wherein the image processing unit adds the moving direction of the subject to the processed image data.
4. An image storage system as described in claim 1 or 2, characterized in that the added location information of the subject includes road attribute information indicating where the subject is located on the road.
5. An image storage method in which each step is performed by a computer, comprising: acquiring three-dimensional information obtained by capturing an image of the surrounding environment using a camera and measuring the distance to an object using an optical distance measuring device that emits and receives light; detecting a face of a subject included in a captured image and deriving a direction of the detected face of the subject from three-dimensional information; a step of masking the face of the subject included in the captured image and adding information indicating the orientation of the face to the captured image data to generate a processed image; and storing data of the generated processed image, The captured image and the three-dimensional information are associated with each other based on time stamps indicating the detection times added to the captured image and the three-dimensional information, respectively; In the deriving step, facial parts of the subject are derived based on the three-dimensional information, and a facial orientation is derived according to the positions of the derived facial parts; In the step of generating the processed image, position information of the subject is added to data of the processed image; An image storage method characterized in that the added position information of the subject is calculated based on the position information of the vehicle and the distance and direction from the vehicle of the subject shown in the three-dimensional information.
Citation Information
Patent Citations
Display processing system for vehicle, display processing method for vehicle, and program
JP2013032949A
Display device, display method, and program
JP2017135695A
Imaging state detection device, imaging state detection method, program, and non-transitory recording medium
JP2018106288A
Method and system for monitoring driving behavior
JP2018528536A
Control method, information processing device, information processing system, and program
JP2021089529A