Video providing system and user terminal
The video providing system uses person re-identification and image processing to anonymize individuals in videos, addressing privacy concerns by accurately tracking and concealing specific persons, thus enhancing privacy protection.
Patent Information
- Application Number
- JP2023075607
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-01
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-05-01
AI Technical Summary
Existing video systems fail to adequately protect the privacy of individuals appearing in captured videos when provided to users, leading to potential identification and misuse of personal information.
A video providing system that utilizes person re-identification technology to track and anonymize individuals in videos based on their position information, performing image processing to conceal specific persons without affecting the visibility of intended viewers, using machine learning-based feature extraction and image processing techniques.
Accurately tracks and anonymizes individuals in videos, ensuring high privacy protection by preventing misidentification of intended and unintended viewers, thereby maintaining privacy while allowing authorized viewing.
Smart Images

Figure 0007782507000001 
Figure 0007782507000002 
Figure 0007782507000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for providing video to a user. [Background technology]
[0002] Patent Document 1 discloses a video surveillance system. A surveillance camera is installed with a viewing angle that overlooks a personal authentication device installed near a security gate. The video surveillance system captures video of the area near the security gate. The video surveillance system extracts people from the video captured by the surveillance camera and identifies the person who has been authenticated for access based on the position and behavior of the extracted person. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-154134 Summary of the Invention [Problem to be solved by the invention]
[0004] When providing video to users, it is desirable to properly protect the privacy of people appearing in the video.
[0005] One object of the present disclosure is to provide a technology that can appropriately protect the privacy of people appearing in a video when the video is provided to a user. [Means for solving the problem]
[0006] One aspect of the present disclosure relates to a video providing system that provides a video to a user. The video providing system includes one or more processors. A concealed person is a person other than a displayed person that the user wishes to see. The one or more processors Acquire one or more images captured by one or more cameras; By performing a person re-identification process, each person appearing in one or more videos is tracked, and person information including position information of each person in one or more videos is acquired; performing image processing on one or more images based on the person information so as to conceal the person to be concealed without concealing the person to be displayed; Displaying one or more images after image processing on the user's user terminal. It is configured as follows. [Effects of the Invention]
[0007] According to the present disclosure, when a video is provided to a user, a person to be anonymized who appears in the video is anonymized. This makes it possible to appropriately protect the privacy of the person to be anonymized. Furthermore, by using a person re-identification process, it becomes possible to stably and accurately track each person appearing in one or more videos. This reduces misrecognition of a person to be displayed and a person to be anonymized. Therefore, it becomes possible to perform image processing with high accuracy to anonymize a person to be anonymized without anonymizing a person to be displayed. As a result, it becomes possible to more appropriately protect the privacy of the person to be anonymized. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a conceptual diagram for explaining an overview of a video providing system according to an embodiment. [Figure 2] 1 is a block diagram illustrating an example of the configuration of a video providing system according to an embodiment. [Figure 3] 1 is a block diagram illustrating an example of a functional configuration of a person information management system according to an embodiment. [Figure 4] FIG. 10 is a conceptual diagram illustrating an example of person information according to the embodiment. [Figure 5] FIG. 10 is a conceptual diagram for explaining an example of processing by the video providing system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1. Overview FIG. 1 is a conceptual diagram for explaining an overview of a video providing system 10 according to this embodiment. The video providing system 10 acquires video VID captured by a camera CAM installed in a predetermined area. The video providing system 10 then provides (presents) the video VID to a user. More specifically, the video providing system 10 includes a user terminal, and displays the video VID on the user terminal. For example, in a monitoring service that allows parents to monitor their children, a video VID showing the child is displayed in real time on the parent's user terminal. The parent can monitor the child by viewing the video VID displayed on the user terminal.
[0010] For example, person P11 is a target that user A wants to watch. The person P11 appears in video VID-1 captured by camera CAM-1. However, the video VID-1 also includes other persons P12 and P13. In this case, when providing (presenting) the video VID-1 to user A, it is desirable to protect the privacy of persons P12 and P13 other than person P11. Therefore, the video providing system 10 performs image processing on the video VID-1 to anonymize the persons P12 and P13. In other words, the video providing system 10 performs image processing on the video VID-1 so that the persons P12 and P13 cannot be identified. For example, image processing is performed so that at least the faces of the persons P12 and P13 are masked. The video providing system 10 then displays the processed video VID-1′ on the user terminal of user A.
[0011] As another example, person P13 is a target that user B wants to watch. In this case, the video providing system 10 performs image processing on the video VID-1 so as to conceal persons P11 and P12 other than person P13. Then, the video providing system 10 displays the processed video VID-1′ on user B's user terminal.
[0012] Furthermore, video VID-2 captured by camera CAM-2 shows people P21 and P22. Here, let's assume that person P13 in video VID-1 and person P21 in video VID-2 are the same person. If the angles of view of cameras CAM-1 and CAM-2 do not overlap, the timing at which person P13 appears in video VID-1 and the timing at which person P21 appears in video VID-2 will be different. If the angles of view of cameras CAM-1 and CAM-2 partially overlap, it is possible that the timing at which person P13 appears in video VID-1 and the timing at which person P21 appears in video VID-2 will be the same person. In either case, video providing system 10 correctly recognizes (determines) that person P13 in video VID-1 and person P21 in video VID-2 are the same person, using a method described below. Then, the video providing system 10 performs image processing on the video VID-2 so as to conceal person P22 other than person P21. Then, the video providing system 10 displays the image-processed video VID-2' on the user terminal of user B. For example, the video providing system 10 switches between displaying the video VID-1' and the video VID-2'.
[0013] As another example, user C monitors the background (e.g., traffic) other than the people in video VID-1. In this case, video providing system 10 performs image processing on video VID-1 so as to conceal all of people P11, P12, and P13. Then, video providing system 10 displays the processed video VID-1' on user C's user terminal.
[0014] More generally, a person that a user wants to see is hereinafter referred to as a "display target person TD." The display target person TD differs for each user. Typically, the display target person TD is specified by the user. There may be cases where there is no display target person TD. On the other hand, a person who is anonymized so that he or she cannot be identified by the user is hereinafter referred to as a "concealment target person TX." The concealment target person TX is a person other than the display target person TD. For example, for user A, person P11 is a display target person TD, and persons P12 and P13 are concealment target persons TX. As another example, for user B, persons P13 and P21 are display target persons TD, and persons P11, P12, and P22 are concealment target persons TX. For user C, persons P11, P12, and P13 are concealment target persons TX.
[0015] The video providing system 10 acquires one or more videos VID-1 to VID-N captured by one or more cameras CAM-1 to CAM-N, where N is an integer equal to or greater than 1. The video providing system 10 tracks each person appearing in the one or more videos VID-1 to VID-N and acquires in-image position information of each person appearing in the one or more videos VID-1 to VID-N. Furthermore, the video providing system 10 performs image processing on the one or more videos VID-1 to VID-N based on the in-image position information so as to conceal the anonymization target person TX without concealing the display target person TD. In other words, the video providing system 10 performs image processing so as to identify the display target person TD but not the anonymization target person TX. The video providing system 10 then displays the one or more videos VID-1′ to VID-N′ after image processing on a user terminal.
[0016] By using this image processing, when providing a video VID to a user, it is possible to appropriately protect the privacy of the person TX to be concealed who appears in the video VID. Even if the person TX to be concealed differs for each user, it is possible to appropriately protect the privacy of the person TX to be concealed. In other words, it is possible to appropriately control privacy for each user.
[0017] In order to ensure the accuracy of image processing that conceals the concealment target person TX without concealing the display target person TD, it is necessary to accurately track the display target person TD and the concealment target person TX. If the concealment target person TX is mistakenly recognized as the display target person TD, or if the display target person TD is mistakenly recognized as the concealment target person TX, the accuracy of image processing will decrease. A decrease in the accuracy of image processing will lead to a decrease in the accuracy of privacy protection.
[0018] Therefore, according to this embodiment, "person re-identification" is used to track each person appearing in one or more videos VID-1 to VID-N. Person re-identification is a technology for identifying the same person from one or more videos VID-1 to VID-N. More specifically, people are detected from the images constituting the video VID, overall feature amounts of the detected people are extracted, and person re-identification processing is performed based on the extracted feature amounts. The feature amounts are extracted based on a person re-identification model based on machine learning.
[0019] Such person re-identification processing makes it possible to stably and accurately track each person appearing in one or more videos VID-1 to VID-N. For example, in FIG. 1, person P13 appearing in video VID-1 and person P21 appearing in video VID-2 are the same person, but this person can also be recognized with high accuracy by the person re-identification processing. In other words, the person re-identification processing suppresses erroneous recognition of the display target person TD and the concealment target person TX. Therefore, it becomes possible to perform image processing with high accuracy to conceal the concealment target person TX without concealing the display target person TD. As a result, it becomes possible to more appropriately protect the privacy of the concealment target person TX.
[0020] A specific example of the video providing system 10 according to this embodiment will be described in detail below.
[0021] 2. Example of a video provision system 2 is a block diagram showing an example of the configuration of a video providing system 10 according to this embodiment. The video providing system 10 includes a video management system 100, a person information management system 200, and a user terminal 300 used by a user.
[0022] 2-1.Video management system The video management system 100 acquires and manages one or more videos VID-1 to VID-N captured by one or more cameras CAM-1 to CAM-N. The video database 130 is a database of one or more videos VID-1 to VID-N acquired during a predetermined period.
[0023] The video management system 100 includes one or more processors 110 (hereinafter simply referred to as "processors 110") and one or more storage devices 120 (hereinafter simply referred to as "storage devices 120"). The processor 110 performs various processes. For example, the processor 110 includes a CPU. The storage device 120 stores various information required for the processes. The video database 130 is also stored in the storage device 120. Examples of the storage device 120 include an HDD, an SSD, a volatile memory, and a non-volatile memory. The functions of the video management system 100 may be realized by cooperation between the processor 110, which executes a computer program, and the storage device 120.
[0024] 2-2.Personal Information Management System The person information management system 200 acquires one or more videos VID-1 to VID-N captured by one or more cameras CAM-1 to CAM-N. The person information management system 200 then acquires and manages person information PSN related to each person appearing in the one or more videos VID-1 to VID-N. For example, the person information management system 200 performs a person re-identification process to track each person appearing in the one or more videos VID-1 to VID-N with high accuracy. The person information PSN includes in-image position information indicating the position of each person in the one or more videos VID-1 to VID-N. The person database 230 is a database of person information PSN acquired during a predetermined period.
[0025] The person information management system 200 includes one or more processors 210 (hereinafter simply referred to as "processor 210") and one or more storage devices 220 (hereinafter simply referred to as "storage device 220"). The processor 210 executes various processes. For example, the processor 210 includes a CPU. The storage device 220 stores various information required for the processes. The person database 230 is also stored in the storage device 220. Examples of the storage device 220 include an HDD, an SSD, a volatile memory, and a non-volatile memory. The functions of the person information management system 200 may be realized by cooperation between the processor 210, which executes a computer program, and the storage device 220.
[0026] 3 is a block diagram showing an example of the functional configuration of the person information management system 200. The person information management system 200 includes, as functional blocks, a person detection unit 211, a tracking unit 212, a face authentication unit 213, a posture estimation unit 214, and a person information provision unit 215.
[0027] One or more videos VID-1 to VID-N are input to the person detection unit 211. Each video VID includes a series of images (frames). The person detection unit 211 performs a person detection process to detect a person in each image. A bounding box indicates the position of the person detected in the image. The person detection unit 211 acquires information about the bounding box of the person in each image. Note that the person detection process is a well-known technique, and the method is not particularly limited. For example, YOLOX is used as the person detection unit 200.
[0028] The tracking unit 212 automatically tracks each person appearing in one or more videos VID-1 to VID-N.
[0029] For example, the tracking unit 212 tracks the same person in a video VID captured by a certain camera CAM. Specifically, the tracking unit 212 tracks bounding boxes representing the same person in a series of images constituting the video VID. Multiple bounding boxes representing the same person at multiple time steps are associated with each other. A "track" is a set of multiple bounding boxes representing the same person at multiple time steps. Such tracking processing is a well-known technique, and the method is not particularly limited. For example, ByteTrack is used.
[0030] Furthermore, the tracking unit 212 tracks the same person appearing in one or more videos VID-1 to VID-N by performing a person re-identification process. In particular, the tracking unit 212 tracks the same person appearing in different videos VID captured by different cameras CAM by performing the person re-identification process. More specifically, the tracking unit 212 extracts person features (hereinafter referred to as "ReID features") for the re-identification process based on an image of a portion of the person. A partial image surrounded by each bounding box in an image corresponds to the image of the person. The tracking unit 212 extracts ReID features from the partial images by using a ReID model based on machine learning. In particular, the tracking unit 212 extracts ReID features for multiple bounding boxes that form tracks. The extracted multiple ReID features are associated with the tracks. The tracking unit 212 then performs matching based on the ReID features to determine whether two or more different tracks are identical. That is, the tracking unit 212 determines whether two or more people appearing in one or more videos VID-1 to VID-N are the same based on the ReID feature. Note that such person re-identification processing is a well-known technique, and the method is not particularly limited. The ReID model may be a model based on a Transformer.
[0031] The face authentication unit 213 holds face images of people that have been registered in advance. The face authentication unit 213 performs face authentication processing based on the pre-registered face images to individually identify a person appearing in one or more videos VID-1 to VID-N. For example, a face image of a display target person TD is registered in advance by a user. The face authentication unit 213 performs face authentication processing based on the pre-registered face images to individually identify a display target person TD appearing in one or more videos VID-1 to VID-N. The individually identified person is associated with the above-mentioned track. Note that such face authentication processing is a well-known technique, and the method thereof is not particularly limited.
[0032] The pose estimation unit 214 performs a pose estimation process to estimate the pose of a person based on an image of a portion of the person. The partial image surrounded by each bounding box in the image corresponds to the image of the person. The pose estimation unit 214 extracts key points from the partial image by using a pose estimation model based on machine learning, and estimates the pose of the person. In particular, the pose estimation unit 214 estimates the pose of the person in each bounding box that constitutes a track. The estimated pose information is associated with the track. Note that the pose estimation process is a well-known technique, and the method is not particularly limited. For example, TransPose is used as the pose estimation unit 214.
[0033] The tracking unit 212 generates person information PSN relating to each person appearing in one or more videos VID-1 to VID-N. For example, the person information PSN is generated for each track.
[0034] FIG. 4 shows an example of person information PSN. In the example shown in FIG. 4, the person information PSN indicates correspondences among track ID, person ID, camera ID, timestamp, in-image position information, posture information, etc. The track ID is identification information of the track. The person ID is identification information of the person associated with the track and identified by the face authentication unit 213. Information on each image including each bounding box constituting the track is obtained from one or more videos VID-1 to VID-N. The camera ID is identification information of the camera CAM that captured each image. The timestamp is the time when each image was captured. The in-image position information is position information of the person (bounding box) in each image, and is obtained by the tracking unit 212 tracking the person. The posture information is information on the posture of the person in each image, and is obtained by the posture estimation unit 214.
[0035] The person database 230 is a database of person information PSN acquired during a predetermined period. The tracking unit 212 generates and updates the person database 230 by accumulating the person information PSN.
[0036] In response to a request from the user terminal 300 , the person information providing unit 215 acquires the necessary person information PSN from the person database 230 and provides the person information PSN to the user terminal 300 .
[0037] 2-3.User terminal The user terminal 300 includes one or more processors 310 (hereinafter simply referred to as "processor 310"), one or more storage devices 320 (hereinafter simply referred to as "storage device 320"), and a user interface 330. The processor 310 executes various processes. For example, the processor 310 includes a CPU. The storage device 320 stores various information required for the processes. Examples of the storage device 320 include an HDD, an SSD, a volatile memory, and a non-volatile memory. The functions of the user terminal 300 may be realized by cooperation between the processor 310, which executes a computer program, and the storage device 320.
[0038] The user interface 330 is an interface that receives information from a user and outputs information to the user. The user interface 330 includes an input device and an output device. Examples of the input device include a touch panel, a keyboard, and buttons. Examples of the output device include a display device and a speaker.
[0039] The user terminal 300 can communicate with the video management system 100 and the person information management system 200. The user terminal 300 acquires the necessary video VID from the video management system 100. The user terminal 300 also acquires the necessary person information PSN from the person information management system 200. The video VID and the person information PSN are stored in the storage device 320. The user terminal 300 then performs image processing on the video VID based on the person information PSN. The user terminal 300 then displays the processed video VID' on a display device, i.e., provides (presents) the processed video VID' to the user. An example of processing by the user terminal 300 will be described in more detail below.
[0040] 3. Example of processing by user terminal 3-1. First example The user specifies the person ID of the person TD to be displayed using the user interface 330 (input device). The user terminal 300 transmits a request including the person ID of the specified person TD to be displayed to the person information management system 200.
[0041] The personal information management system 200 acquires personal information PSN related to the specified display target person TD from the person database 230. For example, the personal information PSN related to the display target person TD includes the personal information PSN of the anonymization target person TX who is captured on the same camera CAM as the display target person TD. The personal information PSN of such an anonymization target person TX has a camera ID that is common to the personal information PSN of the display target person TD. The personal information PSN related to the display target person TD may include the personal information PSN of the display target person TD itself. The personal information management system 200 transmits the acquired personal information PSN to the user terminal 300.
[0042] The user terminal 300 checks the camera ID included in the person information PSN acquired from the person information management system 200. Then, the user terminal 300 acquires the video VID captured by the camera CAM corresponding to the camera ID. For example, the user terminal 300 transmits a request including the camera ID to the video management system 100. The video management system 100 selects the video VID captured by the camera CAM corresponding to the camera ID from the video database 130 and transmits the selected video VID to the user terminal 300. As another example, the user terminal 300 may first acquire videos VID-1 to VID-N from the video management system 100 and then select from them the video VID captured by the camera CAM corresponding to the camera ID.
[0043] The user terminal 300 performs image processing on the video VID to conceal the anonymization target person TX based on the personal information PSN of the anonymization target person TX. More specifically, the user terminal 300 temporally associates the image with the personal information PSN based on the timestamp of each image (frame) included in the video VID and the timestamp included in the personal information PSN. In-image position information of the anonymization target person TX within the image is obtained from the personal information PSN. Therefore, the user terminal 300 can perform image processing to conceal the anonymization target person TX within the image. For example, if the personal information PSN includes posture information, the user terminal 300 may identify the facial position of the anonymization target person PSN based on the posture information and mask the facial portion of the anonymization target person PSN. As another example, the user terminal 300 may mask the entire position (bounding box) of the anonymization target person TX. Note that the display target person TD does not need to be masked.
[0044] The user terminal 300 then displays the processed video VID' on a display device. When the video VID showing the display target person TD is switched, the user terminal 300 switches the video VID' to be displayed in conjunction with the switching.
[0045] 3-2. Second example The user specifies a camera using the user interface 330 (input device). The user terminal 300 transmits a request including the camera ID of the specified camera to the person information management system 200. The person information management system 200 acquires person information PSN including the specified camera ID from the person database 230, and transmits the acquired person information PSN to the user terminal 300. Furthermore, the user terminal 300 acquires video VID captured by the camera CAM corresponding to the camera ID, as in the first example above.
[0046] In the second example, the person to be displayed TD is not specifically designated, so all people appearing in the video VID are the person to be concealed TX. As in the first example, the user terminal 300 performs image processing on the video VID to conceal the person to be concealed TX. The user terminal 300 then displays the processed video VID' on the display device.
[0047] 3-3. Third Example A combination of the first and second examples above is also possible. In this case, the user specifies the camera and the person TD to be displayed.
[0048] 3-4. Fourth Example As described above, the personal information management system 200 acquires the personal information PSN by skipping various processes. Therefore, the sampling rate (acquisition rate) of the personal information PSN may be lower than the frame rate of the video VID. For example, the frame rate of the video VID may be 30 FPS, while the sampling rate of the personal information PSN may be 10 FPS.
[0049] For example, in FIG. 5, the video VID includes frames FR1 to FR10. The timestamps of the frames FR1 to FR10 are ts1 to ts10, respectively. On the other hand, only the timestamps ts1, ts4, ts7, and ts10 of the person information PSN are available. Even in such a case, it is desirable to accurately perform the concealment process based on the person information PSN.
[0050] Therefore, the user terminal 300 selects frames FR1, FR4, FR7, and FR10 corresponding to the timestamps of the person information PSN from among the frames FR1 to FR10 constituting the video VID. Then, the user terminal 300 selectively performs image processing on the selected frames FR1, FR4, FR7, and FR10. After that, the user terminal 300 selectively displays the frames FR1, FR4, FR7, and FR10 after image processing on a display device. The remaining frames are not displayed. Although the frame rate of the displayed video VID' decreases, the anonymization of the anonymization target person TX is performed with high accuracy. In other words, the privacy of the anonymization target person TX is appropriately protected.
[0051] 4. Variations A part of the processing by the user terminal 300 may be executed by the person information management system 200. For example, the person information management system 200 may perform image processing for each user, generate a post-image-processing video VID' for each user, and transmit the post-image-processing video VID' to the user terminal 300 of that user. This also provides the same effect as the above-described embodiment. [Explanation of symbols]
[0052] 10...Video providing system, 100...Video management system, 200...Personal information management system, 300...User terminal, PSN...Personal information, TD...Person to be displayed, TX...Person to be concealed, VID...Video
Claims
1. A video providing system for providing a video to a user, one or more processors; The person to be concealed is a person other than the person to be displayed that the user wishes to see, the one or more processors: Acquire one or more images captured by one or more cameras; By performing a person re-identification process, each person appearing in the one or more videos is tracked, and person information including position information of each person in the one or more videos is acquired; performing image processing on the one or more images based on the person information so as to conceal the person to be concealed without concealing the person to be displayed; Displaying the one or more images after the image processing on a user terminal of the user. It is configured as follows: If the sampling rate of the person information is lower than the frame rate of the one or more videos, the one or more processors select a frame corresponding to a timestamp of the person information from among frames constituting the one or more videos, perform the image processing on the selected frame, selectively display the selected frame after the image processing on the user terminal, and do not display frames other than the selected frame on the user terminal. Video provision system.
2. The video providing system according to claim 1, In the image processing, the one or more processors anonymize the person to be anonymized based on the position information included in the personal information about the person to be anonymized. Video provision system.
3. 3. The video providing system according to claim 2, The one or more processors further estimate a pose of each of the people in the one or more images; the person information includes the position information and the posture information of each person, In the image processing, the one or more processors mask at least a face of the person to be anonymized based on the position information and the posture information included in the personal information on the person to be anonymized. Video provision system.
4. 4. The video providing system according to claim 1, The one or more processors further perform a face recognition process to identify the person to be displayed who appears in the one or more images. Video provision system.
5. A user terminal used by a user, a processor for acquiring one or more images captured by one or more cameras; the person information management system performs a person re-identification process to track each person appearing in the one or more videos and acquire person information including position information of each person in the one or more videos; The concealed person is a person other than the displayed person that the user wants to see, The processor: Acquire the personal information from the personal information management system; performing image processing on the one or more images based on the person information so as to conceal the person to be concealed without concealing the person to be displayed; Displaying the one or more images after the image processing It is configured as follows: If the sampling rate of the person information is lower than the frame rate of the one or more videos, the processor selects a frame corresponding to a timestamp of the person information from among frames constituting the one or more videos, performs the image processing on the selected frame, and selectively displays the selected frame after the image processing, and does not display frames other than the selected frame. User terminal.
Citation Information
Patent Citations
Video processing method, device, electronic equipment and readable storage medium
CN110363172A
Video coding method and device and electronic equipment
CN111614973A
Child nursing method and child nursing system and care service system
JP2002269209A
Image encryption system and method
JP2005229265A
Method and apparatus for automatically blurring faces
JP2005512203A