Video provision system, video provision method and video provision program

The video providing system enhances usability by allowing users to determine and confirm masking processes on personal information, ensuring compliance with privacy regulations and effective utilization of masked video data as training data.

JP2025158872AActive Publication Date: 2025-10-17SAFIE INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024061810
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-06
Publication Date
2025-10-17
Estimated Expiration
2044-04-06

AI Technical Summary

Technical Problem

Existing video delivery systems lack efficient mechanisms for utilizing video data stored on servers as training data due to the need to remove personal information, and there is a lack of usability in the context of masking processes.

Method used

A video providing system that includes a setting screen for determining the type of masking process on personal information, performs the masking process on video data, and provides the masked video to a terminal, with confirmation and deletion options, enhancing usability and compliance with regulations.

Benefits of technology

Improves the usability of the video providing system by ensuring proper masking, enabling effective utilization of masked video data as training data while complying with privacy regulations, reducing storage, and enhancing searchability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025158872000001_ABST
    Figure 2025158872000001_ABST
Patent Text Reader

Abstract

To improve the usability of a video provision system in a context of masking processing to video.SOLUTION: A video provision system 1 stores videos photographed by cameras 2 in a server 3 and provides a terminal with the videos. The video provision system 1 makes a user terminal 4 display a setting screen of masking processing to be executed, determines a kind of the masking processing according to input operation of a user U to the setting screen, generates a second video by executing the determined kind of masking processing to an object showing individual information included in a first video, stores the second video in the server 3, and makes the server 3 provide a company side terminal 5 with the second video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a video providing system, a video providing method, and a video providing program. [Background technology]

[0002] With the development of AI (artificial intelligence) technology, the demand for training data necessary for building machine learning models is increasing. In particular, building highly accurate machine learning models (especially image recognition models using machine learning) requires a large amount of training data. In this regard, in a video delivery system equipped with a camera, a server, and a user terminal, a large amount of video data is accumulated on the server every day, and it is conceivable to effectively utilize the video data accumulated on the server as training data. However, since video data that can be used as training data contains personal information such as people's faces and vehicle license plates, it is necessary to remove the personal information contained in the video data before using the video data as training data in order to comply with laws and regulations such as the Act on the Protection of Personal Information.

[0003] In this regard, Patent Document 1 discloses an image processing technology that identifies a human face area (face area) contained in video data and then performs mosaic processing (an example of masking processing) on ​​the identified face area. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-205835 Summary of the Invention [Problem to be solved by the invention]

[0005] In the disclosure of Patent Document 1, a mosaic process is performed on human faces in video data, thereby removing personal information from the video data, making it possible to utilize the mosaic-processed video data as training data. Meanwhile, in a video delivery system equipped with a camera, a server, and a user terminal, few mechanisms have been proposed at present for efficiently utilizing video data stored in the server as training data. In particular, there is room for further study on new mechanisms for improving the usability of the video delivery system in the context of the masking process performed on the video data.

[0006] In view of the above, the present disclosure aims to improve the usability of a video providing system in the context of masking processing for video. [Means for solving the problem]

[0007] A video provision system according to one embodiment of the present disclosure is a system that stores video captured by a camera on a server and provides the video to a terminal, displays a setting screen for the masking process to be performed on a first terminal, determines the type of masking process in accordance with a user's input operation on the setting screen, and performs the determined type of masking process on objects that indicate personal information contained in the first video to generate a second video, stores the second video on the server, and provides the second video from the server to the second terminal.

[0008] According to the above configuration, a setting screen related to the setting of the masking process (in other words, the anonymization process of personal information) is displayed on the first terminal, and the type of masking process is determined in response to a user's input operation on the setting screen. The determined type of masking process is then performed on objects (e.g., people, vehicles, etc.) that indicate personal information included in the first video, and the second video is then provided to the second terminal. In this way, the user can determine the type of masking process to be performed on the first video through the setting screen, thereby improving the usability of the video providing system in the context of the masking process on the video. Furthermore, since the second video is provided to the second terminal, it is possible to effectively utilize the video stored on the server. For example, the masked video can be effectively used as training data for building a machine learning model (an image recognition model using machine learning).

[0009] In addition, the video providing system may display a confirmation screen on the first terminal to confirm whether the masking process has been properly performed on the second video, determine whether the second video can be used depending on the user's input operation on the confirmation screen, and if it is determined that the second video can be used, store the second video on the server.

[0010] According to the above configuration, a confirmation screen is displayed on the first terminal, and whether or not the second video can be used is determined based on the user's input operation on the confirmation screen. In this way, the user can understand through the confirmation screen whether masking has been properly performed on the second video, thereby improving the usability of the video provision system in the context of the video masking process. Furthermore, the user can objectively confirm that appropriate masking has been performed on the video in compliance with laws and regulations such as the Personal Information Protection Act, thereby ensuring sufficient transparency and reliability of the masking process. Therefore, a third party who receives masked video can use the video with peace of mind. Furthermore, because the masked video is provided to the second terminal, the video stored on the server can be effectively utilized. For example, the masked video can be effectively used as training data for building a machine learning model (an image recognition model using machine learning).

[0011] The video providing system may also display the first video and the second video side by side on the confirmation screen.

[0012] According to the above configuration, the user can clearly see whether the masking process has been properly performed on the second video through the confirmation screen on which the first video and the second video are displayed side by side.

[0013] The video providing system may also delete the first video after storing the second video.

[0014] According to the above configuration, the first video is deleted after the second video is saved, which makes it possible to effectively reduce the amount of data stored in the server and effectively reduce the maintenance and management costs of the server.

[0015] The video providing system may also acquire attribute information of the object and provide the attribute information together with the second video.

[0016] According to the above configuration, the attribute information and the second video are provided to the second terminal. Therefore, even if the attribute of the object cannot be identified from the second video (for example, even if the second video has been subjected to mosaic processing or blurring), the attribute information can be used to identify the attribute of the object. In this way, the masked video can be effectively used as training data for building a machine learning model.

[0017] The object may be a person, and the attribute information may include at least one of gender, age, face angle, and facial expression.

[0018] According to the above configuration, even if the attributes of a person included in the second video cannot be identified, it is possible to identify the attributes of the person based on attribute information including at least one of gender, age, facial angle, and facial expression. In this way, the masked video can be effectively used as training data for building a machine learning model related to human behavior, etc.

[0019] The second video may also be stored with metadata associated with the second video, which may include at least one of identification information of the second video, a time when the second video was captured, a location where the second video was captured, and the masking process performed on the second video.

[0020] According to the above configuration, it is possible to improve the searchability of the second video stored on the server by using metadata including at least one of video identification information, shooting time, shooting location, and masking process.

[0021] The video providing system may also display a search screen on the second terminal for searching for a desired second video from among a plurality of the second videos, determine search conditions for the second video in response to a user's input operation on the search screen, extract at least one second video corresponding to the determined search conditions from among the plurality of second videos stored on the server, and provide the extracted second video to the second terminal.

[0022] According to the above configuration, it is possible to provide the second terminal with a desired second video that matches the search conditions from among a plurality of second videos stored in the server.

[0023] The search conditions may also include at least one of the time of the video, the data size of the video, the masking process performed on the video, the type of object contained in the video, the location where the video was taken, the type of business where the video was taken, the area where the video was taken, the time period when the video was taken, and metadata associated with the video.

[0024] According to the above configuration, the operator of the second terminal can obtain a desired second image that matches at least one of the above information.

[0025] The setting screen may have a masking process selection area in which one of a plurality of types of masking processes can be selected, which may include a first masking process that makes it impossible to identify both the attribute information and personal information of the object, and a second masking process that makes it possible to identify the attribute information of the object but makes it impossible to identify the personal information of the object.

[0026] According to the above configuration, a setting screen having a masking process selection area is displayed on the first terminal, and the masking process is determined in response to a user's input operation on the setting screen. In this way, the user can determine through the setting screen whether to perform a first masking process (e.g., a mosaic process or a blur process) or a second masking process (e.g., a deep masking process) on the first video. This makes it possible to improve the usability of the video providing system in the context of the masking process on the video.

[0027] Another aspect of the present disclosure is a video provision system that stores video captured by a camera on a server and provides the video to a terminal. The system generates a second video by performing a masking process on objects that indicate personal information contained in a first video, displays a confirmation screen on the first terminal to confirm whether the masking process has been performed properly on the second video, determines whether the second video can be used depending on a user's input operation on the confirmation screen, and if it is determined that the second video can be used, stores the second video on the server and provides the second video from the server to the second terminal.

[0028] According to the above configuration, a confirmation screen is displayed on the first terminal, and whether or not the second video can be used is determined based on the user's input operation on the confirmation screen. In this way, the user can understand through the confirmation screen whether masking has been properly performed on the second video, thereby improving the usability of the video provision system in the context of the video masking process. Furthermore, the user can objectively confirm that appropriate masking has been performed on the video in compliance with laws and regulations such as the Personal Information Protection Act, thereby ensuring sufficient transparency and reliability of the masking process. Therefore, a third party who receives masked video can use the video with peace of mind. Furthermore, because the masked video is provided to the second terminal, the video stored on the server can be effectively utilized. For example, the masked video can be effectively used as training data for building a machine learning model (an image recognition model using machine learning).

[0029] Another aspect of the present disclosure is a video provision system that stores video captured by a camera on a server and provides the video to a terminal, generating a second video by performing a masking process on objects that indicate personal information contained in a first video, storing the second video on the server, displaying a search screen on the second terminal for searching for the desired second video from among multiple second videos, determining search conditions for the second video in response to user input operations on the search screen, extracting at least one second video that corresponds to the determined search conditions from among the multiple second videos stored on the server, and providing the extracted second video to the second terminal.

[0030] According to the above configuration, it is possible to provide the second terminal with the desired second video that matches the search criteria from among multiple second videos stored on the server, thereby improving the usability of the video provision system in the context of video searchability.

[0031] A video provision method according to one embodiment of the present disclosure is executed by a video provision system that stores video captured by a camera on a server and provides the video to a terminal, and includes the steps of: displaying a setting screen for masking processing to be performed on a first terminal; determining the type of masking processing in accordance with a user's input operation on the setting screen; generating a second video by performing the determined type of masking processing on an object that indicates personal information contained in the first video; storing the second video on the server; and providing the second video from the server to a second terminal.

[0032] Also provided is a video providing program that causes a video providing system to execute the video providing method. [Effects of the Invention]

[0033] According to the present disclosure, it is possible to improve the usability of a video providing system in the context of masking processing for video. [Brief explanation of the drawings]

[0034] [Figure 1] 1 is a diagram illustrating a video providing system according to an embodiment of the present disclosure (hereinafter referred to as the present embodiment). [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of a camera. [Figure 3] FIG. 2 illustrates an example of a hardware configuration of a server. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a user terminal. [Figure 5] 1 is a flowchart illustrating a series of processes executed by the video providing system according to the present embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a video display screen displayed on a user terminal. [Figure 7] FIG. 10 is a diagram illustrating an example of a setting screen displayed on a user terminal. [Figure 8]1A and 1B are diagrams for explaining masking processing performed on a person's face, where (a) is a diagram showing the original face, (b) is a diagram for explaining mosaic processing, (c) is a diagram for explaining blurring processing, and (d) is a diagram for explaining deep masking processing. [Figure 9] FIG. 10 is a diagram illustrating an example of a masking task list screen. [Figure 10] FIG. 10 is a diagram showing an example of a video comparison screen displayed on a user terminal. [Figure 11] FIG. 10 is a diagram showing an example of a data registration screen displayed on a user terminal. [Figure 12] FIG. 2 is a diagram illustrating an example of video management data. [Figure 13] FIG. 10 is a diagram illustrating an example of a video search screen. DETAILED DESCRIPTION OF THE INVENTION

[0035] (System configuration) A video providing system 1 according to this embodiment will be described below with reference to the drawings. FIG. 1 is a diagram illustrating the video providing system 1 according to this embodiment. As illustrated in FIG. 1, the video providing system 1 includes a camera 2, a server 3, a user terminal 4, and an enterprise terminal 5. These are connected to a communication network 8. Each of the multiple cameras 2 is communicatively connected to the server 3 via the communication network 8. In this example, two cameras 2 are illustrated, but the number of cameras 2 provided in the video providing system 1 is not particularly limited, and three or more cameras 2 may be provided. The server 3 is communicatively connected to the user terminal 4 and the enterprise terminal 5 via the communication network 8. The communication network 8 is configured by at least one of a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a wireless core network. For ease of explanation, the video providing system 1 according to this embodiment illustrates one server 3, one user terminal 4, and one enterprise terminal 5, but the number of these is not particularly limited.

[0036] (Camera 2 configuration) Next, the hardware configuration of camera 2 will be described below. FIG. 2 is a diagram showing an example of the hardware configuration of camera 2. Camera 2 is configured to acquire video data showing its surrounding environment through photography, and may be placed inside or near a store such as a convenience store or restaurant. As shown in FIG. 2, camera 2 includes a control unit 20, a storage device 21, a position information acquisition unit 22, a communication unit 24, an input operation unit 25, an imaging unit 26, and a PTZ mechanism 27, and these elements are connected to a communication bus 28. Camera 2 may also have a built-in battery (not shown). Camera 2 may also be provided with a microphone and a speaker.

[0037] The control unit 20 includes a memory and a processor. The memory is configured to store computer-readable instructions (programs). For example, the memory may include a read-only memory (ROM) storing various programs and a random access memory (RAM) having multiple work areas for storing various programs executed by the processor. The processor may include at least one of a central processing unit (CPU), a micro processing unit (MPU), and a graphics processing unit (GPU). The CPU may include multiple CPU cores. The GPU may include multiple GPU cores. The processor may be configured to load a specified program from various programs stored in the storage device 21 or the ROM onto the RAM and execute various processes in cooperation with the RAM.

[0038] The storage device 21 is a storage device (storage) such as a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc., and is configured to store programs and various data. The position information acquisition unit 22 is configured to acquire position information (longitude, latitude) of the camera 2, and is, for example, a global positioning system (GPS) receiver.

[0039] The communication unit 24 is configured to connect the camera 2 to the communication network 8. The communication unit 24 includes a wireless communication module for wirelessly communicating with external devices such as base stations and wireless LAN routers. The wireless communication module includes a transmitting / receiving antenna and a signal processing circuit. The wireless communication module may be a wireless communication module compatible with short-range wireless communication standards such as Wi-Fi (registered trademark) and Bluetooth (registered trademark), or may be a wireless communication module compatible with an X-generation mobile communication system (for example, a fourth-generation mobile communication system such as LTE) that uses a SIM (Subscriber Identity Module).

[0040] The input operation unit 25 is configured to receive input operations from an operator and to generate operation signals in response to the input operations from the operator. The imaging unit 26 is configured to capture an image of the surrounding environment of the camera 2. In particular, the imaging unit 26 is configured to generate a video signal indicating the surrounding environment of the camera 2, and includes an optical system, an image sensor, and an analog processing circuit. The optical system includes, for example, an optical lens and a color filter. The image sensor is configured by a CCD (Charge-Coupled Device) or a CMOS (Complementary Metal-Oxide Semiconductor) or the like. The analog processing circuit is configured to process the video signal (analog signal) photoelectrically converted by the image sensor, and includes, for example, an amplifier and an AD converter.

[0041] The PTZ mechanism 27 includes a pan mechanism, a tilt mechanism, and a zoom mechanism. The pan mechanism is configured to change the orientation of the camera 2 in the horizontal direction. The tilt mechanism is configured to change the orientation of the camera 2 in the vertical direction. The zoom mechanism is configured to enlarge (zoom in) or reduce (zoom out) an image showing an object to be captured by changing the angle of view of the camera 2. The zoom mechanism may optically change the angle of view of the camera 2 by changing the focal length of an optical lens included in the imaging unit 26, or may digitally change the angle of view of the camera 2. In this embodiment, in response to an input operation by the user U on the user terminal 4, an instruction signal instructing the camera 2 to pan, tilt, and / or tilt is transmitted from the user terminal 4 to the camera 2 via the server 3. In this case, the control unit 20 drives the PTZ mechanism 27 in response to the received instruction signal, thereby realizing the pan, tilt, and zoom functions (PTZ functions) of the camera 2 in real time. In this manner, the PTZ function of the camera 2 can be realized through remote control of the user terminal 4.

[0042] The camera 2 can transmit a video (video data stream) showing the surrounding environment of the camera 2 to the server 3 via the communication network 8 in real time.

[0043] (Server 3 configuration) Next, the hardware configuration of the server 3 will be described below. FIG. 3 is a diagram showing an example of the hardware configuration of the server 3. The server 3 is configured to receive video data from the camera 2 via the communication network 8 and transmit the video data to the user terminal 4 in response to a video transmission request from the user terminal 4. The server 3 may be configured with multiple servers. The server 3 functions as a web server configured to provide a cloud-based video distribution application as a web application. In this regard, the server 3 is configured to transmit data (e.g., HTML files, CSS files, image / video files, program files, etc.) for displaying a video display screen 50 (see FIG. 6) on the web browser of the user terminal 4. In this way, the server 3 functions as a server for providing SaaS (System as a Service). The server 3 may be constructed on-premise or may be a cloud server.

[0044] 3, the server 3 includes a control unit 30, a storage device 31, an input / output interface 32, a communication unit 33, an input operation unit 34, and a display unit 35. These elements are connected to a communication bus 36.

[0045] The control unit 30 includes a memory and a processor. The memory is configured to store computer-readable instructions. In particular, the memory may store a program that causes the processor to execute a series of processes executed by the server 3. The memory is configured with a ROM and a RAM. The processor is configured with at least one of a CPU, an MPU, and a GPU.

[0046] The storage device 31 is, for example, a storage device (storage) such as an HDD, SSD, or flash memory, and is configured to store programs and various data. User management data and camera management data are stored in the storage device 31. The storage device 31 also stores original video data (video data before masking), metadata associated with the original video data, masking task data, and an image recognition model (learning model). The storage device 31 also stores video data after masking (learning data), metadata associated with the video data after masking, and attribute information indicating the attributes of each object (e.g., person, vehicle, etc.) included in the video data. The storage device 31 also stores video management data (video management table) for managing multiple video data after masking.

[0047] The user management data includes management information for each user U who uses the video provision system 1. The camera management data includes management information for each camera 2. Because the original video data contains personal information such as people's faces and vehicle license plates, external access to the video data is restricted in order to comply with the Personal Information Protection Act. In this regard, even the business providing the video provision system 1 is restricted from accessing the original video data, and only the user U who operates the store where the camera 2 is installed (for example, the store manager who owns the camera 2, or a store clerk who has been authorized by the owner to access the camera footage) can access the video data. Multiple pieces of video data captured by multiple cameras 2 are stored in the storage device 31, and each piece of original video data may be deleted from the storage device 31 after a predetermined number of days have passed.

[0048] The metadata associated with the original video data may include a series of information related to the management of the video data. The series of information related to the management of the video data may include, for example, identification information of the video data, shooting time information, shooting location information, user information, and camera information. The masking task data includes information related to the task of the masking process (see, for example, FIG. 9). Here, the masking process is a process for removing personal information contained in the video data. The types and details of the masking process will be described later.

[0049] An image recognition model is a trained model constructed by machine learning. An image recognition model may be composed of multiple image recognition models of different types. An image recognition model is constructed from training data consisting of image data and information about objects (e.g., people, vehicles, etc.) contained in the image data. The training data is prepared by annotating (tagging) the image data. The information about the objects may include information indicating the type, attributes, position, etc. of the objects.

[0050] For example, if the type of object is a person, the information about the object may include information indicating that the object is a person, attribute information of the person, and face area information indicating the position of the person's face. The face area information may be identified by coordinate information of two diagonal vertices out of four vertices that form a rectangular area surrounding the person's face. The attribute information of the person may include facial information such as gender, age, whether or not the person is wearing a mask, facial angle, and facial expression. Furthermore, if the type of object is a vehicle (such as a car or a motorcycle), the information about the object may include information indicating that the object is a vehicle, attribute information of the vehicle, and license plate area information indicating the position of the vehicle's license plate.

[0051] When video data showing multiple objects is input to an image recognition model, information indicating the type, attributes, position, etc. of the objects included in each frame of the video data may be output from the image recognition model. For example, when video data showing multiple people is input to the image recognition model, attribute information (gender, age, whether or not a mask is worn, facial angle, facial expression, etc.) and facial area information for each person may be output from the image recognition model. Furthermore, when video data showing multiple vehicles is input to the image recognition model, license plate area information for each vehicle may be output from the image recognition model.

[0052] In the video data after the masking process (video data on which the masking process has been performed), personal information contained in the video data has been removed through the masking process. Therefore, external access to the video data after the masking process is not restricted. In this way, AI development company K, which develops artificial intelligence (AI), can utilize the video data after the masking process as training data for building a machine learning model (an image recognition model using machine learning). The metadata associated with the video data after the masking process may include a series of information related to the video data after the masking process (e.g., identification information of the video data, shooting time information, shooting location information, masking process information, etc.). Because the video data after the masking process is utilized as training data, the metadata includes information to improve the searchability of the video data.

[0053] When the object included in the video data is a person, the attribute information of the object may include at least one of the person's gender information, age information, facial angle information, and facial expression information. In this regard, the control unit 30 may acquire the attribute information of each object included in the video data by using an image recognition model stored in the storage device 31. In this way, the attribute information of each object acquired through the image recognition model is stored in the storage device 31 in a state associated with the video data after the masking process. Specifically, each object included in the video data is assigned identification information, and the identification information and attribute information of each object may be associated with each other.

[0054] The input / output interface 32 is an interface that enables connection between an external device and the server 3, and includes an interface conforming to a predetermined communication standard such as the USB standard or the HDMI (registered trademark) standard. The communication unit 33 may include various wired communication modules for communicating with external terminals on the communication network 8. The input operation unit 34 is, for example, a touch panel, a mouse, and / or a keyboard, and is configured to accept input operations by an operator and to generate operation signals in response to the input operations by the operator. The display unit 35 is, for example, configured by a video display and a video display circuit.

[0055] (Configuration of user terminal 4) Next, the configuration of the user terminal 4 (an example of a first terminal) will be described below. FIG. 4 is a diagram showing an example of the hardware configuration of the user terminal 4. As shown in FIG. 1, the user terminal 4 is operated by a user U who runs a store in which the camera 2 is installed. The user terminal 4 is communicatively connected to the server 3 via a communication network 8. The user terminal 4 may be, for example, a personal computer, a smartphone, a tablet, or a wearable device worn by the user U. The user terminal 4 has a web browser. A video distribution application provided by the server 3 runs on the web browser of the user terminal 4. Note that, particularly when the user terminal 4 is a smartphone, tablet, or the like, the video distribution application may run on software downloaded to the user terminal 4 instead of the web browser.

[0056] As shown in Fig. 4, the user terminal 4 includes a control unit 40, a storage device 41, an input / output interface 42, a communication unit 43, an input operation unit 44, and a display unit 45. These elements are connected to a communication bus 46. The user terminal 4 can access not only the masked video data but also the data in the restricted access area among the data stored in the storage device 31 of the server 3, if the video was taken with a camera owned by the user. On the other hand, the user terminal 4 cannot access not only the masked video data but also the data in the restricted access area for video taken with a camera owned by someone other than the user.

[0057] The control unit 40 includes a memory and a processor. The memory is configured to store computer-readable instructions. In particular, the memory may store a program for causing the processor to execute a series of processes executed by the user terminal 4. The memory is configured with a ROM and a RAM. The processor is configured with at least one of a CPU, an MPU, and a GPU. The storage device 41 is, for example, a storage device such as an HDD, an SSD, or a flash memory, and is configured to store programs and various data.

[0058] The input / output interface 42 is an interface that enables connection between an external device and the user terminal 4. The communication unit 43 is configured to connect the user terminal 4 to the communication network 8. The communication unit 43 includes, for example, a wireless communication module and a wired communication module for wireless communication with external devices such as base stations and wireless LAN routers. The input operation unit 44 is, for example, a touch panel, a mouse, and / or a keyboard arranged over the video display of the display unit 45, and is configured to accept input operations by the user U and generate operation signals corresponding to the input operations. The display unit 45 is, for example, configured by a video display and a video display circuit that drives and controls the video display.

[0059] (Configuration of company terminal 5) As shown in FIG. 1, the enterprise terminal 5 (an example of a second terminal) is a terminal operated by AI development company K. The enterprise terminal 5 may be, for example, a personal computer, a smartphone, a tablet, or a wearable device worn by a person in charge of AI development company K. The hardware configuration of the enterprise terminal 5 may be the same as the hardware configuration of the user terminal 4. The enterprise terminal 5 is communicatively connected to the server 3 via a communication network 8. The enterprise terminal 5 can access masked video data of footage captured by cameras owned by various users among the data stored in the storage device 31 of the server 3, but cannot access data in access-restricted areas.

[0060] (Flow of a series of processes executed by the video delivery system) Next, a series of processing steps executed by the video providing system 1 according to this embodiment will be described below with reference to FIG. 5. FIG. 5 is a flowchart illustrating a series of processing steps executed by the video providing system 1. As shown in FIG. 5, in step S1, the server 3 receives video data from each camera 2 via the communication network 8 and stores the received video data (original video data) in the storage device 31. Next, in step S2, the server 3 receives a request to transmit video data from the user terminal 4. In step S3, the server 3 transmits the video data to the user terminal 4 in response to the transmission request from the user terminal 4. More specifically, the server 3 transmits data (such as an HTML file, a CSS file, an image / video file, and a program file) for displaying a video display screen 50 (see FIG. 6) displaying the video data to the user terminal 4. Thereafter, the video display screen 50 is displayed on the web browser of the user terminal 4 (step S4).

[0061] As shown in FIG. 6 , the video display screen 50 has a video display area 52 in which video data is displayed, a timeline 82 indicating the playback time, and a slider 83 that can be slid on the timeline 82. This type of screen is called a viewer. The user U can change the playback time of the video data by moving the slider 83 on the timeline 82 of the viewer. For example, if the user U changes the position of the slider 83 to 15:00, the user terminal 4 transmits a transmission request to the server 3 for video data for the time period before and after 15:00. In response to the transmission request, the server 3 transmits video data for the time period before and after 15:00 to the user terminal 4. In this way, the video data for the time period before and after 15:00 is displayed in the video display area 52. Note that, in addition to time, year, month, and day information may be specified using, for example, a calendar. Furthermore, the user U may specify a period (start point / end point) on the timeline 82, allowing the user U to extract that period and create a movie clip.

[0062] Next, in response to an input operation by the user U on the user terminal 4, the user terminal 4 transmits to the server 3 a transmission request for a setting screen 60 (see FIG. 7) related to settings for masking processing to be executed on the video data (step S5). In response to the transmission request from the user terminal 4, the server 3 transmits data for displaying the setting screen 60 to the user terminal 4 (step S6). Thereafter, the user terminal 4 receives the data for displaying the setting screen 60 and displays the setting screen 60 on the web browser (step S7).

[0063] As shown in FIG. 7 , the setting screen 60 is a screen related to the setting of the masking process, particularly a screen for registering a masking task. The setting screen 60 is provided with a masking task name input area 63, a registration date and time display area 64, a video data selection area 65, a masking process selection area 67, and a task registration button 68. The masking task name input area 63 allows the user to input a masking task name. The registration date and time display area 64 displays the date and time when the masking task is to be registered. The video data selection area 65 allows the user to select the video data on which the masking process is to be executed. For example, if multiple cameras 2 associated with user U acquire video data indicating store A, store B, store C, and store D, the video data selection area 65 allows the user to select the video data of stores A to D.

[0064] Furthermore, since the storage device 31 of the server 3 stores consecutive video data (video data from store A to store D) for a predetermined number of days, for the video data to be masked, the setting screen 60 may be provided with a store designation area (not shown) for designating the store that captured the video, a camera designation area (not shown) for designating the camera that captured the video, a period designation area (not shown) for designating the period during which the video was captured, and the like. In this case, for example, when a camera is designated in the camera designation area, the screen shown in FIG. 6 is displayed, and the desired period may be designated there using the aforementioned movie clip function. Also, for example, when a camera is designated in the camera designation area, a list of videos already created with that camera using the movie clip function may be displayed, and the desired video may be selected from among them. Also, when a store is designated in the store designation area before designating a camera, the cameras of that store may be narrowed down and displayed in the camera designation area.

[0065] In the masking process selection area 67, a masking process to be executed on an object (e.g., a person's face) that indicates personal information included in the video data can be selected. In FIG. 7, deep masking process, mosaic process, blurring process, character mosaic process, etc. are selectable as examples of masking processes in the masking process selection area 67. In this example, the user U can select any one of the deep masking process, mosaic process, and blurring process shown in the masking process selection area 67. Furthermore, the user U can additionally select character mosaic process.

[0066] Next, types of masking processing will be described below with reference to FIG. 8. Note that the information in FIG. 8 may be displayed as help in, for example, the masking processing selection area 67. As shown in FIG. 8(a), the original face (an example of the target) before the masking processing enables personal information to be identified. Furthermore, the original face before the masking processing enables AI (image recognition model) to recognize people and also enables attribute information to be identified. Examples of attribute information include the target's age, gender, whether or not they are wearing a mask, facial angle, and facial expression, but the types of attribute information are not particularly limited.

[0067] As shown in Figure 8(b), the face after mosaic processing makes it impossible to identify personal information. Furthermore, the face after mosaic processing makes it impossible for AI to recognize the person, and also makes it impossible to identify attribute information. In this way, mosaic processing (an example of the first masking process) removes not only the personal information of the target object, but also information about the type and attributes of the target object.

[0068] As shown in Figure 8(c), the blurred face makes it impossible to identify personal information. Furthermore, while the blurred face allows AI to recognize people, it makes it impossible to identify attribute information. In this way, the blurring process (an example of the second masking process) removes not only the personal information of the target, but also information about the target's attributes. On the other hand, the blurring process does not remove information about the type of target.

[0069] As shown in Figure 8(d), the face after deep masking processing makes it impossible to identify personal information. Furthermore, the face after deep masking processing enables AI to recognize people and identify attribute information. In this way, deep masking processing (an example of the third masking processing) removes personal information about the target object, but does not remove information about the type and attributes of the target object.

[0070] As such, the amount of information about the object shown in the video data that is removed increases in the order of mosaic processing > blurring processing > deep masking processing. If user U thinks that it is OK for attribute information of people included in the video data after masking processing to be used in machine learning, user U will likely select deep masking processing, which does not delete attribute information of the object, through the masking processing selection area 67. On the other hand, if user U thinks that it is not OK for attribute information of people's faces to be used in machine learning, user U will likely select mosaic processing or blurring processing through the masking processing selection area 67.

[0071] Character mosaic processing is performed on character information that corresponds to personal information. For example, vehicle identification information shown on a vehicle's license plate corresponds to personal information, so mosaic processing is performed on the vehicle identification information. In this case, the user's approach to the type of masking processing is similar to that for a human face.

[0072] Returning to FIG. 5, in step S8, the user terminal 4 transmits a series of information related to the setting of the masking process (more specifically, information related to the masking task name, registration date and time, selection of video data, and selection of masking process) to the server 3 in response to the user U's input operation on the setting screen 60 displayed on the display unit 45 (specifically, the user U's operation on the task registration button 68).

[0073] Next, the server 3 (specifically, the control unit 30 of the server 3) determines the video data to be masked and the masking process to be performed on the video data, based on a series of information related to the settings of the masking process received from the user terminal 4. Thereafter, the server 3 acquires and stores attribute information of each object (particularly, a person) indicating personal information included in the video data to be masked (step S9). More specifically, the server 3 acquires attribute information of each object included in the video data by using an image recognition model, and then stores the acquired attribute information in the storage device 31.

[0074] In step S10, the server 3 performs a masking process on the video data to be selected. For example, if the user U selects deep masking, the server 3 performs the deep masking process on each object representing personal information included in the video data. More specifically, the server 3 uses an image recognition model to identify facial region information and attribute information of people included in each frame of the video data. Next, the server 3 performs the deep masking process on the faces of people representing personal information based on the identified facial region information and attribute information. Furthermore, if the user U selects mosaic processing or blurring, the server 3 performs mosaic processing or blurring on each object representing personal information included in the video data. More specifically, the server 3 uses an image recognition model to identify facial region information of people included in each frame of the video data. Next, the server 3 performs mosaic processing or blurring on the faces of people representing personal information based on the identified facial region information. Note that these masking processes are performed on copies of the original video data, but after the processes are successfully completed, the original video data may be deleted or retained.

[0075] After the masking process is performed on the video data, the server 3 updates the masking task data stored in the storage device 31 (step S11). As shown in FIG. 9, the masking task data includes information about the task of the masking process. As shown in FIG. 9, the masking task data may include information about the masking task name, information about the registration date and time of the masking task, information about the result of the masking process, information about the evaluation of the masking process, and information about the name of the video data on which the masking process is performed. After performing the masking process on the video data, the server 3 registers information about the task of the masking process in the masking task data. For example, if the masking process is completed successfully, information indicating that the masking process has been completed successfully is recorded in the masking task data. On the other hand, if the masking process is completed abnormally, information indicating that the masking process has been completed abnormally (error information) is recorded in the masking task data.

[0076] In step S12, the server 3 generates a masking task list screen 70 (see FIG. 9) based on the updated masking task data, and then transmits data (HTML files, CSS files, image / video files, program files, etc.) for displaying the masking task list screen 70 to the user terminal 4. Thereafter, the masking task list screen 70 is displayed on the web browser of the user terminal 4.

[0077] 9, various information related to masking processing tasks is displayed on the masking task list screen 70. Specifically, information related to the masking task name, information related to the registration date and time of the masking task, information related to the result of the masking processing, information related to the evaluation of the masking processing, and information related to the name of the video data on which the masking processing is to be executed are displayed on the masking task list screen 70. Note that other information such as the store (not shown) specified in the store specification area, the camera (not shown) specified in the camera specification area, and the period specified in the period specification area (not shown) may also be displayed.

[0078] The information regarding the evaluation of the masking process includes evaluation information indicating whether the masking process has been appropriately performed on each object included in the video data. More specifically, the information regarding the evaluation of the masking process includes information indicating that the masking process has been appropriately performed (approval of the masking process) and information indicating that the masking process has not been appropriately performed (non-approval of the masking process). Furthermore, if the masking process for certain video data has not yet been evaluated, the evaluation button 72 may be displayed on the masking task list screen 70 in a state visually associated with the certain video data. In the example shown in FIG. 9, the masking process for the video data of Store A has not yet been evaluated, so the evaluation button 72 is displayed on the masking task list screen 70 in a state visually associated with the video data of Store A.

[0079] In step S13, the user terminal 4 transmits a transmission request for the video comparison screen 80 (see FIG. 10 ) to the server 3 in response to an input operation by the user U on the user terminal 4. Specifically, the user terminal 4 transmits the transmission request for the video comparison screen 80 to the server 3 in response to an operation by the user U on the evaluation button 72 displayed on the masking task list screen 70. In response to the transmission request from the user terminal 4, the server 3 transmits data for displaying the video comparison screen 80 to the user terminal 4 (step S14). Thereafter, the user terminal 4 receives the data for displaying the video comparison screen 80 and displays the video comparison screen 80 on the web browser (step S15).

[0080] As shown in FIG. 10 , a video comparison screen 80 (an example of a video confirmation screen) displays side by side video data V2 on which masking processing has been performed (e.g., video data of store A before masking processing) and video data V1 before masking processing (e.g., video data of store A after masking processing). The video comparison screen 80 has a video display area 87 in which the video data V1 before masking processing is displayed, and a video display area 88 in which the video data V2 after masking processing is displayed. The video display area 87 and the video display area 88 are arranged side by side. The video comparison screen 80 has a timeline 82a and a slider 83a associated with the video display area 87, and a timeline 82b and a slider 83b associated with the video display area 88. The video data V1 and the video data V2 may be played back in conjunction with each other (in time synchronization). In this case, the timelines 82a and 82b, and the sliders 83a and 83b may also be linked to each other. Also, the set of the timeline and slider may be displayed as a single set common to the video data V1 and the video data V2.

[0081] In the example shown in FIG. 10, it is assumed that objects M1 to M4 indicating personal information are present in pre-masking video data V1 displayed in video display area 87. Object M1 is text indicating personal information. Each of objects M2 to M4 is a person (more specifically, a human face). In post-masking video data V2 displayed in video display area 88, text mosaic processing is performed on object M1 indicating text information. Deep masking processing is performed on each of objects M2 to M4 indicating human faces.

[0082] The video comparison screen 80 also has an approval button 84 indicating approval of the masking process, a disapproval button 85 indicating disapproval of the masking process, and a masking correction button 86. In step S16, the user terminal 4 transmits evaluation information on the masking process to the server 3 in response to an input operation by the user U on the user terminal 4. Here, the evaluation information on the masking process is information indicating whether the masking process has been appropriately performed on an object included in the video data. For example, when the user U operates the approval button 84, the user terminal 4 transmits evaluation information indicating that the masking process has been appropriately performed (information indicating approval of the masking process) to the server 3. On the other hand, when the user U operates the disapproval button 85, the user terminal 4 transmits evaluation information indicating that the masking process has not been appropriately performed (information indicating disapproval of the masking process) to the server 3. For example, when the face region of the object is not appropriately masked, the user U will likely determine that the masking process has not been appropriately performed on the object included in the video data.

[0083] In step S17, the server 3 determines whether the masking process has been properly performed on each object included in the video data based on the evaluation information on the masking process received from the user terminal 4, and updates the masking task data. In particular, the server 3 updates information related to the evaluation of the masking process included in the masking task data based on the evaluation information. Here, when the server 3 receives evaluation information indicating that the masking process has been properly performed, the server 3 determines that the masking process has been properly performed on each object included in the video data and updates the masking task data. Thereafter, the server 3 transmits data for displaying a data registration screen 90 (see FIG. 11) to the user terminal 4 (step S18). Thereafter, the user terminal 4 receives the data for displaying the data registration screen 90 and displays the data registration screen 90 on the web browser.

[0084] If the server 3 receives evaluation information indicating that the masking process has not been performed appropriately, the server 3 determines that the masking process has not been performed appropriately for each object included in the video data and updates the masking task data. The server 3 may then re-perform the masking process on the video data. The user U may also operate the masking correction button 86 and then correct the masking process of the video data V2 by inputting an input operation (e.g., mouse operation) through the input operation unit 44. For example, the user U may select M1 to M4 on the video V2 in FIG. 10 to exclude them from the masking target, or may select other faces or characters to add them as masking targets. A UI for selecting the type of masking may also be displayed, allowing the user to change the type of masking by selecting one of the options. In this case, the user terminal 4 transmits information regarding the correction of the masking process performed by the user U's input operation to the server 3. The server 3 may then update the video data after the masking process based on the information regarding the correction of the masking process received from the user terminal 4. The server 3 may then transmit data for displaying the data registration screen 90 to the user terminal 4.

[0085] As shown in FIG. 11 , the data registration screen 90 is a screen for registering video data on which masking processing has been performed. The data registration screen 90 includes a masking task name display area 93, a masking processing display area 94, a video data name display area 95, a metadata availability selection area 97, a video-related information display area 99, and a data registration button 98. The masking task name display area 93 displays the name of the masking task. The masking processing display area 94 displays information related to the masking processing performed on the video data. When deep masking processing and character mosaic processing have been performed on the video data, the deep masking processing and character mosaic processing are selected as the initial state in the masking processing display area 94. The user U can change and / or add the type of masking processing performed on the video data by inputting information into the masking processing display area 94. The video data name display area 95 displays information related to the video data name. In addition, other information may also be displayed, such as the store (not shown) specified in the store specification area, the camera (not shown) specified in the camera specification area, and the period specified in the period specification area (not shown).

[0086] The metadata availability selection area 97 is an area for selecting whether a third party other than the user U (in this example, AI development company K) can use the metadata associated with the video data. The user U can determine whether the metadata associated with the video data can be used by inputting information into the metadata availability selection area 97.

[0087] The video-related information display area 99 displays information about the time of the video (shooting time), information about the data size of the video, information about the type of object contained in the video, information about the location where the video was shot, information about the business where the video was shot, information about the area where the video was shot, information about the time period when the video was shot, identification information (ID) of the video data, and the ID of the metadata linked to the video data.

[0088] The information about the type of object may be information indicating that the object is a person or a vehicle (more specifically, a passenger car, a bicycle, a motorcycle, or a special vehicle). The information about the location may be information indicating that the shooting location is indoors or outdoors. The information about the type of business may be information indicating that the business type is a food and beverage establishment (specifically, an izakaya or restaurant), a retail establishment (specifically, a convenience store, a supermarket, or a department store), or a construction site (specifically, a building, a detached house, or a road). The information about the region may be information indicating the country, prefecture, or city / ward / town / village where the video was shot. In this regard, if the region where the video was shot is domestic (within Japan), the information about the region may indicate information about the prefecture and the city / ward / town / village. If the region where the video was shot is foreign, the information about the region may indicate information about the country and the city. The information about the time of day may be information indicating daytime, nighttime, early morning, or evening. This information may be set based on information preset for each camera, or may be set by analyzing the video data.

[0089] Returning to FIG. 5 , in step S19, the user terminal 4 transmits information indicating an instruction to register the masked video data to the server 3 in response to the user U's input operation on the data registration screen 90 (specifically, the user U's operation on the data registration button 98). The information indicating the instruction to register the video data may include information regarding the masking process selected in the masking process display area 94 (in this example, information indicating deep masking and character mosaic processes) and information regarding whether the metadata is available (in this example, information indicating that the metadata is available). Thereafter, in response to the instruction to register the video data from the user terminal 4, the server 3 stores the masked video data and the metadata in the storage device 31 (step S20). Here, because the original video data (the video data before the masking process) contains personal information, access to the original video data by third parties other than the user U is restricted. On the other hand, because the masked video data does not contain personal information, access to the masked video data by third parties (in this example, AI development company K) is not restricted. Furthermore, the server 3 updates the video management data (video management table) for managing a plurality of video data stored in the storage device 31 (see FIG. 12).

[0090] As shown in Figure 12, the video management data includes the ID of the video data, the time of the video, the data size of the video, the type of object contained in the video, the location where the video was shot (shooting location), the business type where the video was shot, the area where the video was shot, the time period when the video was shot, whether the metadata is available, and the ID of the metadata.

[0091] Furthermore, the metadata associated with the masked video data may include at least one of identification information of the video data, shooting time information of the video data, shooting location information of the video data, and information related to the masking process performed on the video data. The presence of metadata improves the searchability of the masked video data. When multiple masked video data are stored in the storage device 31, the multiple pieces of metadata may be stored in the storage device 31 with each piece of metadata associated with a corresponding one of the multiple masked video data.

[0092] Furthermore, attribute information of each object included in the video data is stored in the storage device 31 in a state associated with the video data after the masking process. If the object included in the video data is a person, the attribute information of the object may include at least one of the person's gender information, age information, facial angle information, and facial expression information. For example, as shown in FIG. 10, identification information and attribute information of objects M1 to M4 included in the video data may be stored in a state associated with the video data. In this case, the identification information of each object M1 to M4 may be displayed on the video data after the masking process in a state visually associated with a corresponding one of the objects M1 to M4.

[0093] In step S21, the server 3 deletes the original video data from the storage device 31. For example, if the video data of store A after the masking process is saved in the storage device 31, the original video data of store A may be deleted from the storage device 31. In this way, the original video data before the masking process is executed is deleted from the server 3, so that the amount of data stored in the server 3 can be appropriately reduced, and the maintenance and management costs of the server 3 can be appropriately reduced. Note that the original video data may not be deleted immediately, but may be retained for at least a certain period (depending on the cloud usage fee / capacity / period, etc. related to the user's contract).

[0094] Next, in step S22, the company terminal 5 transmits a transmission request for the video search screen 100 (see FIG. 13) to the server 3 in response to an input operation by a person in charge of AI development company K. In response to receiving the transmission request, the server 3 transmits the video search screen 100 to the company terminal 5 (step S23).

[0095] In step S24, the enterprise terminal 5 acquires information about search conditions for video data in response to an input operation by the person in charge of AI development company K on the video search screen 100. Thereafter, the enterprise terminal 5 transmits the information about the search conditions to the server 3 (step S25).

[0096] 13, the video search screen 100 displays a search condition specification area 104 and a send button 105. In the search condition specification area 104, the video duration (shooting time), the video data size, the type of masking process, the type of object shown in the video, the location where the video was shot, the type of business where the video was shot, the region where the video was shot, the time period when the video was shot, and whether metadata is available can be specified as search conditions. For example, when a person in charge at AI development company K specifies search conditions for video data through the search condition specification area 104 and then presses the send button 105, the company terminal 5 sends information related to the specified search conditions to the server 3.

[0097] Next, in step S26, the server 3 receives information regarding search conditions for video data from the company terminal 5. Thereafter, based on the received information regarding the search conditions, the server 3 refers to the video management data stored in the storage device 31, and extracts at least one piece of masked video data that matches the search conditions from among the multiple pieces of masked video data stored in the storage device 31. More specifically, the server 3 refers to the video management data to identify the ID of the video data that matches the search conditions, and then acquires the video data corresponding to the ID of the video data from the storage device 31.

[0098] Next, the server 3 transmits the extracted at least one piece of masked video data (i.e., at least one piece of masked video data that matches the search criteria) and attribute information of each object included in the extracted at least one piece of video data to the company terminal 5. In this way, the AI ​​development company K operating the company terminal 5 can acquire desired video data that matches the search criteria through input operations on the video search screen 100. More specifically, the AI ​​development company K can acquire desired video data from the perspectives of information related to the time of the video data, information related to the data size of the video data, information related to the masking process performed on the video data, information related to the type of object included in the video data, information related to the location where the video data was shot, information related to the business type where the video data was shot, information related to the area where the video data was shot, information related to the time period when the video data was shot, and information related to the metadata associated with the video data.

[0099] When mosaic processing or blurring processing is performed on the video data as a masking process, attribute information of each object is deleted from the video data after the masking process, and therefore the attribute information of each object may be transmitted together with the video data to the company terminal 5. In this case, the presence of the attribute information of each object allows the video data after the masking process to be effectively used as learning data for building a machine learning model related to human behavior, etc.

[0100] Furthermore, in step S26, the server 3 may transmit a video list screen showing a list of video data that matches the search conditions to the enterprise terminal 5 before transmitting at least one piece of masked video data that matches the search conditions to the enterprise terminal 5. In this case, the enterprise terminal 5 transmits a transmission request for the desired video data to the server 3 through an input operation by the person in charge on the video list screen. Thereafter, the server 3 transmits the desired video data to the enterprise terminal 5 in response to the transmission request.

[0101] On the other hand, when deep masking processing is performed on the video data as a masking processing, the attribute information of each object remains in the video data after the masking processing, so the attribute information of each object does not need to be sent to the company terminal 5.

[0102] In step S27, the company terminal 5 receives the masked video data and attribute information of each object from the server 3. AI development company K makes effective use of the received video data, etc. as learning data for building a machine learning model (step S28).

[0103] According to this embodiment, a setting screen 60 related to the setting of the masking process (in other words, the anonymization process of personal information) is displayed on the user terminal 4, and the masking process is determined in response to an input operation by the user U on the setting screen 60. The determined masking process is then performed on objects (e.g., people, vehicles, etc.) that indicate personal information included in the video data, and the masked video data is provided to the enterprise terminal 5. In this way, the user U can determine what kind of masking process should be performed on the video data through the setting screen 60, thereby improving the usability of the video providing system 1 in the context of the masking process on the video data. Furthermore, since the masked video data is provided to the enterprise terminal 5, the video data stored on the server 3 can be effectively utilized. For example, the masked video data can be effectively utilized as training data for building a machine learning model (an image recognition model utilizing machine learning).

[0104] Furthermore, according to this embodiment, the video comparison screen 80 is displayed on the user terminal 4, and then, in response to an input operation by the user U on the user terminal 4, it is determined whether masking has been appropriately performed on objects (e.g., people, vehicles, etc.) included in the video data. The masked video data is then provided to the company terminal 5. In this way, the user U can understand through the video comparison screen 80 whether masking has been appropriately performed on the video data, thereby improving the usability of the video providing system 1 in the context of the masking process on the video data. Furthermore, the user U can objectively confirm that appropriate masking in compliance with laws and regulations such as the Personal Information Protection Act has been performed on the video data, thereby ensuring sufficient transparency and reliability of the masking process. Therefore, the AI ​​development company K, which has been provided with the masked video data, can safely use the video data as training data for building a machine learning model.

[0105] Although the embodiments of the present invention have been described above, the technical scope of the present invention should not be construed as being limited by the description of the present embodiments. The present embodiments are merely examples, and it will be understood by those skilled in the art that various modifications of the embodiments are possible within the scope of the invention described in the claims. The technical scope of the present invention should be determined based on the scope of the invention described in the claims and its equivalents. [Explanation of symbols]

[0106] 1: Video provision system, 2: Camera, 3: Server, 4: User terminal, 5: Company terminal, 8: Communication network, 20: Control unit, 21: Storage device, 22: Location information acquisition unit, 24: Communication unit, 25: Input operation unit, 26: Imaging unit, 27: PTZ mechanism, 28: Communication bus, 30: Control unit, 31: Storage device, 32: Input / output interface, 33: Communication unit, 34: Input operation unit, 35: Display unit, 36: Communication bus, 40: Control unit, 41: Storage device, 42: Input / output interface, 43: Communication unit, 44: Input operation unit, 45: Display unit, 46: Communication bus, 50: Video display screen, 52: Video display area, 60: Setting screen, 63: Masking task name input area, 64: Registration date and time display area, 65: Video data selection area, 67: Mask Masking process selection area, 68: Task registration button, 70: Masking task list screen, 72: Evaluation button, 80: Video comparison screen, 82, 82a, 82b: Timeline, 83, 83a, 83b: Slider, 84: Approval button, 85: Disapproval button, 86: Masking correction button, 87, 88: Video display area, 90: Data registration screen, 93: Masking task name display area, 94: Masking process display area, 95: Video data name display area, 97: Metadata availability selection area, 98: Data registration button, 99: Video related information display area, 100: Video search screen, 104: Search condition specification area, 105: Send button, K: AI development company, M1, M2, M3, M4: Object, U: User, V1, V2: Video data

Claims

1. A video providing system that stores video captured by a camera in a server and provides the video to a terminal, displaying a setting screen for the masking process to be executed on the first terminal; determining a type of the masking process in response to a user's input operation on the setting screen; generating a second image by performing the determined type of masking process on an object indicating personal information included in the first image; storing the second image on the server; providing the second video from the server to a second terminal; Video provision system.

2. displaying on the first terminal a confirmation screen for confirming whether the masking process has been properly performed on the second video; determining whether or not the second video is available for use in response to an input operation by a user on the confirmation screen; If it is determined that the second video is available, storing the second video on the server. The video providing system according to claim 1 .

3. the first image and the second image are displayed side by side on the confirmation screen; The video providing system according to claim 2 .

4. After storing the second image, deleting the first image. The video providing system according to claim 1 .

5. acquiring attribute information of the object; providing the attribute information together with the second video; The video providing system according to claim 1 .

6. the object is a person, The attribute information is Gender and Age and Face angle and Facial expressions and including at least one of The video providing system according to claim 5 .

7. storing metadata associated with the second video; The metadata includes: Identification information of the video; The shooting time of the video; The location where the video was shot; the masking process performed on the video; and including at least one of The video providing system according to claim 1 .

8. displaying on the second terminal a search screen for searching for a desired second video from among the plurality of second videos; determining a search condition for the second video in response to an input operation by a user on the search screen; extracting at least one of the second videos corresponding to the determined search criteria from the plurality of second videos stored in the server; providing the extracted second image to the second terminal; The video providing system according to claim 1 .

9. The search conditions are: The time of the video, The data size of the video; a masking process performed on the video; and The type of object included in the video; and The location where the footage was taken; The type of business in which the video was taken; The area where the footage was taken; The time period when the video was taken; metadata associated with the video; and including at least one of The video providing system according to claim 8 .

10. the setting screen has a masking process selection area in which one of a plurality of types of masking processes can be selected; The plurality of types of masking processes include: a first masking process that makes it impossible to identify both the attribute information and personal information of the object; a second masking process that enables identification of attribute information of the object while making it impossible to identify personal information of the object; Including, The video providing system according to claim 1 .

11. A video providing system that stores video captured by a camera in a server and provides the video to a terminal, generating a second image by performing a masking process on an object indicating personal information included in the first image; displaying a confirmation screen on the first terminal for confirming whether the masking process has been properly performed on the second video; determining whether or not the second video is available for use in response to an input operation by a user on the confirmation screen; If it is determined that the second video is available, storing the second video on the server; providing the second video from the server to a second terminal; Video provision system.

12. A video providing system that stores video captured by a camera in a server and provides the video to a terminal, generating a second image by performing a masking process on an object indicating personal information included in the first image; storing the second image on the server; displaying a search screen on the second terminal for searching for a desired second video from among the plurality of second videos; determining a search condition for the second video in response to an input operation by a user on the search screen; extracting at least one of the second videos corresponding to the determined search criteria from the plurality of second videos stored in the server; providing the extracted second image to the second terminal; Video provision system.

13. A video providing method executed by a video providing system that stores video captured by a camera in a server and provides the video to a terminal, comprising: displaying a setting screen for the masking process to be executed on the first terminal; determining the type of the masking process in response to an input operation by a user on the setting screen; generating a second image by performing the determined type of masking process on an object indicating personal information included in the first image; storing the second video on the server; providing the second video from the server to a second terminal; A method for providing video, including:

14. A video providing program that causes a video providing system to execute the video providing method according to claim 13.

Citation Information

Patent Citations

  • Image processing apparatus, image processing system, image processing method, and computer program

    JP2014067131A

  • Learning video selecting device, program and method for selecting, as learning video, shot video with predetermined image region masked

    JP2019079357A

  • Image data processing apparatus, image data processing system, image data processing method, and program

    JP2020068425A

  • Image editing device and image editing method

    JP2020141246A

  • Information processing device, control method of the same, and program

    JP2021027384A