Camera video management system
The camera video management system addresses privacy and data size issues by generating abstract images and captions, reducing data requirements and ensuring privacy protection.
Patent Information
- Application Number
- JP2024081072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Conventional systems face issues with determining the scope of image processing for privacy protection based on score tables, leading to potential privacy breaches, and adding text information to videos increases data size.
A camera video management system that generates abstract images and linguistic information, associating person recognition with scene information, and outputs a reproduction image with captioned abstract spaces to reduce data size and protect privacy.
The system reduces data size and protects privacy by using abstract images and captions, while maintaining scene information integrity.
Smart Images

Figure 2025174595000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a system for managing camera footage. [Background technology]
[0002] Japanese Patent Application Laid-Open Publication No. 2006-295251 discloses a system that processes surveillance camera footage and outputs the processed footage on an external monitor. This conventional system extracts a person from the surveillance camera footage, detects the gestures and actions of the extracted person, and calculates the importance of the detected gestures and actions. The conventional system also determines the range of the surveillance camera footage to be subjected to image processing based on the calculated importance. In the image processing, privacy protection processing such as mosaic processing and smoothing processing is performed on the range to be processed. In other words, in the conventional system, the video that has been subjected to image processing according to the importance of the person's gestures and actions is output from the external monitor.
[0003] Conventional systems also analyze the scenes captured by the surveillance camera, and if the analysis determines that the scene is abnormal, the system adds text information indicating that the scene is abnormal to the processed image.
[0004] In addition to JP 2006-295251 A, examples of documents showing the technical state of the art in the technical field related to the present disclosure include JP 2022-056533 A, JP 2019-144830 A, and JP 2002-024962 A. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-295251 [Patent Document 2] Japanese Patent Publication No. 2022-056533 [Patent Document 3] Japanese Patent Application Publication No. 2019-144830 [Patent Document 4] Japanese Patent Application Laid-Open No. 2002-024962 Summary of the Invention [Problem to be solved by the invention]
[0006] In conventional systems, the importance of a person's gestures and actions is calculated by referencing a score table created in advance. However, depending on the granularity of the information in the score table, the scope of image processing may not be determined correctly, which may result in insufficient privacy protection for people captured on surveillance camera footage.
[0007] In conventional systems, adding text information indicating an abnormal scene to the output video from an external monitor is expected to make it easier to inform the viewer of the external monitor that the video scene is abnormal. However, in order to play back the output video with such text information added, it is necessary to combine the image-processed video with the text information and save it, which poses a problem of increasing the data size of the video.
[0008] One object of the present disclosure is to provide a technology that can reduce the size of data required to play back a scene captured in a camera video while protecting the privacy of people captured in the camera video. [Means for solving the problem]
[0009] The present disclosure relates to a camera video management system and has the following features. The system includes a storage device, a processing circuit, and a display device. The storage device stores the camera image. The processing circuit is configured to perform various processes. The display device is configured to output an image. The processing circuit is configured to generate recognition information of an object appearing in the camera image by performing object recognition processing on the camera image stored in the storage device, generate linguistic information of the scene appearing in the camera image by performing verbalization processing on the camera image stored in the storage device, and if the object recognition information includes person recognition information, generate scene information that associates the person recognition information with the scene linguistic information generated by the verbalization processing on the camera image in which the person was recognized, and store this in the storage device, and perform reproduction processing of the scene appearing in the camera image based on the scene information stored in the storage device. The reproduction process includes rendering an abstract image of the space shown in the camera image based on recognition information of static objects included in the scene information, rendering an abstract image of the person shown in the camera image onto the abstract image of the space based on recognition information of the person included in the scene information, generating caption information for the scene based on linguistic information of the scene included in the scene information, and adding the caption information to the abstract image of the space into which the abstract image of the person has been rendered, to generate a reproduction image to be output from the display device. [Effects of the Invention]
[0010] According to the present disclosure, a scene captured in a camera video is reproduced using an abstract image of a person captured in the camera video, rendered onto an abstract image of the space captured in the camera video, to which linguistic information about the scene captured in the camera video is added as a caption. Here, the abstract image of the space and the person is rendered based on object recognition information generated by object recognition processing of the camera video. Therefore, the data size required to reproduce the video scene can be reduced compared to when raw video data or video data subjected to privacy protection processing using conventional technology is used. Furthermore, the use of an abstract image of the person can also protect the privacy of the person. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram illustrating an example of the overall configuration of a camera video management system according to an embodiment of the present disclosure. [Figure 2] 2 is a block diagram showing an example of a functional configuration of the data processing device shown in FIG. 1. FIG. [Figure 3] 3 is a conceptual diagram illustrating an example of reproduction processing by the reproduction unit shown in FIG. 2. FIG. [Figure 4] 3 is a conceptual diagram illustrating another example of the reproduction process performed by the reproduction unit shown in FIG. 2. FIG. [Figure 5] 1. FIG. 4 is a block diagram showing another example of the functional configuration of the data processing device shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In each drawing, the same or corresponding parts are denoted by the same reference numerals, and the description thereof will be simplified or omitted.
[0013] 1. Overall configuration example Fig. 1 is a block diagram showing an example of the overall configuration of a camera video management system according to an embodiment of the present disclosure. Fig. 1 illustrates a camera 10, a data processing device 20, a display device 30, and an input device 40 as components of the management system according to the embodiment. The camera 10, the display device 30, and the input device 40 communicate with the data processing device 20 via a communication network (not shown). The communication network is not particularly limited, and wired and wireless networks may be used.
[0014] The camera 10 is installed in any indoor or outdoor space. The installation position and shooting range of the camera 10 are known. The camera 10 captures an image of its shooting range. The camera 10 transmits the image captured by the camera 10 (i.e., camera image VD) together with its own identification information to the data processing device 20. The total number of cameras 10 is at least one. The shooting ranges of two or more cameras may overlap in part or in whole.
[0015] The data processing device 20 includes at least one processing circuit 21 and at least one storage device 22. Examples of the processing circuit 21 include a general-purpose processor, a specific-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA). Examples of the storage device 22 include a hard disk drive (HDD), a solid state drive (SSD), a volatile memory, and a non-volatile memory.
[0016] The processing circuit 21 develops various programs stored in the storage device 22 and processes various data stored in the storage device 22. Data processing by the processing circuit 21 includes processing of the camera video VD. In data processing of the camera video VD, a two-dimensional or three-dimensional virtual image is generated from the images constituting the camera video VD and output to the display device 30. For ease of explanation, the video composed of virtual images will hereinafter also be referred to as "video VD_VR." Furthermore, the camera video VD from which the virtual image is generated will also be referred to as "video VD_OR." A detailed example of data processing of the video VD_OR will be described later.
[0017] The display device 30 displays various data. The various data displayed on the display device 30 is provided to a user of the management system. The various data displayed on the display device 30 includes video VD_VR. Examples of the display device 30 include a liquid crystal display, an organic EL display, and a head-up display.
[0018] The input device 40 is operated by a user of the management system. Examples of the input device 40 include a keyboard, a mouse, a touch panel, and a microphone. Input information particularly relevant to the embodiment is search information SAR. The search information SAR is information for searching for scenes shown in the video VD_OR. The search information SAR includes character string information. If the input device 40 is a microphone, audio information input from the microphone is converted into character string information.
[0019] 2. Functional configuration example Fig. 2 is a block diagram showing an example of the functional configuration of the data processing device 20 shown in Fig. 1. Fig. 2 illustrates, as functional blocks of the data processing device 20, an object recognition unit 23, a verbalization unit 24, a scene information generation unit 25, a video recording unit 26, a scene information recording unit 27, and a reproduction unit 28. These functional blocks are realized, for example, by cooperation between the processing circuit 21 and the storage device 22 shown in Fig. 1.
[0020] The object recognition unit 23 performs "object recognition processing" to recognize objects included in images (hereinafter also referred to as "images IMG_OR") that constitute the video VD_OR. In the object recognition processing, objects are detected using, for example, a YOLO (You Only Look Once) network, an SSD (Single Shot multibox Detector) network, or the like. The detection targets in the object recognition processing are static objects SO and dynamic objects MO. Examples of static objects SO include buildings, structures, and natural objects. Examples of dynamic objects MO include people, robots, bicycles, and automobiles.
[0021] In the object recognition process, recognition information REC_SO of a static object SO and recognition information REC_MO of a dynamic object MO are generated. The recognition information REC_MO includes recognition information REC_HM of a person HM and recognition information REC_NHM of a dynamic object NHM other than the person HM (robot, bicycle, automobile, etc.). The recognition information REC_SO and the recognition information REC_MO include information on the type of object and information on the detection time of the object. The recognition information REC_SO and the recognition information REC_MO are transmitted to the scene information generation unit 25.
[0022] In the object recognition process, the two-dimensional pose (2D Pose) and three-dimensional pose (3D Pose) of the person HM are further estimated. The two-dimensional pose and three-dimensional pose are represented by lines connecting parts such as joints, head, hands, and feet. Information on the two-dimensional pose and three-dimensional pose is added to the recognition information REC_HM of the person HM. In the object recognition process, tracking of the person HM may be performed. In the object recognition process, processing to identify the same person between two or more cameras 10 (person re-identification processing) may also be performed. Tracking information and re-identification information are also added to the recognition information REC_HM of the person HM. Note that pose estimation, tracking, and re-identification are well-known techniques, and methods applicable to the present disclosure are not particularly limited.
[0023] The verbalization unit 24 performs a "verbalization process" that assigns linguistic information LAN to the scene shown in the video VD_OR. In the verbalization process, for example, a framework based on the LLM model (Large Language Models) is used to generate text that describes the scene shown in the video VD_OR. The text that describes the scene includes, for example, a description of the environment shown in the video VD_OR, a description of the person and surrounding objects shown in the video VD_OR, and a description of the interaction between the person and the surrounding objects. In another example, a VLM model (Vision Language Models) is used to generate text that describes the scene shown in the video VD_OR.
[0024] The text output from such a learning model is an example of a linguistic information LAN. The linguistic information LAN includes information on the start and end times of the scene described by the text. The linguistic information LAN is sent to the scene information generation unit 25.
[0025] The scene information generation unit 25 generates scene information SCN by associating the information received from the object recognition unit 23 (recognition information REC_SO and recognition information REC_MO) with the information received from the verbalization unit 24 (language information LAN). The scene information SCN is generated by associating these pieces of information based on, for example, scene time information included in the language information LAN and object detection time information included in the recognition information REC_SO or the recognition information REC_MO. The scene information SCN is transmitted to the scene information recording unit 27.
[0026] The video recording unit 26 stores the video VD_OR received from the camera 10 in the storage device 22 shown in Fig. 1. Because the data size of the video VD_OR stored in the storage device 22 is large, the video VD_OR stored in the storage device 22 is compressed or deleted after a certain period of time has passed. The video VD_OR stored in the storage device 22 can be referenced in a scene search, which will be described later, or in updating scene information, which will be described later.
[0027] The scene information recording unit 27 stores the scene information SCN in the storage device 22 shown in FIG. 1. It is desirable that the storage device 22 in which the scene information SCN is stored is a different device from the device in which the video VD_OR is stored. The video scene information SCN stored in the storage device 22 is compressed or deleted after a certain period of time, just like the video VD_OR. However, by storing the video VD_OR and the scene information SCN in separate storage devices 22, even if the scene information SCN is accidentally deleted from the storage device 22, the scene information SCN can be regenerated from the video VD_OR stored in another storage device 22.
[0028] The reproducing unit 28 performs a "reproducing process" to reproduce the scene shown in the video VD_OR based on the scene information SCN stored in the scene information recording unit 27. FIG. 3 is a conceptual diagram illustrating an example of the reproducing process by the reproducing unit 28 shown in FIG. 2. On the left side of FIG. 3, a real space RS shown in the video VD_OR included in the scene information SCN is depicted. This real space RS contains static objects SO and dynamic objects MO (people HM and dynamic objects NHM other than people HM), and these objects are detected by object recognition processing.
[0029] In the reproduction process, a virtual space VS is rendered, which is an abstract representation of the real space RS shown in the video VD_OR. The virtual space VS is expressed in the same world coordinate system (X, Y, Z) as the real space RS. In the virtual space VS, a static object (virtual static object) SO_VR is defined, which corresponds to a static object SO existing in the real space RS and is an abstraction of this static object SO. The configuration of the virtual static object SO_VR (e.g., position, orientation, shape, size, etc.) is defined by the recognition information REC_SO of the static object SO. Therefore, the configuration of the virtual static object SO_VR roughly matches that of the static object SO.
[0030] In the reproduction process, the person HM shown in the video VD_OR is also rendered in the virtual space VS. When rendering the person HM, the configuration (e.g., position, posture, size, etc.) of a person (virtual person) HM_VR that is an abstraction of the person HM is defined based on the recognition information REC_HM of the person HM. The posture of the virtual person HM_VR is defined based on the two-dimensional posture or three-dimensional posture included in the recognition information REC_HM. Therefore, in the example shown in FIG. 3, the virtual person HM_VR is represented by a skeleton shape.
[0031] When a dynamic object NHM other than a person HM appears in the video VD_OR, the reproduction process renders this dynamic object NHM in the virtual space VS simultaneously with the person HM. When rendering the dynamic object NHM, the configuration (e.g., position, orientation, shape, size, etc.) of a dynamic object (virtual dynamic object) NHM_VR that is an abstraction of the dynamic object NHM is defined based on the recognition information REC_NHM of the dynamic object NHM.
[0032] In addition to rendering the virtual space VS, the reproduction process generates caption information CAP for the scene shown in the video VD_OR based on the language information LAN. The caption information CAP is, for example, a description of the person HM shown in the video VD_OR from the language information LAN, or a summary of this description. If a dynamic object NHM other than the person HM is shown in the video VD_OR, the caption information CAP is a description of the interaction between the person HM and the dynamic object NHM shown in the video VD_OR, or a summary of this description. The caption information CAP, together with the virtual space VS, constitutes the video VD_VR. When the video VD_VR is output from the display device 30, the caption information CAP is displayed near the display area of the virtual space VS.
[0033] Returning to Fig. 2, the reproduction process can be performed based on search information SAR input from input device 40. For example, if the search information SAR includes position and time information, the reproduction unit 28, in cooperation with the scene information recording unit 27, identifies scene information SCN whose position and time information matches those of the search information SAR. Then, the reproduction unit 28 performs the reproduction process to generate a video VD_VR from the identified scene information SCN and output it to the display device 30. If the search information SAR further includes keyword information KWD, the reproduction process may highlight an explanation corresponding to the keyword information KWD in the caption information CAP (see Fig. 4).
[0034] Fig. 5 is a block diagram showing another example of the functional configuration of the data processing device 20 shown in Fig. 1. In addition to the functional blocks (such as the object recognition unit 23) shown in Fig. 2, Fig. 5 also shows a scene information update unit 29. These functional blocks are realized, for example, by cooperation between the processing circuit 21 and storage device 22 shown in Fig. 1.
[0035] The scene information update unit 29 performs an "update process" to update the scene information SCN. In the update process, first, in cooperation with the scene information recording unit 27, a specific language is detected from the language information LAN included in the scene information SCN. The specific language is set in advance, for example, by a user of the management system. Examples of the specific language include a language that expresses a state of the person HM that is deemed abnormal from the perspective of maintaining the person's health or traffic safety. Examples of states that are deemed abnormal from the perspective of maintaining the person's health or traffic safety include the person HM falling, crouching, bleeding, and having an abnormal posture.
[0036] If a specific language is detected, in the update process, the video VD_OR from which the language information LAN was generated is identified in cooperation with the video recording unit 26. Then, a verbalization process is performed to re-assign the language information LAN to the scene shown in the identified video VD_OR. In this verbalization process, for example, a VLM model (Vision Language Models) is used to generate text that explains a detailed situation corresponding to the specific language. As a result, detailed language information LAN about the situation related to the specific language included in the scene shown in the identified video VD_OR is generated. If detailed language information LAN is generated, in the update process, the scene information SCN is updated in cooperation with the scene information recording unit 27.
[0037] 3.Effects According to the embodiment described above, a reproduction process of a scene shown in the video VD_OR is performed. In the reproduction process, a virtual space VS that abstractly represents the real space RS shown in the video VD_OR is rendered, and a virtual person HM_VR that abstracts the person HM shown in the video VD_OR is further rendered in this virtual space VS. Therefore, compared to using the video VD_OR directly, the data size required to reproduce a scene from the video VD_OR can be reduced. Furthermore, since the virtual person HM_VR is used, the privacy of the person HM can also be protected. [Explanation of symbols]
[0038] 10...camera, 20...data processing device, 21...processing circuit, 22...storage device, 23...object recognition unit, 24...verbalization unit, 25...scene information generation unit, 26...video recording unit, 27...scene information recording unit, 28...reproduction unit, 29...scene information update unit, 30...display device, 40...input device, HM...person, MO...dynamic object, SO...static object, NHM...dynamic object other than person, HM_VR...virtual person, SO_VR...virtual static object, NHM_VR...virtual dynamic object, VD_OR, VD_VR...video, CAP...caption information, KWD...keyword information, LAN...language information, SAR...retrieval information, SCN...scene information, REC_MO, REC_SO...recognition information
Claims
1. A system for managing camera footage, a storage device in which the camera image is stored; a processing circuit for performing various processes; a display device that outputs an image; Equipped with the processing circuitry generating recognition information of an object captured in the camera image by performing object recognition processing on the camera image stored in the storage device; generating linguistic information of the scene shown in the camera image by performing a linguistic process on the camera image stored in the storage device; If the object recognition information includes person recognition information, scene information is generated that associates the person recognition information with linguistic information of a scene generated by the linguistic processing of the camera image in which the person is recognized, and the generated scene information is stored in the storage device; performing a reproduction process of the scene captured in the camera image based on the scene information stored in the storage device; The reproduction process Rendering an abstract image of the space captured in the camera image based on static object recognition information included in the scene information; Rendering an abstract image of the person captured in the camera image on the abstract image of the space based on the person recognition information included in the scene information; generating caption information for the scene based on language information for the scene included in the scene information; adding the caption information to the abstract image of the space into which the abstract image of the person has been rendered to generate a representation for output on the display device; A camera image management system comprising:
2. 10. The system of claim 1, the processing circuitry If the object recognition information includes recognition information of a dynamic object other than a person, add the recognition information of the dynamic object to the scene information; The reproduction process further comprises: Rendering an abstract image of the dynamic object on the abstract image of the space based on recognition information of the dynamic object included in the scene information; A camera image management system comprising:
3. 3. The system according to claim 1 or 2, The processing circuitry further comprises: detecting a predetermined specific language by referring to language information of the scene included in the scene information; When linguistic information of the scene including information of the specific language is detected, detailed linguistic information of the scene shown in the camera image is re-generated by performing a linguistic processing again on the camera image from which the detected linguistic information of the scene was generated; The scene information including the information of the specific language is updated based on detailed language information of the scene. A camera image management system characterized by:
4. 3. The system according to claim 1 or 2, further comprising an input device for inputting information; The processing circuitry further comprises: When search information for a scene captured in the camera image is input from the input device, scene information that matches the search information is identified based on the search information and language information for the scene included in the scene information stored in the storage device; The reproduction process is performed based on scene information that matches the search information. A camera image management system characterized by:
5. 3. The system according to claim 1 or 2, the abstract image of the space includes a three-dimensional image that abstracts the space captured in the camera image, The abstract image of the person includes a three-dimensional image that abstracts the person captured in the camera image. A camera image management system characterized by:
Citation Information
Patent Citations
Method for reporting state on site its system and image pickup unit using it
JP2002024962A
Imaging apparatus for monitor and its control method
JP2006295251A
Program, device, and method for recognizing actions of persons using a plurality of recognition engines
JP2019144830A
Information processing apparatus, information processing system, and information processing method
JP2022056533A