Summary image generation device, summary image generation method, and computer program
The summary image generating device addresses the inadequacy of single thumbnail images by extracting and representing objects of interest and background images, enhancing the representation of video content.
Patent Information
- Application Number
- JP2024082782
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-12-04
AI Technical Summary
Existing video recording systems fail to adequately represent the content of a recording file using a single thumbnail image, as multiple objects of interest may not be displayed, leading to missed important content.
A summary image generating device that extracts an object of interest and background image from video, generating a summary image based on these elements to represent the content effectively.
Enables the generation of summary images that appropriately represent the content of a video, ensuring important objects are displayed and facilitating easier identification of relevant recording files.
Smart Images

Figure 2025176548000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a summary image generating device, a summary image generating method, a computer program, and the like. [Background technology]
[0002] In video recording systems that record and manage camera footage, such as security camera systems, a technique is known that displays a list of thumbnail images of each recording file, allowing monitors to find the recording files containing content they need to check from the large number of stored recording files.
[0003] The video recording system displays thumbnail images of each recording file in a tiled format on the monitor operated by the monitor, and the monitor can select one of the displayed thumbnail images to play the recording file linked to that thumbnail image.
[0004] To determine the image to be displayed as the thumbnail image for a recorded file, any one video frame within the video, such as the first video frame or a video frame with a lot of movement within the video, is extracted, and the obtained image is resized to a size suitable for display and used as the thumbnail image.
[0005] For example, Patent Document 1 discloses a configuration in which important video frames are extracted from a video and used as thumbnail images of the video. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-8508 Summary of the Invention [Problem to be solved by the invention]
[0007] However, in a video recording system like the one in Patent Document 1, which uses a video frame in a recording file as a thumbnail image, there is a problem in that a single thumbnail image cannot adequately represent the contents of the recording file. For example, if multiple objects that deserve attention appear in a video at different times, one of the objects of interest may not be displayed in the thumbnail image, and the monitor may miss a recording file with content that they should check.
[0008] The present invention has been made in view of the above-mentioned problems, and one of its objects is to provide a summary image generating device capable of generating a summary image that appropriately represents the content of a video. [Means for solving the problem]
[0009] In order to solve the above problems, the summary image generating device according to the present invention comprises: an extraction means for extracting an object of interest that satisfies a predetermined condition from a video and extracting a background image from the video excluding at least the object of interest; a summary image generating means for generating a summary image based on the object image generated corresponding to the target object and the background image; The present invention is characterized by comprising: [Effects of the Invention]
[0010] According to the present invention, it is possible to realize a summary image generating device capable of generating a summary image that appropriately represents the content of a video. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram showing an example of the configuration of a video recording system according to a first embodiment of the present invention. [Figure 2] 2 is a diagram illustrating an example of the hardware configuration of a video analysis server 20B according to the first embodiment. FIG. [Figure 3] 3 is a functional block diagram showing an example of the functional configuration of a video analysis server 20B according to the first embodiment. FIG. [Figure 4] FIG. 3 is a diagram showing a table illustrating an example of video information according to the first embodiment. [Figure 5] FIG. 3 is a diagram showing a table illustrating an example of object information according to the first embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of detailed movement trajectory information of movement trajectory ID: B1. [Figure 7] FIG. 10 is a diagram showing an example of a condition setting screen for selecting an object of interest to be used in a summarized thumbnail image according to the first embodiment. [Figure 8] 10 is a flowchart showing an example of a process of a summary image generation method executed by a video analysis server 20B according to the first embodiment. [Figure 9] 10 is a flowchart showing an example of processing executed by a video analysis server 20B according to the second embodiment. [Figure 10] FIG. 10 is a diagram showing an example of an object identification table 1000 for identifying a target object based on display area information according to the second embodiment. [Figure 11] FIG. 11 is a diagram showing an example of a GUI for making settings related to thumbnail image generation according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.
[0013] (Embodiment 1) In the first embodiment, an example of a system that performs image analysis on a recording file of video captured by a surveillance camera, extracts objects, and generates thumbnail images will be described.
[0014] In this embodiment, only the set target object and the background are displayed in the thumbnail image, and unnecessary objects are not displayed, so that even small thumbnail images can be displayed that are easy to view.
[0015] In the following, this embodiment will be described using as an example a video recording system in which a network camera is used as the imaging device, but this embodiment is not limited to this case and can also be applied to other thumbnail image display applications.
[0016] FIG. 1 is a diagram showing an example of the configuration of a video recording system according to a first embodiment of the present invention, and shows an example of a network configuration when the video recording system is applied to a network camera system.
[0017] In FIG. 1, a network camera system 100 includes cameras (for example, network cameras) 10A and 10B as imaging devices, a video recording server 20A, a video analysis server 20B, and an operation terminal 20C.
[0018] The video recording server 20A, the video analysis server 20B, and the operation terminal 20C function as information processing devices. The cameras 10A and 10B are connected to the video recording server 20A, the video analysis server 20B, and the operation terminal 20C via a network 30.
[0019] The network 30 may be a LAN (Local Area Network), the Internet, a WAN (Wide Area Network), or a combination of these.
[0020] The network 30 may be configured in any communication standard, scale, or configuration as long as it allows communication between the cameras 10A and 10B and the video recording server 20A, video analysis server 20B, and operation terminal 20C.
[0021] The physical connection to the network 30 may be wired or wireless. Furthermore, the numbers of cameras 10A and 10B and information processing devices such as the video recording server 20A, the video analysis server 20B, and the operation terminal 20C connected to the network 30 are not limited to those shown in FIG.
[0022] Cameras 10A and 10B are imaging devices such as surveillance cameras, each capturing an image of a predetermined subject within a monitored space at a predetermined angle of view. Cameras 10A and 10B can transmit the captured images to at least one of video recording server 20A, video analysis server 20B, and operation terminal 20C via network 30.
[0023] The cameras 10A and 10B may transmit video in response to requests from the video recording server 20A, the video analysis server 20B, and the operation terminal 20C, or may actively transmit video to the video recording server 20A, the video analysis server 20B, and the operation terminal 20C. For example, the video recording server 20A, the video analysis server 20B, and the operation terminal 20C may be configured as physically independent devices, or may be configured as the same device.
[0024] The video recording server 20A is configured with a storage device that receives, saves, and accumulates the video transmitted from the camera 10A via the network 30. In addition, the video recording server 20A transmits the stored video to each device in response to requests received from the video analysis server 20B as a video analysis server and the operation terminal 20C.
[0025] The video analysis server 20B is an information processing device that receives the video recorded in the video recording server 20A via the network 30 and performs video analysis processing. In this embodiment, the video analysis server 20B functions as a summary image generation device.
[0026] In this embodiment, a thumbnail image is generated as the summary image, but the summary image does not have to be a thumbnail image and may be, for example, an image of a normal size. Specific processing of the video analysis processing will be described later.
[0027] The operation terminal 20C is configured by a terminal device that can be operated by an operator, such as a personal computer (PC), a smartphone, a tablet PC, etc. The operation terminal 20C has a display unit configured by a display or the like, and has a display control function for displaying images and playing videos.
[0028] The operation terminal 20C can display a list of thumbnail images of the video received from the camera 10A, the video recorded in the video recording server 20A, and the recorded video received from the video analysis server 20B.
[0029] Furthermore, the operation terminal 20C has a user interface and an input unit, and is configured to be able to accept settings of extraction parameters by the operator, selection of recorded video to be played from a list of thumbnail images of recorded video, etc. The operation terminal 20C transmits the settings accepted from the operator to the video analysis server 20B, and the video analysis server 20B performs video analysis processing based on the settings received from the operation terminal 20C.
[0030] Fig. 2 is a diagram showing an example of the hardware configuration of the video analysis server 20B according to the first embodiment, and the hardware configurations of the video recording server 20A and the operation terminal 20C may be similar to the configuration shown in Fig. 2. The video recording server 20A, the video analysis server 20B, and the operation terminal 20C may have the same hardware configuration.
[0031] The video analysis server 20B includes a CPU 201 as a computer, a ROM 202, a RAM 203, an external memory 204, an input device 205, an output device 206, a communication I / F 207, and a system bus 208. Note that the video analysis server 20B may further include components other than those described above.
[0032] The CPU 201 controls the overall operation of the video analysis server 20B. The ROM 202 is a non-volatile memory that stores computer programs and data required for the CPU 201 to execute processing. Note that the computer programs may be stored in the external memory 204 or a removable storage medium (not shown).
[0033] The RAM 203 functions as a main memory, a work area, etc. for the CPU 201. When executing processing, the CPU 201 loads necessary computer programs, etc. from the ROM 202 into the RAM 203 and executes the computer programs, etc. to realize various functional operations.
[0034] The external memory 204 is a non-volatile storage device such as a hard disk drive (HDD), flash memory, or SD card, and may be detachable. The external memory 204 is used as a permanent storage area for the OS, various programs, and various data, and can also be used as a short-term storage area for various data.
[0035] The input device 205 is an I / O device that can be operated by an operator, including a keyboard, a mouse, or other pointing device. The output device 206 has a monitor such as a liquid crystal display (LCD) and can present information to the operator by displaying images, playing videos, etc. The communication I / F 207 transmits and receives data to and receives data from the cameras 10A and 10B, the video recording server 20A, and the operation terminal 20C via the network 30.
[0036] Some or all of the functions of the video analysis server 20B are realized by the CPU 201 executing a computer program stored in the ROM 202 or the external memory 204. However, at least some of the functions of the video analysis server 20B may be performed by dedicated hardware. In this case, the dedicated hardware operates under the control of the CPU 201.
[0037] Here, we have explained the case where the video recording server 20A, the video analysis server 20B, and the operation terminal 20C have the same hardware configuration, but the video recording server 20A and the video analysis server 20B do not need to be equipped with the input device 205 and the output device 206.
[0038] The hardware configuration of cameras 10A and 10B can also be similar to the hardware configuration shown in Fig. 2. However, cameras 10A and 10B are provided with an imaging unit instead of output device 206.
[0039] The imaging unit captures an image of a subject and generates a captured image. The imaging unit includes an imaging element such as a complementary metal oxide semiconductor (CMOS) or a charge coupled device (CCD), an A / D converter, a development processing unit, and the like.
[0040] In the case of cameras 10A and 10B, input device 205 is composed of a power button, setting button, etc., and operators of cameras 10A and 10B can give instructions to cameras 10A and 10B via input device 205.
[0041] Some or all of the functions of cameras 10A and 10B are realized by a CPU in cameras 10A and 10B corresponding to CPU 201 executing a computer program. However, at least some of the functions of cameras 10A and 10B may be performed by dedicated hardware. In this case, the dedicated hardware operates under the control of the CPU.
[0042] In this embodiment, the video analysis server 20B operates as an information processing device that performs video analysis processing to generate thumbnail images. However, the video recording server 20A, the operation terminal 20C, or a general PC connected to the cameras 10A and 10B so as to be able to communicate with them may operate as the information processing device that performs the video analysis processing, or the cameras 10A and 10B may operate as the information processing device.
[0043] Fig. 3 is a functional block diagram showing an example of the functional configuration of the video analysis server 20B according to the first embodiment, and shows an example of the functional configuration of the video analysis server 20B. Note that some of the functional blocks shown in Fig. 3 are realized by causing a CPU or the like serving as a computer included in the video analysis server to execute a computer program stored in a memory serving as a storage medium for the video analysis server.
[0044] However, some or all of these functions may be implemented by hardware, which may be a dedicated circuit (ASIC) or a processor (reconfigurable processor, DSP).
[0045] Furthermore, the functional blocks shown in FIG. 3 do not have to be housed in the same housing, and may be configured as separate devices connected to each other via signal paths.
[0046] The video analysis server 20B includes a control unit 301 , a receiving unit 302 , an acquiring unit 303 , an extracting unit 304 , an object information acquiring unit 305 , an identifying unit 306 , and a generating unit 307 .
[0047] In this embodiment, the video analysis server 20B is described as having the functions shown in Fig. 3, but some of the functions may be provided by other devices. For example, some of the functions may be provided by the camera 10A, or by other information processing devices including the video recording server 20A.
[0048] The control unit 301 controls the receiving unit 302 , the acquiring unit 303 , the extracting unit 304 , the object information acquiring unit 305 , the identifying unit 306 , and the generating unit 307 .
[0049] The receiving unit 302 acquires video to be processed. Here, the video acquired by the receiving unit 302 may be recorded video captured and recorded by the camera 10A or the camera 10B. The receiving unit 302 may acquire recorded video stored in the external memory 204, or may acquire recorded video on the network 30 via the communication I / F 207.
[0050] The acquisition unit 303 acquires extraction conditions for an object of interest and a video analysis target range. The extraction unit 304 extracts an object of interest and a background image from the video acquired by the receiving unit 302, based on the extraction conditions for an object of interest and the video analysis target range acquired by the acquisition unit 303.
[0051] There are several methods for extracting foreground objects such as moving objects from video, but here we use a method that combines background subtraction and inter-frame subtraction. However, other methods for extracting objects may also be used.
[0052] The background image is an image obtained by excluding at least the object of interest extracted by the extraction unit 304 from the video. Here, the extraction unit 304 functions as extraction means that extracts an object of interest that satisfies a predetermined condition from the video and extracts a background image obtained by excluding at least the object of interest from the video.
[0053] The object information acquisition unit 305 acquires movement trajectory information indicating the behavior of the object extracted by the extraction unit 304 and attribute information of the object, and generates object information. The identification unit 306 identifies the target object from the object information acquired by the object information acquisition unit 305, based on the extraction conditions for the target object acquired by the acquisition unit 303.
[0054] The generation unit 307 generates thumbnail images of the video within the video analysis range based on the object information of the target object identified by the identification unit 306 and the background information acquired by the extraction unit 304. The generated thumbnail images are stored in the external memory 204, and can be displayed in a list together with other thumbnail images on the output device 206, such as the display of the operation terminal 20C.
[0055] When the operator selects one thumbnail image from the multiple thumbnail images displayed in a list, the recording file associated with the selected thumbnail image can be obtained from the video recording server 20A and played on the output device 206.
[0056] In addition, by placing the mouse cursor on an object in a desired thumbnail image among the thumbnail images displayed in a list, the image of the object may be enlarged. When an object is selected and playback of the recorded video is started, playback of the recorded video may start from the time the object appears.
[0057] In this embodiment, the object information acquisition unit 305 extracts a person or a vehicle as an example of the object. However, the object to be extracted is not limited to this, and specific objects such as animals or suspicious objects can also be extracted.
[0058] In this embodiment, the extraction unit 304 extracts an object from the video acquired by the receiving unit 302, and the object information acquisition unit 305 acquires movement trajectory information indicating the behavior of the object and attribute information of the object to generate object information. However, the receiving unit 302 may receive the extracted object information.
[0059] For example, the camera 10A, the video recording server 20A, or another information processing device may extract an object from a video and generate object information, and the receiving unit 302 may receive the extracted object information and the video. In this case, the video analysis server 20B can omit the extraction process in the extraction unit 304 and the object information generation process in the object information acquisition unit 305.
[0060] The thumbnail image generation process in this embodiment will be specifically described below.
[0061] 4 is a diagram showing a table representing an example of video information according to the first embodiment, in which the video information table 400 represents video information relating to recorded video within the analysis target range acquired by the acquisition unit 303. The video information table 400 includes an identifier 401, a source ID 402, a start time 403, an end time 404, an object information ID 405, a background image 406, and a thumbnail image 407.
[0062] Videos 408 to 410 are identified by identifier 401, which is an ID for identifying video information. Source ID 402 is an ID for identifying the video source.
[0063] The start time 403 is the start time of the video, and the end time 404 is the end time of the video. The start time 403 and the end time 404 are acquired by the acquisition unit 303. For the video specified by the source ID 402, the analysis target range is from the start time 403 to the end time 404 of the video.
[0064] The acquisition unit 303 periodically acquires recorded video from the video recording server 20A and performs thumbnail image generation processing before a thumbnail image display request is made, thereby speeding up the display of thumbnail images. For video 408 and video 409 in Figure 4, thumbnail image generation processing is performed every hour for source ID: S001.
[0065] The object information ID 405 includes the object information ID of the object extracted from the video of the analysis range target. The background image 406 includes a background image for the recorded video of the analysis range extracted by the extraction unit 304. The thumbnail image 407 includes a thumbnail image for the recorded video of the analysis range generated by the generation unit 307.
[0066] FIG. 5 is a diagram showing a table representing an example of object information according to embodiment 1, and object information table 500 represents a table of object information contained in the recorded video of the analysis target range acquired by extraction unit 304 and object information acquisition unit 305.
[0067] 5, object information table 500 includes identifier 501, type 502, start time 503, movement trajectory ID 504, and attribute information 505. Objects 506 to 510 are examples of objects included in recorded video within the analysis target range and identified by identifier 501. Note that object information table 500 in FIG. 5 shows a detailed example of object information for object information ID: OL001 of video 408 in FIG. 4.
[0068] Identifier 501 is an ID for identifying object information, and a unique ID is assigned to each object extracted from video by extraction unit 304. Type 502 includes the object type obtained by performing image analysis of the extracted object by object information acquisition unit 305. In Fig. 5, for example, objects 506 and 507 are classified as human bodies.
[0069] Start time 503 is the time when an object appears in a video, and is a relative time from the start time of the video. In Fig. 5, object 506 appears when 0 seconds have elapsed since the start of playback of video 408. Movement trajectory ID 504 is the ID of the movement trajectory information indicating the behavior of the object, acquired by extraction unit 304. For example, the movement trajectory ID of object 506 is B1.
[0070] The attribute information 505 includes attribute information of the object acquired by the object information acquisition unit 305 through video analysis of the object. The attribute information is detected by applying a classifier that has undergone machine learning to identify a desired detection target object, such as a fallen state or an ambulance, to an image of the object.
[0071] 5, for example, the attributes of object 506 are stored as a fall, and the attribute of object 510 is stored as an ambulance. The attribute information 505 may include information other than the type of object. In addition to the method using a machine-learned classifier, attributes may be assigned based on tracking of a detected object, estimating the posture of a detected human body, or posture changes. Alternatively, object information obtained from a source other than video analysis, such as a depth sensor or temperature sensor, may be included.
[0072] Fig. 6 is a diagram illustrating an example of detailed trajectory information of trajectory ID: B1. That is, trajectory information 600 in Fig. 6 shows a detailed example of trajectory information of object 506 having trajectory ID: B1.
[0073] The movement trajectory information 600 in FIG. 6 represents information indicating the behavior of the object 506 acquired by the extraction unit 304. The movement trajectory information 600 includes an identifier 601, a relative time 602, a circumscribing rectangle 603, and attribute information 604, and information on each video frame is stored in chronological order.
[0074] An identifier 601 indicates the identifier of each video frame. A relative time 602 indicates the relative time from when the object 506 appears in the video. This table records the movement trajectory of the object 506, whose identifier is OBJ001, from time 0 to 79.
[0075] The circumscribing rectangle 603 represents the circumscribing rectangle (x coordinate, y coordinate, width, height) of the object 506 for each video frame. Here, the object range is a circumscribing rectangle, but any format that allows the range to be specified may be used. For example, it may be polygonal data composed of multiple coordinates, or it may be in the form of mask information that represents the object position as bitmap data in pixel units.
[0076] Attribute information 604 represents attribute information of object 506 in each video frame. The attribute information 604 is detected using the same method as the attribute information detected in attribute information 505, but stores attribute information that can be detected on a video frame-by-video frame basis.
[0077] 7 is a diagram showing an example of a condition setting screen for extracting an object of interest to be used in a summarized thumbnail image according to embodiment 1. In this embodiment, a summarized thumbnail image is a thumbnail image in which only the image of the extracted object of interest is combined with a background image to clearly display the presence, position, etc. of the image of interest together with the background image.
[0078] The GUI 700 is displayed on the display of the operation terminal 20C, and receives input from the operator. The input result is acquired by the acquisition unit 303 as a condition for extracting an object of interest, and is used by the identification unit 306.
[0079] The above extraction conditions can be set using the GUI 700 at any time before the thumbnail images are displayed. The GUI 700 has multiple check boxes for selecting conditions for extracting an object of interest from multiple objects. The conditions selectable using the GUI 700 are the same as the object types and attributes that can be acquired by the object information acquisition unit 305.
[0080] In the initial state, the GUI 700 allows the types "People," "Man," "Woman," and "Child" (check boxes 701 to 704) to be selected.
[0081] In addition, in this embodiment, the GUI 700 is set to set, for example, a male fallen person (type "Man", state "Fall") and an ambulance vehicle (type "Car" or "Truck", vehicle type "Ambulance") as specific objects.
[0082] To do this, in GUI 700, for example, all check boxes for types selected in the initial state are unchecked, new check boxes 702, 705, 706, 707, and 708 are selected, and the fallen man and the ambulance are set as the object of interest conditions.
[0083] After setting the target object conditions in the GUI 700, clicking the apply button 709 closes the GUI 700. The target object conditions set in the GUI 700 are used by the identification unit 306 as conditions for extracting the target object.
[0084] In this embodiment, the fallen man and the ambulance are set as the object of interest conditions in the GUI 700. Therefore, the generation unit 307 generates an image in which the object 506 and the object 510 are superimposed on the background image BG1, and this image is recorded as a thumbnail image I1 of the object 506 and used when displaying the thumbnail image of the video 408.
[0085] Next, the operation of the video analysis server 20B will be described with reference to a flowchart, in which a method for generating thumbnail images using the video information table 400 in FIG. 4 and the object information table 500 in FIG.
[0086] Fig. 8 is a flowchart showing an example of a process of a summary image generation method executed by the video analysis server 20B according to embodiment 1. The process flow shown in Fig. 8 starts, for example, when the receiving unit 302 receives recorded video from the camera 10A or the camera 10B.
[0087] The video analysis server 20B implements each process shown in the flowchart of FIG. 8 by the CPU 201 serving as the computer in FIG. 2 reading and executing the necessary computer programs.
[0088] First, in step S801, the acquisition unit 303 acquires the extraction conditions for the target object and the analysis target range of the video via the reception unit 302, and records them in the video information table 400 of Fig. 4. The extraction conditions for the target object are set by receiving input from the operator using the GUI 700 displayed on the display of the operation terminal 20C, as described in Fig. 7.
[0089] Here, step S801 functions as an extraction condition acquisition step in which the GUI 700 as an extraction condition acquisition means acquires extraction conditions for extracting an object of interest.
[0090] The extraction conditions for the target object set in the GUI 700 are used by the identification unit 306. Information about the analysis target range of the acquired video is stored in the source ID 402, start time 403, and end time 404 in FIG.
[0091] In this embodiment, the extraction conditions for the target object are set using the GUI 700, but this is not limiting. A history of extraction conditions for the target object that were previously set using the GUI 700 may be stored, and when receiving video, the selected extraction conditions may be read from the past setting history and applied as setting information. In this case, the setting information history may be applied commonly to all recorded video, or a history may be stored and applied for each video source.
[0092] In step S802, the target video is acquired. That is, the acquisition unit 303 acquires, via the reception unit 302, the recorded video of the analysis target range acquired in step S801.
[0093] In step S803, the video analysis server 20B determines whether or not to generate a summary thumbnail image.
[0094] If extraction conditions have been set for the background image in step S801, it is determined that a summary thumbnail is to be generated (Yes in step S803), and the process proceeds to step S804. On the other hand, if it is determined that a summary thumbnail is not to be generated (No in step S803), the process proceeds to step S807.
[0095] In step S804, the extraction unit 304 extracts an object and background image from the target recorded video acquired in step S802, and acquires object information. Here, step S804 functions as an object information acquisition step (object information acquisition means) that acquires object information, which is information about the target object.
[0096] 5, including type 502, start time 503, movement trajectory ID 504, and attribute information 505. Here, "fall" is stored as the attribute information of object 506, and "ambulance" is stored as the attribute information of object 510. Note that the background image may be extracted by removing at least the target object from the video in step S805, instead of step S804.
[0097] In step S805, the identification unit 306 identifies the target object from the object information acquired in step S804 based on the target object extraction conditions acquired in step S801. In this case, the fallen man and the ambulance are set as the specified objects, so all extracted objects (objects 506 to 510) are scanned, and object 506 and object 510 that match the specified object extraction conditions acquired in step S801 are identified as the specified objects.
[0098] As described above, the extraction of the background image in step S804 may be performed after the extraction of the target object. Here, steps S804 and S805 function as an extraction step (extraction means) that extracts the target object that satisfies predetermined conditions from the video and extracts a background image from the video excluding at least the target object.
[0099] In step S806, generation unit 307 acquires an image of the object identified in step S805. The object image is the object image of the video frame with the largest object area among the video frames in which the object exists. In this way, in step S806, the object image is generated based on the object information, and also based on the target range of the video and the extraction conditions for the target object.
[0100] In step S807, if a summary thumbnail image is to be generated (if Yes in step S803), the generation unit 307 generates the summary thumbnail image by superimposing the object image acquired in step S806 on the background image extracted in step S804.
[0101] Here, step S807 functions as a summary image generating step (summary image generating means) that generates a summary image based on the object image generated corresponding to the target object and the background image.
[0102] If a summary thumbnail image is not to be generated (No in step S803), the first video frame of the recorded video acquired in step S802 is acquired as the first thumbnail image. The thumbnail image generated in step S807 is saved as thumbnail image 407 in FIG.
[0103] If there are multiple specific objects in the summary thumbnail image, the object image with the earliest application time in Figure 5, i.e., the start time 503, is superimposed on the background image. Then, the object image with the latest start time 503 is superimposed so as to be positioned in front of the earlier object image.
[0104] However, although the above describes an example in which an object image that appears later is superimposed so as to be positioned in front, it is also possible to, for example, increase the transparency of the object image superimposed in front so that the object image of the object behind it can be seen. Alternatively, the display position of the object image of the next target object may be shifted to a position where it does not obscure the previously placed object image. In other words, the display positions of the object images may be determined so as to reduce the overlap of multiple object images present on the same screen.
[0105] Alternatively, each object image may be reduced in size so that all object images of the target object can be displayed, and then arranged within the summary thumbnail image. Alternatively, if the target object size is too small to be easily recognized in the summary thumbnail image, the extracted object image may be enlarged to a predetermined size or larger and then superimposed on the background image.
[0106] Furthermore, in this embodiment, the object image is extracted from the video frame with the largest object area in step S806, but it is also possible to extract object images by changing the video frame from which they are cut out for each target object so as to minimize overlap of object images of multiple target objects.
[0107] For example, a pattern that minimizes the overlapping area of object images of multiple target objects may be detected from all patterns of combinations of cut-out images of all video frames of all target objects, and then the object image of each target object may be extracted from each video frame and superimposed to achieve this pattern.
[0108] In this embodiment, if a summary thumbnail image is not generated (No in step S803), the first video frame of the recorded video is used as the thumbnail image, but this is not limited to this. The video frame with the most moving objects or the largest moving object area may be selected from the video. Alternatively, a video frame after a predetermined time (e.g., 5 seconds) from the start may be selected as the thumbnail image.
[0109] Next, in step S808, it is determined whether or not all thumbnail generation processes have been completed. That is, the video analysis server 20B determines whether or not to continue the thumbnail image generation process of Fig. 8. If thumbnail image generation processes for the analysis target range of all videos acquired in step S801 have been completed (Yes in step S808), the thumbnail image generation process of Fig. 8 ends.
[0110] On the other hand, if the thumbnail image generation process for the analysis target range of all videos has not been completed (No in step S808), the process returns to step S802, and the thumbnail image generation process for the remaining videos continues.
[0111] As described above, in this embodiment, recorded video is analyzed, an object of interest and a background image are extracted, and a condensed thumbnail image is generated, so that only the object of interest is displayed in the condensed thumbnail image, and unnecessary objects are not displayed. Therefore, it is possible to display a condensed thumbnail image that is easy to view.
[0112] (Embodiment 2) In the first embodiment, the conditions for extracting an object of interest used to generate a summarized thumbnail image are set using the GUI 700 .
[0113] However, for example, when multiple thumbnail images are displayed on a screen, the size of each thumbnail image becomes smaller as the number of thumbnail images displayed increases. As a result, the size of the objects displayed in each thumbnail image also becomes smaller, making it difficult to visually identify what objects are displayed in each thumbnail image, and there is a risk of missing a recorded file that should be checked.
[0114] Therefore, in the second embodiment, in addition to the extraction conditions set in the first embodiment, control is performed so that the generated summary thumbnail images are changed based on area information when a list of multiple summary thumbnail images is displayed.
[0115] Fig. 9 is a flowchart showing an example of processing executed by the video analysis server 20B according to embodiment 2. Note that the operation of each step in the flowchart of Fig. 9 is performed sequentially by a CPU or the like serving as a computer in the video analysis server executing a computer program stored in memory.
[0116] In FIG. 9, steps S901 to S904 and step S909 are the same as steps S801 to S804 and step S808 in FIG. 8, respectively, and therefore description thereof will be omitted.
[0117] In step S905, the acquiring unit 303 acquires display area information of the display device of the operation terminal 20C on which thumbnail images are displayed in a list. The display area information is information used when displaying thumbnail images in a list, such as the display size (width, height) of each thumbnail image. Here, step S905 functions as a display area information acquiring step (display area information acquiring means) that acquires display area information used when displaying thumbnail images in a list.
[0118] When thumbnail images are displayed in a list in a tiled arrangement on the screen, the number of rows and columns of thumbnail images displayed in a list may be acquired as the display area information. Alternatively, the display area information may be information on the total number of thumbnail images displayed on the screen or the zoom magnification when the thumbnail images are displayed.
[0119] That is, the display area information in this embodiment includes at least one of the display size (width, height) of each thumbnail image, the number of rows and columns of thumbnail images displayed in a list, the number of thumbnail images displayed in a list on the screen, and the zoom magnification when displaying the thumbnail images.
[0120] In step S906, the identification unit 306 identifies the object of interest based on the extraction conditions for the object of interest acquired in step S901 and the display area information acquired in step S905.
[0121] 10 is a diagram showing an example of an object identification table 1000 for identifying an object of interest based on display area information according to embodiment 2. The object identification table 1000 includes parameters such as a thumbnail image display size 1001, a maximum number of objects to be displayed 1002, an extracted object size 1003, and a type 1004.
[0122] It should be noted that the parameters (1001 to 1004) in FIG. 10 are defined in advance taking into consideration the display size for displaying thumbnail images, but the parameters (1001 to 1004) can be changed arbitrarily.
[0123] 10, a thumbnail image display size 1001 corresponds to the display area information acquired in step S905. The object extraction conditions of the parameters (1002 to 1004) are switched according to the thumbnail image display size 1001 as the acquired display area information.
[0124] For example, if the thumbnail image display size 1001 is 320 pixels wide and 240 pixels high, the maximum number of objects displayed in each thumbnail image, that is, the maximum number of objects displayed in each thumbnail image, 1002, is 4. By limiting the maximum number of objects displayed, it is possible to improve the visibility of the objects displayed in each thumbnail image.
[0125] Extracted object size 1003 indicates the threshold for whether or not to display the extracted object based on the object size in the recorded video, and in the above case, indicates that only objects with a size of 60 pixels wide and 60 pixels high or larger can be selected as extracted objects. By limiting the display of small objects, small objects are no longer displayed in thumbnail images, improving visibility.
[0126] Type 1004 indicates the type of object that can be identified, and in the above case, only a human body or a vehicle can be identified as an object of interest. By determining whether or not to display an object based on Type 1004, it is possible to prevent object types that do not need to be checked by an observer from being displayed in thumbnail images, thereby improving the visibility of thumbnail images.
[0127] In this manner, in this embodiment, at least one of the number of object images included in a thumbnail image, the size of the object image, and the type of the target object is determined according to the display area information.
[0128] 9, in step S907, the generation unit 307 acquires an image of the object identified in step S906. That is, in step S907, an object image is generated based on the display area information.
[0129] In step S908, if a summary thumbnail image is to be generated (if Yes in step S903), the generation unit 307 generates the summary thumbnail image by superimposing the object image acquired in step S907 on the background image acquired in step S904.
[0130] If a summary thumbnail image is not to be generated (No in step S903), a thumbnail image is generated in step S908 from, for example, the first video frame of the recorded video acquired in step S902, and saved as thumbnail image 407 in FIG.
[0131] In the second embodiment, it is desirable to save thumbnail images for each thumbnail image display size. This makes it possible to display the already saved thumbnail image without performing the thumbnail image generation process again if the display size of the thumbnail image is the same as the saved thumbnail image display size. After that, step S909 is performed, and the thumbnail image generation process flow in FIG. 9 ends.
[0132] As described above, in the second embodiment, by determining a specific object based on display area information when displaying a list of thumbnail images and generating thumbnail images, it is possible to display thumbnail images that are easy to view according to the display size.
[0133] (Embodiment 3) In the first and second embodiments, the generation unit 307 generates a summary thumbnail image using the image of the object identified by the identification unit 306. In contrast, in the second embodiment, instead of extracting an object image from the recorded video, a summary thumbnail image is generated using a predetermined pseudo image generated based on object information as the object image.
[0134] The pseudo-image in this embodiment may be, for example, an illustration generated by CG (Computer Graphics), or an avatar obtained from a predetermined website, or may be an image obtained by editing or processing a predetermined natural image, or may be generated by a generation AI.
[0135] By generating summary thumbnail images using pseudo images, it is possible to prevent the identification of individuals in the summary thumbnail images and protect privacy. In addition, by using generated images that succinctly represent the characteristics of an object, unnecessary object information can be omitted, making it easier to visually recognize the characteristics of an object even in small thumbnail images.
[0136] In the second embodiment, an image generation model trained by machine learning that finds latent patterns from large amounts of data through repeated calculations is used as the image generator for generating the pseudo-images.
[0137] 11 is a diagram showing an example of a GUI for making settings related to the generation of summarized thumbnail images according to the third embodiment. The GUI 1100 is displayed on the display of the operation terminal 20C, and is used to receive input from the operator and make settings related to the generation of summarized thumbnail images. The conditions are set using the GUI 1100 at any timing before the display of summarized thumbnail images.
[0138] GUI 1100 has check boxes and radio buttons for specifying settings for generating summary thumbnail images. Check box 1101 is for selecting whether or not to display a summary thumbnail. Checking check box 1101 generates a summary thumbnail image, and a Yes determination is made in step S803 of FIG. 8.
[0139] On the other hand, if the check box 1101 is not checked, the determination in step S803 in FIG. 8 is No, and as explained in FIG. 8, a thumbnail image is generated using, for example, the first video frame.
[0140] Radio buttons 1102 and 1103 are buttons for selecting whether the operator will manually specify specific condition settings for the target object or whether they will be set automatically. When radio button 1102 is selected, a detailed settings button 1106 is enabled. Clicking the detailed settings button 1106 displays the GUI 700 described in FIG. 7, allowing settings to be made to select the conditions for the target object to be used in generating a thumbnail image.
[0141] On the other hand, when the radio button 1103 is selected, predetermined extraction conditions are applied as the target object extraction conditions. Also, when the radio button 1103 is selected, the check box 1104 is enabled, and the past setting history is used with priority.
[0142] In other words, the settings made in the past when target object extraction conditions were set are given priority, and target object extraction conditions are set using these settings. In this way, extraction conditions may be acquired based on extraction conditions used in the past.
[0143] Checkbox 1105 is a checkbox for selecting whether or not to protect the privacy of the human body. If checkbox 1105 is selected, a predetermined pseudo image automatically generated from object information will be used as the object image displayed in the thumbnail image, instead of an object image extracted from the recorded video.
[0144] The above-mentioned predetermined pseudo image is generated based on the object information acquired by the object information acquisition unit 305. For example, in the case of object 506 in Fig. 5, it can be seen from Fig. 5 that the type is a human body and the attribute information is a fallen state, and the moving direction and size of the object can be acquired from Fig. 6.
[0145] The generation unit 307 inputs the object position information and the object information acquired above as prompts, executes a process of generating a predetermined pseudo image using a generation AI or the like, and uses the pseudo image as the object image.
[0146] In this case, the object information used in the image generation process is acquired by the object information acquisition unit 305, and by acquiring more attribute information about the object by the object information acquisition unit 305, the reproduction accuracy of the specified pseudo-image generated in the image generation process can be improved.
[0147] Examples of information acquired by the object information acquisition unit 305 include color information and position of an object, the color of the upper and lower clothes and mask in the case of a human body, the presence or absence of a backpack and the color of the backpack, and the posture of the human body based on skeleton estimation.
[0148] Attribute information 505 includes attribute information of an object acquired by video analysis of the object in object information acquisition unit 305. The attribute information is detected by applying a classifier that has learned a desired detection target object, such as a fallen state or an ambulance, to an image of the object.
[0149] In the example of Figure 5, the attributes of object 506 and object 510 are stored as a fallen object and an ambulance, respectively. The attribute information may include information other than the object type. In addition to the method using a trained classifier, attributes may be assigned based on tracking of a detected object, estimating the posture of a detected human body, or posture changes. Alternatively, object information obtained from a source other than video analysis, such as a depth sensor or temperature sensor, may be included.
[0150] After configuring the summary thumbnail image generation settings in the GUI 1100, clicking the Apply button 1107 closes the GUI 1100.
[0151] As described above, by using a predetermined pseudo-image generated from object information to display a summary thumbnail image, it is possible to prevent the identification of individuals in the summary thumbnail image and protect privacy. Furthermore, by replacing the summary thumbnail image with a predetermined pseudo-image that succinctly represents only the characteristics of the object, unnecessary object information can be removed, making the characteristics of the object easier to see even in small thumbnail images.
[0152] (Modification of the third embodiment) In the third embodiment, a summary thumbnail image is generated using a predetermined pseudo-image generated from object information as the object image in the object image acquisition in the first embodiment. However, it is also possible to use an image generator and input an extraction condition for the recorded video as a prompt to generate a summary thumbnail image.
[0153] In this case, the image generator in this modification has the functions of the acquisition unit 303, extraction unit 304, object information acquisition unit 305, identification unit 306, and generation unit 307 in the first embodiment.
[0154] The image generator in this modification executes a pseudo-image generation process using the recorded video of the analysis target range and a prompt including the extraction object conditions as input data. By performing this pseudo-image generation process, a summarized thumbnail image of the recorded video similar to that in the first embodiment is generated.
[0155] As described above, in the modification of the third embodiment, it is possible to generate a digest thumbnail image that is easy to view from recorded video.
[0156] In the above-described first to third embodiments, an example has been described in which an extracted image of a target object or a pseudo image is synthesized with a background image to generate a summary thumbnail image that clearly displays the presence, position, etc. of the target image together with the background image. However, the above-described first to third embodiments are not limited to generating small images such as summary thumbnail images, and may also generate summary images of a desired size.
[0157] The present invention has been described above in detail based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible based on the spirit of the present invention, and these are not excluded from the scope of the present invention.
[0158] The present invention also includes those that realize the functions of the above embodiments using, for example, at least one processor such as a CPU, memory, or circuit (for example, ASIC). Also, multiple processors may be used to perform distributed processing.
[0159] In order to realize some or all of the control in the above embodiments, a computer program that realizes the functions of the above embodiments may be supplied to a thumbnail image generating device or the like via a network or various storage media. Then, a computer (or a CPU, MPU, or the like) in the thumbnail image generating device or the like may read and execute the program. In this case, the program and the storage medium storing the program constitute the present invention. The present invention also includes the following combinations.
[0160] (Configuration 1) A summary image generating device comprising: an extraction means for extracting an object of interest that satisfies predetermined conditions from a video and extracting a background image from the video excluding at least the object of interest; and a summary image generating means for generating a summary image based on an object image generated corresponding to the object of interest and the background image.
[0161] (Configuration 2) The summary image generating device according to configuration 1, further comprising extraction condition acquisition means for acquiring extraction conditions for extracting the target object.
[0162] (Configuration 3) The summary image generating device according to Configuration 2, wherein the summary image generating means generates the object image based on the target range of the video and the extraction conditions.
[0163] (Configuration 4) The summary image generating device according to configuration 2 or 3, wherein the extraction condition acquisition means acquires the extraction conditions based on the extraction conditions used in the past.
[0164] (Configuration 5) An object information acquisition means for acquiring object information that is information about the target object; 5. The abstract image generating device according to any one of configurations 1 to 4, wherein the object image is generated based on the object information.
[0165] (Configuration 6) The summary image generating device according to any one of configurations 1 to 5, wherein the summary image is a thumbnail image.
[0166] (Configuration 7) A summary image generating device according to Configuration 6, characterized in that a display area information acquiring means acquires display area information when displaying the thumbnail images in a list, and the summary image generating means generates the object image based on the display area information.
[0167] (Configuration 8) The summary image generating device according to Configuration 7, wherein the display area information includes at least one of the display size of the thumbnail images, the number of thumbnail images displayed in a list, the number of rows and columns of the thumbnail images, and the zoom magnification of the thumbnail images.
[0168] (Configuration 9) The summary image generating device according to Configuration 7 or 8, wherein the summary image generating means determines at least one of the number of object images included in the thumbnail image, the size of the object image, and the type of the target object, according to the display area information.
[0169] (Configuration 10) The summary image generating device according to any one of configurations 1 to 9, characterized in that the summary image generating means determines the display position of the object image so as to reduce overlap of multiple object images present on the same screen.
[0170] (Configuration 11) The summary image generating device according to any one of configurations 1 to 10, wherein the summary image generating means uses a predetermined pseudo image as the object image.
[0171] (Method) A summary image generation method comprising: an extraction step of extracting an object of interest that satisfies predetermined conditions from a video and extracting a background image from the video excluding at least the object of interest; and a summary image generation step of generating a summary image based on an object image generated corresponding to the object of interest and the background image.
[0172] (Program) A computer program for controlling each means of the summary image generating device according to any one of configurations 1 to 11 by a computer. [Explanation of symbols]
[0173] 100...Network camera system 10A, 10B...Camera 20A...Video recording server 20B...Video analysis server 20C...Operation terminal 301...Control unit 302...Receiver 303…Acquisition Department 304...Extraction part 305…Object information acquisition unit 306…Specific part 307...Generation section
Claims
1. an extraction means for extracting an object of interest that satisfies a predetermined condition from a video and extracting a background image from the video excluding at least the object of interest; a summary image generating means for generating a summary image based on the object image generated corresponding to the target object and the background image; A summary image generating device comprising:
2. 2. The apparatus for generating a summarized image according to claim 1, further comprising extraction condition acquisition means for acquiring extraction conditions for extracting the target object.
3. The summary image generating means Based on the target range of the video and the extraction conditions, 3. The apparatus for generating a summarized image according to claim 2, wherein the apparatus generates the object image.
4. The extraction condition acquisition means 3. The apparatus according to claim 2, wherein the extraction conditions are acquired based on the extraction conditions used in the past.
5. an object information acquisition means for acquiring object information that is information about the target object; 2. The apparatus according to claim 1, wherein the object image is generated based on the object information.
6. 2. The apparatus according to claim 1, wherein the summary image is a thumbnail image.
7. a display area information acquisition means for acquiring display area information when displaying the thumbnail images in a list; 7. The apparatus according to claim 6, wherein said abstract image generating means generates said object image based on said display area information.
8. The display area information is 8. The apparatus according to claim 7, wherein the display information includes at least one of a display size of the thumbnail images, a number of thumbnail images to be displayed in a list, a number of rows and columns of the thumbnail images, and a zoom magnification of the thumbnail images.
9. The summary image generating means 8. The summary image generating device according to claim 7, wherein at least one of the number of object images included in the thumbnail image, the size of the object image, and the type of the target object is determined according to the display area information.
10. The summary image generating means 2. The apparatus for generating a summarized image according to claim 1, wherein the display positions of the object images are determined so as to reduce overlapping of the object images present on the same screen.
11. The summary image generating means 2. The apparatus for generating an abstract image according to claim 1, wherein a predetermined pseudo image is used as the object image.
12. an extraction step of extracting an object of interest that satisfies a predetermined condition from the video and extracting a background image from the video excluding at least the object of interest; a summary image generating step of generating a summary image based on the object image generated corresponding to the target object and the background image; A method for generating a summary image, comprising:
13. A computer program for controlling each means of the abstract image generating apparatus according to any one of claims 1 to 11 by a computer.
Citation Information
Patent Citations
Significant information extraction method and device
JP2011008508A