Video analysis device, video analysis method, and program

JP7913290B2Active Publication Date: 2026-09-01NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022108648
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2026-09-01
Estimated Expiration
2042-07-05

AI Technical Summary

Benefits of technology

【0013】 本発明の一態様によれば、複数の映像を解析した結果を活用することが可能になる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007913290000001
    Figure 0007913290000001
  • Figure 0007913290000002
    Figure 0007913290000002
  • Figure 0007913290000003
    Figure 0007913290000003
Patent Text Reader

Abstract

To utilize a result of analyzing a plurality of videos.SOLUTION: A video analyzer 100 includes a type acceptance part 110, an acquisition part 111 and an integration part 112. The type acceptance part 110 accepts a selection of a kind of engine for detecting a detection object included in a plurality of videos through analyzing the plurality of videos. The acquisition part 111 acquires results of analyzing the plurality of videos using the selected type of engine, from among results of analyzing the plurality of videos using a plurality of kinds of engines, respectively. The integration part 112 integrates the results of analyzing the plurality of videos thus acquired.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video analysis apparatus, a video analysis method, and a program. Background Art

[0002] Patent Document 1 discloses a distributed object tracking system for tracking an object by concatenating analysis results from a plurality of image analysis apparatuses. This distributed object tracking system comprises a plurality of image analysis apparatuses and a cluster management service apparatus.

[0003] Each of the plurality of image analysis apparatuses described in Patent Document 1 is connected to at least one corresponding camera apparatus, and analyzes an object in at least one corresponding real-time video stream transmitted by the at least one corresponding camera apparatus to generate an analysis result of the object. Patent Document 1 discloses that the object includes a person or a suitcase, and the analysis result includes features of a person's face or characteristics of a suitcase.

[0004] The cluster management service apparatus described in Patent Document 1 is a cluster management service apparatus connected to a plurality of image analysis apparatuses, and concatenates the analysis results generated by each of the plurality of image analysis apparatuses to generate a trajectory of the object.

[0005] Note that Patent Document 2 describes a technique of calculating a feature amount for each of a plurality of key points of a human body included in an image, searching for images including human bodies with similar postures or similar movements based on the calculated feature amounts, and collectively classifying those having similar postures or movements. Non-Patent Document 1 describes a technique related to human skeleton estimation. Prior Art Documents Patent Documents

[0006] Patent Document 1 Japanese Unexamined Patent Application Publication No. 2020-184292 Patent Document 2 International Publication No. 2021 / 084677 [Non-patent literature]

[0007] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, [Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields];, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 [Overview of the project] [Problems that the invention aims to solve]

[0008] Generally, by analyzing video, it is possible to detect various features related to the appearance of the object being tracked, not limited to the features of a person's face or the characteristics of a suitcase. Although the distributed object tracking system described in Patent Document 1 can track an object within a real-time video stream, it is difficult to utilize the results of analyzing multiple videos for purposes other than tracking the object. Furthermore, Patent Document 2 and Non-Patent Document 1 do not disclose any technology that utilizes the results of analyzing multiple videos.

[0009] One example of the object of the present invention is to provide a video analysis device, a video analysis method, and a program that solve the problem of utilizing the results of analyzing multiple videos, in view of the above-mentioned problems. [Means for solving the problem]

[0010] According to one aspect of the present invention, A type receiving means that accepts the selection of the type of engine for analyzing each of multiple videos and detecting the target object contained in each of the multiple videos, acquisition means for acquiring, among results obtained by analyzing each of the plurality of videos using each of a plurality of types of said engines, results obtained by analyzing each of the plurality of videos using said selected type of said engine; and integrating means for integrating the acquired results of analysis of said plurality of videos 、 The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. The aforementioned integration means is A grouping unit that groups the detection targets included in the plurality of videos based on the similarity of the appearance features of the detection targets, and generates integrated information that associates the detection targets with the group to which the detection targets belong and the shooting identification information, Upon receiving the designation of the group, the statistical processing unit includes a function that counts the number of detection targets belonging to the designated group included in the results of the analysis of the multiple images for each shooting identification piece, and calculates the number of appearances for each shooting device. there is provided a video analysis apparatus.

[0011] According to one aspect of the present invention, a computer: accepts a selection of a type of engine that analyzes each of a plurality of videos to detect detection targets included in each of the plurality of videos, acquires, among results obtained by analyzing each of the plurality of videos using each of a plurality of types of said engines, results obtained by analyzing each of the plurality of videos using said selected type of said engine, and integrates the acquired results of analysis of said plurality of videos This includes, The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. Integrating the above results means Based on the similarity of the appearance features of the detection targets, the detection targets included in the plurality of videos are grouped, and integrated information is generated that associates the detection targets with the group to which the detection targets belong and the shooting identification information. Upon receiving the designation of the group, the process includes counting the number of detection targets belonging to the designated group included in the results of the analysis of the multiple videos for each shooting identification piece, and calculating the number of appearances for each shooting device. there is provided a video analysis method.

[0012] According to one aspect of the present invention, a computer is caused to: accept a selection of a type of engine that analyzes each of a plurality of videos to detect detection targets included in each of the plurality of videos, acquire, among results obtained by analyzing each of the plurality of videos using each of a plurality of types of said engines, results obtained by analyzing each of the plurality of videos using said selected type of said engine, execute integration of the acquired results of analysis of said plurality of videos 、 The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. Integrating the above results means Based on the similarity of the appearance features of the detection targets, the detection targets included in the plurality of videos are grouped, and integrated information is generated that associates the detection targets with the group to which the detection targets belong and the shooting identification information. Upon receiving the designation of the group, the number of detection targets belonging to the designated group included in the results of the analysis of the multiple videos is counted for each of the shooting identification pieces, and the number of appearances for each shooting device is calculated. A program is provided. Effects of the Invention

[0013] According to one aspect of the present invention, it becomes possible to utilize results obtained by analyzing a plurality of videos. Brief Description of the Drawings

[0014] [Figure 1] It is a diagram showing an outline of a video analysis device according to an embodiment. [Figure 2] It is a diagram showing an outline of a video analysis system according to an embodiment. [Figure 3] It is a flowchart showing an example of video analysis processing according to an embodiment. [Figure 4] It is a diagram showing a detailed example of the configuration of a video analysis system according to an embodiment. [Figure 5] It is a diagram showing a configuration example of video information according to an embodiment. [Figure 6] It is a diagram showing a configuration example of analysis information according to an embodiment. [Figure 7] It is a diagram showing a detailed example of the functional configuration of a video analysis device according to an embodiment. [Figure 8] It is a diagram showing a configuration example of integrated information according to an embodiment. [Figure 9] It is a diagram showing a physical configuration example of a video analysis device according to an embodiment. [Figure 10] It is a flowchart showing an example of analysis processing according to an embodiment. [Figure 11] An example of a start screen according to an embodiment is shown. [Figure 12] It is a flowchart showing a detailed example of integration processing according to an embodiment. [Figure 13] It is a diagram showing an example of an integration result screen according to an embodiment. [Figure 14]This figure shows an example of a screen displaying the number of appearances according to one embodiment. [Modes for carrying out the invention]

[0015] Hereinafter, one embodiment of the present invention will be described with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted where appropriate.

[0016] <Embodiment> Figure 1 shows an overview of a video analysis device 100 according to one embodiment. The video analysis device 100 includes a type receiving unit 110, an acquisition unit 111, and an integration unit 112.

[0017] The type reception unit 110 accepts the selection of the type of engine to analyze each of the multiple videos and detect the detection target contained in each of the multiple videos. The acquisition unit 111 acquires the results of analyzing each of the multiple videos using the selected type of engine from among the results of analyzing each of the multiple videos using each of the multiple types of engines. The integration unit 112 integrates the results of analyzing the acquired multiple videos.

[0018] This video analysis device 100 makes it possible to utilize the results of analyzing multiple videos.

[0019] Figure 2 shows an overview of a video analysis system 120 according to one embodiment. The video analysis system 120 comprises a video analysis device 100, a plurality of shooting devices 121_1 to 121_K, and an analysis device 122. Here, K is an integer of 2 or more, and the same applies hereafter.

[0020] Multiple imaging devices 121_1 to 121_K are devices for capturing multiple images. The analysis device 122 analyzes each of the multiple images using multiple types of engines.

[0021] This video analysis system 120 makes it possible to utilize the results of analyzing multiple videos.

[0022] Figure 3 is a flowchart showing an example of video analysis processing according to one embodiment.

[0023] The type receiving unit 110 accepts the selection of the type of engine for analyzing each of the multiple video streams and detecting the detection target contained in each of the multiple video streams (step S101).

[0024] The acquisition unit 111 acquires the results of analyzing each of the multiple videos using the selected type of engine from among the results of analyzing each of the multiple videos using each of the multiple types of engines (step S102).

[0025] The integration unit 112 integrates the results of analyzing the multiple acquired video images (step S103).

[0026] This video analysis process makes it possible to utilize the results of analyzing multiple videos.

[0027] The following describes a detailed example of a video analysis system 120 according to one embodiment.

[0028] Figure 4 shows a detailed example of the configuration of the video analysis system 120 according to this embodiment.

[0029] The video analysis system 120 comprises a video analysis device 100, K-unit shooting devices 121_1 to 121_K, and an analysis device 122.

[0030] The video analysis device 100, each of the shooting devices 121_1 to 121_K, and the analysis device 122 are connected to each other via a communication network N, which is configured as wired, wireless, or a combination thereof. The video analysis device 100, each of the shooting devices 121_1 to 121_K, and the analysis device 122 send and receive information to each other via the communication network N.

[0031] (Configuration of imaging devices 121_1~121_K) Each of the imaging devices 121_1 to 121_K is a device for capturing images.

[0032] Each of the imaging devices 121_1 to 121_K is a camera installed to photograph a predetermined imaging area within a predetermined range. The predetermined range may be a building, facility, city, town, or prefecture, or it may be a range that is appropriately determined from among these. The imaging areas of each of the imaging devices 121_1 to 121_K may partially overlap with each other, or they may be different areas from each other.

[0033] The imaging device 121_i captures a predetermined imaging area, for example, at a predetermined frame rate. The imaging device 121_i then generates video information 124a_i, which includes video, by capturing the predetermined imaging area. The video consists of multiple frame images in a time series. Here, i is an integer between 1 and K, and the same applies below. That is, imaging device 121_i means any one of the imaging devices 121_1 to 121_K.

[0034] The imaging device 121_i transmits video information 124a_i, which represents the captured video, to the analysis device 122 via the communication network N. The timing at which the imaging device 121_i transmits the video information 124a_i to the analysis device 122 varies. For example, the imaging device 121_i may transmit the video information 124a_i to the analysis device 122 at any given time, or it may transmit the video information 124a_i to the analysis device 122 all at once at a predetermined time (for example, a predetermined time each day).

[0035] Figure 5 shows an example of the structure of video information 124a_i. Video information 124a_i is information that includes video composed of multiple frame images. In detail, for example as shown in Figure 5, video information 124a_ i is The video ID, recording device ID, recording time, and video (frame image group) are associated.

[0036] The video ID is information (video identification information) for identifying each of multiple videos. The camera ID is information (shooting identification information) for identifying each of the shooting devices 121_1 to 121_K. The shooting time is information indicating the time when the video was shot. The shooting time may include, for example, the start timing and end timing of the shooting. The shooting time may further include the frame shooting timing when each frame image was captured. The start timing, end timing, and frame shooting timing may each consist of, for example, a date and a time.

[0037] In video information 124a_i, a video ID is associated with the video identified using that video ID. Furthermore, in video information 124a_i, the video ID is associated with the recording device ID of the recording device 121_i that captured the video identified using that video ID, and the recording time (start timing, end timing) indicating the time when the video identified using that video ID was captured. In addition, video information 124a _i Then, a video ID is associated with each frame image that makes up the video identified using that video ID, along with its shooting time (frame shooting timing).

[0038] (Functions of the analysis device 122) The analysis device 122 analyzes multiple video images captured by each of the imaging devices 121_1 to 121_K by analyzing each frame image captured by each of the imaging devices 121_1 to 121_K. As shown in Figure 4, the analysis device 122 comprises an analysis unit 123 and an analysis storage unit 124.

[0039] The analysis unit 123 acquires video information 124a_1 to 124a_K from each of the imaging devices 121_1 to 121_K and stores the acquired video information 124a_1 to 124a_K in the analysis storage unit 124. The analysis unit 123 analyzes the multiple videos contained in each of the acquired video information 124a_1 to 124a_K. Specifically, for example, the analysis unit 123 analyzes the multiple frame images contained in each of the multiple video information 124a_1 to 124a_K.

[0040] The analysis unit 123 generates analysis information 124b showing the results of analyzing multiple video images and stores it in the analysis storage unit 124. The analysis unit 123 also transmits the multiple video information 124a_1 to 124a_K and the analysis information 124b to the video analysis device 100 via the communication network N.

[0041] The analysis unit 123 has the function of analyzing images using multiple types of engines. Each type of engine has the function of analyzing an image and detecting the object to be detected contained in the image. That is, the analysis unit 123 according to this embodiment analyzes the frame images (i.e., video) contained in each of the video information 124a_1 to 124a_K using multiple types of engines and generates analysis information 124b.

[0042] In this embodiment, the object to be detected is a person. However, the object to be detected may be a predetermined object such as a car or a bag.

[0043] Examples of engine types include (1) object detection engine, (2) face analysis engine, (3) human figure analysis engine, (4) posture analysis engine, (5) behavior analysis engine, (6) appearance attribute analysis engine, (7) gradient feature analysis engine, (8) color feature analysis engine, and (9) movement path analysis engine. The analysis device 122 may be equipped with at least two engines from the types of engines exemplified here and other types of engines.

[0044] (1) The object detection engine detects people and objects from an image. The object detection function can also determine the location of people and objects within an image. One example of a model applied to object detection processing is YOLO (You Only Look Once).

[0045] (2) The face analysis engine detects human faces from images, extracts facial features from the detected faces, and classifies the detected faces. The face analysis engine can also determine the position of a face within an image. The face analysis engine can also determine the identity of people detected from different images based on the similarity between facial features of people detected from different images.

[0046] (3) The human figure analysis engine extracts human physical characteristics of people contained in an image (for example, values ​​indicating overall characteristics such as body shape, height, and clothing), and classifies (categorizes) the people contained in the image. The human figure analysis engine can also identify the position of a person within an image. The human figure analysis engine can also determine the identity of people contained in different images based on the human physical characteristics of people contained in those different images.

[0047] (4) The posture analysis engine generates posture information that indicates a person's posture. This posture information includes, for example, a person's posture estimation model. The posture estimation model is a model that connects the joints of a person estimated from an image. The posture estimation model consists of multiple model elements, such as joint elements corresponding to joints, trunk elements corresponding to the torso, and bone elements corresponding to the bones connecting the joints. The posture analysis function, for example, detects the joint points of a person from an image and connects the joint points to create a posture estimation model.

[0048] The posture analysis engine then uses information from the posture estimation model to estimate a person's posture, extracts the estimated posture features, and classifies the people in the image. The posture analysis engine can also determine the identity of people in different images based on the posture features of people in those different images.

[0049] For example, the techniques disclosed in Patent Document 2 and Non-Patent Document 1 can be applied to the attitude analysis engine.

[0050] (5) The behavioral analysis engine can estimate human movement using information from a posture estimation model and changes in posture, extract movement features, and classify people in an image. The behavioral analysis engine can also estimate a person's height using information from a stick figure model and identify a person's position in an image. For example, the behavioral analysis engine can estimate actions such as changes or transitions in posture and movement (changes or transitions in position) from an image and extract movement features related to those actions.

[0051] (6) The appearance attribute analysis engine can recognize appearance attributes associated with a person. The appearance attribute analysis engine extracts features related to the recognized appearance attributes (appearance attribute features), classifies people included in the image, and so on. Appearance attributes are attributes of appearance, and include one or more, such as the color of clothing, the color of shoes, hairstyle, whether a hat or tie is worn, or whether glasses are worn or not.

[0052] (7) The gradient feature analysis engine extracts gradient features (gradient features) from the image. Techniques such as SIFT, SURF, RIFF, ORB, BRISK, CARD, and HOG can be applied to the gradient feature analysis engine.

[0053] (8) The color feature analysis engine can detect objects from an image, extract color features from the detected objects, and classify the detected objects. Color features include, for example, a color histogram. The color feature analysis engine can, for example, detect people and objects contained in an image.

[0054] (9) The movement analysis engine can determine the movement paths (trajectories of movement) of people included in the video by using, for example, the results of identity determination performed by one or more of the engines described above. Specifically, for example, the movement paths of people can be determined by connecting people who have been determined to be the same across images that are different in time series. Alternatively, for example, the movement analysis engine can determine movement features that indicate the direction and speed of movement of a person. The movement features may be either the direction or speed of movement of a person.

[0055] The movement path analysis engine can also determine movement paths that span across multiple images captured in different shooting areas, such as when images are acquired from multiple shooting devices 121_2 to 121_K that capture different shooting areas.

[0056] Furthermore, each of the engines (1) through (9) can calculate the confidence level of the feature it requires.

[0057] Furthermore, each of the engines (1) to (9) may appropriately utilize the results of analyses performed by other engines. The video analysis device 100 may also include an analysis unit that has the functions of the analysis device 122.

[0058] The analysis memory unit 124 is a memory unit for storing various types of information, such as video information 124a_1 to 124a_K and analysis information 124b.

[0059] Figure 6 shows an example of the structure of analysis information 124b. Analysis information 124b associates the video ID, shooting device ID, shooting time, and analysis results.

[0060] The video ID, recording device ID, and recording time associated in analysis information 124b are the same as those associated in video information 124a_i.

[0061] The analysis results are information that shows the results of analyzing the video identified using the associated video ID. In analysis information 124b, the analysis results are associated with a video ID used to identify the video that was analyzed in order to obtain the analysis results.

[0062] The analysis results, for example, correlate the detected target ID, engine type, appearance features, and confidence level.

[0063] The detection target ID is information for identifying the detection target (detection target identification information). In this embodiment, as described above, the detection target is a person. Therefore, the detection target ID is information for identifying a person detected by the analysis device 122 when it analyzes each of the multiple frame images. In this embodiment, the detection target ID is information for identifying each of the images showing a person (person image) detected from each of the multiple frame images, regardless of whether the detection targets are the same person or not.

[0064] Furthermore, the detection target ID may be information used to identify each person represented by the person image detected in each of multiple frame images. In this case, the detection target ID will be the same if the detection targets are the same person, and different if the detection targets are different people.

[0065] In analysis information 124b, the detection target ID is information used to identify the detection target contained in the video, which is identified using the associated video ID.

[0066] The engine type indicates the type of engine used to analyze the video.

[0067] Appearance features represent features related to the appearance of the object being detected. Examples of appearance features include the detection results of objects in object detection functions, facial features, human body features, posture features, motion features, appearance attribute features, gradient features, color features, and movement features.

[0068] In the analysis results of analysis information 124b, the appearance features represent the features obtained using the associated type of engine for the detection target indicated by the associated detection target ID.

[0069] The confidence level indicates the confidence level of the appearance feature. In the analysis results of analysis information 124b, the confidence level indicates the confidence level of the appearance feature associated with it.

[0070] For example, when the analysis device 122 obtains appearance features using each of the engines (1) to (9) described above, the analysis results associate a common detection target ID with an engine type indicating the type of engine (1) to (9). Then, for each engine type, the analysis results associate the appearance features obtained using the engine of the type indicated by that engine type with the appearance features.

[0071] (Functions of the video analysis device 100) Figure 7 shows a detailed example of the functional configuration of the video analysis device 100 according to this embodiment. The video analysis device 100 includes a storage unit 108, a receiving unit 109, a type receiving unit 110, an acquisition unit 111, an integration unit 112, a display control unit 113, and a display unit 114. The video analysis device 100 may also include an analysis unit 123, in which case the video analysis system 120 does not need to include an analysis device 122.

[0072] The memory unit 108 is a memory unit for storing various types of information.

[0073] The receiving unit 109 receives various types of information, such as video information 124a_1 to 124a_K and analysis information 124b, from the analysis device 122 via the communication network N. The receiving unit 109 may receive the video information 124a_1 to 124a_K and analysis information 124b from the analysis device 122 in real time, or it may receive them as needed, such as when used for processing by the video analysis device 100.

[0074] The receiving unit 109 stores the received information in the storage unit 108. In other words, in this embodiment, the information stored in the storage unit 108 includes video information 124a_1 to 124a_K and analysis information 124b.

[0075] The receiving unit 109 may receive video information 124a_1 to 124a_K from each of the imaging devices 121_1 to 121_K via the communication network N and store the received information in the storage unit 108. Furthermore, the receiving unit 109 may receive video information 124a_1 to 124a_K and analysis information 124b from the analysis device 122 via the communication network N as needed, such as when used for processing by the video analysis device 100. In this case, the video information 124a_1 to 124a_K and analysis information 124b do not need to be stored in the storage unit 108. Moreover, for example, if the receiving unit 109 receives all of the video information 124a_1 to 124a_K and analysis information 124b from the analysis device 122 and stores them in the storage unit 108, the analysis device 122 does not need to retain the video information 124a_1 to 124a_K and analysis information 124b.

[0076] The type reception unit 110 receives, for example, a selection from the user regarding the type of engine used by the analysis device 122 to analyze the video. The type reception unit 110 is not limited to accepting only one type of engine, but may accept multiple types.

[0077] In detail, for example, the type receiving unit 110 receives information indicating one of the following types: (1) object detection engine, (2) face analysis engine, (3) human figure analysis engine, (4) posture analysis engine, (5) behavior analysis engine, (6) appearance attribute analysis engine, (7) gradient feature analysis engine, (8) color feature analysis engine, and (9) movement path analysis engine.

[0078] The selection of the engine type may also be performed by selecting the result of analyzing each of the multiple video clips. In this case, for example, the type receiving unit 110 may receive the selection of the result of analyzing each of the multiple video clips and identify the type of engine used to obtain the selected result.

[0079] The acquisition unit 111 acquires analysis information 124b from the storage unit 108, which indicates the results of analyzing each of the multiple videos using a selected type of engine, i.e., the type of engine accepted by the type acceptance unit 110, from among the results of analyzing each of the multiple videos using each of the multiple types of engines. The acquisition unit 111 may also receive the analysis information 124b from the analysis device 122 via the communication network N.

[0080] The results of analyzing each of the multiple videos are included in the analysis information 124b. Therefore, the results of analyzing each of the multiple videos include, for example, the appearance features of the object to be detected contained in each of the multiple videos. Also, for example, the results of analyzing each of the multiple videos include the shooting device ID (shooting identification information) for identifying the shooting devices 121_1 to 121_K that shot the video containing the object to be detected. Furthermore, for example, the results of analyzing each of the multiple videos include the shooting time when the video containing the object to be detected was shot. This shooting time may include at least one of the start and end timings of the video containing the object to be detected and the frame shooting timing of the frame image containing the object to be detected.

[0081] Here, the multiple videos that are the subject of analysis for generating the analysis information 124b acquired by the acquisition unit 111 are videos that are related in terms of location and time. That is, in this embodiment, these multiple videos are videos obtained by taking pictures of multiple locations within a predetermined range at different times within a predetermined period (for example, one day, one week, one month).

[0082] Furthermore, the multiple videos contained in each of the multiple video information 124a_1 to 124a_K are not limited to videos that are geographically and temporally related, but may be videos that are geographically or temporally related. In other words, the videos that are the subject of analysis for generating the analysis information 124b acquired by the acquisition unit 111 may be videos obtained by taking pictures of the same location at different times within a predetermined period, or they may be videos obtained by taking pictures of multiple locations within a predetermined range at the same time.

[0083] The integration unit 112 integrates the analysis results acquired by the acquisition unit 111. That is, the integration unit 112 integrates the results of analyzing each of the multiple videos using the selected type of engine, i.e., the type of engine accepted by the type acceptance unit 110. More specifically, for example, the integration unit 112 integrates the results of analyzing each of the multiple videos using the same type of engine.

[0084] In addition, multiple types of engines may be selected. In this case, the integration unit 112 may integrate the results of analyzing each of the multiple videos using each of the selected engines, for each selected engine type. That is, when multiple types of engines are selected, the integration unit 112 may integrate the results of analyzing each of the multiple videos using the same type of engine for each selected engine type.

[0085] In this embodiment, the integration unit 112 integrates the analysis results by grouping the detection targets based on the visual characteristics of the detection targets detected by the analysis.

[0086] In detail, for example, the integration unit 112 includes a grouping unit 112a and a statistical processing unit 112b, as shown in Figure 7.

[0087] The grouping unit 112a groups the detection targets included in multiple videos based on the similarity of the appearance features of the detection targets, and generates integrated information 108a that associates the detection targets with the groups to which they belong. The grouping unit 112a stores the generated integrated information 108a in the storage unit 108.

[0088] More specifically, the grouping unit 112a accepts the designation of the videos to be integrated, for example, based on user input or pre-set default values. The grouping unit 112a then groups the detection targets detected using the designated videos based on the similarity of the appearance features of the detection targets.

[0089] The videos to be integrated are specified, for example, using a combination of the recording devices 121_1 to 121_K that captured the videos to be integrated, and the recording period, which is the period during which the videos were recorded. The recording period is specified, for example, using a combination of time zones and dates. The grouping unit 112a identifies the videos captured by the specified recording devices 121_1 to 121_K during the specified recording period, and groups the detection targets included in the identified videos.

[0090] The grouping unit 112a may also group together detection targets included in all images captured by all imaging devices 121_1 to 121_K. Alternatively, the grouping unit 112a may also group together detection targets included in multiple images captured by all imaging devices 121_1 to 121_K during a specified time period.

[0091] The grouping unit 112a acquires grouping conditions for grouping detection targets based on, for example, user input or pre-set default values. The grouping unit 112a retains the grouping conditions. Based on the grouping conditions, the grouping unit 112a groups the detection targets included in multiple videos.

[0092] The grouping criteria include at least one of the following: a first threshold for the confidence level of the appearance features, a second threshold for the similarity level of the appearance features, and the number of groups. Note that the grouping criteria only need to include at least one of the first threshold, the second threshold, and the number of groups.

[0093] The grouping unit 112a may, for example, extract detection targets associated with appearance features whose confidence level is equal to or greater than a first threshold, based on grouping conditions. Then, the grouping unit 112a may group the extracted detection targets based on appearance features.

[0094] For example, the grouping unit 112a may group the detection targets such that the similarity of the appearance features is equal to or greater than the second threshold, and the detection targets with a similarity of appearance features less than the second threshold are placed in different groups, based on the grouping conditions.

[0095] Furthermore, for example, the grouping unit 112a may group the detection targets such that the number of groups into which the detection targets are divided equals the number of groups included in the grouping condition.

[0096] The grouping unit 112a may use common grouping conditions for grouping the detection targets regardless of the user of the video analysis device 100, or it may use grouping conditions specified by the user from among multiple grouping conditions for grouping the detection targets.

[0097] Furthermore, the grouping unit 112a may store grouping conditions in association with user identification information for identifying users. In this case, the grouping unit 112a may use grouping conditions associated with, for example, user identification information for identifying logged-in users, or user identification information entered by users, to group the detection targets. This allows the grouping unit 112a to group the detection targets included in multiple videos based on grouping conditions defined for each user.

[0098] Figure 8 shows an example of the configuration of integrated information 108a. Integrated information 108a, for example, associates the integration target and group information.

[0099] The data to be integrated is information used to identify multiple video footages that are to be integrated. In the example shown in Figure 8, the data to be integrated associates the camera ID, shooting period, shooting time, and engine type.

[0100] The recording device ID and recording period are the recording devices 121_1 to 121_K and the recording period, respectively, designated to identify the video subject to the specified integration. The recording time is the recording time of the video captured within the recording period. The recording device ID and recording time included in the integration target can be associated with the recording device ID and recording time contained in the video information 124a_i to identify the video ID and video.

[0101] The engine type indicates the type of engine selected. In other words, the engine type indicates the type of engine used to determine the feature quantities to be detected from the multiple screens that are subject to integration (analysis of those multiple screens).

[0102] Group information is information that shows the results of grouping, and associates the group ID and the detection target ID. The group ID is information used to identify a group (group identification information). In the group information, the group ID is associated with the detection target ID of the detection target belonging to the group identified using that group ID.

[0103] The statistical processing unit 112b uses the integrated information 108a to count the number of detection targets included in multiple videos and determine the number of appearances of the detection targets. More specifically, for example, the statistical processing unit 112b uses the integrated information 108a to count the number of detection targets belonging to a user-specified group included in multiple videos taken by, for example, designated shooting devices 121_1 to 121_K during a specified shooting period, and determines the number of appearances of the detection targets belonging to that group.

[0104] The number of appearances includes at least one of the following: total number of appearances, number of appearances by time slot, etc.

[0105] The total number of appearances is calculated by counting the number of times a detection target belonging to a user-specified group is included in all of the multiple videos filmed during the shooting period.

[0106] The number of appearances per time period is obtained by counting the number of times a detection target belonging to a user-specified group is included in multiple videos taken during each time period into which the shooting period is divided. These time periods may be determined according to predetermined time lengths, such as every hour, or they may be specified by the user.

[0107] The display control unit 113 displays various types of information on the display unit 114. For example, the display control unit 113 displays the results of the integration by the integration unit 112 on the display unit 114. The integrated results include, for example, the detection targets for each group, the shooting device IDs of the shooting devices 121_1 to 121_K that captured the video in which the detection targets were detected, the shooting time of the video in which the detection targets were detected, and the number of times the targets appeared.

[0108] For example, when a user specifies a time period, the display control unit 113 will display one or more videos taken during that specified time period on the display unit 114.

[0109] (Physical configuration of the video analysis device 100) Figure 9 shows an example of the physical configuration of the video analysis device 100 according to this embodiment. The video analysis device 100 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, and a user interface 1060.

[0110] Bus 1010 is a data transmission path for the processor 1020, memory 1030, storage device 1040, network interface 1050, and user interface 1060 to send and receive data to and from each other. However, the method of connecting the processor 1020 and other components to each other is not limited to bus connection.

[0111] The 1020 processor is a processor implemented in components such as the CPU (Central Processing Unit) and GPU (Graphics Processing Unit).

[0112] Memory 1030 is a main memory device implemented using RAM (Random Access Memory), etc.

[0113] The storage device 1040 is an auxiliary storage device implemented as an HDD (Hard Disk Drive), SSD (Solid State Drive), memory card, or ROM (Read Only Memory). The storage device 1040 stores program modules for realizing the functions of the video analysis device 100. The processor 1020 loads each of these program modules into memory 1030 and executes them, thereby realizing the functions corresponding to those program modules.

[0114] The network interface 1050 is an interface for connecting the video analysis device 100 to the communication network N.

[0115] The user interface 1060 includes touch panels, keyboards, mice, etc., as interfaces for the user to input information, and liquid crystal panels, organic EL (Electro-Luminescence) panels, etc., as interfaces for presenting information to the user.

[0116] The analysis device 122 should be physically configured similarly to the video analysis device 100 (see Figure 9). Therefore, a diagram showing the physical configuration of the analysis device 122 is omitted.

[0117] (Operation of video analysis system 120) From here, the operation of the video analysis system 120 will be explained with reference to the diagram.

[0118] (Analysis process) Figure 10 is a flowchart showing an example of the analysis process according to this embodiment. The analysis process is for analyzing the images captured by the imaging devices 121_1 to 121_K. The analysis process is repeatedly executed, for example, while the imaging devices 121_1 to 121_K and the analysis unit 123 are in operation.

[0119] The analysis unit 123 acquires video information 124a_1 to 124a_K from each of the imaging devices 121_1 to 121_K, for example in real time, via the communication network N (step S201).

[0120] The analysis unit 123 stores the multiple video information pieces 124a_1 to 124a_K acquired in step S201 in the analysis storage unit 124, and analyzes the video included in the multiple video information pieces 124a_1 to 124a_K (step S202).

[0121] For example, as described above, the analysis unit 123 analyzes the frame images contained in each video using multiple types of engines to detect the target object. The analysis unit 123 also uses each type of engine to determine the appearance features of the detected target object and the confidence level of those appearance features. By performing such analysis, the analysis unit 123 generates analysis information 124b.

[0122] The analysis unit 123 stores the analysis information 124b generated by the analysis performed in step S202 in the analysis storage unit 124 and transmits it to the video analysis device 100 via the communication network N (step S203). At this time, the analysis unit 123 may also transmit the video information 124a_1 to 124a_K acquired in step S201 to the video analysis device 100 via the communication network N.

[0123] The receiving unit 109 receives the analysis information 124b transmitted in step S203 via the communication network N (step S204). At this time, the receiving unit 109 may also receive the video information 124a_1 to 124a_K transmitted in step S203 via the communication network N.

[0124] The receiving unit 109 stores the analysis information 124b received in step S204 in the storage unit 108 (step S205), and terminates the analysis process. At this time, the receiving unit 109 may also receive the video information 124a_1 to 124a_K received in step S204 via the communication network N.

[0125] (Video analysis processing) As explained with reference to Figure 3, the video analysis process is a process for integrating the results of video analysis. The video analysis process starts, for example, when a user logs in, and the display control unit 113 displays the start screen 131 on the display unit 114. The start screen 131 is a screen for receiving user specifications.

[0126] Figure 11 shows an example of the start screen 131 according to this embodiment. The start screen 131 shown in Figure 11 includes input fields for specifying or selecting the imaging device and imaging period associated with the integration target, the type of engine, and the first threshold, second threshold, and number of groups associated with the grouping conditions.

[0127] Figure 11 shows an example where "all" of imaging devices 121_1 to 121_K are entered in the input field associated with "imaging device". In this input field, for example, the imaging device IDs of one or more imaging devices 121_1 to 121_K may be entered.

[0128] Figure 11 shows an example where "April 1, 2022 0:00 - April 2, 2022 0:00" is entered in the input field associated with "Shooting Period". Any appropriate period should be entered in this input field.

[0129] Figure 11 shows an example where "Appearance Attribute Analysis Engine" is entered in the input field associated with "Engine Type". This input field should contain the type of engine used to calculate appearance features. Furthermore, multiple types of engines used to calculate appearance features may be entered in this input field.

[0130] Figure 11 shows an example where "0.35", "0.25", and "3" are entered into the input fields corresponding to "First Threshold", "Second Threshold", and "Number of Groups", respectively. These input fields are initially set with grouping conditions associated with the user identification information of the logged-in user, for example, and these initial values ​​may be changed by the user as needed.

[0131] When the user presses the integration start button 131a, the video analysis device 100 starts the video analysis process shown in Figure 3.

[0132] As described above with reference to Figure 3, the type receiving unit 110 accepts the selection of the type of engine for analyzing the video and detecting the object to be detected contained in the video (step S101).

[0133] At this time, the type reception unit 110 receives information specified on the start screen 131 in addition to the engine type. As explained with reference to Figure 11, this information is, for example, information for specifying the imaging device, imaging period, first threshold, second threshold, and number of groups.

[0134] As described above, the acquisition unit 111 acquires the results of analyzing each of the multiple videos using the type of engine selected in step S101 (step S102).

[0135] In detail, for example, the acquisition unit 111 acquires analysis information 124b regarding multiple images to be integrated from the storage unit 108 based on the engine type indicating the type of engine selected, the specified imaging device ID, and the shooting period. Here, the acquisition unit 111 acquires analysis information 124b from the storage unit 108 that includes the engine type indicating the type of engine selected, the specified imaging device ID, and the shooting time within the specified shooting period.

[0136] The integration unit 112 integrates the results obtained in step S102 (step S103). In other words, it integrates the analysis information 124b obtained in step S102.

[0137] Figure 12 is a flowchart showing a detailed example of the integrated processing (step S103) according to this embodiment.

[0138] The grouping unit 112a groups the detection targets included in multiple videos based on the similarity of the appearance features included in the analysis information 124b acquired in step S102 (step S103a). As a result, the grouping unit 112a generates integrated information 108a and stores it in the storage unit 108.

[0139] The display control unit 113 displays the grouped results from step S103a on the display unit 114 (step S103b).

[0140] Figure 13 shows an example of the integrated results screen 132, which displays the results of grouping. The integrated results screen 132 displays a list of the camera IDs of the camera devices 121_1 to 121_K that captured the video in which the detected target belonging to the group was detected, for each group.

[0141] In the example shown in Figure 13, Group 1, Group 2, and Group 3 represent the group IDs of three groups corresponding to the specified number of groups. In the example shown in Figure 13, the imaging device IDs "Imaging Device 1" and "Imaging Device 2," corresponding to imaging devices 121_1 to 121_2, are associated with Group 1. The imaging device IDs "Imaging Device 2" and "Imaging Device 3," corresponding to imaging devices 121_2 to 121_3, are associated with Group 2. The imaging device ID "Imaging Device 4," corresponding to imaging device 121_4, is associated with Group 3.

[0142] The integrated results screen 132 is not limited to this; for example, it may display a list of video IDs of videos in which the detected target belonging to the group was detected, for each group.

[0143] The statistical processing unit 112b accepts the designation of a group (step S103c).

[0144] For example, each of the "Group 1," "Group 2," and "Group 3" on the integrated results screen 132 illustrated in Figure 13 is selectable. When the user selects one of "Group 1," "Group 2," or "Group 3," the statistical processing unit 112b accepts the group specification.

[0145] The statistical processing unit 112b counts the number of detection targets that belong to the group specified in step S103c and determines the number of occurrences of the detection targets belonging to that group (step S103d).

[0146] In detail, for example, the statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to the group specified in step S103c that are included in the analysis information 124b obtained in step S102. This makes it possible to count the number of detection targets belonging to the user-specified group included in multiple images captured by the specified imaging devices 121_1 to 121_K during the specified imaging period.

[0147] The statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to the specified group that are included in the total analysis information 124b obtained in step S102, and calculates the total number of occurrences.

[0148] The statistical processing unit 112b divides the analysis information 124b acquired in step S102 into time zones based on the shooting time included in the analysis information 124b. The statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to a specified group that are included in the analysis information 124b for each time zone, and determines the number of appearances for each time zone.

[0149] The statistical processing unit 112b may count the number of detection targets (detection target IDs) belonging to a specified group that are included in the total analysis information 124b for each imaging device ID, and determine the total number of appearances for each imaging device. Alternatively, the statistical processing unit 112b may count the number of detection targets (detection target IDs) belonging to a specified group that are included in the analysis information 124b for each time period, for each imaging device ID, and determine the number of appearances for each time period and each imaging device.

[0150] The display control unit 113 displays the number of occurrences obtained in step S103d on the display unit 114 (step S103e), and then terminates the video analysis process (see Figure 3).

[0151] Figure 14 shows an example of an appearance count display screen 133, which shows the number of appearances. The appearance count display screen 133 illustrated in Figure 14 is an example of a screen that shows the number of appearances for Group 1 by time period and by shooting device in a line graph.

[0152] For example, it may be possible to select a time to indicate each time period, and when a time period is specified by such selection, the display control unit 113 may display one or more videos taken during the specified time period on the display unit 114. More specifically, for example, the display control unit 113 may identify a video ID corresponding to a video containing a group of frame images taken during the specified time period, based on the shooting time included in the analysis information 124b acquired in step S102. The display control unit 113 may then display the image associated with the identified video ID on the display unit 114, based on the video information 124a_1 to 124a_K.

[0153] Furthermore, the screen displaying the number of appearances 133 may not be limited to line graphs; pie charts, bar graphs, or other methods may be used to show the number of appearances.

[0154] By performing video analysis processing, detection targets contained in multiple videos can be grouped based on appearance features obtained using a selected type of engine. This allows for the grouping of detection targets with similar appearance features.

[0155] Furthermore, by referring to the integrated results screen 132, users can view the grouped results. Additionally, by referring to the occurrence count display screen 133, users can view the occurrence count of the detection targets classified based on their appearance features. This allows users to understand the trends in the appearance of detection targets with similar appearance features, such as when, where, and to what extent similar detection targets with similar appearance features are present.

[0156] (Effects / Actions) As described above, according to this embodiment, the video analysis device 100 comprises a type receiving unit 110, an acquisition unit 111, and an integration unit 112. The type receiving unit 110 accepts the selection of the type of engine for analyzing each of a plurality of videos and detecting the detection target contained in each of the plurality of videos. The acquisition unit 111 acquires the results of analyzing each of the plurality of videos using the selected type of engine from among the results of analyzing each of the plurality of videos using each of the plurality of engines. The integration unit 112 integrates the results of analyzing the acquired plurality of videos.

[0157] This allows for the acquisition of integrated information by analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0158] According to this embodiment, the type of engine is selected by selecting the result of analyzing each of the multiple images.

[0159] This allows for the acquisition of integrated information by analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0160] According to this embodiment, the integration unit 112 integrates the results of analyzing each of the multiple images using the same type of engine.

[0161] This allows for the acquisition of integrated information from the analysis of multiple videos using the same type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0162] According to this embodiment, the results of analyzing multiple videos include the appearance features of the detection targets contained in each of the multiple videos. The integration unit 112 groups the detection targets contained in the multiple videos based on the similarity of the appearance features of the detection targets and generates integrated information 108a that associates the detection targets with the groups to which they belong.

[0163] This allows for the acquisition of integrated information 108a by integrating the results of analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0164] According to this embodiment, the integration unit 112 groups the detection targets included in multiple videos based on grouping conditions for grouping the detection targets.

[0165] This allows for grouping of detection targets using grouping conditions. Consequently, it becomes possible to utilize the results of analyzing multiple videos.

[0166] According to this embodiment, the grouping condition includes at least one of a first threshold for the confidence level of the appearance features, a second threshold for the similarity level of the appearance features, and the number of groups.

[0167] This allows for grouping the detection targets based on at least one of the following conditions: a first threshold, a second threshold, and the number of groups. Consequently, it becomes possible to utilize the results of analyzing multiple videos.

[0168] According to this embodiment, the integration unit 112 groups the detection targets included in multiple videos based on grouping conditions defined for each user.

[0169] This allows for grouping detection targets using grouping criteria tailored to the user. Consequently, it becomes possible to utilize the results of analyzing multiple videos.

[0170] According to this embodiment, the results of analyzing multiple images further include shooting identification information for identifying the shooting devices 121_1 to 121_K that captured images containing the detection target. The integrated information 108a further associates the shooting identification information.

[0171] This allows for the analysis of integrated information 108a for each imaging device. Consequently, it becomes possible to utilize the results of analyzing multiple video images.

[0172] According to this embodiment, the integration unit 112 further counts the number of detection targets included in the multiple videos and determines the number of times the detection targets appear.

[0173] This allows the detection count to be obtained by integrating the results of analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0174] According to this embodiment, the results of analyzing multiple videos further include the shooting time when the video containing the detection target was taken. The integration unit 112 further counts the number of detection targets included in the multiple videos for each time period in which the video was taken, and determines the number of times the detection target appears for each time period.

[0175] This allows for the acquisition of the number of occurrences of the target object by time period, by integrating the results of analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0176] According to this embodiment, the video analysis device 100 further includes a display control unit 113 that displays the integrated results on a display unit 114.

[0177] This allows the user to view the display unit 114 and see the integrated results of analyzing multiple videos using the selected type of engine. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0178] According to this embodiment, when a time period is specified, the display control unit 113 causes one or more videos captured during the specified time period to be displayed on the display unit 114.

[0179] This allows users to easily view the video footage used to obtain the analysis results, as needed. Therefore, it becomes possible to utilize the results of analyzing multiple videos.

[0180] According to this embodiment, the multiple images are images captured using multiple shooting devices 121_1 to 121_K.

[0181] This makes it possible to utilize the results of analyzing multiple videos taken in different locations.

[0182] According to this embodiment, the multiple images are images that are related in terms of location or time.

[0183] This makes it possible to utilize the results of analyzing multiple videos that are related in terms of location or time.

[0184] According to this embodiment, the multiple images are images obtained by shooting the same shooting area at different times within a predetermined period, or images obtained by shooting multiple shooting areas within a predetermined range at the same time or at different times within a predetermined period.

[0185] This makes it possible to utilize the results of analyzing multiple videos that are related in terms of location or time.

[0186] The embodiments and modifications of the present invention have been described above with reference to the drawings, but these are merely examples of the present invention, and various other configurations can also be adopted.

[0187] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in the embodiment is not limited to the order in which they are described. In the embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above-described embodiments and modifications can be combined to the extent that their content does not conflict.

[0188] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0189] 1. A type receiving means that accepts the selection of the type of engine for analyzing each of multiple video files and detecting the target object contained in each of the multiple video files, An acquisition means for acquiring the results of analyzing each of the multiple images using a selected type of engine from among the results of analyzing each of the multiple images using each of the multiple types of engines, The system includes an integration means for integrating the results of analyzing the multiple acquired video images. Video analysis device. 2. The selection of the engine type is made by selecting the result of analyzing each of the multiple images. The video analysis device described in item 1 above. 3. The integration means integrates the results of analyzing each of the multiple images using the same type of engine. The video analysis device described in item 1 or 2 above. 4. The results of the analysis of the multiple videos include the appearance features of the target to be detected, contained in each of the multiple videos. The integration means groups the detection targets included in the plurality of videos based on the similarity of the appearance features of the detection targets, and generates integrated information that associates the detection targets with the groups to which they belong. A video analysis device as described in any one of the above 1 to 3. 5. The integration means further groups the detection targets included in the plurality of videos based on grouping conditions for grouping the detection targets. The video analysis device described in item 4 above. 6. The grouping condition includes at least one of the following: a first threshold for the confidence level of the appearance features, a second threshold for the similarity level of the appearance features, and the number of groups. The video analysis device described in item 5 above. 7. The integration means groups the detection targets included in the plurality of videos based on the grouping conditions defined for each user. The video analysis device described in item 5 or 6 above. 8. The results of the analysis of the multiple videos further include shooting identification information for identifying the shooting device that took the video including the detected target, The integrated information further relates to the image identification information. A video analysis device as described in any one of items 4 through 7 above. 9. The integration means further counts the number of detection targets included in the plurality of videos and determines the number of times the detection targets appear. A video analysis device as described in any one of items 1 through 8 above. 10. The results of the analysis of the multiple videos further include the time at which the video containing the detection target was taken, The integration means further counts the number of detection targets included in the plurality of videos according to the time period in which each of the plurality of videos was filmed, and determines the number of appearances of the detection targets for each time period. The video analysis device described in item 9 above. 11. The system further comprises a display control means for displaying the integrated result on a display means. A video analysis device as described in any one of items 1 through 10 above. 12. When a time period is specified, the display control means causes one or more of the images captured during the specified time period to be displayed on the display means. The video analysis device described in item 11 above. 13. The aforementioned multiple images are images captured using multiple cameras. A video analysis device as described in any one of the above 1 to 12. 14. The aforementioned images are spatially or temporally related. The video analysis device described in item 13 above. 15. The aforementioned multiple images are images obtained by photographing the same shooting area at different times within a predetermined period, or images obtained by photographing multiple shooting areas within a predetermined range at the same time or at different times within a predetermined period. The video analysis device described in item 13 or 14 above. 16. A video analysis device as described in any one of items 1 to 15 above, Multiple shooting devices for capturing the aforementioned multiple images, The system includes an analysis device that analyzes each of the multiple images using multiple types of engines. Video analysis system. 17. Computers, The system accepts the selection of the type of engine that analyzes each of multiple videos and detects the target object contained in each of those videos. From the results obtained by analyzing each of the multiple images using each of the multiple types of engines, the results obtained by analyzing each of the multiple images using the selected type of engine are acquired. The results of analyzing the multiple acquired video images are integrated. Video analysis methods. 18. To the computer, The system accepts the selection of the type of engine that analyzes each of multiple videos and detects the target object contained in each of those videos. From the results obtained by analyzing each of the multiple images using each of the multiple types of engines, the results obtained by analyzing each of the multiple images using the selected type of engine are acquired. A program for integrating the results of analyzing the multiple acquired video images. 19. To the computer, The system accepts the selection of the type of engine that analyzes each of multiple videos and detects the target object contained in each of those videos. From the results obtained by analyzing each of the multiple images using each of the multiple types of engines, the results obtained by analyzing each of the multiple images using the selected type of engine are acquired. A recording medium containing a program for performing the integration of the results of analyzing the multiple acquired video images. [Explanation of symbols]

[0190] 100 Video Analysis Devices 108 Storage section 108a Integrated Information 109 Receiving Unit 110 types of reception desk 111 Acquisition Department 112 Integration Department 112a Grouping section 112b Statistical Processing Unit 113 Display Control Unit 114 Display section 120 Video Analysis System 121_2~121_K Imaging device 122 Analysis equipment 123 Analysis Department 124 Analysis storage section 124a_1~124a_K Video Information 124b Analysis information 131 Start screen 131a Integration Start Button 132. Integrated Results Screen 133 Screen showing the number of appearances

Claims

1. A type receiving means that accepts the selection of the type of engine for analyzing each of multiple videos and detecting the target object contained in each of the multiple videos, An acquisition means for acquiring the results of analyzing each of the multiple images using a selected type of engine from among the results of analyzing each of the multiple images using each of the multiple types of engines, The system includes an integration means for integrating the results of analyzing the multiple acquired video images, The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. The aforementioned integration means is A grouping unit that groups the detection targets included in the plurality of videos based on the similarity of the appearance features of the detection targets, and generates integrated information that associates the detection targets with the group to which the detection targets belong and the shooting identification information, Upon receiving the designation of the group, the statistical processing unit includes a function that counts the number of detection targets belonging to the designated group included in the results of the analysis of the multiple videos for each shooting identification piece, and calculates the number of appearances for each shooting device. Video analysis device.

2. The selection of the engine type is performed by selecting the result of analyzing each of the multiple images. The video analysis apparatus according to claim 1.

3. The integration means integrates the results of analyzing each of the multiple images using the same type of engine. The video analysis apparatus according to claim 1.

4. The integration means further groups the detection targets included in the plurality of videos based on grouping conditions for grouping the detection targets. The video analysis apparatus according to any one of claims 1 to 3.

5. The results of the analysis of the aforementioned multiple videos further include the time at which the video containing the detection target was taken. The integration means further counts the number of detection targets included in the plurality of videos according to the time period in which each of the plurality of videos was filmed, and determines the number of appearances of the detection targets for each time period. The video analysis apparatus according to any one of claims 1 to 3.

6. Computers The system accepts the selection of the type of engine that analyzes each of multiple videos and detects the target object contained in each of those videos. From the results obtained by analyzing each of the multiple images using each of the multiple types of engines, the results obtained by analyzing each of the multiple images using the selected type of engine are acquired. This includes integrating the results of analyzing the multiple videos that have been acquired, The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. Integrating the above results means Based on the similarity of the appearance features of the detection targets, the detection targets included in the plurality of videos are grouped, and integrated information is generated that associates the detection targets with the group to which the detection targets belong and the shooting identification information. Upon receiving the designation of the group, the process includes counting the number of detection targets belonging to the designated group included in the results of the analysis of the multiple videos for each shooting identification piece, and calculating the number of appearances for each shooting device. Video analysis methods.

7. On the computer, The system accepts the selection of the type of engine that analyzes each of multiple videos and detects the target object contained in each of those videos. From the results obtained by analyzing each of the multiple images using each of the multiple types of engines, the results obtained by analyzing each of the multiple images using the selected type of engine are acquired. The system then performs the process of integrating the results of analyzing the multiple videos that were acquired. The results of analyzing the aforementioned multiple videos include the visual characteristics of the object to be detected contained in each of the multiple videos, and shooting identification information for identifying the shooting device that took the video containing the object to be detected. Integrating the above results means Based on the similarity of the appearance features of the detection targets, the detection targets included in the plurality of videos are grouped, and integrated information is generated that associates the detection targets with the group to which the detection targets belong and the shooting identification information. A program that, upon receiving a designation of the aforementioned group, counts the number of detection targets belonging to the designated group included in the results of the analysis of the multiple images, for each of the shooting identification pieces, and calculates the number of appearances for each shooting device.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2019016098A

  • Dispersion-type target tracking system

    JP2020184292A

  • System and method for improving speed of similarity based searches

    US20200082212A1

  • Data processing device, data processing method, and program

    WO2017077902A1

  • Image processing device, image processing method, and non-transitory computer-readable medium having image processing program stored thereon

    WO2021084677A1