Video analyzer, video analysis method, and program

JP2024007263A5Active Publication Date: 2025-06-17NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022108648
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-06-17
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Existing video analysis systems struggle to effectively utilize the results of analyzing multiple videos for purposes beyond object tracking, limiting the application of these results.

Method used

A video analysis device and method that allows for the selection and integration of analysis results from multiple videos using various engines, such as object detection, face analysis, and posture estimation, to extract and integrate features like appearance attributes and movement patterns across videos.

Benefits of technology

Enables the utilization of comprehensive video analysis results, facilitating the grouping and counting of detection targets based on similarity and reliability, enhancing the understanding and visualization of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To utilize a result of analyzing a plurality of videos.SOLUTION: A video analyzer 100 includes a type acceptance part 110, an acquisition part 111 and an integration part 112. The type acceptance part 110 accepts a selection of a kind of engine for detecting a detection object included in a plurality of videos through analyzing the plurality of videos. The acquisition part 111 acquires results of analyzing the plurality of videos using the selected type of engine, from among results of analyzing the plurality of videos using a plurality of kinds of engines, respectively. The integration part 112 integrates the results of analyzing the plurality of videos thus acquired.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a video analysis device, a video analysis method, and a program. [Background technology]

[0002] Patent Document 1 discloses a distributed object tracking system for linking analysis results of image analysis devices to track an object. This distributed object tracking system includes a plurality of image analysis devices and a cluster management service device.

[0003] Each of the multiple image analysis devices described in Patent Document 1 is connected to at least one corresponding camera device, and analyzes an object in at least one corresponding real-time video stream transmitted by the at least one corresponding camera device to generate an analysis result of the object. Patent Document 1 discloses that the object includes a person or a suitcase, and the analysis result includes facial features of the person or characteristics of the suitcase.

[0004] The cluster management service device described in Patent Document 1 is a cluster management service device that is connected to multiple image analysis devices, and concatenates the analysis results generated by each of the multiple image analysis devices to generate a trajectory of an object.

[0005] Patent Document 2 describes a technology for calculating the feature amount of each of a plurality of key points of a human body included in an image, searching for images including human bodies with similar postures or movements based on the calculated feature amount, and classifying images with similar postures or movements together. Non-Patent Document 1 describes a technology related to human skeleton estimation. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2020-184292 A [Patent Document 2] International Publication No. 2021 / 084677 [Non-patent literature]

[0007] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, [Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields];, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]

[0008] In general, by analyzing a video, various feature quantities related to the appearance of a detection target can be detected, including but not limited to facial features of a person and characteristics of a suitcase. According to the distributed target tracking system described in Patent Document 1, even if it is possible to track a target in a real-time video stream, it is difficult to utilize the results of analyzing multiple videos for purposes other than tracking the target. Note that Patent Document 2 and Non-Patent Document 1 also do not disclose a technology that utilizes the results of analyzing multiple videos.

[0009] In view of the above-mentioned problems, one example of an object of the present invention is to provide a video analysis device, a video analysis method, and a program that solve the problem of utilizing the results of analyzing a plurality of videos. [Means for solving the problem]

[0010] According to one aspect of the present invention, a type receiving means for receiving a selection of a type of engine for analyzing each of the plurality of videos and detecting a detection target included in each of the plurality of videos; an acquisition means for acquiring a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; and an integration means for integrating the results of analyzing the plurality of acquired images. A video analysis device is provided.

[0011] According to one aspect of the present invention, The computer Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; The results of analyzing the acquired multiple images are integrated. A method for video analysis is provided.

[0012] According to one aspect of the present invention, On the computer, Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; A program is provided to execute the step of integrating the results of analyzing the captured images. Effect of the Invention

[0013] According to one aspect of the present invention, it becomes possible to utilize the results of analyzing multiple videos. [Brief description of the drawings]

[0014] [Figure 1] 1 is a diagram showing an overview of a video analysis device according to an embodiment; [Diagram 2]1 is a diagram showing an overview of a video analysis system according to an embodiment; [Diagram 3] 11 is a flowchart illustrating an example of a video analysis process according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating a detailed example of a configuration of a video analysis system according to an embodiment. [Diagram 5] FIG. 2 is a diagram showing an example of the configuration of video information according to an embodiment; [Figure 6] FIG. 13 is a diagram illustrating an example of a configuration of analysis information according to an embodiment. [Figure 7] FIG. 2 is a diagram illustrating a detailed example of a functional configuration of a video analysis device according to an embodiment. [Figure 8] FIG. 13 is a diagram illustrating an example of a configuration of integration information according to an embodiment. [Figure 9] FIG. 1 is a diagram illustrating an example of a physical configuration of a video analysis device according to an embodiment. [Figure 10] 11 is a flowchart illustrating an example of an analysis process according to an embodiment. [Figure 11] 13 illustrates an example of a start screen according to an embodiment. [Figure 12] 11 is a flowchart illustrating a detailed example of an integration process according to an embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of an integrated result screen according to an embodiment. [Figure 14] FIG. 13 is a diagram showing an example of an appearance frequency display screen according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0016] <Embodiment> 1 is a diagram showing an overview of a video analysis device 100 according to an embodiment. The video analysis device 100 includes a type receiving unit 110, an acquisition unit 111, and an integration unit 112.

[0017] The type receiving unit 110 receives a selection of a type of engine for analyzing each of a plurality of videos and detecting a detection target included in each of the plurality of videos. The acquiring unit 111 acquires a result of analyzing each of the plurality of images using the selected type of engine from among the results of analyzing each of the plurality of videos using each of the plurality of types of engine. The integrating unit 112 integrates the acquired results of analyzing the plurality of videos.

[0018] According to this video analysis device 100, it becomes possible to utilize the results of analyzing a plurality of videos.

[0019] 2 is a diagram showing an overview of a video analysis system 120 according to an embodiment. The video analysis system 120 includes a video analysis device 100, a plurality of imaging devices 121_1 to 121_K, and an analysis device 122. Here, K is an integer equal to or greater than 2, and the same applies hereinafter.

[0020] The plurality of image capturing devices 121_1 to 121_K are devices for capturing a plurality of images. The analysis device 122 uses a plurality of types of engines to analyze each of the plurality of images.

[0021] According to this video analysis system 120, it becomes possible to utilize the results of analyzing a plurality of videos.

[0022] FIG. 3 is a flowchart illustrating an example of a video analysis process according to an embodiment.

[0023] The type receiving unit 110 receives a selection of a type of engine for analyzing each of a plurality of videos and detecting a detection target included in each of the plurality of videos (step S101).

[0024] The acquisition unit 111 acquires the results of analyzing each of the multiple videos using a selected type of engine from among the results of analyzing each of the multiple videos using each of the multiple types of engines (step S102).

[0025] The integration unit 112 integrates the results of analyzing the multiple acquired videos (step S103).

[0026] This video analysis process makes it possible to utilize the results of analyzing multiple videos.

[0027] A detailed example of the video analysis system 120 according to an embodiment will be described below.

[0028] FIG. 4 is a diagram showing a detailed example of the configuration of the video analysis system 120 according to this embodiment.

[0029] The video analysis system 120 includes a video analysis device 100, K image capturing devices 121_1 to 121_K, and an analysis device 122.

[0030] The video analysis device 100, each of the imaging devices 121_1 to 121_K, and the analysis device 122 are connected to each other via a communication network N configured by wired or wireless communication or a combination of these. The video analysis device 100, each of the imaging devices 121_1 to 121_K, and the analysis device 122 transmit and receive information to and from each other via the communication network N.

[0031] (Configuration of the imaging devices 121_1 to 121_K) Each of the imaging devices 121_1 to 121_K is a device for capturing an image.

[0032] Each of the photographing devices 121_1 to 121_K is, for example, a camera installed to photograph a predetermined photographing area within a predetermined range. The predetermined range may be a building, a facility, a city, town, village, prefecture, etc., or may be a range appropriately determined within these. The photographing areas of the photographing devices 121_1 to 121_K may partially overlap each other, or may be different areas.

[0033] The image capturing device 121_i captures an image of a predetermined image capturing area at, for example, a predetermined frame rate. The image capturing device 121_i captures an image of the predetermined image capturing area to generate image information 124a_i including an image. The image is composed of a plurality of frame images in a time series. Here, i is an integer between 1 and K, and the same applies hereinafter. That is, the image capturing device 121_i means any one of the image capturing devices 121_1 to 121_K.

[0034] The image capturing device 121_i transmits image information 124a_i indicating the captured image to the analysis device 122 via the communication network N. The image capturing device 121_i may transmit the image information 124a_i to the analysis device 122 at various times. For example, the image capturing device 121_i may transmit the image information 124a_i to the analysis device 122, or may transmit the image information 124a_i to the analysis device 122 collectively at a predetermined time (for example, a predetermined time of day).

[0035] Fig. 5 is a diagram showing an example of the configuration of the video information 124a_i. The video information 124a_i is information including a video composed of a plurality of frame images. In detail, for example, as shown in Fig. 5, the video information 124a_i,j associates a video ID, a photographing device ID, a photographing time, and a video (a group of frame images).

[0036] The video ID is information (video identification information) for identifying each of the multiple videos. The imaging device ID is information (imaging identification information) for identifying each of the imaging devices 121_1 to 121_K. The imaging time is information indicating the time when the video was captured. The imaging time may include, for example, the start timing and end timing of imaging. The imaging time may further include the frame imaging timing when each frame image was captured. Each of the start timing, end timing, and frame imaging timing may be composed of, for example, a date and a time.

[0037] In the video information 124a_i, a video ID is associated with a video identified by the video ID. In addition, in the video information 124a_i, a video ID is associated with a photographing device ID of the photographing device 121_i that photographed the video identified by the video ID, and a photographing time (start timing, end timing) indicating the time when the video identified by the video ID was photographed. Furthermore, in the video information 124a, a video ID is associated with each of the frame images constituting the video identified by the video ID and a photographing time (frame photographing timing).

[0038] (Functions of analysis device 122) Analysis device 122 analyzes each of the frame images captured by each of image capturing devices 121_1 to 121_K, thereby analyzing a plurality of videos captured by each of image capturing devices 121_1 to 121_K. Analysis device 122 includes an analysis unit 123 and an analysis storage unit 124, as shown in FIG.

[0039] The analysis unit 123 acquires the video information 124a_1-124a_K from each of the imaging devices 121_1-121_K, and stores the acquired plurality of pieces of video information 124a_1-124a_K in the analysis storage unit 124. The analysis unit 123 analyzes the plurality of images included in each of the acquired plurality of pieces of video information 124a_1-124a_K. In detail, for example, the analysis unit 123 analyzes the plurality of frame images included in each of the plurality of pieces of video information 124a_1-124a_K.

[0040] The analysis unit 123 generates analysis information 124b indicating the results of analyzing the multiple videos, and stores the analysis information 124b in the analysis storage unit 124. Furthermore, the analysis unit 123 transmits the multiple video information 124a_1 to 124a_K and the analysis information 124b to the video analysis device 100 via the communication network N.

[0041] The analysis unit 123 has a function of analyzing an image using a plurality of types of engines. The various engines have a function of analyzing an image and detecting a detection target included in the image. That is, the analysis unit 123 according to the present embodiment uses a plurality of types of engines to analyze frame images (i.e., videos) included in each of the video information 124a_1 to 124a_K, and generates the analysis information 124b.

[0042] The detection target according to the present embodiment is a person, but the detection target may be a predetermined object such as a car or a bag.

[0043] Examples of engine types include (1) an object detection engine, (2) a face analysis engine, (3) a human shape analysis engine, (4) a posture analysis engine, (5) a behavior analysis engine, (6) an appearance attribute analysis engine, (7) a gradient feature analysis engine, (8) a color feature analysis engine, and (9) a movement line analysis engine. The analysis device 122 may include at least two of the engines of the types exemplified here and other types of engines.

[0044] (1) The object detection engine detects people and objects from images. The object detection function can also determine the location of people and objects within an image. An example of a model that can be applied to object detection processing is YOLO (You Only Look Once).

[0045] (2) The face analysis engine detects human faces from images, extracts the features of the detected faces, and classifies the detected faces. The face analysis engine can also determine the position of the face within the image. The face analysis engine can also determine the identity of people detected from different images based on the similarity between the facial features of people detected from different images.

[0046] (3) The human morphology analysis engine extracts the human features of people in an image (for example, values ​​that indicate overall characteristics such as whether the body is fat or thin, height, and clothing) and classifies (classifies) people in the image. The human morphology analysis engine can also identify the position of a person in an image. The human morphology analysis engine can also determine the identity of people in different images based on the human features of the people in different images.

[0047] (4) The posture analysis engine generates posture information indicating the posture of a person. The posture information includes, for example, a posture estimation model of a person. The posture estimation model is a model in which the joints of a person are connected, estimated from an image. The posture estimation model is composed of a plurality of model elements corresponding to, for example, joint elements corresponding to joints, trunk elements corresponding to the torso, and bone elements corresponding to bones connecting the joints. The posture analysis function, for example, detects the joint points of a person from an image and creates a posture estimation model by connecting the joint points.

[0048] Then, the posture analysis engine uses the information of the posture estimation model to estimate the posture of a person, extracts features of the estimated posture (posture features), classifies people included in the image, etc. The posture analysis engine can also determine the identity of people included in different images based on the posture features of people included in different images.

[0049] For example, the techniques disclosed in Patent Document 2 and Non-Patent Document 1 can be applied to the posture analysis engine.

[0050] (5) The behavior analysis engine can estimate human movements using posture estimation model information, posture changes, etc., extract features of human movements (movement features), and classify (classify) people included in images. The behavior analysis engine can also estimate a person's height and identify a person's position in an image using stick figure model information. The behavior analysis engine can estimate behaviors such as posture changes or transitions, and movement (position changes or transitions) from images, and extract movement features related to the behaviors.

[0051] (6) The appearance attribute analysis engine can recognize appearance attributes associated with a person. The appearance attribute analysis engine extracts features related to the recognized appearance attributes (appearance attribute features) and classifies (classifies) people included in the image. Appearance attributes are attributes of appearance, and include, for example, one or more of the color of clothing, the color of shoes, hairstyle, whether a hat or tie is worn, and whether glasses are worn or not.

[0052] (7) The gradient feature analysis engine extracts gradient features in an image. For example, technologies such as SIFT, SURF, RIFF, ORB, BRISK, CARD, and HOG can be applied to the gradient feature analysis engine.

[0053] (8) The color feature analysis engine can detect objects from an image, extract color features of the detected objects, and classify the detected objects. Color features are, for example, color histograms. The color feature analysis engine can detect, for example, people and objects included in an image.

[0054] (9) The movement line analysis engine can determine the movement line (movement trajectory) of a person included in a video by using, for example, the result of the identity determination performed by one or more of the above-mentioned engines. In detail, for example, by connecting a person determined to be the same between chronologically different images, the movement line of the person can be determined. Also, for example, the movement line analysis engine can determine a movement feature amount indicating the movement direction and movement speed of the person. The movement feature amount may be either one of the movement direction and movement speed of the person.

[0055] When images captured by a plurality of image capturing devices 121_2 to 121_K capturing different imaging areas are acquired, the flow line analysis engine can also find a flow line spanning a plurality of images capturing different imaging areas.

[0056] In addition, the engines (1) to (9) can calculate the reliability of the features they each require.

[0057] Each of the engines (1) to (9) may use the results of analysis performed by the other engines as appropriate. Video analysis device 100 may include an analysis unit having the functions of analysis device 122.

[0058] The analysis storage unit 124 is a storage unit for storing various types of information such as the video information 124a_1 to 124a_K and the analysis information 124b.

[0059] 6 is a diagram showing an example of the configuration of the analysis information 124b. The analysis information 124b associates a video ID, a shooting device ID, a shooting time, and an analysis result.

[0060] The video ID, the photographing device ID, and the photographing time associated in the analysis information 124b are the same as the video ID, the photographing device ID, and the photographing time associated in the video information 124a_i, respectively.

[0061] The analysis result is information indicating the result of analyzing a video identified by the associated video ID. In the analysis information 124b, the analysis result is associated with a video ID for identifying the video that was the subject of analysis to obtain the analysis result.

[0062] The analysis results are associated with, for example, the detection target ID, engine type, appearance features, and reliability.

[0063] The detection target ID is information for identifying the detection target (detection target identification information). In this embodiment, as described above, the detection target is a person. Therefore, the detection target ID is information for identifying a person detected by the analysis device 122 analyzing each of the multiple frame images. The detection target ID in this embodiment is information for identifying each image (human image) showing a person detected from each of the multiple frame images, regardless of whether the detection target is the same person.

[0064] The detection target ID may be information for identifying each person represented by a human image detected from each of the multiple frame images. In this case, the detection target ID will be the same when the detection targets are the same person, and different detection target IDs when the detection targets are different people.

[0065] In the analysis information 124b, the detection target ID is information for identifying a detection target included in a video identified using the video ID associated therewith.

[0066] The engine type indicates the type of engine used to analyze the video.

[0067] The appearance feature indicates a feature related to the appearance of the detection target, such as the result of object detection by the object detection function, a face feature, a human body feature, a posture feature, a movement feature, an appearance attribute feature, a gradient feature, a color feature, a movement feature, etc.

[0068] In the analysis results of the analysis information 124b, the appearance feature amount indicates a feature amount obtained for the detection target indicated by the detection target ID associated therewith, by using the type of engine associated therewith.

[0069] The reliability indicates the reliability of the appearance feature amount. In the analysis result of the analysis information 124b, the reliability indicates the reliability of the appearance feature amount associated therewith.

[0070] For example, when analysis device 122 obtains appearance feature amounts using each of the above-mentioned engines (1) to (9), the analysis result associates an engine type indicating the type of each of engines (1) to (9) with a common detection target ID. Then, the analysis result associates each engine type with the appearance feature amount obtained using the type of engine indicated by the engine type and the reliability of the appearance feature amount.

[0071] (Functions of the video analysis device 100) 7 is a diagram showing a detailed example of the functional configuration of the video analysis device 100 according to this embodiment. The video analysis device 100 includes a storage unit 108, a receiving unit 109, a type receiving unit 110, an acquiring unit 111, an integrating unit 112, a display control unit 113, and a display unit 114. Note that the video analysis device 100 may include an analyzing unit 123, in which case the video analysis system 120 does not need to include the analyzing device 122.

[0072] The storage unit 108 is a storage unit for storing various types of information.

[0073] The receiving unit 109 receives various information such as the video information 124a_1 to 124a_K and the analysis information 124b from the analysis device 122 via the communication network N. The receiving unit 109 may receive the video information 124a_1 to 124a_K and the analysis information 124b from the analysis device 122 in real time, or may receive them as needed, such as when they are used for processing in the video analysis device 100.

[0074] The receiving unit 109 stores the received information in the storage unit 108. That is, in this embodiment, the information stored in the storage unit 108 includes the video information 124a_1 to 124a_K and the analysis information 124b.

[0075] The receiving unit 109 may receive the video information 124a_1 to 124a_K from each of the imaging devices 121_1 to 121_K via the communication network N and store the received information in the storage unit 108. The receiving unit 109 may also receive the video information 124a_1 to 124a_K and the analysis information 124b from the analysis device 122 via the communication network N as necessary, such as when the information is used for processing in the video analysis device 100. In this case, the video information 124a_1 to 124a_K and the analysis information 124b may not be stored in the storage unit 108. Furthermore, for example, when the receiving unit 109 receives all of the video information 124a_1 to 124a_K and the analysis information 124b from the analysis device 122 and stores them in the storage unit 108, the analysis device 122 may not hold the video information 124a_1 to 124a_K and the analysis information 124b.

[0076] Type receiving unit 110 receives, for example, from a user, a selection of the type of engine used to analyze a video by analysis device 122. The type of engine received by type receiving unit 110 is not limited to one, and may be multiple.

[0077] In detail, for example, the type receiving unit 110 receives information indicating any one of the types of engines, such as (1) object detection engine, (2) face analysis engine, (3) human shape analysis engine, (4) posture analysis engine, (5) behavior analysis engine, (6) appearance attribute analysis engine, (7) gradient feature analysis engine, (8) color feature analysis engine, and (9) movement line analysis engine.

[0078] The engine type may be selected by selecting the results of analyzing each of the multiple videos. In this case, for example, the type receiving unit 110 may receive the selection of the results of analyzing each of the multiple videos and specify the type of engine used to obtain the selected result.

[0079] The acquiring unit 111 acquires, from the storage unit 108, analysis information 124b indicating the results of analyzing each of the multiple videos using a selected type of engine, i.e., the type of engine accepted by the type accepting unit 110, among the results of analyzing each of the multiple videos using each of the multiple types of engines. Note that the acquiring unit 111 may receive the analysis information 124b from the analysis device 122 via the communication network N.

[0080] The result of analyzing each of the multiple videos is information included in the analysis information 124b. Therefore, the result of analyzing each of the multiple videos includes, for example, an appearance feature amount of the detection target included in each of the multiple videos. In addition, for example, the result of analyzing each of the multiple videos includes a camera device ID (camera identification information) for identifying the camera devices 121_1 to 121_K that captured the video including the detection target. Furthermore, for example, the result of analyzing each of the multiple videos includes a capture time when the video including the detection target was captured. This capture time may include at least one of the start timing and end timing of the video including the detection target and the frame capture timing of the frame image including the detection target.

[0081] Here, the multiple images that are the subject of analysis for generating the analysis information 124b acquired by the acquisition unit 111 are images that are related in terms of location and time. That is, in this embodiment, the multiple images are images obtained by shooting multiple locations within a predetermined range at different times within a predetermined period (for example, one day, one week, one month).

[0082] The multiple images included in each of the multiple image information 124a_1 to 124a_K are not limited to images related in terms of location and time, but may be images related in terms of location or time. That is, the images analyzed to generate the analysis information 124b acquired by the acquisition unit 111 may be images obtained by shooting the same place at different times within a predetermined period, or may be images obtained by shooting multiple places within a predetermined range at the same time.

[0083] The integration unit 112 integrates the analysis results acquired by the acquisition unit 111. That is, the integration unit 112 integrates the results of analyzing each of the multiple videos using the selected type of engine, i.e., the type of engine accepted by the type acceptance unit 110. In detail, for example, the integration unit 112 integrates the results of analyzing each of the multiple videos using the same type of engine.

[0084] Note that multiple types of engines may be selected, in which case the integration unit 112 may integrate the results of analyzing each of the multiple videos using each selected type of engine for each selected engine type. In other words, when multiple types of engines are selected, the integration unit 112 may integrate the results of analyzing each of the multiple videos using the same type of engine for each selected engine type.

[0085] In this embodiment, the integrating unit 112 integrates the analysis results by grouping the detection targets based on the appearance feature amounts of the detection targets detected by the analysis.

[0086] In detail, for example, the integration unit 112 includes a grouping unit 112a and a statistical processing unit 112b as shown in FIG.

[0087] The grouping unit 112a groups the detection targets included in the multiple videos based on the similarity of the appearance feature amounts of the detection targets, and generates integrated information 108a that associates the detection targets with the groups to which the detection targets belong. The grouping unit 112a stores the generated integrated information 108a in the storage unit 108.

[0088] More specifically, the grouping unit 112a accepts designation of videos to be integrated based on, for example, a user's input or a preset value, and groups the detection targets detected using the designated videos based on the similarity of the appearance feature amounts of the detection targets.

[0089] The images to be integrated are specified, for example, by using a combination of the image capturing devices 121_1-121_K that captured the images to be integrated and a capture period during which the images were captured. The capture period is specified, for example, by a combination of a time period, a date, etc. The grouping unit 112a identifies the images captured by the designated image capturing devices 121_1-121_K during the designated capture period, and groups the detection targets included in the identified images.

[0090] The grouping unit 112a may group detection targets included in all the images captured by all the image capturing devices 121_1 to 121_K. The grouping unit 112a may also group detection targets included in a plurality of images captured by all the image capturing devices 121_1 to 121_K in a specified time period.

[0091] The grouping unit 112a acquires grouping conditions for grouping detection targets based on, for example, a user input or a preset value. The grouping unit 112a holds the grouping conditions. The grouping unit 112a groups detection targets included in multiple videos based on the grouping conditions.

[0092] The grouping condition includes at least one of a first threshold value related to the reliability of the appearance feature amount, a second threshold value related to the similarity of the appearance feature amount, and the number of groups. Note that the grouping condition may include at least one of the first threshold value, the second threshold value, and the number of groups.

[0093] The grouping unit 112a may extract detection objects associated with appearance features whose reliability is equal to or greater than a first threshold based on the grouping conditions, and may then group the extracted detection objects based on the appearance features.

[0094] Further, for example, the grouping unit 112a may perform grouping based on the grouping condition such that detection objects whose similarity in appearance feature amount is equal to or greater than a second threshold value are placed in the same group, and detection objects whose similarity in appearance feature amount is less than the second threshold value are placed in a different group.

[0095] Furthermore, for example, the grouping unit 112a may group the detection targets so that the number of groups into which the detection targets are grouped is the number of groups included in the grouping condition.

[0096] Grouping section 112a may use common grouping conditions for grouping detection targets regardless of the user of video analysis device 100, or may use grouping conditions designated by the user from among a plurality of grouping conditions for grouping detection targets.

[0097] Furthermore, the grouping unit 112a may hold grouping conditions in association with user identification information for identifying a user. In this case, the grouping unit 112a may use grouping conditions associated with user identification information for identifying a logged-in user, user identification information input by a user, or the like, for grouping the detection targets. This allows the grouping unit 112a to group the detection targets included in multiple videos based on the grouping conditions defined for each user.

[0098] 8 is a diagram showing an example of the configuration of the integration information 108a. The integration information 108a associates, for example, integration targets with group information.

[0099] The integration target is information for identifying a plurality of videos to be integrated. In the example shown in Fig. 8, the integration target is associated with the image capture device ID, the image capture period, the image capture time, and the engine type.

[0100] The photographing device ID and the photographing period are the photographing devices 121_1 to 121_K and the photographing period, respectively, designated to identify the images to be integrated. The photographing time is the photographing time of the images captured within the photographing period. The photographing device ID and the photographing time included in the integration target can be used to identify the image ID and the image by associating them with the photographing device ID and the photographing time included in the image information 124a_i.

[0101] The engine type is information indicating the type of the selected engine. In other words, the engine type indicates the type of engine used to obtain the feature amount of the detection target detected from the multiple screens to be integrated (analysis of the multiple screens).

[0102] The group information is information indicating the result of grouping, and associates a group ID with a detection target ID. The group ID is information for identifying a group (group identification information). In the group information, the group ID is associated with the detection target ID of the detection target belonging to the group identified using the group ID.

[0103] The statistical processing unit 112b uses the integrated information 108a to count the number of detection targets included in a plurality of videos, and obtains the number of times the detection targets appear. In detail, for example, the statistical processing unit 112b uses the integrated information 108a to count the number of detection targets included in a group designated by the user, for a plurality of videos captured by the designated image capturing devices 121_1 to 121_K during a designated image capturing period, and obtains the number of times the detection targets appear in the group.

[0104] The number of appearances includes at least one of the total number of appearances, the number of appearances by time period, and the like.

[0105] The total appearance count is the appearance count obtained by counting the number of detection targets belonging to the group designated by the user that are included in all of the multiple videos captured during the capture period.

[0106] The appearance count by time period is obtained by counting the number of detection targets belonging to a group designated by the user in multiple videos captured during each time period into which the shooting period is divided, by time period. This time period may be determined according to a predetermined length of time, such as every hour, or may be designated by the user.

[0107] The display control unit 113 causes various pieces of information to be displayed on the display unit 114. For example, the display control unit 113 causes the display unit 114 to display the results integrated by the integration unit 112. The integrated results include, for example, the detection targets for each group, the imaging device IDs of the imaging devices 121_1 to 121_K that captured the video in which the detection targets were detected, the capture time of the video in which the detection targets were detected, the number of times the detection targets appeared, and the like.

[0108] Furthermore, for example, when a time period is designated by the user, the display control unit 113 causes the display unit 114 to display one or more images captured during the designated time period.

[0109] (Physical Configuration of Video Analysis Device 100) 9 is a diagram showing an example of the physical configuration of the video analysis device 100 according to this embodiment. The video analysis device 100 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, and a user interface 1060.

[0110] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, the memory 1030, the storage device 1040, the network interface 1050, and the user interface 1060. However, the method of connecting the processor 1020 and the like to each other is not limited to a bus connection.

[0111] The processor 1020 is implemented by a central processing unit (CPU) or a graphics processing unit (GPU).

[0112] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.

[0113] The storage device 1040 is an auxiliary storage device realized by a hard disk drive (HDD), a solid state drive (SSD), a memory card, a read only memory (ROM), etc. The storage device 1040 stores program modules for realizing the functions of the video analysis device 100. The processor 1020 reads each of these program modules into the memory 1030 and executes them to realize the function corresponding to the program module.

[0114] The network interface 1050 is an interface for connecting the video analysis device 100 to the communication network N.

[0115] The user interface 1060 includes a touch panel, a keyboard, a mouse, etc., as interfaces for the user to input information, and a liquid crystal panel, an organic EL (Electro-Luminescence) panel, etc., as interfaces for presenting information to the user.

[0116] Analysis device 122 may be physically configured in the same manner as video analysis device 100 (see FIG. 9). Therefore, a diagram showing the physical configuration of analysis device 122 will be omitted.

[0117] (Operation of video analysis system 120) Now, the operation of the video analysis system 120 will be described with reference to the drawings.

[0118] (Analysis processing) 10 is a flowchart showing an example of the analysis process according to the present embodiment. The analysis process is a process for analyzing the video images captured by the image capturing devices 121_1 to 121_K. The analysis process is repeatedly executed while the image capturing devices 121_1 to 121_K and the analysis unit 123 are in operation, for example.

[0119] The analysis unit 123 acquires the video information 124a_1 to 124a_K from each of the imaging devices 121_1 to 121_K via the communication network N, for example, in real time (step S201).

[0120] The analysis unit 123 stores the plurality of pieces of video information 124a_1 to 124a_K acquired in step S201 in the analysis storage unit 124, and analyzes the videos included in the plurality of pieces of video information 124a_1 to 124a_K (step S202).

[0121] For example, as described above, the analysis unit 123 uses multiple types of engines to analyze frame images included in each video to detect the detection target. Furthermore, the analysis unit 123 uses each type of engine to obtain the appearance feature amount of the detected detection target and the reliability of the appearance feature amount. The analysis unit 123 generates analysis information 124b by performing such analysis.

[0122] The analysis unit 123 stores the analysis information 124b generated by performing the analysis in step S202 in the analysis storage unit 124, and transmits the analysis information 124b to the video analysis device 100 via the communication network N (step S203). At this time, the analysis unit 123 may transmit the video information 124a_1 to 124a_K acquired in step S201 to the video analysis device 100 via the communication network N.

[0123] The receiving unit 109 receives the analysis information 124b transmitted in step S203 via the communication network N (step S204). At this time, the receiving unit 109 may receive the video information 124a_1 to 124a_K transmitted in step S203 via the communication network N.

[0124] The receiving unit 109 stores the analysis information 124b received in step S204 in the storage unit 108 (step S205), and ends the analysis process. At this time, the receiving unit 109 may receive the video information 124a_1 to 124a_K received in step S204 via the communication network N.

[0125] (Video analysis processing) The video analysis process is a process for integrating the results of analyzing the video, as described with reference to Fig. 3. The video analysis process is started, for example, when a user logs in, and the display control unit 113 causes the display unit 114 to display a start screen 131. The start screen 131 is a screen for receiving a user's designation.

[0126] Fig. 11 shows an example of a start screen 131 according to the present embodiment. The start screen 131 shown in Fig. 11 includes input fields for specifying or selecting the imaging device and imaging period associated with the integration target, the type of engine, and the first threshold, the second threshold, and the number of groups associated with the grouping condition.

[0127] 11 shows an example in which "all" of the camera devices 121_1 to 121_K is input in an input field corresponding to "camera device." For example, the camera device ID of one or more of the camera devices 121_1 to 121_K may be input in this input field.

[0128] 11 shows an example in which "April 1, 2022, 0:00-April 2, 2022, 0:00" is entered in an input field associated with "shooting period." An appropriate period may be entered in this input field.

[0129] 11 shows an example in which "appearance attribute analysis engine" is entered in an input field associated with "engine type." The type of engine used to obtain appearance features may be entered in this input field. Also, multiple types of engines used to obtain appearance features may be entered in this input field.

[0130] 11 shows an example in which "0.35", "0.25", and "3" are entered in the input fields corresponding to "first threshold", "second threshold", and "number of groups", respectively. In these input fields, for example, grouping conditions associated with the user identification information of the logged-in user are set as initial values, and these initial values ​​may be changed by the user as necessary.

[0131] When the user presses integration start button 131a, video analysis device 100 starts the video analysis process shown in FIG.

[0132] As described above with reference to FIG. 3, the type receiving unit 110 receives a selection of the type of engine for analyzing a video and detecting a detection target included in the video (step S101).

[0133] At this time, the type receiving unit 110 receives, in addition to the engine type, information specified on the start screen 131. As described with reference to Fig. 11, this information is, for example, information for specifying the imaging device, the imaging period, the first threshold, the second threshold, and the number of groups.

[0134] As described above, the acquisition unit 111 acquires the results of analyzing each of the multiple videos using the type of engine selected in step S101 (step S102).

[0135] In detail, for example, the acquisition unit 111 acquires, based on the engine type indicating the type of the selected engine, the specified image capture device ID, and the image capture period, analysis information 124b regarding the multiple videos to be integrated from the storage unit 108. Here, the acquisition unit 111 acquires, from the storage unit 108, analysis information 124b including the engine type indicating the type of the selected engine, the specified image capture device ID, and the image capture time within the specified image capture period.

[0136] The integration unit 112 integrates the results acquired in step S102 (step S103), that is, integrates the analysis information 124b acquired in step S102.

[0137] FIG. 12 is a flowchart showing a detailed example of the integration process (step S103) according to this embodiment.

[0138] The grouping unit 112a groups the detection targets included in the multiple videos based on the similarity of the appearance feature amount included in the analysis information 124b acquired in step S102 (step S103a). As a result, the grouping unit 112a generates the integrated information 108a and stores it in the storage unit 108.

[0139] The display control unit 113 causes the display unit 114 to display the results of grouping in step S103a (step S103b).

[0140] 13 is a diagram showing an example of an integrated result screen 132 which is a screen showing the grouping result. The integrated result screen 132 displays, for each group, a list of the image capturing device IDs of the image capturing devices 121_1 to 121_K which captured the video in which the detection object belonging to the group is detected.

[0141] 13, group 1, group 2, and group 3 indicate group IDs of three groups according to the designation of the number of groups. In the example shown in FIG. 13, the camera device IDs "camera device 1" and "camera device 2" corresponding to the camera devices 121_1 to 121_2 are associated with group 1. The camera device IDs "camera device 2" and "camera device 3" corresponding to the camera devices 121_2 to 121_3 are associated with group 2. The camera device ID "camera device 4" corresponding to the camera device 121_4 is associated with group 3.

[0142] Note that the integrated result screen 132 is not limited to this, and may display, for example, a list of image IDs of images in which detection objects belonging to a group have been detected, for each group.

[0143] The statistical processing unit 112b accepts the designation of the group (step S103c).

[0144] For example, each of "Group 1", "Group 2", and "Group 3" on the integrated result screen 132 illustrated in Fig. 13 is selectable. When the user selects one of "Group 1", "Group 2", and "Group 3", the statistical processing unit 112b accepts the group designation.

[0145] The statistical processing unit 112b counts the number of detection targets belonging to the group designated in step S103c, and obtains the appearance frequency of the detection targets belonging to that group (step S103d).

[0146] In detail, for example, the statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to the group specified in step S103c that are included in the analysis information 124b acquired in step S102. This makes it possible to count the number of detection targets that belong to the group specified by the user among the multiple videos captured by the specified imaging devices 121_1 to 121_K during the specified imaging period.

[0147] The statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to the specified group included in the entire analysis information 124b acquired in step S102, to obtain the total number of appearances.

[0148] The statistical processing unit 112b divides the analysis information 124b acquired in step S102 into time periods based on the shooting times included in the analysis information 124b. The statistical processing unit 112b counts the number of detection targets (detection target IDs) belonging to the specified group included in the analysis information 124b by time period, and obtains the number of appearances by time period.

[0149] The statistical processing unit 112b may count the number of detection targets (detection target IDs) belonging to a specified group included in the entire analysis information 124b for each imaging device ID to obtain the total number of appearances for each imaging device. The statistical processing unit 112b may also count the number of detection targets (detection target IDs) belonging to a specified group included in the analysis information 124b for each time period for each imaging device ID to obtain the number of appearances for each time period and each imaging device.

[0150] The display control unit 113 causes the appearance count calculated in step S103d to be displayed on the display unit 114 (step S103e), and ends the video analysis process (see FIG. 3).

[0151] Fig. 14 is a diagram showing an example of an appearance count display screen 133 which is a screen showing the appearance count. The appearance count display screen 133 shown in Fig. 14 is an example of a screen showing the appearance count by time period and by imaging device for group 1 in a line graph.

[0152] For example, the time indicating each time period may be selectable, and when a time period is designated by the selection, the display control unit 113 may display one or more images captured during the designated time period on the display unit 114. In detail, for example, the display control unit 113 may identify an image ID corresponding to an image including a group of frame images captured during the designated time period based on the shooting time included in the analysis information 124b acquired in step S102. The display control unit 113 may display an image associated with the identified image ID on the display unit 114 based on the image information 124a_1 to 124a_K.

[0153] It should be noted that the appearance count display screen 133 is not limited to a line graph, and may display the appearance count using a pie chart, a bar graph, or the like.

[0154] By executing the video analysis process, it is possible to group the detection targets included in multiple videos based on the appearance feature values ​​obtained using the selected type of engine, thereby making it possible to group detection targets having similar appearance features.

[0155] Also, the user can check the grouped results by referring to the integrated result screen 132. Furthermore, the user can check the appearance count of the detection targets classified based on the appearance feature amount by referring to the appearance count display screen 133. This allows the user to know the tendency of detection targets with similar appearance characteristics to appear, such as when, where, and how many detection targets with similar appearance characteristics are present.

[0156] (Action and effect) As described above, according to this embodiment, the video analysis device 100 includes a type receiving unit 110, an acquisition unit 111, and an integration unit 112. The type receiving unit 110 accepts a selection of a type of engine for analyzing each of a plurality of videos and detecting a detection target contained in each of the plurality of videos. The acquisition unit 111 acquires a result of analyzing each of the plurality of videos using the selected type of engine from among results of analyzing each of the plurality of videos using each of the plurality of types of engines. The integration unit 112 integrates the acquired results of analyzing the plurality of videos.

[0157] This makes it possible to obtain information that integrates the results of analyzing multiple videos using the selected type of engine, making it possible to utilize the results of analyzing multiple videos.

[0158] According to this embodiment, the type of engine is selected by selecting the results of analyzing each of a plurality of videos.

[0159] This makes it possible to obtain information that integrates the results of analyzing multiple videos using the selected type of engine, making it possible to utilize the results of analyzing multiple videos.

[0160] According to this embodiment, the integration unit 112 integrates the results of analyzing each of the multiple videos using the same type of engine.

[0161] This makes it possible to obtain information that integrates the results of analyzing multiple videos using the same type of engine, making it possible to utilize the results of analyzing multiple videos.

[0162] According to the present embodiment, the result of analyzing the multiple videos includes appearance feature amounts of the detection targets included in each of the multiple videos. The integration unit 112 groups the detection targets included in the multiple videos based on the similarity of the appearance feature amounts of the detection targets, and generates integration information 108a that associates the detection targets with the groups to which the detection targets belong.

[0163] As a result, it is possible to obtain integrated information 108a as a result of integrating the results of analyzing a plurality of videos using the selected type of engine. This makes it possible to utilize the results of analyzing a plurality of videos.

[0164] According to this embodiment, the integration unit 112 groups the detection targets included in a plurality of videos based on grouping conditions for grouping the detection targets.

[0165] This allows detection targets to be grouped using grouping conditions, making it possible to utilize the results of analyzing multiple videos.

[0166] According to this embodiment, the grouping conditions include at least one of a first threshold value related to the reliability of the appearance feature amount, a second threshold value related to the similarity of the appearance feature amount, and the number of groups.

[0167] This allows detection targets to be grouped using at least one of the first threshold, the second threshold, and the number of groups as a condition, making it possible to utilize the results of analyzing multiple videos.

[0168] According to this embodiment, the integration unit 112 groups detection targets included in a plurality of videos based on grouping conditions defined for each user.

[0169] This allows detection targets to be grouped using grouping conditions suited to the user, making it possible to utilize the results of analyzing multiple videos.

[0170] According to the present embodiment, the result of analyzing a plurality of videos further includes shooting identification information for identifying the camera devices 121_1 to 121_K that shot the videos including the detection target. The integrated information 108a further associates the shooting identification information.

[0171] This allows the integrated information 108a to be analyzed for each imaging device, making it possible to utilize the results of analyzing a plurality of videos.

[0172] According to this embodiment, the integration unit 112 further counts the number of detection targets included in a plurality of videos to obtain the number of appearances of the detection targets.

[0173] This allows the number of appearances of the detected target to be obtained as a result of integrating the results of analyzing multiple videos using the selected type of engine, making it possible to utilize the results of analyzing multiple videos.

[0174] According to this embodiment, the result of analyzing the multiple videos further includes the shooting time when the video including the detection target was shot. The integration unit 112 further counts the number of detection targets included in the multiple videos by time period when each video was shot, and obtains the number of times the detection target appears by time period.

[0175] This allows the number of times the target object appears by time period to be obtained by integrating the results of analyzing multiple videos using the selected type of engine, making it possible to utilize the results of analyzing multiple videos.

[0176] According to this embodiment, video analysis device 100 further includes a display control unit 113 that causes display unit 114 to display the integrated result.

[0177] This allows the user to know the integrated results of analyzing multiple videos using the selected type of engine by looking at the display unit 114. This makes it possible to utilize the results of analyzing multiple videos.

[0178] According to this embodiment, when a time period is specified, the display control unit 113 causes the display unit 114 to display one or more images captured during the specified time period.

[0179] This allows the user to easily view the video used to obtain the analysis results, as necessary, making it possible to utilize the results of analyzing multiple videos.

[0180] According to this embodiment, the multiple images are images captured by using multiple image capturing devices 121_1 to 121_K.

[0181] This makes it possible to utilize the results of analyzing multiple videos taken in different locations.

[0182] According to this embodiment, the multiple images are images related to each other in terms of location or time.

[0183] This makes it possible to utilize the results of analyzing multiple videos that are related in terms of location or time.

[0184] According to this embodiment, the multiple images are images obtained by photographing the same shooting area at different times within a predetermined period, or images obtained by photographing multiple shooting areas within a predetermined range at the same or different times within a predetermined period.

[0185] This makes it possible to utilize the results of analyzing multiple videos that are related in terms of location or time.

[0186] Although the embodiment and modified examples of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various configurations other than those described above can also be adopted.

[0187] In addition, in the multiple flowcharts used in the above description, multiple steps (processing) are described in order, but the execution order of the steps performed in the embodiments is not limited to the order described. In the embodiments, the order of the steps shown in the figures can be changed to the extent that the content is not affected. In addition, the above-mentioned embodiments and modified examples can be combined to the extent that the content is not contradictory.

[0188] A part or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes.

[0189] 1. A type receiving means for receiving a selection of a type of engine for analyzing each of a plurality of videos and detecting a detection target included in each of the plurality of videos; an acquisition means for acquiring a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; and an integration means for integrating the results of analyzing the plurality of acquired images. Video analysis equipment. 2. The type of engine is selected by selecting the result of analyzing each of the plurality of videos. 1. The video analysis device according to claim 1. 3. The integration means integrates the results of analyzing each of the plurality of videos using the same type of engine. 3. The video analysis device according to claim 1 or 2. 4. The result of analyzing the plurality of images includes an appearance feature amount of the detection target included in each of the plurality of images; The integration means groups the detection targets included in the plurality of videos based on a similarity in appearance feature amounts of the detection targets, and generates integration information that associates the detection targets with the groups to which the detection targets belong. 4. A video analysis device according to any one of 1 to 3 above. 5. The integration means further groups the detection objects included in the plurality of videos based on grouping conditions for grouping the detection objects. 5. The video analysis device according to claim 4. 6. The grouping condition includes at least one of a first threshold value related to a reliability of the appearance feature amount, a second threshold value related to a similarity of the appearance feature amount, and a number of groups. 6. The video analysis device according to claim 5. 7. The integration means groups the detection targets included in the plurality of videos based on the grouping conditions defined for each user. 7. The video analysis device according to claim 5 or 6. 8. The result of analyzing the plurality of images further includes image capture identification information for identifying the image capture device that captured the images including the detection target; The integrated information further associates the photography identification information. 8. A video analysis device according to any one of claims 4 to 7. 9. The integration means further counts the number of the detection target included in the plurality of videos to obtain the number of appearances of the detection target. 9. A video analysis device according to any one of 1 to 8 above. 10. The result of analyzing the plurality of images further includes a shooting time when the image including the detection target was shot, The integration means further counts the number of the detection target included in the plurality of videos by time period in which each of the plurality of videos was shot, to obtain the number of appearances of the detection target by time period. 10. The video analysis device according to claim 9. 11. The display control means further includes a display means for displaying the integrated result. 11. A video analysis device according to any one of claims 1 to 10. 12. When a time period is specified, the display control means causes the display means to display one or more of the images captured during the specified time period. 12. The video analysis device according to claim 11. 13. The multiple images are images taken using multiple image capture devices. 13. A video analysis device according to any one of claims 1 to 12. 14. The multiple images are images related in location or time. 14. The video analysis device according to claim 13. 15. The multiple images are images obtained by photographing the same shooting area at different times within a predetermined period, or images obtained by photographing multiple shooting areas within a predetermined range at the same time or at different times within a predetermined period. 15. The video analysis device according to claim 13 or 14. 16. The video analysis device according to any one of 1 to 15 above, A plurality of image capturing devices for capturing the plurality of images; and an analysis device that analyzes each of the plurality of videos using a plurality of types of the engine. Video analysis system. 17. The computer Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; The results of analyzing the acquired multiple images are integrated. Video analysis methods. 18. To the computer: Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; A program for executing the step of integrating the results of analyzing the multiple acquired images. 19. To the computer: Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; A recording medium having a program recorded thereon for executing the process of integrating the results of analyzing the multiple acquired images. [Explanation of symbols]

[0190] 100 Image Analysis Equipment 108 Storage section 108a Integrated Information 109 Receiving section 110 Type Reception 111 Acquisition Department 112 Integration Department 112a Grouping section 112b Statistical processing section 113 Display control unit 114 Display section 120 Video Analysis System 121_2~121_K Imaging Device 122 Analyzer 123 Analysis Department 124 Analysis storage section 124a_1~124a_K Video information 124b Analysis information 131 Start screen 131a Integration start button 132 Integrated Results Screen 133 Appearance count display screen

Claims

1. a type receiving means for receiving a selection of a type of engine for analyzing each of the plurality of videos and detecting a detection target included in each of the plurality of videos; an acquisition means for acquiring a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; and an integration means for integrating the results of analyzing the plurality of acquired images. Video analysis equipment.

2. The engine type is selected by selecting a result of analyzing each of the plurality of videos. The video analysis device according to claim 1 .

3. The integration means integrates the results of analyzing each of the plurality of videos using the engines of the same type. The video analysis device according to claim 1 .

4. a result of analyzing the plurality of images includes an appearance feature amount of the detection target included in each of the plurality of images; The integration means groups the detection targets included in the plurality of videos based on a similarity in appearance feature amounts of the detection targets, and generates integration information that associates the detection targets with the groups to which the detection targets belong. The video analysis device according to any one of claims 1 to 3.

5. The integration means further groups the detection targets included in the plurality of videos based on a grouping condition for grouping the detection targets.

5. The video analysis device according to claim 4.

6. The result of analyzing the plurality of images further includes image capture identification information for identifying an image capture device that captured the images including the detection target, The integrated information further associates the photography identification information.

5. The video analysis device according to claim 4.

7. The integration means further counts the number of the detection target included in the plurality of videos to obtain the number of appearances of the detection target.

5. The video analysis device according to claim 4.

8. The result of analyzing the plurality of images further includes a shooting time when the image including the detection target was shot, The integration means further counts the number of the detection target included in the plurality of videos by time period in which each of the plurality of videos was shot, to obtain the number of appearances of the detection target by time period. The video analysis device according to claim 7.

9. The computer Accepting a selection of a type of engine that analyzes each of a plurality of videos and detects a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; The results of analyzing the acquired multiple images are integrated. Video analysis methods.

10. On the computer, Accepting a selection of a type of engine for analyzing each of the plurality of videos and detecting a detection target included in each of the plurality of videos; obtaining a result of analyzing each of the plurality of videos using a selected type of the engine from among results of analyzing each of the plurality of videos using each of the plurality of types of the engine; A program for executing the step of integrating the results of analyzing the multiple acquired images.