Parameter viewing system, parameter adjustment system, server, and server control program
Patent Information
- Application Number
- JP2021146640
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-09
AI Technical Summary
Conventional systems face difficulties in setting appropriate condition parameters for object detection in camera frames, requiring labor-intensive manual input and verification of numerical coordinates, which is cumbersome and inaccurate.
A parameter browsing system that statistically processes superimposed videos to generate visualization statistical information, allowing detection of similar past data for condition setting parameters, facilitating easy input and verification of accuracy.
Enables users to easily input and verify condition setting parameters, reducing labor intensity and improving accuracy in setting object detection frames.
Smart Images

Figure 0007695695000001 
Figure 0007695695000002 
Figure 0007695695000003
Abstract
Description
Technical Field
[0001] The present invention relates to a parameter browsing system, a parameter adjustment system, a server, and a server control program.
Background Art
[0002] Conventionally, there have been known devices and systems for performing analysis processing such as object detection on frame images captured by cameras such as AI cameras using a learned neural network model (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In devices and systems for performing analysis processing such as object detection on frame images captured by conventional cameras as described above, it is difficult to set conditions such as object detection frames for the camera to be set. For example, when setting a detection frame for an in-store customer or the like within the viewing angle of the camera to be set, if one tries to input it in a form that can be understood by a computer, it is necessary to input the numerical coordinate information of the four vertices of the detection frame within the viewing angle of the camera to be set. However, since the viewing angles of each camera, the way people enter (into the viewing angle), and the way people and objects appear are various, it is difficult for a person (user) such as the system operation administrator to appropriately input condition setting parameters such as the numerical coordinate information of the above detection frame.
[0005] Also, even if condition setting parameters such as the numerical coordinate information of the detection frame can be appropriately input, in order to confirm how correct (appropriate) the condition setting by the input of the above parameters is, it is necessary for a user such as the system operation administrator to visually compare and inspect the analysis result of the captured image (frame image) using the above condition setting parameters and the actual video. Therefore, there is a problem that the work for confirming the accuracy of the condition setting is very labor-intensive.
[0006] The present invention solves the above problems, and an object thereof is to provide a parameter browsing system, a parameter adjustment system, a server, and a server control program that enable a user such as a system operation administrator to easily input appropriate condition setting parameters. Another object is to provide a parameter adjustment system that enables a user to easily confirm the accuracy of the condition setting by the input of the condition setting parameters.
Means for Solving the Problems
[0007] In order to solve the above problems, a parameter browsing system according to a first aspect of the present invention statistically processes a superimposed video in which a display means, a captured image of a camera on the edge side, and a result of recognition processing by an image analysis device on the edge side for a frame image included in the captured image are superimposed to generate visualization statistical information. A visualization statistical information generation means, a statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means, and from the past visualization statistical information stored in the statistical information storage means, similar to the visualization statistical information based on the captured image of the camera to be set this time. Detection means for detecting past visualization statistical information, and parameter output means for outputting to the display means a screen including condition setting parameters for captured image analysis applied to the past captured image that is the source of the past visualization statistical information detected by the detection means.
[0008] In this parameter browsing system, the detection means compares a vector corresponding to the moving direction of a person obtained from visualization statistical information based on a captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from past visualization statistical information stored in the statistical information storage means, and may detect past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time.
[0009] In this parameter browsing system, the detection means includes display control means for controlling to display, on the display means, visualization statistical information based on a captured image of the camera to be set this time and a plurality of pieces of past visualization statistical information stored in the statistical information storage means. The parameter browsing system may further include selection input means for selecting, from among the plurality of pieces of past visualization statistical information displayed on the display means, the past visualization statistical information most similar to the visualization statistical information based on the captured image of the camera to be set this time.
[0012] According to an aspect of the present invention, Second the server includes visualization statistical information generation means for statistically processing a superimposed video obtained by superimposing a captured image of an edge-side camera and a result of recognition processing of a frame image included in the captured image by an edge-side image analysis device to generate visualization statistical information, statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means, detection means for detecting, from among the past visualization statistical information stored in the statistical information storage means, past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time, and parameter output means for outputting, to a display means, a screen including condition setting parameters for analyzing a captured image, which were applied to the past captured image that was the source of the past visualization statistical information detected by the detection means.
[0013] In this server, the detection means compares a vector corresponding to the moving direction of a person obtained from visualization statistical information based on a captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from past visualization statistical information stored in the statistical information storage means, and may detect past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time.
[0014] According to an aspect of the present invention Third The server control program is a server control program for controlling the server so as to perform a process for finding condition setting parameters for analyzing a captured image suitable for the captured image of the camera to be set this time, and the server includes a superimposed video in which a captured image of an edge-side camera and a result of recognition processing of a frame image included in the captured image by an edge-side image analysis device are superimposed, and performs statistical processing to generate visualization statistical information, a statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means, a detection means for detecting past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from among the past visualization statistical information stored in the statistical information storage means, and functions as a parameter output means for outputting to a display means a screen including condition setting parameters for analyzing a captured image applied to the past captured image that is the source of the past visualization statistical information detected by the detection means.
[0015] In this server control program, the detection means compares a vector corresponding to the moving direction of a person obtained from visualization statistical information based on a captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from past visualization statistical information stored in the statistical information storage means, and may detect past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time.
Advantages of the Invention
[0016] The parameter browsing system according to the first aspect of the present invention SecondAccording to the server and Third the server control program according to the embodiment, past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time (hereinafter abbreviated as "the current camera") is detected, and the shooting image analysis conditions applied to the past captured image that was the basis of the detected past visualization statistical information are detected. A screen including the setting parameters is output to the display means so that a user such as an operation administrator of the system can view it. Here, the past visualization statistical information similar to the visualization statistical information based on the captured image of the current camera is likely to be visualization statistical information using the result of performing the same recognition process as the recognition process performed on the captured image of the current camera for the captured image (frame image in) of the moving range of the same person as the captured image of the current camera. For this reason, the user refers to the condition setting parameters applied to the captured image that was the basis of the past visualization statistical information (similar to the visualization statistical information based on the captured image of the current camera) viewed above, and inputs the condition setting parameters for the captured image analysis of the current camera. As a result, a user such as an operation administrator of the system can easily input appropriate condition setting parameters (for example, a human detection frame).
Brief Description of Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Mode for Carrying Out the Invention
[0019] Hereinafter, a parameter browsing system, a parameter adjustment system, a server, and a server control program according to an embodiment embodying the present invention will be described with reference to the drawings. FIG. 1 is a block configuration diagram showing a schematic configuration of an image analysis system 10 (corresponding to the "parameter browsing system" and the "parameter adjustment system" in the claims) according to the present embodiment. In the present embodiment, a plurality of fixed cameras (monitoring cameras) 2, which are network cameras for monitoring that photograph a predetermined photographing area, an analysis box 3 that analyzes the video from the fixed cameras 2, and a plurality of signage 5, which are tablet terminals for digital signage equipped with built-in cameras 4 (or web cameras), will be described as an example of the case where they are arranged in a store S such as a chain store. As shown in FIG. 1, the image analysis system 10 includes, in addition to the plurality of fixed cameras 2, the analysis box 3, the plurality of signage 5, a WiFi AP (WiFi Access Point) 6, a hub 7, and a router 8 arranged in the store S, a management server 1 (corresponding to the "server" in the claims) arranged on the cloud C, and a personal computer (personal computer) 9 used by a user such as a system operation administrator. Note that the camera on the signage 5 side may be a built-in camera 4 disposed inside the housing of the signage 5 or a web camera attached to the signage 5. In the following description, an example in which the camera on the signage 5 side is a built-in camera 4 will be mainly described.
[0020] The above-mentioned fixed cameras 2 have IP addresses and can be directly connected to the network. As shown in FIG. 1, the analysis box 3 is connected to the plurality of fixed cameras 2 via a LAN (Local Area Network) and a hub 7, and analyzes the images input from each of these fixed cameras 2. Specifically, the analysis box 3 performs object detection processing (detection processing of customers and customer faces) on the images input from each of the fixed cameras 2, and object recognition processing (recognition processing such as estimation of customer attributes (gender and age (age group))) on the images of the objects detected by this object detection processing.
[0021] The signage 5 is mainly installed on the product shelves in the store S, and displays content such as advertisements for customers who visit the store S on its touch panel display 14 (see Figure 2). At the same time, it performs object detection processing (detection processing of customers and customer faces) on the frame images from its built-in camera 4, object recognition processing on the images of the objects detected by this object detection processing, and the like.
[0022] The above management server 1 is installed in the management department (head office, etc.) of the store including the store S, and manages a large number of fixed cameras 2, analysis boxes 3, and signages 5 distributed to each store. Specifically, the management server 1 controls the installation of application packages to the analysis boxes 3 and signages 5 in each store, and the startup and stop of the fixed cameras 2 connected to the analysis boxes 3. In addition, the management server 1 receives from the edge-side image analysis device (analysis box 3 or signage 5) the captured images of the fixed camera 2 or the built-in camera 4 and the analysis results including the results of the recognition processing for the captured images, and uses the received captured images and the above analysis results to perform a process for finding the condition setting parameters for photograph image analysis suitable for the captured images of the camera (fixed camera 2 or built-in camera 4) to be set this time. As shown in Figure 1, a personal computer 9 used by a user such as an operation administrator of the image analysis system 10 is connected to the management server 1, and the user can use the personal computer 9 to access the management server 1 (portal) and perform a process (such as viewing past condition setting parameters) for finding the condition setting parameters for photograph image analysis suitable for the camera (fixed camera 2 or built-in camera 4) to be set this time.
[0023] Next, referring to FIG. 2, the hardware configuration of the above-described tablet-type signage 5 will be described. In addition to the built-in camera 4 described above, the signage 5 includes a System-on-a-Chip (SoC) 11, a touch panel display 14, a speaker 15, various programs including an app package, a memory 16 that stores various data including condition setting parameters, a communication unit 17, a secondary battery 18, a charging terminal 19, a storage 20 such as an SD memory card that stores the captured image of the built-in camera 4 and an analysis (result) superimposed video (see FIG. 6 etc.) described later. The SoC 11 includes a CPU 12 that controls the entire device and performs various calculations, and a GPU 13 that is used for inference processing of various learned Deep Neural Networks (DNN) models. The GPU 13 corresponds to the DNN calculation chip 69 shown in FIGS. 6 and 7.
[0024] The programs stored in the memory 16 include the learned DNN models 70 (various learned inference models) shown in FIGS. 6 and 7. The communication unit 17 includes a communication IC and an antenna. The signage 5 is connected to the management server 1 on the cloud C via the communication unit 17 and the network. The secondary battery 18 is a battery such as a lithium-ion battery that can be repeatedly used by charging, and stores the power from the commercial power supply converted to DC power by an AC / DC converter and supplies it to each part of the signage 5.
[0025] Next, with reference to FIG. 3, the hardware configuration of the analysis box 3 will be described. The analysis box 3 includes a CPU 21 that controls the entire device and performs various operations, a hard disk 22 that stores various data and programs, a RAM (Random Access Memory) 23, inference chips (hereinafter abbreviated as "chips") 24a to 24h that are DNN (Deep Neural Networks) inference processors, and a communication control IC 25. The CPU 21 is a general-purpose CPU or a CPU designed to enhance parallel processing performance for simultaneously processing a large number of video streams. In addition, the data stored in the hard disk 22 includes video data after decoding the video stream (data) input from each of the fixed cameras 2, the data of the analysis result superimposed video described later with reference to FIG. 6 and the like, and analysis data 58 (the same as the video metadata 74 in FIG. 7). Further, the programs stored in the hard disk 22 include pre-trained DNN models for inference processing such as human detection processing, face detection processing, and attribute recognition processing (various pre-trained DNN models for inference processing: corresponding to the DNN model 57 in FIGS. 6 and 7), an analysis box OS program, a VMS (Video Management System) program, a video analysis program, and an analysis superimposed video generation program.
[0026] Next, with reference to FIG. 4, the hardware configuration of the management server 1 will be described. The management server 1 includes a CPU 31 that controls the entire device and performs various operations, a hard disk 32 (corresponding to the "statistical information storage means" in the claims) that stores various data and programs, a RAM (Random Access Memory) 33, a display 34, an operation unit 35, and a communication unit 36 (corresponding to the "reception means" in the claims).
[0027] The programs stored in the hard disk 32 include a server control program 37, a portal (for use) program 38, and a dashboard 39. The data stored in the hard disk 32 includes statistical data 40. The server control program 37 is a program for controlling the management server 1. This server control program 37 performs processes such as finding condition setting parameters for analyzing captured images suitable for the captured images of the camera to be set this time (hereinafter, may be abbreviated as "the current camera") (such as viewing the accumulated past condition setting parameters), and processes for adjusting (editing) the condition setting parameters of the current camera. The portal 38 is a type of so-called enterprise portal (software for aggregating and displaying various information and applications scattered within the enterprise on the computer screen in order to efficiently search for and use these information and applications). The processes that can be performed using this portal 38 will be described in the explanations of FIGS. 6 and 7. The dashboard 39 is software for aggregating and visualizing statistical information such as customer attributes, stay times, and tracking results of customer behavior. The dashboard 39 is used for the process of extracting the counts of visitors, passersby, lingerers, etc. from the statistical data 40 based on a request from a user using the personal computer 9 and displaying them on the display 50b (see FIG. 5) of the personal computer 9.
[0028] Figure 5 shows the functional blocks of the above-described management server 1 and personal computer 9. The CPU 31 of the management server 1 includes, as functional blocks, a visualization statistical information generation unit 41, a detection unit 42, a parameter output unit 43, a parameter display control unit 45, a reanalysis unit 46, and a reanalysis result display control unit 47. Further, the above-described detection unit 42 includes a display control unit 44. The above-described visualization statistical information generation unit 41, detection unit 42, parameter output unit 43, display control unit 44, parameter display control unit 45, reanalysis unit 46, and reanalysis result display control unit 47 respectively correspond to the visualization statistical information generation means, detection means, parameter output means, display control means, parameter display control means, reanalysis means, and reanalysis result display control means in the claims. Also, the personal computer 9 includes an operation unit 50a (corresponding to the "adjustment input means" and "selection input means" in the claims) and a display 50b (corresponding to the "display means" in the claims).
[0029] The above-described visualization statistical information generation unit 41 statistically processes a superimposed video obtained by superimposing a captured image of a camera on the edge side (fixed camera 2 or built-in camera 4) and a result of recognition processing of the frame image included in the captured image by an image analysis device on the edge side (analysis box 3 or signage 5) to generate visualization statistical information. This visualization statistical information is an image that is the result of performing some statistical processing on a superimposed video composed of each frame image included in the captured image for a predetermined time and each superimposed image obtained by superimposing the result of recognition processing on these frame images. For example, the face detection aggregation image 89 (a type of visualization statistical information) shown in FIG. 13 is a heat map created based on each superimposed image obtained by superimposing each frame image included in the captured image of a certain camera for one hour and the face detection results for these frame images. In this example, the portions with a high face detection frequency in the captured image of the corresponding camera for one hour are shown in a dark color (in the actual image, a dark red, etc.).
[0030] The statistical data 40 including the visualized statistical information 48 generated by the above-described visualized statistical information generation unit 41 is stored in the hard disk 32. The above statistical data 40 includes the condition setting parameters 49 (such as detection frames) of each camera (fixed camera 2 or built-in camera 4) on the edge side.
[0031] The detection unit 42 in FIG. 5 detects past visualized statistical information similar to the visualized statistical information based on the captured image of the camera (fixed camera 2 or built-in camera 4) to be set this time from the past visualized statistical information stored in the hard disk 32. Further, the parameter output unit 43 outputs a screen including the condition setting parameters (such as detection frames) for analyzing the captured image, which were applied to the past captured image from which the past visualized statistical information detected by the detection unit 42 was derived, to the display 50b of the personal computer 9. The display control unit 44 controls to display, on the display 50b of the personal computer 9, the visualized statistical information based on the captured image of the camera (to be set this time) (in the case of the example in FIG. 16, the face detection aggregation image) and a plurality of past visualized statistical information stored in the hard disk 32, as shown in FIG. 16, based on the instruction operation of the user of the personal computer 9.
[0032] The above parameter display control unit 45 reads the current condition setting parameters 49 (such as detection frames) for analyzing the captured image of the camera to be set this time from the hard disk 32, superimposes them on the visualized statistical information based on the captured image of the camera to be set this time, and controls to display them on the display 50b of the personal computer 9. In the example shown in FIG. 10, the parameter display control unit 45 superimposes the current condition setting parameters (detection frame) for analyzing the captured image of the camera to be set this time on the visualized statistical information (movement trajectory) based on the captured image of the camera to be set this time and displays it on the display 50b of the personal computer 9.
[0033] When the user uses the operation unit 50a of the personal computer 9 to perform adjustment input on the condition setting parameters displayed on the display 50b, the above-described reanalysis unit 46 re-analyzes the captured image of the camera to be set this time (received by the communication unit 36) using the condition setting parameters after adjustment (input). Further, the reanalysis result display control unit 47 controls to display the analysis result by the reanalysis unit 46 on the display 50b of the personal computer 9 used by the user.
[0034] Next, with reference to FIG. 6, the flow of generation and display of the analysis result superimposed video performed in this image analysis system 10 will be described. First, the analysis result superimposed video generation process on the analysis box 3 side will be described. The CPU 21 of the analysis box 3 decodes the video stream input from the fixed camera 2 using the VMS 51 in FIG. 6, and stores the decoded video (data) as collection data 52 (including personal information) in the hard disk 22. Next, the CPU 21 of the analysis box 3 uses the video analysis program 53 to divide the collection data 52 (including personal information) into frame images 55 (including personal information) and outputs them to the main memory 54 (corresponding to the RAM 23 in FIG. 3). Then, the CPU 21 of the analysis box 3 uses the video analysis program 53 to input each of the above frame images 55 into the (trained) DNN model 57 for various image recognition, and the DNN model 57 performs image recognition processing (processing such as human detection and attribute analysis) on each frame image 55 and outputs analysis data 58 including the results of the above recognition processing. The DNN calculation chip 56 (corresponding to the (inference) chips 24a to 24h in FIG. 3) is used for the image recognition processing by the above DNN model 57. Further, the above frame image 55 is immediately deleted from the main memory 54 after being input to the DNN model 57. As shown by the dashed line in FIG. 6, the video analysis program 53 includes the above DNN model 57.
[0035] The above analysis data 58 is text data indicating which camera, at which time frame image, at which position, and what kind of person is present, and does not contain personal information. This analysis data 58 includes the count results of visitors and passers-by obtained by the above video analysis program 53 using the person detection results by the above DNN model 57 and the condition setting parameters 59.
[0036] When the output process of the analysis data 58 using the above video analysis program 53 is completed, the CPU 21 of the analysis box 3 generates an analysis result superimposed video using the analysis superimposed video generation program 60. Specifically, the CPU 21 of the analysis box 3 reads the above (personal information-including) collection data 52 from the hard disk 22 using the VMS 51 according to the analysis superimposed video generation program 60, and generates an analysis result superimposed video, which is a video obtained by superimposing the recognition results (person detection results and attribute analysis results) by the DNN model 57 on this collection data 52. Then, the CPU 21 of the analysis box 3 stores the above analysis result superimposed video in the hard disk 22 as a part of the (personal information-including) collection data 52 using the VMS 51 according to the analysis superimposed video generation program 60. Therefore, the above collection data 52 includes the video (data) after decoding the video stream input from the fixed camera 2 and the above analysis result superimposed video. An example of the above analysis result superimposed video is the analysis result superimposed video 63 shown in FIG. 8. This analysis result superimposed video 63 is a video obtained by superimposing the person detection results by the DNN model 57 on the collection data 52 (video data of the entrance camera).
[0037] The analysis result superimposed video generation process on the signboard 5 side is basically the same as the analysis result superimposed video generation process on the analysis box 3 side. The collected data 65, video analysis program 66, main memory 67, frame image 68, DNN calculation chip 69, DNN model 70, analysis data 71, condition setting parameter 73, and analysis superimposed video generation program 72 on the signboard 5 side respectively correspond to the collected data 52, video analysis program 53, main memory 54, frame image 55, DNN calculation chip 56, DNN model 57, analysis data 58, condition setting parameter 59, and analysis superimposed video generation program 60 on the analysis box 3 side. Also, the storage 20 on the signboard 5 side serves the same role as the hard disk 22 on the analysis box 3 side. Although not described in FIGS. 6 and 7, the signboard 5 also has a program with the same function as the VMS 51 on the analysis box 3 side.
[0038] Next, the processing on the management server 1 side and the personal computer 9 side in FIG. 6 will be described. The management server 1 has an analysis result superimposed video display program 61 and a parameter adjustment program 62. These programs are included in the server control program 37 in FIG. 4. When a user such as an operation administrator on the personal computer 9 side (a user of the portal 38 shown in FIG. 4) uses the operation unit 50a of the personal computer 9 to instruct the management server 1 to display the above analysis result superimposed video, the CPU 31 of the management server 1 uses the above analysis result superimposed video display program 61 to access the analysis box 3 and uses the VMS 51 to read the analysis result superimposed video included in the collected data 52 from the hard disk 22 of the analysis box 3, or accesses the signage 5 and reads the analysis result superimposed video included in the collected data 65 from the storage 20 of the signage 5. Then, a screen including this analysis result superimposed video (for example, the current count status confirmation screen 81 shown in FIG. 8) is output (displayed) to the display 50b of the personal computer 9. As a result, a user such as an operation administrator can view a video (analysis result superimposed video) in which the analysis result is superimposed on the video data (collected data) input from the camera (fixed camera 2 or built-in camera 4) and confirm the current analysis status. Also, although details will be described later, when a user such as an operation administrator uses the operation unit 50a of the personal computer 9 to perform a parameter adjustment input operation on the management server 1, the CPU 31 of the management server 1 uses the above parameter adjustment program 62 to rewrite the condition setting parameter 59 on the analysis box 3 side or the condition setting parameter 73 on the signage 5 side. As a result, a user such as an operation administrator (including workers of the system development outsourcing company and the system operation outsourcing company) can easily adjust the condition setting parameters (condition setting parameter 59 or condition setting parameter 73) of the camera to be set this time for analysis improvement.
[0039] Next, with reference to FIG. 7, the statistical processing of the image analysis results (analysis data) performed in this image analysis system 10 and the process of viewing and editing the detection conditions (condition setting parameters) superimposed on the video (visualized statistical information or analysis result superimposed video) will be described. The processing program 76, video search and display program 78, detection condition data extraction program 79, and program 80 for viewing and editing the detection conditions superimposed on the video (hereinafter abbreviated as the "detection condition viewing and editing program") in FIG. 7 are all programs included in the portal 38 shown in FIG. 4.
[0040] First, the statistical processing of the above image analysis results will be described. The CPU 31 of the management server 1 uses the above processing program 76 to perform processing (aggregation and statistical processing) on the analysis data 58 or analysis data 71 (in FIG. 7, "video metadata 74" or "video metadata 75") including the recognition processing results by the DNN model 57 or DNN model 70 described in the above description of FIG. 6, and stores the processing result in the database 77 in the hard disk 32 (see FIG. 4) as statistical data 40. Note that for the video metadata 74 on the analysis box 3 side and the video metadata 75 on the signage 5 side, the data after being processed and aggregated (summarized) for each person, rather than the analysis (recognition) results themselves in frame image units, are sent to the management server 1 and become the targets of processing by the processing program 76. The above statistical data 40 is, for example, data such as the number of customers visiting a certain store at 9 o'clock and the gender and age distribution, and even when combined with the collected data 52 and 65, it is non-personal information that cannot identify an individual. As shown in FIG. 5, this statistical data 40 includes visualized statistical information 48 and the condition setting parameters 49 (detection frames, etc.) of each camera (fixed camera 2 or built-in camera 4) on the edge side.
[0041] In this image analysis system 10, as described above, personal information (collection data 52 and collection data 65) is stored only on the edge side (analysis box 3 and signage 5), and only non-personal statistical data 40 (including condition setting parameters 49) is stored on the cloud side (management server 1). This enables the operation of the cloud side (management server 1) to be entrusted to operators other than the store operation company (such as system development outsourcing contractors and system operation outsourcing contractors).
[0042] Next, the process of viewing and editing the detection conditions (condition setting parameters) superimposed on the video in this image analysis system 10 will be described. When a user such as an operation administrator uses the operation unit 50a of the personal computer 9 to give an instruction operation for viewing and editing the detection conditions (condition setting parameters such as detection frames) to the management server 1, the CPU 31 of the management server 1 first uses the video search and display program 78 to search for the videos (including analysis result superimposed videos) of the past collection data 52 and 65 stored in the analysis box 3 and the signage 5, and the visualization statistical information included in the statistical data 40 stored in the database 77, and extracts the videos (including visualization statistical information) of the search results. Next, the CPU 31 of the management server 1 uses the detection condition data extraction program 79 to extract the current detection conditions (condition setting parameters 49 such as detection frames) applied to the extracted search result videos from the statistical data 40. Then, the CPU 31 of the management server 1 uses the detection condition viewing and editing program 80 to display a video or image with the above-extracted current detection conditions superimposed on the above search result videos (including visualization statistical information) on the display 50b of the personal computer 9, and accepts the editing operation of the detection conditions (such as detection frames) using the operation unit 50a of the personal computer 9 by the user.
[0043] Next, with reference to FIGS. 8 to 16, the user interface employed in this image analysis system 10 will be described. FIG. 8 shows the current count status confirmation screen 81. This screen is a screen including the analysis result superimposed video 63, which is a kind of the analysis result superimposed video described in the explanation of FIG. 6. This analysis result superimposed video 63 is a video obtained by superimposing the person detection result by the DNN model 57 in FIG. 6 on the video data from 9:00 to 10:00 captured by the entrance camera of a certain store (corresponding to the collected data 52 in FIG. 6). In the analysis result superimposed video 63, a bounding box 82 corresponding to each person detected by the DNN model 57 is displayed. Further, in the count result display column 81a on the current count status confirmation screen 81, the store visitor count result based on the current store visitor count condition of the above entrance camera is displayed. By using the current count status confirmation screen 81 as described above, the detection status and count result of store visitors and the like based on the current count condition can be visualized in an easy-to-understand manner.
[0044] In addition, the user interface adopted in this image analysis system 10 includes a screen for simulating the setting of condition setting parameters (such as detection frames). FIG. 9 is a flowchart of the simulation process for setting condition setting parameters (such as detection frames) using this simulation screen. FIG. 10 shows an example of the above simulation screen, the customer count area setting simulation screen 83. In this customer count area setting simulation screen 83, the detection frame (set detection frame 85) of the customer count area corresponds to the condition setting parameters in FIG. 9. On this customer count area setting simulation screen 83, the set detection frame 85 (a type of condition setting parameter) of the current customer count area of the camera to be set this time is superimposed and displayed on the visualization statistical information of a predetermined time period of the camera to be set this time (the movement trajectory aggregation image 84, face detection aggregation image 89 (see FIG. 13), or gender and age recognition aggregation image shown in FIG. 10). Here, in the example of FIG. 10, the above movement trajectory aggregation image 84 is an image showing the movement trajectory created based on the person detection results for the captured images of the entrance camera from 9:00 to 10:00. Also, in the example of FIG. 13, the above face detection aggregation image 89 is a heat map created based on the face detection results (to be exact, the face detection frequency in each area within the captured image) for the captured images of the entrance camera of store XX from 9:00 to 10:00, and the above gender and age recognition aggregation image is a heat map created based on the gender and age recognition results (to be exact, the gender and age recognition frequency in each area within the captured image) for the captured images of the camera to be set this time in a predetermined time period. Note that in the above face detection aggregation image 89 (see FIG. 13) and gender and age recognition aggregation image, the background image (the image obtained by removing moving objects (people) from the captured image) in the original captured image is displayed on the back side of the heat map.
[0045] Taking the case of the customer count area setting simulation screen 83 in Fig. 10 as an example, the simulation process of setting condition setting parameters (such as detection frames) using the above simulation screen will be described according to the flowchart in Fig. 9. When the user uses the operation unit 50a of the personal computer 9 to instruct the management server 1 to display the customer count area setting simulation screen 83, the CPU 31 of the management server 1 displays the customer count area setting simulation screen 83. Then, when the user operates the camera selection button 83a, time zone selection button 83b, and statistical information selection button 83c in Fig. 10 using the operation unit 50a (such as a mouse) of the personal computer 9 to select the camera to be set (analyzed), the acquisition time zone of the inference result (analysis result), and the type of visualized statistical information to be displayed (S1), the CPU 31 of the management server 1 superimposes the current condition setting parameters (set detection frame 85) of the selected camera on the visualized statistical information corresponding to the selected camera, acquisition time zone, and type of visualized statistical information (in the example of Fig. 10, the movement trajectory aggregation image 84), and displays it on the display 50b of the personal computer 9. At the same time, the analysis result (customer count result) by the current set detection frame 85 is displayed on the display 50b (count result display column 83h) (S2). Note that the CPU 31 of the management server 1 extracts the visualized statistical information corresponding to the selected camera, acquisition time zone, and type of visualized statistical information, and the current condition setting parameters (set detection frame 85) of the selected camera from the above statistical data 40 (see Fig. 7, etc.) in the process of the above S2.
[0046] When the process of S2 is completed, a user such as an operation manager checks the current condition setting parameters (setting detection frame 85) displayed on the display 50b and the analysis result using these condition setting parameters (visitor count result using the current setting detection frame 85) (S3). As a result of this check, when the user makes an adjustment input to the condition setting parameters (setting detection frame 85) using the operation unit 50a (such as a mouse) of the personal computer 9 (YES in S4), the CPU 31 of the management server 1 acquires the past captured images of the camera to be set this time (captured images corresponding to the camera selected in S1 and the acquisition time zone) stored in the analysis box 3 and the signage 5 (S5), and re-analyzes the captured images of the camera to be set this time acquired in S5 using the condition setting parameters (setting detection frame 85) after the adjustment input in S4 (performs re-analysis) (S6). Then, the visitor count result of the re-analysis result is displayed on the display 50b (count result display column 83h) of the personal computer 9 (S7).
[0047] The method of adjusting and inputting the condition setting parameters (setting detection frame 85) in S4 will be described more specifically. First, the user can use a mouse or the like of the personal computer 9 to drag the vertex 85a of the setting detection frame 85 on the visitor count area setting simulation screen 83, thereby adjusting the setting detection frame 85 (the frame of the setting detection area (visitor count area)). Also, the user can use a keyboard or the like of the personal computer 9 to directly input coordinates into the detection area adjustment input field 83f on the visitor count area setting simulation screen 83, thereby also adjusting the detection area, that is, adjusting the setting detection frame 85. Further, on the movement trajectory aggregation image 84 of the visitor count area setting simulation screen 83. In addition to the current setting detection frame 85 of the camera to be set this time, a recommended detection frame 86 is displayed. When the user uses a mouse or the like of the personal computer 9 to click the recommended detection area setting button 83g on the visitor count area setting simulation screen 83, the x-coordinate and y-coordinate values of each vertex 86a of the recommended detection frame 86 are input into the input fields of each vertex of the above detection area adjustment input field 83f. After the adjustment and input of the setting detection frame 85 by various methods as described above, when the user uses a mouse or the like of the personal computer 9 to click the calculation button 83i, the visitor count result based on the visitor count condition (including the setting detection frame 85) after the adjustment and input for the camera to be set this time is displayed in the count result display column 81h.
[0048] On the above customer arrival count area setting simulation screen 83, the user can also adjust and input the count type and moving direction among the customer arrival count conditions by operating the count type selection button 83d and the moving direction selection button 83e using a mouse or the like of the personal computer 9. After the user uses a mouse or the like of the personal computer 9 to adjust and input the customer arrival count conditions (setting detection frame 85 (setting detection area), count type, and moving direction) for the camera to be set this time, if the user clicks the calculation button 83i and is satisfied with the count result (accuracy represented) displayed in the count result display column 81h, the user can click the apply button 83j to apply the customer arrival count conditions after the above adjustment and input. By clicking the above apply button 83j, the customer arrival count conditions (setting detection frame (setting detection area), count type, and moving direction) for the camera to be set this time included in the statistical data 40 in the database 77 in FIG. 7 are rewritten.
[0049] In addition, the user interface adopted in this image analysis system 10 has a screen for displaying the ranking of the accuracy of various counts (for example, counting of customers, passers-by, and stationary persons). FIG. 11 shows a customer arrival count accuracy ranking display screen 87 which is an example of the screen for displaying the ranking of the above count accuracy. This customer arrival count accuracy ranking display screen 87 is a screen that displays the ranking of the setting conditions for customer arrival count set in the past (particularly, the setting detection frame) in which the count accuracy in the setting conditions is high. On this customer arrival count accuracy ranking display screen 87, in addition to the accuracy, detection conditions, and (detection) uses for each condition, the setting detection frames 87a and 87b are displayed. In this image analysis system 10, by using a screen for displaying the ranking of the count accuracy such as the customer arrival count accuracy ranking display screen 87, among the setting conditions for (human) counting set in the past, the setting conditions with high count accuracy under the same detection conditions and uses can be referred to (viewed).
[0050] Next, with reference to the flowchart of FIG. 12, a method for setting condition setting parameters such as a setting detection frame to be applied to the captured image of the camera to be set this time by using visualization statistical information based on the captured images of the cameras that have been set in the past will be described. When the CPU 31 of the management server 1 receives, via the communication unit 36, from the edge side (analysis box 3 or signage 5), a superimposed video in which the captured image of the camera and the result of the recognition process for the frame image included in the captured image are superimposed (S11), it uses the visualization statistical information generation unit 41 (see FIG. 5) to statistically process the superimposed video for a predetermined period of time to generate visualization statistical information (S12). Then, as shown in FIG. 5, the CPU 31 of the management server 1 stores the generated visualization statistical information 48 in the hard disk 32 as part of the statistical data 40 (S13).
[0051] When the user searches for past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from among the past visualization statistical information stored in the hard disk 32 of the management server 1, for example, the similar visualization statistical information search screen 88 shown in FIG. 13 is used. When the user displays the above-mentioned similar visualization statistical information search screen 88 on the display 50b of the personal computer 9 and uses the operation unit 50a of the personal computer 9 to select the store selection button 88a, the camera selection button 88b, the time selection button 88c, and the visualization statistical information selection button 88d to select the type of desired visualization statistical information based on the captured image of the camera to be set this time in a desired time zone, the CPU 31 of the management server 1 displays on the display 50b of the personal computer 9 the visualization statistical information corresponding to the selected store, camera, time zone, and type of visualization statistical information (in the example of FIG. 13, the face detection aggregated image 89). This face detection aggregated image 89 is, in the example of FIG. 13, a heat map created based on the face detection results (more precisely, the face detection frequencies in each area within the captured image) for the frame images included in the captured images of the entrance camera of store XX from 9:00 to 10:00.
[0052] In a state where visualization statistical information such as the face detection aggregated image 89 is displayed on the above-described similar visualization statistical information search screen 88, when the user searches for past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time, the operation unit 50a (such as a mouse) of the personal computer 9 is used to click the similar visualization statistical information search button 88e on the similar visualization statistical information search screen 88. By this click operation, when the user instructs to search for similar past visualization statistical information (YES in S14), the CPU 31 (the detection unit 42 (see FIG. 5)) of the management server 1 detects (searches) past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from the past visualization statistical information stored in the hard disk 32 (S15).
[0053] There are roughly two methods for searching for the above-described similar visualization statistical information. The first method is a method in which the CPU 31 (the detection unit 42) of the management server 1 automatically selects the past visualization statistical information most similar to the visualization statistical information based on the captured image of the camera to be set this time and presents it to the user. The second method is a method in which the CPU 31 of the management server 1 arranges and displays the visualization statistical information based on the captured image of the camera to be set this time and a plurality of past visualization statistical information stored in the hard disk 32 on the display 50b of the personal computer 9, and the user uses the operation unit 50a (such as a mouse) of the personal computer 9 to select the past visualization statistical information most similar to the visualization statistical information based on the captured image of the camera to be set this time from among the plurality of past visualization statistical information displayed on the display 50b.
[0054] As an example of the first method above (the method in which the CPU 31 of the management server 1 automatically detects visualization statistical information similar to the visualization statistical information of the camera to be set this time), in the case of the face detection aggregation image 89 in FIG. 13, the CPU 31 of the management server 1 automatically detects from the past face detection aggregation images stored in the hard disk 32 a face detection aggregation image with a heat map pattern similar to the pattern of the heat map in the face detection aggregation image 89 currently displayed on the similar visualization statistical information search screen 88, and presents it to the user. Specifically, for example, when the user clicks the similar visualization statistical information search button 88e on the similar visualization statistical information search screen 88 shown in FIG. 13 using the operation unit 50a (such as a mouse) of the personal computer 9, the CPU 31 (detection unit 42) of the management server 1, as shown in FIG. 14, compares the vector 92 corresponding to the moving direction of people (from the upper left to the lower right) obtained from the face detection aggregation image 89 (the same as the face detection aggregation image 89 in FIG. 13) of the camera to be set this time (based on the photographed image) with the vector corresponding to the moving direction of people obtained from the past face detection aggregation images stored in the hard disk 32, and detects a past face detection aggregation image 91 having a vector 93 similar to the vector 92 obtained from the face detection aggregation image 89 of the camera to be set this time. Then, the CPU 31 of the management server 1 displays the detected (similar) past face detection aggregation image 91 on the similar visualization statistical information search result display screen 90 side by side with the face detection aggregation image 89 of the camera to be set this time described above.
[0055] As a method for detecting a past face detection aggregation image 91 having a vector 93 similar to the vector 92 of the face detection aggregation image 89 of the camera to be set this time, a method can be considered in which the inner product of the vector 92 of the face detection aggregation image 89 of the camera to be set this time and the vector 93 of the past face detection aggregation image 91 is obtained, and the past face detection aggregation image 91 corresponding to the vector 93 when the value (scalar quantity) of this inner product is large is selected.
[0056] On the above-described similar visualization statistical information search result display screen 90, when the user clicks (selects) a past face detection aggregated image 91 (which has a vector 93 similar to the face detection aggregated image 89 of the camera to be set this time) using the operation unit 50a (such as a mouse) of the personal computer 9, the CPU 31 of the management server 1 reads the condition setting parameters for photograph image analysis (mainly, the coordinates of each vertex of the detection frame) applied to the past photographed image that was the source of the selected past face detection aggregated image 91 from the statistical data 40 on the hard disk 32. Then, as shown in FIG. 15, the CPU 31 (parameter output unit 43) of the management server 1 outputs (displays) a screen (similar pattern image setting display screen 94) including the above-described past face detection aggregated image 91 (the same as the face detection aggregated image 91 in FIG. 14) and the condition setting parameters (the coordinates of each vertex of the detection frame displayed in the detection area display column 94f) applied to the past photographed image that was the source of this past face detection aggregated image 91 to the display 50b of the personal computer 9 (S16 in FIG. 12).
[0057] The user can, as described above, refer to (browse) the condition setting parameters (the coordinates of each vertex of the detection frame) applied to the photographed image that was the source of the past face detection aggregated image 91 (which has a vector 93 similar to the face detection aggregated image 89 of the camera to be set this time) displayed on the display 50b of the personal computer 9, thereby knowing the values of the condition setting parameters suitable for the photographed image of the camera to be set this time and making an input.
[0058] In addition, on the similar pattern image setting display screen 94 shown in FIG. 15, in addition to the above-described detection area display column 94f, a store display column 94a, a camera display column 94b, a time zone display column 94c, a count type display column 94d, a movement direction display column 94e, and a count result display column 94g are provided. In the store display column 94a and the camera display column 94b, information indicating which store's and which camera took the captured image that was the basis of the above-described past face detection aggregated image 91 is displayed. In the time zone display column 94c, the time zone when the captured image that was the basis of the above-described past face detection aggregated image 91 was taken is displayed. Also, in the count type display column 94d and the movement direction display column 94e, the count type (the type of what kind of state of people to count) and the movement direction (information on what movement direction of people to count) applied to the captured image that was the basis of the above-described past face detection aggregated image 91 are displayed. Further, in the count result display column 94g, using the customer count conditions displayed in the count type display column 94d, the movement direction display column 94e, and the detection area display column 94f, the number of people as the result of performing the customer count process on the captured image that was the basis of the above-described past face detection aggregated image 91 is displayed.
[0059] Next, a specific example of the above second method (a method of arranging and displaying the visualization statistical information of the camera to be set this time and the visualization statistical information of a plurality of past times on the display 50b of the personal computer 9, and allowing the user to select the past visualization statistical information most similar to the visualization statistical information of the camera to be set this time) will be described. Also in the case of adopting this method, the user uses the operation unit 50a (such as a mouse) of the personal computer 9 to select the store selection button 88a, the camera selection button 88b, the time selection button 88c, and the visualization statistical information selection button 88d on the similar visualization statistical information search screen 88 shown in FIG. 13, and selects the type of desired visualization statistical information based on the captured images of the camera to be set this time in a predetermined time period. Then, the CPU 31 of the management server 1 displays the visualization statistical information corresponding to the selected store, camera, time period, and type of visualization statistical information (in the example of FIG. 13, the face detection aggregated image 89) on the display 50b of the personal computer 9. In this state, when the user clicks the similar visualization statistical information search button 88e on the similar visualization statistical information search screen 88 using a mouse or the like of the personal computer 9, the CPU 31 (display control unit 44) of the management server 1, as shown in the similar visualization statistical information selection screen 97 of FIG. 16, arranges and displays the face detection aggregated image 89 based on the captured image of the camera to be set this time and a plurality of past face detection aggregated images 96a to 96i stored in the hard disk 32 on the display 50b (candidate image display area 95) of the personal computer 9. Note that in the face detection aggregated images 96a to 96i displayed in the candidate image display area 95 of FIG. 16, only the heat map part is shown by omitting the background image part in these images.
[0060] Among the plurality of past face detection aggregated images 96a to 96i displayed in the candidate image display area 95 described above, it is preferable that they are those with the same camera, usage, and detection conditions as the current setting target among the plurality of past face detection aggregated images stored in the hard disk 32. Further, among the plurality of past face detection aggregated images 96a to 96i displayed in the candidate image display area 95 described above, it is desirable that they are face detection aggregated images of a heat map in a pattern close to the face detection aggregated image 89 based on the captured image of the camera that is the current setting target among the plurality of past face detection aggregated images stored in the hard disk 32.
[0061] When the user uses the operation unit 50a (such as a mouse) of the personal computer 9 to click (select) the past face detection aggregated image that is most similar to the face detection aggregated image 89 based on the captured image of the camera that is the current setting target among the plurality of past face detection aggregated images 96a to 96i displayed in the candidate image display area 95 described above, the CPU 31 of the management server 1 performs the same processing as when a past face detection aggregated image 91 is clicked on the similar visualization statistical information search result display screen 90 shown in FIG. 14 above. That is, the CPU 31 of the management server 1 reads the condition setting parameters for photograph image analysis (coordinates of each vertex of the detection frame) applied to the past captured image that is the source of the selected past face detection aggregated image (any one of 96a to 96i) from the statistical data 40 of the hard disk 32, and as shown in FIG. 15, the past face detection aggregated image 91 selected in FIG. 16 and the condition setting parameters (coordinates of each vertex of the detection frame displayed in the detection area display column 94f) applied to the past captured image that is the source of this past face detection aggregated image 91 are included in a screen (similar pattern image setting display screen 94) and output (displayed) to the display 50b of the personal computer 9.
[0062] As described above, according to the image analysis system 10, the management server 1, and the server control program 37 of the present embodiment, past visualization statistical information (such as the face detection aggregation image 91, etc.) similar to the visualization statistical information (such as the face detection aggregation image 89, etc.) based on the captured image of the camera to be set this time (hereinafter abbreviated as "the current camera") (fixed camera 2 or built-in camera 4) is detected, and the captured image analysis condition setting parameters (coordinates of each vertex of the detection frame (detection area) shown in FIG. 15, etc.) applied to the past captured image that is the source of the detected past visualization statistical information are included in a screen (similar pattern image setting display screen 94 in FIG. 15) and output (displayed) on the display 50b of the personal computer 9 so that a user such as an operation administrator of the system can view it. Here, the past visualization statistical information similar to the visualization statistical information based on the captured image of the current camera is likely to be visualization statistical information using the result of performing the same recognition process as the recognition process performed on the captured image of the current camera for the captured image (frame image in it) of the movement range of the same people as the captured image of the current camera. Therefore, by referring to the condition setting parameters (for example, coordinates of each vertex of the detection frame (detection area) shown in FIG. 15, etc.) applied to the captured image that is the source of the above-viewed (similar to the visualization statistical information based on the captured image of the current camera) past visualization statistical information (for example, the past face detection aggregation image 91) and inputting the condition setting parameters for the captured image analysis of the current camera, a user such as an operation administrator of the system can easily input appropriate condition setting parameters (for example, coordinates of each vertex of the person detection frame (detection area), etc.).
[0063] Also, according to the image analysis system 10 of the present embodiment, the condition setting parameters for analyzing the current captured image of the current (camera to be set) (for example, the setting detection frame 85 of a person in FIG. 10) are superimposed on the visualization statistical information based on the captured image of the current camera (for example, the movement trajectory aggregation image 84 in FIG. 10) and displayed on the display 50b of the personal computer 9, so that a user such as an operation administrator of the system can perform an adjustment input of the displayed condition setting parameters. As a result, the user can perform an adjustment input of the condition setting parameters while looking at the current condition setting parameters (such as the setting detection frame 85) superimposed on the visualization statistical information (such as the movement trajectory aggregation image 84, etc.) based on the captured image of the current camera. Therefore, for the current camera, appropriate condition setting parameters can be easily input.
[0064] Also, according to the image analysis system 10, the management server 1, and the server control program 37 of the present embodiment, as shown in FIG. 14, the CPU 31 (detection unit 42) of the management server 1 compares the vector 92 corresponding to the moving direction of a person obtained from the visualization statistical information (for example, the face detection aggregation image 89) based on the captured image of the current camera to be set with the vector 93 corresponding to the moving direction of a person obtained from the past visualization statistical information (for example, the past face detection aggregation image 91) stored in the hard disk 32, and detects the past visualization statistical information similar to the visualization statistical information based on the captured image of the current camera to be set. As a result, the past visualization statistical information (such as the face detection aggregation image 91, etc.) similar to the visualization statistical information (such as the face detection aggregation image 89, etc.) based on the captured image of the current camera to be set can be easily detected.
[0065] Further, according to the image analysis system 10 of the present embodiment, the CPU 31 (display control unit 44) of the management server 1 displays visualization statistical information (for example, the face detection aggregation image 89 in FIG. 16) based on the captured image of the camera to be set this time and a plurality of past visualization statistical information (for example, the face detection aggregation images 96a to 96i in FIG. 16) stored in the hard disk 32 on the display 50b of the personal computer 9, so that the user can use a mouse or the like of the personal computer 9 to select, from among the plurality of past visualization statistical information (for example, the face detection aggregation images 96a to 96i) displayed on the display 50b, the past visualization statistical information most similar to the visualization statistical information (for example, the face detection aggregation image 89) based on the captured image of the camera to be set this time. Thereby, it becomes possible to search for the past visualization statistical information most similar to the visualization statistical information based on the captured image of the camera to be set this time with a simple configuration.
[0066] Also, according to the image analysis system 10 of the present embodiment, after the user adjusts using the operation unit 50a of the personal computer 9, the condition setting parameter (for example, the (person's) setting detection frame 85 in FIG. 10) is used to re-analyze the captured image of the camera to be set this time (re-analyze), and the analysis result after the re-analysis (for example, the customer count result in the customer count result display column 83h in FIG. 10) is displayed on the display 50b of the personal computer 9. Thereby, for the camera to be set this time, it is possible to simulate the difference in the analysis result (for the captured image of this camera) according to the adjustment content of the condition setting parameter. Therefore, the user can easily confirm the accuracy of the condition setting (for example, the setting detection frame 85 in FIG. 10) set by the adjustment input of the condition setting parameter.
[0067] Modification example: Note that the present invention is not limited to the configurations of the above-described embodiments, and various modifications are possible without changing the gist of the invention. Next, modification examples of the present invention will be described.
[0068] Modification example 1: In the above embodiment, on the simulation screen of the condition setting parameters (for example, the customer count area setting simulation screen 83 in FIG. 10), the current condition setting parameters of the camera to be set this time (for example, the setting detection frame 85 in FIG. 10) are superimposed on the visualization statistical information based on the captured image of the camera to be set this time (for example, the moving trajectory aggregation image 84 in FIG. 10), and then displayed on the display 50b of the personal computer 9. However, on the simulation screen of the condition setting parameters, (1) the current condition setting parameters (such as the setting detection frame) of the camera to be set this time may be superimposed on the captured images (videos) of the camera to be set this time for a predetermined time and then displayed on the display of the personal computer, or (2) the current condition setting parameters (such as the setting detection frame) of the camera to be set this time may be superimposed on the analysis result superimposed video (for example, the analysis result superimposed video 63 in FIG. 8) of the camera to be set this time for a predetermined time and then displayed on the display of the personal computer.
[0069] Also, in the above embodiment, on the simulation screen of the condition setting parameters (for example, the customer count area setting simulation screen 83 in FIG. 10), the current condition setting parameters of the camera to be set this time (for example, the setting detection frame 85 in FIG. 10) are superimposed on the moving trajectory aggregation image 84 based on the captured image of the camera to be set this time, and then displayed on the display 50b of the personal computer 9. However, on the simulation screen of the condition setting parameters, the visualization statistical information on which the current condition setting parameters of the camera to be set this time are superimposed is not limited to the above moving trajectory aggregation image, and may be a face detection aggregation image 89 as shown in FIG. 13, or the above gender and age recognition aggregation image (a heat map generated by statistically processing each superimposed image obtained by superimposing each frame image included in the captured images of the camera to be set this time for a predetermined time period and the gender and age recognition results for these respective frame images).
[0070] Modification Example 2: In the above-described embodiment, as shown in FIGS. 13 to 16, an example was shown in which past visualization statistical information similar to the face detection aggregated image 89 based on the captured image of the camera to be set this time was detected (searched) from the past face detection aggregated images stored in the hard disk 32. However, it is also possible to detect (search) a past movement trajectory aggregated image having a pattern similar to the pattern of the movement trajectory in the movement trajectory aggregated image (based on the captured image of the camera to be set this time) from the past movement trajectory aggregated images stored in the hard disk 32, or to detect (search) a past gender and age recognition aggregated image having a pattern similar to the pattern of the heat map in the past gender and age recognition aggregated image stored in the hard disk 32.
[0071] Modification Example 3: In the above-described embodiment, an example was shown in which the "condition setting parameter for analyzing a captured image" in the claims is a detection frame (or the coordinates of each vertex of the detection frame) of a person. However, the condition setting parameter (for analyzing a captured image) in the present invention is not limited to this, and may be, for example, the reliability of object detection or object recognition (for example, person detection or face detection of a person) for the captured image. For example, by setting the condition setting parameter to the reliability of person detection, the amount (range of analysis data) of the analysis data (corresponding to the video metadata 74 in FIG. 7) sent to the management server 1 on the cloud C side can be adjusted by adjusting this reliability.
[0072] Modification Example 4: In the above embodiment, an example in which the image analysis system 10 includes only the management server 1 on the cloud C has been shown. However, the configuration of the image analysis system is not limited to this. For example, the image analysis system may include an AI analysis server on the cloud in addition to the management server. The AI analysis server analyzes the behavior of people in each store based on the object recognition results from, for example, the analysis box and the signage, and converts the information of the analysis results into data that is easy to use for various applications such as marketing and security, and outputs it. Further, the image analysis system may include a management server for managing the fixed camera and the analysis box and a signage management server for managing the signage on the cloud.
[0073] Modification Example 5: In the above embodiment, an example in which the image analysis system 10 includes the fixed camera 2, the analysis box 3, and the signage 5 in the store S has been shown. However, the image analysis system may include only the fixed camera and the analysis box in the store and not include the signage, or may include only the signage in the store and not include the fixed camera and the analysis box.
[0074] Modification Example 6: In the above embodiment, an example in which the CPU 31 of the management server 1 includes the visualization statistical information generation unit 41 has been shown. However, the analysis box and the signage may include the visualization statistical information generation unit to generate the visualization statistical information.
[0075] Further, in the above embodiment, an example in which the display means in the claims is the display 50b of the personal computer 9 and the selection input means and the adjustment input means in the claims are mainly the operation unit 50a of the personal computer 9 has been shown. However, the display means in the claims may be the display of the management server, and the selection input means and the adjustment input means in the claims may be the operation unit 35 of the management server.
Explanation of Reference Numerals
[0076] 1 Management server (server) 2 Fixed camera (part of the edge-side camera) 3 Analysis box (part of the edge-side image analysis device) 4 Built-in camera (part of the edge-side camera) 5 Signage (part of the edge-side image analysis device) 10 Image analysis system (parameter viewing system, parameter adjustment system) 32 Hard disk (statistical information storage means) 36 Communication unit (receiving means) 37 Server control program 41 Visualization statistical information generation unit (visualization statistical information generation means) 42 Detection unit (detection means) 43 Parameter output unit (parameter output means) 44 Display control unit (display control means) 45 Parameter display control unit (parameter display control means) 46 Re-analysis unit (re-analysis means) 47 Re-analysis result display control unit (re-analysis result display control means) 48 Visualization statistical information 49 Condition setting parameter 50a Operation unit (adjustment input means, selection input means) 50b Display (display means) 92 Vector (vector corresponding to the moving direction of people obtained from the visualization statistical information based on the captured image of the camera to be set this time) 93 Vector (vector corresponding to the moving direction of people obtained from the past visualization statistical information)
Claims
1. A display means, a visualization statistical information generation means for statistically processing a superimposed video obtained by superimposing a captured image of a camera on the edge side and a result of recognition processing by an image analysis device on the edge side for a frame image included in the captured image, and generating visualization statistical information; a statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means; a detection means for detecting past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from among the past visualization statistical information stored in the statistical information storage means; a parameter output means for outputting to the display means a screen including condition setting parameters for analyzing a captured image, which were applied to the past captured image that was the basis of the past visualization statistical information detected by the detection means A parameter viewing system comprising the same.
2. The detection means compares a vector corresponding to the moving direction of a person obtained from the visualization statistical information based on the captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from the past visualization statistical information stored in the statistical information storage means, and detects past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time. The parameter viewing system according to Claim 1, characterized in that.
3. The detection means includes display control means for controlling to display on the display means the visualization statistical information based on the captured image of the camera to be set this time and a plurality of past visualization statistical information stored in the statistical information storage means, The parameter viewing system further includes a selection input means for selecting, from among the plurality of past visualization statistical information displayed on the display means, the past visualization statistical information most similar to the visualization statistical information based on the captured image of the camera to be set this time. The parameter viewing system according to Claim 1, characterized in that.
4. a visualization statistical information generation means for statistically processing a superimposed video obtained by superimposing a captured image of a camera on the edge side and a result of recognition processing by an image analysis device on the edge side for a frame image included in the captured image, and generating visualization statistical information; a statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means; A detection means for detecting past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from among the past visualization statistical information stored in the statistical information storage means; A parameter output means for outputting to the display means a screen including condition setting parameters for analyzing a captured image, which were applied to the past captured image that was the basis of the past visualization statistical information detected by the detection means; A server comprising the above.
5. The detection means compares a vector corresponding to the moving direction of a person obtained from the visualization statistical information based on the captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from the past visualization statistical information stored in the statistical information storage means, and detects past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time. The server according to claim 4, characterized in that.
6. A server control program for controlling a server so as to perform a process for finding condition setting parameters for analyzing a captured image suitable for the captured image of the camera to be set this time, The server is A visualization statistical information generation means for statistically processing a superimposed video in which a captured image of an edge-side camera and a result of recognition processing of a frame image included in the captured image by an edge-side image analysis device are superimposed to generate visualization statistical information; A statistical information storage means for storing the visualization statistical information generated by the visualization statistical information generation means; A detection means for detecting past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time from among the past visualization statistical information stored in the statistical information storage means; A server control program for causing the server to function as a parameter output means for outputting to the display means a screen including condition setting parameters for analyzing a captured image, which were applied to the past captured image that was the basis of the past visualization statistical information detected by the detection means.
7. The detection means compares a vector corresponding to the moving direction of a person obtained from the visualization statistical information based on the captured image of the camera to be set this time with a vector corresponding to the moving direction of a person obtained from the past visualization statistical information stored in the statistical information storage means, and detects past visualization statistical information similar to the visualization statistical information based on the captured image of the camera to be set this time. The server control program according to claim 6, characterized in that.
Citation Information
Patent Citations
Activity status analysis device, activity status analysis system and activity status analysis method
JP2016076893A
Surveillance camera and detection method
JP6573297B1