System, control device and program
The system addresses the challenge of tracking a subject across imaging devices with different positions and directions by using feature information generation and comparison, ensuring accurate subject recognition and control.
Patent Information
- Application Number
- JP2024079718
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-09-07
- Filing Date
- 2024-05-15
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-05-15
AI Technical Summary
Existing tracking systems using multiple image capture devices struggle to identify a specific subject when the devices have different imaging positions and directions, leading to decreased similarity and difficulty in recognizing the same subject across devices.
A system comprising a first imaging device with a wide-angle view and a second imaging device with variable view, controlled by a first and second control device, which uses feature information generation and comparison to track a subject across devices with different positions and directions.
Enables effective tracking of a specific subject using multiple imaging devices with diverse setups by maintaining similarity in feature information, allowing accurate subject recognition and control.
Smart Images

Figure 0007805396000013 
Figure 0007805396000014 
Figure 0007805396000015
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system for tracking a specific subject using a plurality of image capturing devices with different image capturing positions and directions. [Background technology]
[0002] There is a technology that allows a specific subject to be tracked using an imaging device that can automatically control the pan / tilt / zoom (PTZ) from a remote location. This automatic tracking control automatically controls the PTZ so that the subject to be tracked is positioned at a desired position within the shooting angle of view.
[0003] Patent document 1 describes a technology in which, when detecting a specific subject through image recognition processing, the parameters used in the image recognition processing are changed according to the zoom magnification so that the specific subject does not go undetected due to a change in the zoom magnification.
[0004] Patent document 2 also describes a technology in which, when a subject to be tracked moves near the boundary of the imaging range of a first imaging device, template data of the subject to be tracked generated by the first imaging device is transmitted to a second imaging device, and the second imaging device takes over the tracking target. Patent Document 3 describes a technique for expanding the shooting range to search for a specific subject when the specific subject cannot be detected from a captured image. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-142181 [Patent Document 2] Japanese Patent Application Laid-Open No. 2002-290962 [Patent Document 3] Japanese Patent Application Laid-Open No. 2015-61239 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in Patent Documents 1 and 2, the object to be tracked is identified by template matching, so when tracking a specific object using multiple image capture devices, the multiple image capture devices need to be arranged so that their image capture positions and image capture directions are close to each other. Therefore, if the image capture positions and image capture directions of the multiple image capture devices are arranged far apart, it becomes difficult to track a specific object using the multiple image capture devices.
[0007] Furthermore, if the imaging areas of the tracking target subject are different among the multiple imaging devices, the similarity of the tracking target subject may decrease among the multiple imaging devices. For example, if the imaging angle of view of a first imaging device covers the entire subject and the imaging angle of view of a second imaging device covers only a portion of the subject, or if a portion of the subject is outside the imaging angle of view, it may be difficult for the multiple imaging devices to recognize the same subject.
[0008] Furthermore, in Patent Documents 2 and 3, if the size of the image of the subject to be tracked differs between the multiple imaging devices, the similarity of the subject to be tracked between the multiple imaging devices may decrease, making it difficult for the multiple imaging devices to recognize the same subject. The present invention has been made in view of the above-mentioned problems, and an object of the present invention is to realize a system that makes it possible to track a specific subject using a plurality of imaging devices with different imaging positions and directions. [Means for solving the problem]
[0009] In order to solve the above problems and achieve the object, the present invention provides a system including a first imaging device and a second imaging device having different shooting directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, wherein the first control device has a first determination means that determines the predetermined subject from the subjects included in the first image, an image processing means that cuts out a first region of the predetermined subject included in the first image based on a shooting range of the second image acquired from the second imaging device, a first generation means that generates first feature information of the first region of the predetermined subject, and a first control means that controls the second imaging device to track the predetermined subject, and the second control device has a second generation means that generates second feature information of the subject included in the second image, a second determination means that determines the predetermined subject based on the first feature information and the second feature information acquired from the first control device, and a second control means that controls the second imaging device to track the predetermined subject. [Effects of the Invention]
[0010] According to the present invention, it is possible to track a specific subject using a plurality of imaging devices with different imaging positions and imaging directions. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration according to a first embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of the hardware configuration of devices constituting the system of the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of the functional configuration of the apparatuses that constitute the system of the first embodiment. [Figure 4] 3 is a flowchart illustrating the basic operation of the devices constituting the system of the first embodiment. [Figure 5] 3A to 3C are diagrams for explaining a coordinate conversion method for a captured image according to the first embodiment. [Figure 6]3A to 3C are diagrams for explaining a subject detection method and a coordinate conversion method according to the first embodiment. [Figure 7] 5A to 5C are diagrams for explaining pan control according to the first embodiment. [Figure 8] 5A to 5C are diagrams for explaining tilt control according to the first embodiment. [Figure 9] 4 is a flowchart illustrating a control process according to the first embodiment. [Figure 10] 5A to 5C are diagrams for explaining a method for determining a subject to be tracked according to the first embodiment. [Figure 11] 4A and 4B are diagrams illustrating the relationship between a subject to be tracked and a photographing angle of view. [Figure 12] 4 is a flowchart illustrating a trimming process according to the first embodiment. [Figure 13] 5A to 5C are diagrams for explaining the trimming process according to the first embodiment. [Figure 14] 10 is a flowchart illustrating a control process according to a third embodiment. [Figure 15] 10 is a flowchart illustrating zoom control according to the third embodiment. [Figure 16] FIG. 10 is a diagram illustrating an example of the functional configuration of devices constituting a system according to a fourth embodiment. [Figure 17] 10 is a flowchart illustrating a control process according to the fourth embodiment. [Figure 18] 10 is a flowchart illustrating a control process according to the fourth embodiment. [Figure 19] 10 is a flowchart illustrating a control process according to a fifth embodiment. [Figure 20] 13 is a flowchart illustrating a control process according to a sixth embodiment. [Figure 21] FIG. 13 is a diagram illustrating an example of a system configuration according to a seventh embodiment. [Figure 22] FIG. 20 is a diagram illustrating the roles and contents that can be set in the imaging device of the seventh embodiment. [Figure 23] 13 is a flowchart illustrating a control process according to the eighth embodiment. [Figure 24] 13 is a flowchart illustrating a letter adding process according to the eighth embodiment. [Figure 25] FIG. 10 is a diagram illustrating letter assignment in the eighth embodiment. [Figure 26] 13A to 13C are diagrams illustrating skeleton estimation processing according to the ninth embodiment. [Figure 27] FIG. 20 is a diagram illustrating processing using the skeleton estimation processing of the ninth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0013] [Embodiment 1] <System configuration> First, the system configuration of the first embodiment will be described with reference to FIG.
[0014] The system of this embodiment includes a first control device 100, a second control device 200, a first imaging device 300, and a second imaging device 400. In the system of this embodiment, the second imaging device 400 is controlled by either the first control device 100 or the second control device 200 to track a specific subject. In this embodiment, the specific subject is, for example, a person, but may also be an animal or an object.
[0015] The first control device 100 detects a subject to be tracked from an overhead image captured by the first imaging device 300, and controls the second imaging device 400 based on the detection result. The first control device 100 is also called a workstation. The subject to be tracked is set, for example, by a user operation or automatically.
[0016] The second control device 200 controls the second imaging device 400 based on the result of object recognition of the tracking target from the overhead image captured by the first imaging device 300 and the result of object recognition of the tracking target from the sub-image captured by the second imaging device 300. The second control device 200 is also called an edge box.
[0017] The first imaging device 300 has a fixed wide-angle shooting angle of view and can capture an overhead image including all of subjects A, B, and C. The first imaging device 300 is also called an overhead camera. The second imaging device 400 has a variable shooting angle of view and can capture images of at least one of subjects A, B, and C. The second imaging device 400 is called a sub-camera. The first imaging device 300 and the second imaging device 400 are placed at positions separated from each other so that their shooting positions and / or shooting directions are different.
[0018] The first control device 100, the second control device 200, the first imaging device 300, and the second imaging device 400 are communicably connected via a network 600 such as a LAN (Local Area Network). In this embodiment, an example is described in which the first control device 100, the second control device 200, the first imaging device 300, and the second imaging device 400 are connected via the network 600, but they may be connected via a connection cable (not shown). In this embodiment, an example is described in which there is one second imaging device 400, but there may be two or more. When there are multiple second imaging devices 400, a second control device 200 is provided for each second imaging device 400.
[0019] Next, the basic functions of the system of this embodiment will be described.
[0020] The first imaging device 300 captures an overhead image and transmits the overhead image to the first control device 100 via the network 600.
[0021] The second imaging device 400 captures a sub-image including the subject to be tracked (tracked subject) and transmits the sub-image to the second control device 200 via the network 600. The second imaging device 400 has a PTZ function. The PTZ function is a function that can control the pan, tilt, and zoom of the imaging device. PTZ is an abbreviation of the initials of pan, tilt, and zoom. Pan is the horizontal movement of the optical axis of the imaging device. Tilt is the vertical movement of the optical axis of the imaging device. Zoom is zoom up (telephoto) and zoom out (wide angle). Pan and tilt are functions that change the shooting direction of the imaging device. Zoom is a function that changes the shooting range (shooting angle of view) of the imaging device.
[0022] The first control device 100 determines a tracking subject from subjects detected from the overhead image received from the first imaging device 300, and calculates first feature information of the tracking subject from the overhead image. The first control device controls the second imaging device 400 to change the imaging direction and imaging range of the second imaging device 400 to the imaging direction and imaging range of the tracking subject based on the first feature information of the tracking subject.
[0023] After changing the shooting direction and shooting range of the second imaging device 400 to the shooting direction and shooting range of the tracked subject, the first control device 100 transmits the first feature information of the tracked subject calculated from the overhead image to the second control device 200.
[0024] The second control device 200 detects the subject from the sub-image received from the second imaging device 400 and calculates second feature information of the detected subject. The second control device 200 compares the second feature information of the subject detected from the sub-image with the first feature information of the tracking subject received from the first control device 100.
[0025] If the similarity between the first feature information of the tracked subject and the second feature information of the subject detected from the sub-image is low, the first control device 100 controls the second imaging device 400 to change the shooting direction and shooting range of the second imaging device 400 to the shooting direction and shooting range of the tracked subject based on the first feature information of the tracked subject.
[0026] In addition, if there is a high degree of similarity between the first feature information of the tracked subject and the second feature information of the subject detected from the sub-image, the second control device 200 controls the second imaging device 400 to change the shooting direction and shooting range of the second imaging device 400 to the shooting direction and shooting range of the tracked subject based on the second feature information of the subject detected from the sub-image that has a high degree of similarity to the first feature information of the tracked subject.
[0027] The feature information is information that can identify the same subject when the same subject is photographed by multiple imaging devices with different photographing positions and / or photographing directions. The feature information is an inference result that is output by performing image recognition through inference processing using a trained model, using multiple images of the same subject photographed by multiple imaging devices with different photographing positions and / or photographing directions as input. When an inference result that the subjects are the same subject is obtained, it is possible to identify that the subjects included in the multiple images photographed by multiple imaging devices with different photographing positions and / or photographing directions are the same subject.
[0028] In the following description, the first control device 100 will be referred to as a workstation (WS), the second control device 200 as an edge box (EB), the first imaging device 300 as an overhead camera, and the second imaging device 400 as a sub-camera.
[0029] <Device configuration> Next, the hardware configuration of the WS 100, the EB 200, the overhead camera 300, and the sub-camera 400 will be described in detail with reference to FIG.
[0030] First, the configuration of the WS100 will be described.
[0031] The WS 100 comprises a control unit 101, a volatile memory 102, a non-volatile memory 103, an inference unit 104, a communication unit 105, and an operation unit 106, and each unit is connected via an internal bus 110 so that data can be sent and received.
[0032] The control unit 101 has a processor (CPU) that performs arithmetic processing and control processing of the WS 100 , and controls each component of the WS 100 by executing a control program stored in the nonvolatile memory 103 .
[0033] The volatile memory 102 is a main storage device such as a RAM. Constants and variables for the operation of the control unit 101, as well as control programs and inference programs read from the nonvolatile memory 103, are loaded into the volatile memory 102. The volatile memory 102 also stores information such as image data and inference programs received from external devices via the communication unit 105. The volatile memory 102 also stores overhead image data received from the overhead camera 300. The volatile memory 102 has a storage capacity sufficient to hold this information.
[0034] The nonvolatile memory 103 is an auxiliary storage device such as an EEPROM, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The nonvolatile memory 103 stores an OS (operating system) which is basic software executed by the control unit 101, a control program including applications which cooperate with the OS to realize applied functions, an inference program used by the inference unit 104 for inference processing, etc.
[0035] The inference unit 104 executes inference processing using a trained inference model and inference parameters in accordance with an inference program. The inference unit 104 executes inference processing to estimate the presence or absence and position of a specific subject and feature information of the subject from the overhead image received from the overhead camera 300. The inference processing in the inference unit 104 can be executed by a processing device specialized for image processing and inference processing, such as a GPU (Graphics Processing Unit). A GPU is a processor capable of performing a large number of product-sum operations and has the processing power to perform matrix operations of a neural network in a short period of time. The inference processing in the inference unit 104 may also be implemented by a reconfigurable logic circuit, such as an FPGA (Field-Programmable Gate Array). Note that the inference processing may be performed by a CPU and GPU in the control unit 101 working together, or by either the CPU or the GPU in the control unit 101.
[0036] The communication unit 105 is an interface (I / F) that complies with a wired communication standard such as Ethernet (registered trademark) or an interface that complies with a wireless communication standard such as Wi-Fi (registered trademark). The communication unit 105 is connected to external devices such as the EB 200, the overhead camera 300, and the sub-camera 400 via a network 600 such as a wired LAN or a wireless LAN, and can transmit and receive data to and from the external devices. The control unit 101 realizes communication with the external devices by controlling the communication unit 105. Note that the communication method is not limited to Ethernet (registered trademark) or Wi-Fi (registered trademark), and a communication standard such as IEEE 1394 may also be used.
[0037] The operation unit 106 is an operation member such as various switches, buttons, a touch panel, etc. that accepts various operations by the user and outputs operation information to the control unit 101. The operation unit 106 also provides a user interface for the user to operate the WS100.
[0038] The display unit 111 displays an overhead image, subject recognition results, and a GUI (Graphical User Interface) for interactive operation. The display unit 111 is a display device such as a liquid crystal display or an organic EL display. The display unit 111 may be an integral part of the WS 100 or an external device connected to the WS 100.
[0039] Next, the configuration of the EB200 will be described.
[0040] The EB 200 includes a control unit 201, a volatile memory 202, a nonvolatile memory 203, an inference unit 204, and a communication unit 205, and each unit is connected via an internal bus 210 so as to be able to send and receive data.
[0041] The control unit 201 has a processor (CPU) that performs arithmetic processing and control processing of the EB 200, and controls each component of the EB 200 by executing a control program stored in the nonvolatile memory 203.
[0042] The volatile memory 202 is a main storage device such as a RAM. Constants and variables for the operation of the control unit 201, as well as control programs and inference programs read from the non-volatile memory 203, are loaded into the volatile memory 202. The volatile memory 202 also stores information such as image data and inference programs received from external devices via the communication unit 205. The volatile memory 202 also stores sub-image data received from the sub-camera 400. The volatile memory 202 has a sufficient storage capacity to hold this information.
[0043] The nonvolatile memory 203 is an auxiliary storage device such as an EEPROM, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The nonvolatile memory 203 stores an OS (operating system) which is basic software executed by the control unit 201, a control program including an application that cooperates with the OS to realize applied functions, an inference program used by the inference unit 204 for inference processing, etc.
[0044] The inference unit 204 executes inference processing using a trained inference model and inference parameters in accordance with an inference program. The inference unit 204 executes inference processing to estimate the presence or absence and position of a specific subject and feature information of the subject from a sub-image received from the sub-camera 400. The inference processing in the inference unit 204 can be executed by a processing device specialized for image processing and inference processing, such as a GPU (Graphics Processing Unit). A GPU is a processor capable of performing a large number of product-sum operations and has the processing power to perform matrix operations of a neural network in a short period of time. The inference processing in the inference unit 204 may also be implemented by a reconfigurable logic circuit, such as an FPGA (Field-Programmable Gate Array). Note that the inference processing may be performed by a CPU and GPU of the control unit 201 working together, or by either the CPU or the GPU of the control unit 201.
[0045] The communication unit 205 is an interface (I / F) that complies with a wired communication standard such as Ethernet (registered trademark) or an interface that complies with a wireless communication standard such as Wi-Fi (registered trademark). The communication unit 205 is connected to external devices such as the WS 100 and the sub-camera 400 via a network 600 such as a wired LAN or a wireless LAN, and can transmit and receive data to and from the external devices. The control unit 201 controls the communication unit 205 to realize communication with the external devices. Note that the communication method is not limited to Ethernet (registered trademark) or Wi-Fi (registered trademark), and a communication standard such as IEEE 1394 may also be used.
[0046] Next, the configuration of the overhead camera 300 will be described.
[0047] The overhead camera 300 comprises a control unit 301, a volatile memory 302, a non-volatile memory 303, a communication unit 305, an imaging unit 306, and an image processing unit 307, and each unit is connected via an internal bus 310 so that data can be sent and received.
[0048] The control unit 301 performs overall control of the overhead camera 300 under the control of the WS 100. The control unit 301 has a processor (CPU) that performs calculation and control processing for the overhead camera 300, and controls each component of the overhead camera 300 by executing a control program stored in the nonvolatile memory 303.
[0049] Volatile memory 302 is a main storage device such as a RAM. Constants and variables for the operation of control unit 301, as well as control programs and inference programs read from nonvolatile memory 303, are loaded into volatile memory 302. Volatile memory 302 also stores overhead image data captured by imaging unit 306 and processed by image processing unit 307. Volatile memory 302 has a storage capacity sufficient to hold this information.
[0050] The nonvolatile memory 303 is an auxiliary storage device such as an EEPROM, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The nonvolatile memory 303 stores an OS (operating system), which is basic software executed by the control unit 301, and control programs including applications that cooperate with the OS to realize applied functions.
[0051] The imaging unit 306 has an image sensor configured with a CCD (charge-coupled device), a CMOS (complementary metal-oxide semiconductor) element, etc., and converts an optical image of the subject into an electrical signal. In this embodiment, the overhead camera 300 has a fixed shooting angle of view so that it can capture an overhead image including multiple subjects, including the tracking subject.
[0052] The image processing unit 307 performs various types of image processing on image data output from the imaging unit 306 or image data read from the volatile memory 302. The various types of image processing include, for example, image processing such as noise removal, edge enhancement, and enlargement / reduction, image correction processing such as contrast correction, brightness correction, and color correction, and trimming or cropping processing to cut out part of the image data. The image processing unit 307 converts the image data that has been subjected to image processing into an image file in a predetermined format (for example, JPEG) and records it in the non-volatile memory 303. The image processing unit 307 also performs predetermined calculation processing using the image data, and the control unit 301 performs AF (autofocus) processing and AE (autoexposure) processing based on the calculation results.
[0053] The communication unit 305 is an interface (I / F) that complies with a wired communication standard such as Ethernet (registered trademark) or an interface that complies with a wireless communication standard such as Wi-Fi (registered trademark). The communication unit 305 is connected to an external device such as the WS 100 via a network 600 such as a wired LAN or a wireless LAN, and can transmit and receive data to and from the external device. The control unit 301 realizes communication with the external device by controlling the communication unit 305. Note that the communication method is not limited to Ethernet (registered trademark) or Wi-Fi (registered trademark), and a communication standard such as IEEE 1394 may also be used.
[0054] Next, the configuration of the sub-camera 400 will be described.
[0055] The sub-camera 400 comprises a control unit 401, a volatile memory 402, a non-volatile memory 403, a communication unit 405, an imaging unit 406, an image processing unit 407, an optical unit 408 and a PTZ drive unit 409, and each unit is connected via an internal bus 410 so that data can be sent and received.
[0056] The control unit 401 performs overall control of the sub-camera 400 under the control of the WS 100 or the EB 200. The control unit 401 has a processor (CPU) that performs arithmetic processing and control processing of the sub-camera 400, and controls each component of the sub-camera 400 by executing a control program stored in the non-volatile memory 403.
[0057] The volatile memory 402 is a main storage device such as a RAM. Constants and variables for the operation of the control unit 401, as well as control programs and inference programs read from the nonvolatile memory 403, are loaded into the volatile memory 402. The volatile memory 402 also stores overhead image data captured by the imaging unit 406 and processed by the image processing unit 407. The volatile memory 402 has a storage capacity sufficient to hold this information.
[0058] The nonvolatile memory 403 is an auxiliary storage device such as an EEPROM, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The nonvolatile memory 403 stores an OS (operating system), which is basic software executed by the control unit 401, and control programs including applications that cooperate with the OS to realize applied functions.
[0059] The imaging unit 406 has an image sensor configured with a CCD (charge coupled device), a CMOS (complementary metal oxide semiconductor) element, or the like, and converts an optical image of a subject into an electrical signal.
[0060] The image processing unit 407 performs various types of image processing on image data output from the imaging unit 406 or image data read from the volatile memory 402. The various types of image processing include, for example, image processing such as noise removal, edge enhancement, and enlargement / reduction, image correction processing such as contrast correction, brightness correction, and color correction, and trimming or cropping processing to cut out part of the image data. The image processing unit 407 converts the image data that has been subjected to image processing into an image file in a predetermined format (for example, JPEG) and records it in the non-volatile memory 403. In addition, the image processing unit 407 performs predetermined calculation processing using the image data, and the control unit 401 performs AF (autofocus) processing and AE (autoexposure) processing based on the calculation results.
[0061] The communication unit 405 is an interface (I / F) that complies with a wired communication standard such as Ethernet (registered trademark) or an interface that complies with a wireless communication standard such as Wi-Fi (registered trademark). The communication unit 405 can connect to an external device such as the EB200 via a network 600 such as a wired LAN or a wireless LAN, and can exchange data with the external device. The control unit 401 realizes communication with the external device by controlling the communication unit 405. Note that the communication method is not limited to Ethernet (registered trademark) or Wi-Fi (registered trademark), and a communication standard such as IEEE1394 may also be used.
[0062] The optical unit 408 includes a group of lenses including a zoom lens and a focus lens, a shutter with an aperture function, and a mechanism for driving these optical elements. The optical unit 408 drives the optical elements to at least one of rotate the shooting direction of the sub-camera 400 around the pan (P) axis (horizontal direction) or the tilt (T) axis (vertical direction), and change the shooting range (shooting angle of view) of the sub-camera 400 along the zoom (Z) axis (enlargement / reduction direction).
[0063] The PTZ driving unit 409 includes actuators such as mechanical elements and motors for driving the optical unit 408 in the PTZ direction, and drives the optical unit 408 in the PTZ direction under the control of the control unit 401 .
[0064] The zoom function of this embodiment is not limited to optical zoom, which changes the focal length by moving a zoom lens, but may also be digital zoom, which cuts out and enlarges a portion of the captured image data, or a combination of optical zoom and digital zoom.
[0065] [Control processing] Next, referring to Figures 3 to 10, we will explain the control process for tracking a tracking subject by switching between a mode in which WS100 controls the sub-camera 400 based on an overhead image and a mode in which EB200 controls the sub-camera 400 based on a sub-image.
[0066] First, the functional configuration of the WS 100 and EB 200 for realizing the control processing of this embodiment will be described with reference to FIGS.
[0067] The functions of the WS 100 and the EB 200 are implemented by hardware and / or software. If the functional units shown in Figure 3 are implemented by hardware instead of software, it is sufficient to have a circuit configuration corresponding to each functional unit shown in Figure 3.
[0068] The WS 100 includes an image recognition unit 121, an object of interest determination unit 122, a tracking target determination unit 123, a control information generation unit 124, a feature information determination unit 125, and a tracking state determination unit 126. Software that realizes these functions is stored in the non-volatile memory 103, and is loaded into the volatile memory 102 by the control unit 101 for execution.
[0069] The EB 200 includes an image recognition unit 221, a tracking target determination unit 222, and a control information generation unit 223. These pieces of software are stored in the non-volatile memory 203, and the control unit 201 loads them into the volatile memory 202 and executes them.
[0070] Fig. 4(a) is a flowchart showing the basic operation of the WS 100. Fig. 4(b) is a flowchart showing the basic operation of the EB 200. Fig. 4(c) is a flowchart showing the operation of the overhead camera 300. Fig. 4(d) is a flowchart showing the operation of the sub-camera 400.
[0071] First, the software functions and basic operations of the WS 100 will be described with reference to FIGS.
[0072] In step S101, the control unit 101 transmits a shooting command to the overhead camera 300 using a predetermined protocol via the communication unit 105, receives an overhead image from the overhead camera 300, stores it in the volatile memory 102, and proceeds to step S102.
[0073] In step S102, the control unit 101 executes the function of the image recognition unit 121 in FIG. 3, and the process proceeds to step S103.
[0074] The image recognition unit 121 controls the inference unit 104, the volatile memory 102, and the non-volatile memory 103, and performs the following object recognition processing.
[0075] The image recognition unit 121 receives the overhead-view image IMG of the overhead-view camera 300 read from the volatile memory 102 and reference position information REF_POSI of the overhead-view camera 300. The position information REF_POSI of the overhead-view camera 300 includes information on the position of the overhead-view camera 300 and marker coordinates. The image recognition unit 121 detects the subject and calculates feature information based on the overhead-view image IMG of the overhead-view camera 300 and the reference position information REF_POSIREF_POSI. The image recognition unit 121 then outputs coordinate information POSITION[n] indicating the position of the detected subject, ID[n] indicating identification information of the detected subject, and STAT[n] indicating feature information of the detected subject. The position of the overhead camera 300 is a position in a coordinate space when the area captured by the overhead camera 300 is viewed from directly above, and is known in advance by user operation or by measurement using a sensor (not shown). The marker coordinates are position information of a marker placed in a coordinate space when the area captured by the overhead camera 300 is viewed from directly above in order to calculate a homography transformation matrix (described later), and are known values measured in advance manually or by a sensor (not shown). The marker is like a mark with a color different from the color of the floor, ground, etc., and may be any type that can be measured by user operation or a sensor (not shown). For example, if the sensor (not shown) is a camera, the marker position is acquired by extracting the color of the marker from an image captured using a mark of any color. The position of the overhead camera 300 and marker coordinates may also be input by the user via the operation unit 106 of the WS 100, and the control unit 101 may store them in the volatile memory 102. The reference position information REF_POSI and the subject coordinate information POSITION[n] are expressed in a coordinate system converted into a coordinate space in which the shooting area of the overhead camera 300 is viewed from directly above. n is an index indicating the number of subjects detected; for example, if the inference unit 104 detects three people, the POSITION, ID, and STAT for all three people are output as the inference result. The control unit 101 stores the subject recognition result obtained by the image recognition unit 121 in the volatile memory 102. Details of the subject detection process and the feature information calculation process will be described later.
[0076] Here, a method for calculating the coordinate information POSITION of the subject by the image recognition unit 121 will be described.
[0077] First, with reference to FIG. 5, the relationship between the coordinate system of the overhead image of the overhead camera 300 and the coordinate system when the shooting area of the overhead camera 300 is viewed from directly above will be described.
[0078] In order to calculate a pan value such that the shooting direction of the sub-camera 400 is the direction of the subject being tracked, the calculation is simplified by calculating the angle in a plane coordinate space perpendicular to the axis along which the sub-camera 400 pans. For example, if the sub-camera 400 is installed perpendicular to a ground surface (reference position) such as the floor or ground, the coordinate space perpendicular to the axis along which the sub-camera 400 pans is a coordinate space parallel to the reference position (a coordinate space in which the space in which the sub-camera 400 and the subject are located is viewed from directly above) as shown in Fig. 5(b). In this embodiment, it is assumed that the sub-camera 400 is installed perpendicular to the reference position, and the pan value is calculated in a coordinate system in which the shooting area of the overhead camera 300 is viewed from directly above. That is, the subject position detected in the coordinate system of the overhead image of the overhead camera 300 shown in Fig. 6(a) (hereinafter referred to as the overhead camera coordinate system) is transformed into a coordinate system in which the shooting area of the overhead camera 300 is viewed from directly above (hereinafter referred to as the planar coordinate system) shown in Fig. 5(b). The coordinate transformation is performed using a homography transformation matrix H according to the following equation 1. (Formula 1) TIFF0007805396000001.tif19155In Equation 1, x and y are the horizontal and vertical coordinates in the overhead camera coordinate system, and X and Y are the horizontal and vertical coordinates in the plane coordinate system.
[0079] The control unit 101 reads out the reference position information REF_POSI from the volatile memory 102 and calculates the homography transformation matrix H by substituting the marker coordinates Mark_A to Mark_D shown in FIGS. 5(a) and 5(b) included in the reference position information REF_POSI into Equation 1. Note that the marker coordinates are values in a planar coordinate system. By using Equation 1, any coordinate in the overhead camera coordinate system of FIG. 4(a) can be mapped to any coordinate in the planar coordinate system of FIG. 4(b). In the example of FIG. 4, the control unit 101 can determine the positions of subjects A, B, and C included in the overhead image IMG of the overhead camera 300 in the planar coordinate system of FIG. 4(b). The control unit 101 stores the homography transformation matrix H calculated using Equation 1 in the volatile memory 102.
[0080] Next, a method for detecting the subject position using an inference model for subject detection and a method for converting the position into a planar coordinate system will be described.
[0081] In this embodiment, subject detection is performed by performing image recognition processing using a trained inference model for subject detection created by machine learning such as deep learning.
[0082] The inference model for subject detection takes an overhead image as input and outputs coordinate information on the image of the subject included in the overhead image.
[0083] The control unit 101 detects subjects by performing image recognition processing using an inference model for subject detection with the overhead image IMG captured by the overhead camera 300 as input, via the inference unit 104. Fig. 6(a) shows an example in which subjects detected by the inference unit 104 are displayed in rectangular frames. As shown in Fig. 6(a), the coordinates of rectangular portions circumscribing subjects A, B, and C detected from the overhead image are detected as subject positions. The control unit 101 stores coordinate information of the subject detected from the overhead image in the volatile memory 102. Note that, although the present embodiment has described an example in which subject detection is performed by inference processing using a trained model, the present invention is not limited to this. For example, a method called the SIFT method, which detects by matching local feature points in an image, or a template matching method, which detects by calculating the similarity with a template image, may also be used.
[0084] Furthermore, the control unit 101 converts the bottom edge of the rectangular part of the subject detected in the overhead camera coordinate system shown in Fig. 6(a) into the planar coordinate system shown in Fig. 6(b) as the subject detection position (the coordinates of the person's feet in the example of Fig. 6). For example, the control unit 101 reads out the homography transformation matrix H from the volatile memory 102 and substitutes the coordinates of the feet (xa, ya) of subject A in the overhead camera coordinate system into x and y in Equation 1, thereby enabling conversion to the coordinates of the feet (XA, YA) in the planar coordinate system. Similarly, it is possible to calculate the coordinates (XB, YB) of the feet of subject B and the coordinates (XC, YC) of the feet of subject C in the plane coordinate system for the coordinates (xb, yb) of the feet of subject B and the coordinates (xc, yc) of the feet of subject C. The control unit 101 writes the coordinates of the feet into the volatile memory 102 as the position coordinates POSITION of the subjects.
[0085] Next, a method for generating identification information ID and feature information STAT of a subject by image recognition unit 121 will be described.
[0086] The control unit 101 inputs the subject coordinate information POSITION, which is the inference result of the inference model for subject detection, and the overhead image taken by the overhead camera 300 into the trained inference model for subject identification created by the inference unit 104 through machine learning such as deep learning, and performs inference processing to output identification information ID and feature information STAT. The inference model for subject identification is different from the inference model for subject detection.
[0087] Here, we will explain the inference model for identifying the subject.
[0088] The inference model for subject identification in this embodiment is a trained model that has been trained to increase the similarity of feature information for images of the same subject using training data that is a collection of data correlating a set of images of a specific subject taken from multiple different shooting directions with information that can identify the specific subject, for the number of subjects.By inputting an image of the subject cut out based on the coordinate position POSITION of the subject, which is the output of the inference model for subject detection, into the inference model for subject identification, feature information STAT is output. When an image of the same subject taken with a different camera is input, the output feature information has a higher similarity to the feature information STAT compared to when an image of a different subject is input. Feature information can be a multidimensional vector of the response of the convolution layer of a convolutional neural network. Similarity will be discussed later.
[0089] The inference model for subject detection and the inference model for subject identification are stored in the non-volatile memory 103 before the control process of this embodiment is started.
[0090] The image recognition unit 121 also assigns identification information ID to the subject corresponding to the feature information, which is the inference result of the inference model for subject identification. Furthermore, the image recognition unit 121 inputs the images of the current frame and past frames, and calculates the similarity of the feature information of each subject image obtained by inputting the images of each subject detected by the inference model for subject detection into the inference model for subject identification. The similarity is calculated using cosine similarity. The more similar the multidimensional vectors, which are the feature information of each object image, the closer the cosine similarity is to 1, and the more different they are, the closer the cosine similarity is to 0. The same ID is assigned to objects with the closest similarity between the past and current frames. Note that the method for calculating the similarity is not limited to this, and any method may be used as long as it outputs a higher value when the feature information is closer and a lower value when the feature information is more distant. In this embodiment, feature information is used to assign IDs, but this is not limiting. A method may be used in which rectangular information of a subject obtained by an inference model for subject detection is used to compare the position and size of the rectangular information of the detected subject between the current frame and past frames, and the same ID is assigned to the subject that is closest to the position of the rectangular information. Another method may be used in which the position of rectangular information in the current frame is predicted using a Kalman filter or the like based on the transition of the position of rectangular information for the same ID in several past frames, and the same ID is assigned to the subject that is closest to the position of the predicted rectangular information. IDs may also be assigned using a combination of these methods. By using this method, it is possible to improve the accuracy of ID assignment when a subject with a similar appearance suddenly enters the shooting field of view.
[0091] As described above, image recognition unit 121 receives overhead image 300 as input, performs inference processing using an inference model for subject detection, and outputs the coordinate position POSITION of the subject, which is then stored in volatile memory 102. Image recognition unit 121 also inputs subject coordinate information POSITION, which is the inference result of the inference model for subject detection, and the overhead image captured by overhead camera 300, into an inference model for subject identification, and performs inference processing. Image recognition unit 121 then outputs identification information ID and feature information STAT as the results of the inference processing, which are then stored in volatile memory 102.
[0092] Returning to the description of FIG. 4, in step S103, the control unit 101 executes the function of the target subject determination unit 122 in FIG. 3, and the process proceeds to step S104.
[0093] The target subject determining unit 122 determines a target subject MAIN_SUBJECT from operation information input by the user via the operation unit 106 and coordinate information of the subject, which is the subject recognition result read from the volatile memory 102 by the image recognition unit 121 .
[0094] The control unit 101 displays the overhead image captured by the overhead camera 300 and the object recognition results in the volatile memory 102 on the display unit 111 of the WS 100. The control unit 101 allows the user to select an object of interest from among the objects displayed as the object recognition results via the operation unit 106. For example, if the operation unit 106 is a mouse, the user can click and select one of the objects displayed on the display unit 111. The control unit 101 saves the identification information ID corresponding to the object of interest selected by the user in the volatile memory 102 as the object of interest MAIN_SUBJECT.
[0095] In step S104, the control unit 101 executes the function of the tracking target determination unit 123 in FIG. 3, and the process proceeds to step S105.
[0096] The tracking target determining unit 123 determines the tracking subject SUBJECT_ID of the sub camera 400 from the target subject MAIN_SUBJECT determined by the target subject determining unit 122 .
[0097] Here, a method for determining the subject to be tracked by the sub-camera 400 will be described.
[0098] The control unit 101 reads out the target subject MAIN_SUBJECT determined by the target subject determination unit 122 from the volatile memory 102, and determines the target subject MAIN_SUBJECT as the tracking subject SUBJECT_ID of the sub camera 400. In this way, by setting the same subject as the target subject MAIN_SUBJECT selected by the user as the tracking subject SUBJECT_ID of the sub camera 400, the sub camera 400 can be controlled with the user-selected subject as the tracking target.
[0099] The method for determining the subject to be tracked is not limited to the above method, and may be determined using, for example, information on the subject of interest MAIN_SUBJECT and identification information ID read from volatile memory 102. For example, if multiple subjects are included in the overhead image of overhead camera 300 and multiple sub-cameras 400 are installed, one method is to have one sub-camera track the same subject as the subject of interest, and another sub-camera track a subject different from the subject of interest. By determining the subject to be tracked in this way, it is possible for each sub-camera to comprehensively track multiple subjects included in the overhead image of overhead camera 300. Another method is to read out REF_POSI, which includes the subject's coordinate information POSITION, identification information ID, and sub-camera position, from the volatile memory 102, and determine the subject closest to the sub-camera among the subjects detected from the overhead image of the overhead camera 300, as the subject to be tracked. By determining the subject to be tracked in this manner, it is possible to set the subject that is easiest to fit into the angle of view from the position of the sub-camera as the tracking target. The control unit 101 saves the tracking subject SUBJECT_ID decided as described above in the volatile memory 102, and also saves the identification ID of the tracking subject before saving in the volatile memory 102 as the past tracking subject ID.
[0100] In step S105, the control unit 101 executes the function of the feature information determination unit 125 to transmit feature information corresponding to the subject being tracked by the sub camera 400 to the EB 200. The control unit 101 also executes the function of the tracking state determination unit 126 to update the tracking state information STATE, store it in the volatile memory 102, and proceeds to step S106.
[0101] The tracking state information STATE includes information of either "Tracking by WS100" or "Tracking by EB200." "Tracking by WS100" indicates a state in which the WS100 is tracking the tracking subject by controlling the sub-camera 400. "Tracking by EB200" indicates a state in which the EB200 is tracking the tracking subject by controlling the sub-camera 400. Details of the processing in step S105 will be described later.
[0102] In step S106, the control unit 101 reads the tracking state information STATE from the volatile memory 102, and determines whether "tracking by WS100" or "tracking by EB200" is in progress based on the tracking state information STATE. If the control unit 101 determines that "tracking by WS100" is in progress, the process proceeds to step S107; if the control unit 101 determines that "tracking by EB200" is in progress, the process returns to step S101.
[0103] In step S107, the control unit 101 executes the function of the control information generating unit 124 in FIG. 3, and the process proceeds to step S108.
[0104] The control information generation unit 124 calculates pan / tilt values PT_VALUE of the sub camera 400 for tracking the tracking subject SUBJECT_ID determined by the tracking target determination unit 123 by the sub camera 400. The control unit 101 reads out coordinate information of the sub camera 400 in a planar coordinate system included in the reference position information REF_POSI and coordinate information POSITION of the detected subject from the volatile memory 102. Then, the control unit 101 calculates pan / tilt values from the coordinate information of the subject corresponding to the tracking subject SUBJECT_ID such that the shooting direction of the sub camera 400 is the direction of the tracking subject.
[0105] Here, a method for calculating the pan value will be described with reference to FIG.
[0106] As shown in FIG. 7, the angle θ formed by the line extending from the center of the optical axis of the sub-camera 400 and the line connecting the sub-camera 400 and the tracking subject SUBJECT_ID can be calculated by the following formula 2. (Formula 2) TIFF0007805396000002.tif19155In Equation 2, px and py are the horizontal and vertical coordinates of the position of the tracked subject, and subx and suby are the horizontal and vertical coordinates of the position of the sub-camera 400. px and py can be obtained from the coordinate information POSITION of the detected subject by referencing the coordinate information corresponding to the tracked subject SUBJECT_ID.
[0107] The control information generator 124 calculates the pan value of the sub camera 400 based on the angle θ.
[0108] Next, a method for calculating the tilt control value will be described with reference to FIG.
[0109] As shown in Figure 8, when the height h1 of the optical axis of the sub-camera 400 is taken as the angle ρ formed by a line extending from the center of the optical axis of the sub-camera 400 and a line extending toward the height h2 of a specific part of the tracked subject (in the case of a person, the height of the face), the angle can be calculated using the following equations 3 and 4. (Formula 3) TIFF0007805396000003.tif13155 (Formula 4) TIFF0007805396000004.tif19155 In Equation 4, h1 is the height of the sub-camera 400 from the ground surface, and h2 is the height from the ground surface to a predetermined part of the tracking subject (the face in the case of a person). h1 and h2 may be stored in advance in the volatile memory 102, or may be measured in real time using a sensor (not shown).
[0110] The control information generator 124 calculates a tilt control value for the sub camera 400 based on the angle ρ.
[0111] The pan value / tilt value may be used as a velocity value for directing the sub camera 400 toward a tracking subject. The pan value / tilt value is calculated as follows: first, the control unit 101 acquires the current pan value / tilt value of the sub camera 400 from the EB 200. Next, the control unit 101 calculates a pan angular velocity proportional to the difference from the pan value θ read from the volatile memory 102. The control unit 101 also calculates a tilt angular velocity proportional to the difference from the tilt control value ρ read from the volatile memory 102. Then, the control unit 101 stores the calculated control values in the volatile memory 102.
[0112] In step S108, the control unit 101 reads the pan value / tilt value from the volatile memory 102, converts it into a control command according to a predetermined protocol for controlling the sub-camera 400, stores it in the volatile memory 102, and proceeds to step S109.
[0113] In step S109, the control unit 101 transmits a control command corresponding to the pan value / tilt value calculated in step S108 to the sub-camera 400 via the communication unit 105, and returns the process to step S101.
[0114] The above is the basic operation of the WS100.
[0115] Next, the function and basic operation of the EB 200 will be described with reference to FIG. 3 and FIG. 4(b).
[0116] In step S201, the control unit 201 transmits a shooting command to the sub-camera 400 via the communication unit 205, receives a captured sub-image from the sub-camera 400, stores the captured sub-image in the volatile memory 202, and proceeds to step S202.
[0117] In step S202, the control unit 201 executes the function of the image recognition unit 221 in FIG. 3, and the process proceeds to step S203.
[0118] The image recognition unit 221 has the same functions as the image recognition unit 121 of the WS 100. The control unit 201 inputs the sub-image of the sub-camera 400 read from the volatile memory 202 by the inference unit 204 into a trained model created by machine learning such as deep learning, and performs inference processing. The inference results include coordinate information POSITION of the subject detected from the sub-image of the sub-camera 400, feature information STAT_SUB[m], and identification information ID for each subject, and are saved in the volatile memory 202. Note that the trained model used for the inference processing of the image recognition unit 221 is a common model (inference model for subject detection, inference model for subject identification) with the trained model used by the image recognition unit 121 of the WS 100.
[0119] In step S203, the control unit 201 receives the feature information STAT of the subject from the WS 100 via the communication unit 205, and compares it with the feature information STAT_SUB calculated from the sub-image of the sub-camera 400 using the function of the tracking target determination unit 222 in Fig. 3. If a subject with a high degree of similarity between the feature information STAT and the feature information STAT_SUB is present within the shooting angle of view of the sub-camera 400, the control unit 201 determines the identification information ID of the subject as the identification information ID of the subject to be tracked by the sub-camera 400 = SUBJECT_ID, stores it in the volatile memory 102, and proceeds to step S204. The method of calculating the similarity will be described in detail later.
[0120] In step S204, the control unit 201 checks the communication state for the WS 100 to stop tracking or continue tracking via the communication unit 205, performs processing according to the communication content, and proceeds to step S205. Details of the processing in step S204 will be described later.
[0121] In step S205, the control unit 201 determines whether or not information on the SUBJECT_ID of the tracking subject is stored in the volatile memory 202. If the control unit 201 determines that information on the SUBJECT_ID of the tracking subject is stored in the volatile memory 202, that is, that the identification information ID of the tracking subject of the sub camera 400 is stored in the volatile memory 102, the process proceeds to step S206. If the control unit 201 determines that information on the SUBJECT_ID of the tracking subject is not stored in the volatile memory 202, that is, that the identification information ID of the tracking subject of the sub camera 400 is not stored in the volatile memory 102, the process returns to step S201.
[0122] In step S206, the control unit 201 reads out the identification information ID for each subject, which is the subject recognition result of step S202, from the volatile memory 202, and determines whether or not the tracking subject SUBJECT_ID exists in the sub-image of the sub-camera 400. If the control unit 201 determines that the tracking subject SUBJECT_ID exists in the sub-image, the process proceeds to step S207, and if it determines that the tracking subject SUBJECT_ID does not exist, the process returns to step S201.
[0123] In step S207, the control unit 101 executes the function of the control information generation unit 223 in FIG. 3, and proceeds to step S208.
[0124] The control information generation unit 223 has a function of calculating pan values / tilt values of the sub-camera 400. The control unit 201 reads the subject coordinate information POSITION and the tracking subject SUBJECT_ID from the volatile memory 202, and identifies the current position of the tracking subject corresponding to the tracking subject SUBJECT_ID. The control unit 201 reads the past positions of the tracking subject within the shooting angle of view from the volatile memory 202, and calculates a larger angular velocity for panning if there is a large difference in the horizontal direction between the current position of the tracking subject and the past position of the tracking subject, and calculates a larger angular velocity for tilting if there is a large difference in the vertical direction. The control unit 201 saves the pan values / tilt values in the volatile memory 202.
[0125] In step S208, the control unit 201 converts the pan / tilt values read from the volatile memory 202 into control commands in accordance with a predetermined protocol for controlling the sub-camera 400, stores the control commands in the volatile memory 202, and proceeds to step S209.
[0126] In step S209, the control unit 201 transmits a control command corresponding to the pan value / tilt value calculated in step S208 to the sub-camera 400 via the communication unit 205, and returns the process to step S101.
[0127] The above is the basic operation of the EB200.
[0128] As described above, the WS100 performs image recognition processing on the overhead image of the overhead camera 300, and controls the panning / tilting operations of the sub camera 400 when the tracking state information STATE is "tracking by the WS100." It does not control the panning / tilting operations of the sub camera 400 when it is "tracking by the EB200." The EB200 performs image recognition processing on the sub image of the sub camera 400, and controls the panning / tilting operations of the sub camera 400 when a tracking subject is set and detected from the sub image. If a tracking subject is not set, it does not control the panning / tilting operations of the sub camera 400. 9, it is possible to switch between the WS100 and the EB200 controlling the sub-camera 400 by updating the tracking state information STATE and the settings of the tracked subject. Note that by transmitting the pan / tilt values only from one device controlling the sub-camera 400 and not transmitting them when the other device is in control, it is possible to reduce the amount of communication traffic compared to transmitting the pan / tilt values for each process shown in FIGS. 4(a) and 4(b).
[0129] Next, the operation of the overhead camera 300 when a shooting command is received from the WS 100 will be described with reference to FIG. 4(c).
[0130] In step S301, the control unit 301 receives a shooting command from the WS 100 via the communication unit 305, and the process proceeds to step S302.
[0131] In step S302, the control unit 301 starts the photographing process in response to receiving the photographing command via the communication unit 305, and proceeds to step S303. The control unit 301 captures an image using the imaging unit 306, performs predetermined image processing on the image using the image processing unit 307, and stores the generated image data in the volatile memory 302.
[0132] In step S303, the control unit 301 reads the image data from the volatile memory 302 and transmits it to the WS 100 via the communication unit 305.
[0133] The above is the operation of the overhead camera 300.
[0134] Next, the operation of the sub-camera 400 that receives a control command from the WS 100 or the EB 200 will be described with reference to FIG.
[0135] In step S401, the control unit 401 receives a control command via the communication unit 405, stores the control command in the volatile memory 402, and proceeds to step S402.
[0136] In step S402, the control unit 401 reads out the pan / tilt values from the volatile memory 402 in response to receiving a control command from the communication unit 405, and the process proceeds to step S403.
[0137] In step S403, the control unit 401 calculates drive parameters for controlling panning / tilting in a desired direction at a desired speed based on the panning / tilting values read from the nonvolatile memory 403, and proceeds to step S404. The drive parameters are parameters for controlling the actuators for the panning and tilting directions included in the PTZ drive unit 409, and the panning / tilting values included in the control command are converted into drive parameters by referring to a conversion table stored in the nonvolatile memory 403.
[0138] In step S404, the control unit 401 controls the optical unit 408 by the PTZ driving unit 409 based on the driving parameters calculated in step S403, and changes the shooting direction of the sub-camera 400. The PTZ driving unit 409 changes the shooting direction of the sub-camera 400 by driving the optical unit 408 in the pan / tilt direction based on the driving parameters.
[0139] The above is the operation of the sub-camera 400.
[0140] Next, the control process of the WS 100 will be described with reference to FIG.
[0141] FIG. 9(a) shows the control process of the WS 100, and shows the detailed process of step S105 in FIG. 4(a).
[0142] Part of the processing in FIG. 9(a) is realized by the control unit 101 executing the function of the tracking state determination unit 126 in FIG.
[0143] The tracking state determination unit 126 has a function of updating the tracking state information STATE stored in the volatile memory 102 .
[0144] In step S110, the control unit 101 reads out the SUBJECT_ID of the tracking subject of the sub camera 400 calculated in step S104 of Fig. 4A and the identification information ID indicating the previous tracking subject from the volatile memory 102. Then, the control unit 101 compares the identification information read out from the volatile memory 102 to determine whether the tracking subject of the sub camera 400 has changed. If the control unit 101 determines that the tracking subject of the sub camera 400 has changed, the process proceeds to step S111; if it determines that the tracking subject has not changed, the process proceeds to step S113.
[0145] In step S111, the control unit 101 transmits a tracking stop command to the EB 200 via the communication unit 105, and the process proceeds to step S112.
[0146] In step S112, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 3, and changes the tracking state information STATE to "tracking by WS 100."
[0147] If the subject being tracked by the sub-camera 400 has changed, it is highly likely that the subject being tracked will no longer be within the angle of view of the sub-camera 400. In this case, by performing the processes of steps S111 and S112, the WS 100 controls the sub-camera 400 based on the overhead image from the overhead camera 300 instead of the sub-camera 400.
[0148] In step S113, the control unit 101 reads the tracking state information STATE from the volatile memory 102, and determines whether "tracking by WS100" or "tracking by EB200" based on the tracking state information STATE. If the control unit 101 determines that "tracking by WS100" is occurring, the process proceeds to step S117; if the control unit 101 determines that "tracking by EB200" is occurring, the process proceeds to step S114.
[0149] In step S114, the control unit 101 sends a tracking continuation confirmation request to the EB200 via the communication unit 105, inquiring whether or not the EB200 can continue tracking the tracked subject. The response from the EB200 is either "Continue tracking OK" or "Continue tracking NG." If the control unit 101 receives a notification from the EB200 that "Continue tracking OK," the process returns to step S101, and if the control unit 101 receives a notification from the EB200 that "Continue tracking NG," the process proceeds to step S115.
[0150] In step S115, the control unit 101 transmits a tracking stop command to the EB 200 via the communication unit 105, and the process proceeds to step S116.
[0151] In step S116, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 3 to update the tracking state information STATE to "tracking by WS 100", and ends the process.
[0152] By performing the processes from step S114 to S116, if the tracking state is "tracking by EB200", tracking can be continued by WS100 even if the EB200 becomes unable to perform tracking.
[0153] In step S117, the control unit 101 determines whether or not the tracking subject is present within the shooting angle of view of the sub camera 400. If the control unit 101 determines that the tracking subject is present within the shooting angle of view of the sub camera 400, the process proceeds to step S118. If the control unit 101 determines that the tracking subject is not present within the shooting angle of view of the sub camera 400, the process ends. Whether or not a tracking subject is present within the shooting angle of view of the sub-camera 400 can be determined by comparing the current pan / tilt values acquired by the control unit 101 from the sub-camera 400 with the new pan / tilt values calculated in step S107 of Figure 4(a). If the current pan value / tilt value is sufficiently close to the new pan value / tilt value, it can be determined that the tracking subject is present within the shooting angle of view of the sub camera 400. Alternatively, if the pan / tilt speed value calculated in step S108 is sufficiently small, it can be determined that the tracking subject is present within the shooting angle of view of the sub camera 400 because the current pan value / tilt value is approaching the new pan value / tilt value.
[0154] In step S118, control unit 101 acquires shooting information from sub-camera 400 to determine whether or not to trim the overhead image and the subject area to be trimmed, and then proceeds to step S119. Details of the trimming process will be described later. Here, for simplicity, the description will continue assuming that trimming is not performed.
[0155] In step S119, the control unit 101 executes the function of the feature information determination unit 125 in FIG. 3, and the process proceeds to S120.
[0156] The feature information determination unit 125 has a function of determining feature information of the subject being tracked by the sub camera 400, i.e., feature information of the subject to be transmitted to the EB 200. The feature information determination unit 125 reads out feature information STAT[n] of the subject detected by the image recognition unit 121 from the overhead image of the overhead camera 300 from the volatile memory 102. The feature information determination unit 125 also reads out identification information SUBJECT_ID of the subject to be tracked determined by the tracking target determination unit 123 from the volatile memory 102. The feature information determination unit 125 then determines feature information STAT[i] corresponding to the subject to be tracked from the feature information STAT[n] and stores it in the volatile memory 102, where i is an index indicating the subject to be tracked.
[0157] In step S120, the control unit 101 transmits a tracking start command and feature information STAT[i] of the tracking subject to the EB 200 via the communication unit 105, and the process proceeds to step S121.
[0158] By the processing of steps S117 to S120, a tracking start command and feature information of the tracking subject can be sent to the EB200 only when there is a high possibility that the tracking subject is present within the shooting angle of view of the sub-camera 400. This makes it possible to reduce the amount of communication traffic compared to when information is sent for each process of Figures 4(a) and 9(a).
[0159] In step S121, the control unit 101 receives the matching result of the subject from the EB 200 via the communication unit 105. If the control unit 101 receives match information from the EB 200 indicating that the subjects match, the process proceeds to step S122, and if the control unit 101 receives mismatch information indicating that the subjects do not match, the process ends.
[0160] In step S122, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 3, changes the tracking state information STATE to "Tracking by EB200", and ends the process.
[0161] Next, the control process of the EB 200 will be described with reference to FIGS. 9(b), 9(c), and 10. FIG.
[0162] FIG. 9(b) shows the control process of the EB 200, and shows the detailed process of step S203 in FIG. 4(b).
[0163] In step S210, control unit 201 determines whether or not it has received, via communication unit 205, from WS 100 a tracking start command and feature information STAT[i] of the tracking subject obtained from the overhead image of overhead camera 300. If control unit 201 has received a tracking start command and feature information STAT[i] of the tracking subject from WS 100, it proceeds to step S211;
[0164] In steps S211 to S214, the control unit 201 executes the function of the tracking target determination unit 222 in Figure 3 and determines whether the feature information STAT[i] received from WS100 and the feature information STAT_SUB[m] obtained from the sub-image of the sub-camera 400 satisfy predetermined conditions.
[0165] The tracking target determination unit 222 has a function of calculating the similarity between the feature information STAT[i] received from the WS 100 and the feature information STAT_SUB[m] obtained from the sub-image of the sub-camera 400. The tracking target determination unit 222 also has a function of comparing the similarity of the feature information with a threshold value stored in the volatile memory 102, and storing the comparison result in the volatile memory 102. For example, if two people are present in the sub-image of the sub-camera 400, the tracking target determination unit 222 calculates the similarity between the feature information of the two people (STAT_SUB[1], STAT_SUB[2]) and the feature information STAT[i] received from the WS 100. The similarity is calculated as the cosine similarity between the feature information vectors, and a value between 1 and 0 is obtained as the similarity. The control unit 201 stores the similarities calculated for the m subjects in the volatile memory 202.
[0166] In step S211, the control unit 201 executes the function of the tracking target determination unit 222 in FIG. 3 to perform a process of matching feature information, and then the process proceeds to step S212.
[0167] In step S212, control unit 201 determines whether or not a subject with highly similar feature information exists, based on the comparison result of step S211. The existence of a subject with highly similar feature information means that the same subject has been photographed by overhead camera 300 and sub-camera 400. If control unit 201 determines that a subject with highly similar feature information exists, control proceeds to step S214; if control unit 201 determines that a subject with highly similar feature information does not exist, control proceeds to step S213.
[0168] The control unit 201 reads a predetermined threshold value from the volatile memory 202, and determines that a subject with high similarity in feature information exists if the similarity is equal to or greater than the threshold value as a predetermined condition, or if a subject with a higher similarity exists, or if the subjects match, and stores the identification information ID of the subject in the volatile memory 202. Furthermore, the control unit 201 updates information MATCH, which indicates whether or not there is a subject with high similarity in feature information, and stores the information in the volatile memory 202. In this embodiment, if the value of MATCH is 0, there is no subject with high similarity in feature information, i.e., the subjects do not match between the overhead-view camera 300 and the sub-camera 400. If the value of MATCH is 1, there is a subject with high similarity in feature information, i.e., the subjects match between the overhead-view camera 300 and the sub-camera 400. If there is a subject with high similarity in feature information, the control unit 201 stores MATCH=1 in the volatile memory 202 and proceeds to step S214. If there is no subject with high similarity in feature information, the control unit 201 stores MATCH=0 in the volatile memory 202 and proceeds to step S213.
[0169] Here, the similarity of the feature information of the subject detected from the overhead image captured by the overhead camera 300 and the sub-image captured by the sub-camera 400 will be described with reference to FIG.
[0170] Fig. 10(a) shows the positional relationship between the shooting position and shooting direction of overhead camera 300 and the shooting position and shooting direction of sub-camera 400. Fig. 10(b) shows the subject detected from the overhead image of overhead camera 300 and the tracked subject.
[0171] Assume that subjects A, B, and C are detected from the overhead image of overhead camera 300, and the subject being tracked by sub-camera 400 is subject C. The feature information of the subject being tracked by sub-camera 400 transmitted from sub-camera 400 to WS100 is information corresponding to subject C. Figures 10(c) and (e) show the sub-images of sub-camera 400, and Figures 10(d) and (f) show the similarity between the feature information of the subject being tracked by sub-camera 400 and the feature information of the subject detected from the sub-images.
[0172] As shown in Figure 10(c), when sub camera 400 is capturing images of subjects A and B, the similarity is calculated between the feature information of subject C in the overhead image captured by overhead camera 300 and the feature information of subject A or subject B in the sub image captured by sub camera 400. As shown in Figure 10(d), the similarity between the feature information of subject C in the overhead image captured by overhead camera 300 and the feature information of subject A or subject B in the sub image captured by sub camera 400 is low. In this case, for example, if the threshold for subject similarity is 0.7, the result will be that subjects A and B do not match.
[0173] 10(e), when sub-camera 400 is capturing images of subjects B and C, the similarity is calculated between the feature information of subject C in the overhead image captured by overhead camera 300 and the feature information of subject B or subject C in the sub-image captured by sub-camera 400. Subject C in the overhead image captured by overhead camera 300 and subject C in the sub-image captured by sub-camera 400 have different camera positions and directions, and therefore different shapes in the images. For example, if subject C is facing its face or body toward overhead camera 300, subject C will face forward in the overhead image from overhead camera 300 and will be facing closer to the side in the sub-image from sub-camera 400. The inference models for identifying subjects in image recognition unit 121 of WS100 and image recognition unit 221 of EB200 are models that learn from images of the same subject captured from multiple different directions. For this reason, the same subject captured by multiple cameras with different shooting positions and directions will have a high degree of similarity in feature information even if the morphology in each captured image is different. 10(f), the similarity between the feature information of subject C in the overhead image captured by overhead camera 300 and the feature information of subject C in the sub-image captured by sub camera 400 is high. As a result, for example, when the threshold value for subject similarity is 0.7, subject B is determined to be a mismatch and subject C is determined to be a match, and subject C can be determined to be the same subject.
[0174] Returning to the explanation of FIG. 9B, in step S213, the control unit 201 reads MATCH=0 from the volatile memory 202, transmits it to the WS 100 via the communication unit 205, and ends the process.
[0175] In step S214, the control unit 201 reads from the volatile memory 102 the identification information ID of the subject for which the highest similarity has been calculated, stores it in the volatile memory 102 as the tracking subject SUBJECT_ID, and proceeds to step S215. By selecting the subject for which the highest similarity has been calculated, even if there are subjects wearing similar clothing, for example, it is possible to set the most likely subject among them as the tracking target.
[0176] In step S215, the control unit 201 reads MATCH=1 from the volatile memory 202, transmits it to the WS 100 via the communication unit 205, and ends the process.
[0177] FIG. 9(c) shows the control process of the EB 200, and shows the detailed process of step S204 in FIG. 4(b).
[0178] In step S220, the control unit 201 determines whether or not a tracking stop command has been received from the WS 100 via the communication unit 205. If the control unit 201 has received a tracking stop command from the WS 100, the process proceeds to step S221; if not, the process proceeds to step S223.
[0179] In step S221, the control unit 201 transmits a control command to stop the panning / tilting operation to the sub-camera 400 via the communication unit 205, and the process proceeds to step S222.
[0180] In step S222, the control unit 201 deletes the tracking subject SUBJECT_ID stored in the volatile memory 202, and the process proceeds to step S201.
[0181] In step S223, the control unit 201 determines whether or not a tracking continuation confirmation request has been received from the WS 100 via the communication unit 205. If the control unit 201 has received a tracking continuation confirmation request from the WS 100, the process proceeds to step S224;
[0182] In step S224, the control unit 201 reads out the subject recognition result by the image recognition unit 221 from the volatile memory 202, and determines whether or not the tracking subject SUBJECT_ID has been detected. If the control unit 201 determines that the tracking subject SUBJECT_ID has been detected by the image recognition unit 221, the process proceeds to step S226, and if not, the process proceeds to step S225.
[0183] In step S225, the control unit 201 transmits "Continue tracking NG" to the WS 100 via the communication unit 205, and the process returns to step S201.
[0184] In step S226, the control unit 201 transmits "Continue tracking OK" to the WS 100 via the communication unit 205, and ends the process.
[0185] The above is the detailed control process of the EB200.
[0186] Next, the trimming process in step S118 in Fig. 9(a) will be described. Fig. 11 is a diagram illustrating the relationship between the subject to be tracked and the shooting angle of view, and describes a method for calculating the amount of trimming of the overhead image so that the size of the subject to be tracked in the overhead image matches the size of the subject to be tracked in the sub-image. In this embodiment, an example will be described in which control is performed based on the vertical size of the subject to be tracked.
[0187] As shown in FIG. 11(a), in the overhead image of overhead camera 300, the entire tracking subject C is within the shooting angle of view of overhead camera 300, whereas in the sub-image of sub-camera 400 shown in FIG. 11(c), only a portion of tracking subject C is within the shooting range (part of the subject is outside the shooting angle of view (the subject is cut off)), there is a possibility that the similarity between the feature information of the tracking subject extracted from the overhead image of overhead camera 300 and the feature information of the tracking subject extracted from the sub-image of sub-camera 400 will decrease. For example, this is the case when feature information of the entire body of a person as the tracking subject is extracted from the overhead image of overhead camera 300, while feature information of only the face of the tracking subject is extracted from the sub-image of sub-camera 400.
[0188] If there is a large difference between the size of the tracked subject in the overhead image of the overhead camera 300 (hereinafter referred to as the tracked subject size in the overhead image) and the size of the tracked subject in the sub-image of the sub-camera 400 (hereinafter referred to as the tracked subject size in the sub-image), the similarity of the feature information described in Figure 10(f) will decrease, and there is a possibility that subjects that actually match will be determined to be mismatched.
[0189] Therefore, in this embodiment, when the difference between the size of the tracked subject in the overhead image and the size of the tracked subject in the sub-image is large, the amount of trimming of the overhead image is controlled so that the size of the tracked subject in the overhead image and the size of the tracked subject in the sub-image become closer. This makes it less likely that the subject will be cut off, and therefore reduces the decrease in similarity of the feature information due to the subject being cut off.
[0190] FIG. 11(b) shows the relationship between the tracking subject and the angle of view of the overhead camera 300. FIG. 11(d) shows the relationship between the tracking subject and the angle of view of the sub-camera 400. In FIG. 11, hcam denotes the camera height, φ denotes the vertical angle of view, H denotes the vertical field of view, L denotes the distance to the tracking subject, and hobj denotes the overall height of the subject (the actual height of the subject in real space). In FIG. 11(b), the vertical field of view H includes the ground at the subject's position, so if the subject's overall height hobj is less than the vertical field of view H, no part of the subject will fall outside the angle of view. In FIG. 11(d), the vertical field of view H does not reach the ground at the subject's position, so no part of the subject will fall outside the angle of view. In FIG. 11, the height hview of the subject within the vertical field of view, which is the size of the tracking subject for the overhead camera 300 and the sub-camera 400, can be calculated using the following equation 5. (Formula 5) TIFF0007805396000005.tif26155In this embodiment, the amount of trimming is determined from the relationship between the size information hview of the subject in the vertical field of view H and the overall size information hobj of the subject in the above formula 5. This makes it possible to determine the subject area for calculating feature information according to the size information of the subject in the vertical field of view H.
[0191] FIG. 12 illustrates the trimming process in step S118 of FIG. 9(a).
[0192] In step S123, the control unit 101 reads out subject information from the volatile memory 102. The subject information is the overall height hobj of the subject described in FIG. 11. The subject information is set in advance by a user operation via the operation unit 106 or a setting instruction via the communication unit 105, and is saved in the volatile memory 102. Note that the subject information may be read out from the volatile memory 102 or the non-volatile memory 103 so as to be automatically selected depending on the type of subject. The control unit 101 saves the subject information in the volatile memory 102, and the process proceeds to step S124.
[0193] In step S124, the control unit 101 acquires shooting information from the overhead camera 300 and the sub-camera 400 via the communication unit 105. The shooting information is information for calculating the vertical field of view H shown in FIG. 11. The vertical field of view H can be calculated using the following equations 6 and 7. The distance L can be calculated using the above equation 3, and the vertical angle of view φ can be calculated from the vertical size s of the image sensors of the overhead camera 300 and the sub-camera 400 and the focal length f at the time the information was acquired. Therefore, the control unit 101 can calculate the vertical field of view H by acquiring information about the vertical size s and focal length f of the image sensors from the overhead camera 300 and the sub-camera 400. (Formula 6) TIFF0007805396000006.tif13155 (Formula 7) TIFF0007805396000007.tif13155 The control unit 101 stores the subject information acquired from the overhead camera 300 and the sub-camera 400 in the volatile memory 102, and the process proceeds to step S125.
[0194] In step S125, control unit 101 compares overall size information hobj of the subject with size information hview of the subject in the vertical field of view H. Control unit 101 reads the information saved in steps S123 and S124 from volatile memory 102 and calculates the vertical field of view H. Next, control unit 101 calculates the height hview of the subject within the field of view using the vertical field of view H. Control unit 101 then calculates the ratio of the subject height hview to the subject height hobj as the degree of match r with the overhead image using the following equation 8, saves the calculated degree of match r with the overhead image in volatile memory 102, and proceeds to step S126. (Formula 8) TIFF0007805396000008.tif13155 In step S126, the control unit 101 determines the amount of trimming based on the table read out from the nonvolatile memory 103 in accordance with the degree of match r with the overhead image.
[0195] Here, with reference to FIG. 13, a table showing the relationship between the degree of coincidence r with the overhead image and the amount of trimming, and a method for determining the amount of trimming, will be described.
[0196] In FIG. 13(a), if the degree of match r with the overhead image is greater than or equal to 1 / 2 and less than 1, it is determined that more than half of the entire subject is within the field of view, and the amount of trimming is set to 0. If the degree of match r with the overhead image is greater than or equal to 1 / 3 and less than 1 / 2, it is determined that more than half of the entire subject is outside the field of view, and the amount of trimming is set to up to half of the entire subject, as shown in FIG. 13(b). If the degree of match r with the overhead image is less than 1 / 3, it is determined that more than two-thirds of the subject is outside the field of view, and the amount of trimming is set to up to two-thirds of the entire subject. In this embodiment, since sub-camera 400 captures a person's face as the subject, the direction of trimming is set vertically upward, starting from the feet.
[0197] Note that multiple cropping directions may be set taking into account the horizontal width of the subject depending on the subject's characteristics and the shooting method. The cropping amount may be dynamically calculated from the subject area by defining it as a ratio according to the degree of matching r with the overhead image, rather than the degree of matching r with the overhead image. For example, the cropping amount may be calculated using the following equation 9, where T is the cropping amount: (Formula 9) TIFF0007805396000009.tif7155 The control unit 101 stores the trimming amount determined as described above in the volatile memory 102, and the process proceeds to step S127.
[0198] In step S127, the control unit 101 determines whether or not to perform trimming based on the result of the trimming amount determination process in step S126. If the trimming amount determined in step S126 is not stored in the volatile memory 102, the control unit 101 determines that trimming will not be performed, terminates the process of Fig. 12, and proceeds to step S119 in Fig. 9(a). If the trimming amount determined in step S126 is stored in the volatile memory 102, the control unit 101 determines that trimming will be performed, and proceeds to step S128.
[0199] In step S128, control unit 101 causes inference unit 104 to perform subject detection using the inference model for subject detection described above. That is, as shown in Fig. 6(a), coordinates of a circumscribing rectangle of the subject are acquired from the overhead image captured by overhead camera 300. Control unit 101 causes image processing unit 307 to perform a trimming process on the circumscribing rectangle based on the trimming amount determined in step S126, as shown in Fig. 13.
[0200] The image processing unit 307 generates a subject image by cutting out the portion of the subject that is out of the field of view from the entire subject as shown in Fig. 11(c), and inputs the generated subject image to the image recognition unit 121. The control unit 101 stores the coordinate information of the subject in the subject image that has been subjected to the trimming process in the volatile memory 102, ends the process of Fig. 12, and proceeds to step S119 of Fig. 9(a). In this embodiment, the description will proceed assuming that the amount of trimming determined in step S126 is 1 / 2.
[0201] 9(a), in step S119, the control unit 101 reads from the volatile memory 102 the subject image that has been subjected to the trimming process in step S128, if any, and executes the function of the feature information determination unit 125. The trimmed subject image is input, and the control unit 101 reads from the volatile memory 102 the subject feature information STAT[n] and the identification information SUBJECT_ID of the subject to be tracked determined by the tracking target determination unit 123. The control unit 101 determines STAT[i] corresponding to the subject to be tracked from the feature information STAT[n], stores it in the volatile memory 102, and proceeds to step S120.
[0202] The processing of steps S123 to S128 in Fig. 12 and step S119 in Fig. 9(a) makes it possible to calculate feature information of a portion of the subject that has been trimmed. The control unit 201 of the EB200 receives the feature information of the portion of the subject that has been trimmed from the WS100 and performs the same processing as in Fig. 4. The control unit 201 of the EB200 then compares this with the feature information of the subject detected by the sub-camera 400 when part of the subject is outside the field of view of the sub-camera 400, as shown in Fig. 11(c).
[0203] For example, if the overhead image captured by the overhead camera 300 is shown in Fig. 13(c) and the sub-image captured by the sub-camera 400 is shown in Fig. 13(d), the subject portion above the dashed line is cut out by performing a trimming process on the overhead image in Fig. 13(c). This results in a state similar to that of the subject image in Fig. 13(d), and then feature information about the subject is calculated and compared with the feature information about the subject in the sub-image in Fig. 13(d).
[0204] In the example of Figure 13(e), when no trimming processing is performed, the similarity of the feature information is 0.4, but when trimming processing is performed with the trimming amount reduced to 1 / 2, the similarity of the feature information becomes 0.8, and it is determined that the subjects match between the overhead image and the sub-image.
[0205] According to the first embodiment described above, the same subject can be recognized by the multiple cameras 300 and 400 with different shooting positions and shooting directions. Therefore, a specific subject can be tracked by appropriately switching between control of the sub-camera 400 by the WS 100 and control of the sub-camera 400 by the EB 200.
[0206] When the subject to be tracked does not exist in the sub-image of the sub-camera 400, the WS 100 controls the sub-camera 400, and when the subject to be tracked exists in the angle of view of the sub-camera 400, the control of the sub-camera 400 can be handed over from the WS 100 to the EB 200. Furthermore, when the subject to be tracked moves at high speed and is lost, or when the subject to be tracked is changed, the WS 100 controls the sub-camera 400, allowing tracking to continue.
[0207] In the first embodiment, the WS 100 and the EB 200 switch whether or not to transmit pan / tilt values to the sub-camera 400. However, this is not limiting. For example, the WS 100 and the EB 200 may transmit pan / tilt values to the sub-camera 400 regardless of the tracking state, and the sub-camera 400 may control which device the pan / tilt values received from to perform the pan / tilt operation. In this case, the WS 100 may omit step S106 in FIG. 4(a) and add a process in which the control unit 101 transmits tracking state information STATE to the sub-camera 400 before step S107 in FIG. 4(a). The EB 200 may omit steps S205 and S206 in FIG. 4(b) and step S221 in FIG. 9(c).
[0208] If the tracking state information STATE received from the WS 100 is "tracking by the EB 200", the sub camera 400 performs control to perform panning / tilting operations in accordance with the control command received from the EB 200. If the tracking state information STATE received from the WS 100 is "tracking by the WS 100", the sub camera 400 performs control to perform panning / tilting operations in accordance with the control command received from the WS 100.
[0209] [Embodiment 2] In the first embodiment, the feature information of the subject transmitted from the WS 100 to the EB 200 is one piece of feature information of the tracked subject. In the second embodiment, an example will be described in which the WS 100 generates multiple feature information candidates and transmits them to the EB 200, and the feature information is compared in the EB 200.
[0210] In the first embodiment, the trimming amount is determined by the process of Fig. 12, and feature information is calculated for the subject image that has been subjected to the trimming process in step S119. In the second embodiment, a plurality of subject images that have been subjected to the trimming process based on a plurality of preset trimming amounts are created, and feature information is calculated from each of the subject images.
[0211] The following description will focus on the differences from the first embodiment.
[0212] First, the operation of the WS100 will be described.
[0213] The WS 100 executes the trimming process in step S118 in FIG. 9A regardless of the subject information and the shooting information of the sub-camera 400 acquired in steps S123 and S124 in FIG.
[0214] In step S118, the control unit 101 creates subject images according to the three patterns of trimming amounts for the table shown in FIG.
[0215] In step S119, the control unit 101 calculates feature information for the three patterns of subject images created in step S118. In the first embodiment, one piece of feature information STAT[i] was determined for the tracking subject, but in the second embodiment, three pieces of feature information, STAT[i][0], STAT[i][1], and STAT[i][2], are determined for one subject. Note that the trimming amount is not limited to three patterns, and any number of patterns may be set. For example, if there are p patterns of trimming amount and the number of detected subjects is n, then there will be n×p patterns of feature information. Therefore, for a specific tracking subject SUBJECT_ID=i, feature information STAT[i][0] to STAT[i][p-1] is calculated.
[0216] In step S120, the control unit 101 transmits a tracking start instruction and the feature information STAT[i][0] to STAT[i][2] to the EB 200 via the communication unit 105.
[0217] Next, the operation of the EB200 will be described.
[0218] The second embodiment differs in that a plurality of pieces of characteristic information are received and collated in the process of FIG. 9(b).
[0219] In step S210, control unit 201 determines whether or not it has received, via communication unit 205, from WS 100, a tracking start command and characteristic information STAT[i][0] to STAT[i][2] of the tracking subject in the overhead image of overhead camera 300. If control unit 201 determines that it has received, from WS 100, a tracking start command and characteristic information STAT[i][0] to STAT[i][2] of the tracking subject in the overhead image of overhead camera 300, it proceeds to step S211;
[0220] In step S211, the control unit 201 executes the function of the tracking target determination unit 222 in Fig. 3. Unlike the first embodiment, the tracking target determination unit 222 compares the feature information STAT[i][0] to STAT[i][2] received in step S210 with the feature information STAT_SUB.
[0221] In step S212, the control unit 201 determines whether or not a subject with highly similar feature information exists based on the result of the comparison in step S211, as shown in Fig. 13(e). If the control unit 201 determines that a subject with highly similar feature information exists, the control unit 201 proceeds to step S214, and if the control unit 201 determines that a subject with highly similar feature information does not exist, the control unit 201 proceeds to step S213.
[0222] According to the second embodiment described above, by calculating feature information according to a plurality of trimming amount patterns without using subject information or shooting information of the sub-camera 400, it is possible to obtain the same effect as in the first embodiment.
[0223] [Embodiment 3] In the first and second embodiments, the overhead image is subjected to a trimming process to bring the feature information of the subject detected from the overhead image of the overhead camera 300 closer to the feature information of the subject detected from the sub-image of the sub camera 400.
[0224] In the third embodiment, an example will be described in which, when the feature information of the tracked subject between the overhead-view camera 300 and the sub-camera 400 does not match, zoom control is performed so that part of the subject does not fall outside the shooting angle of view of the sub-camera 400. In this way, by changing the shooting angle of view of the sub-camera 400 and recalculating the feature information of the subject, it is possible to obtain feature information of the entire subject from the sub-image. This control is effective when priority is placed on tracking the subject rather than changing the shooting angle of view.
[0225] The control process of the third embodiment will be described with reference to FIGS.
[0226] The processing from steps S110 to S122 in FIG. 14 is the same as the processing from steps S110 to S122 in FIG.
[0227] In step S121, if the control unit 101 receives mismatch information from the EB 200 indicating that the subjects do not match, the control unit 101 advances the process to step S130.
[0228] Here, the zoom control of the sub-camera 400 in step S130 of FIG. 14 will be described with reference to FIG.
[0229] In step S134, the control unit 101 reads the subject information from the volatile memory 102, similar to step S123, and the process proceeds to step S135.
[0230] In step S135, the control unit 101 acquires shooting information from the overhead camera 300 and the sub-camera 400 via the communication unit 105, and the process proceeds to step S136.
[0231] In step S136, the control unit 101 calculates a zoom value for the sub-camera 400 so that the angle of view of the sub-camera 400 approaches the angle of view of the overhead camera 300 capturing the tracked subject. The control unit 101 calculates the vertical field of view H of the overhead camera 300 based on the information acquired in steps S134 and S135.
[0232] Next, the control unit 101 calculates the height hview of the subject within the field of view of the overhead camera 300 based on the vertical field of view H of the overhead camera 300. Using the above formula 5, the control unit 101 calculates the vertical field of view H of the sub camera 400 such that the difference between the size information hview of the tracked subject in the vertical field of view H of the overhead camera 300 and the size information hview of the subject in the sub camera 400 matches.
[0233] Furthermore, the control unit 101 calculates the focal length f from the vertical field of view H of the sub-camera 400 and the above equations 6 and 7, and calculates the zoom value corresponding to the focal length f. The control unit 101 stores the zoom value in the volatile memory 102, and the process proceeds to step S137.
[0234] The height hview of the subject within the field of view of the sub-camera 400 does not necessarily have to be equal to the height hview of the subject within the field of view of the overhead camera 300, and may be closer (so that the difference is smaller) than before the start of processing. Also, if there are multiple subjects within the shooting field of view of the sub-camera 400, the control unit 101 selects the subject closest to the center of the field of view as the tracking subject.
[0235] In step S137, the control unit 101 reads from the volatile memory 102 the zoom value calculated in step S136, the focal length f of the sub camera 400 acquired in step S135, and a threshold value previously stored in the volatile memory 102. The control unit 101 calculates the difference between the zoom value corresponding to the focal length f of the sub camera 400 acquired in step S135 and the zoom value calculated in step S136. If the control unit 101 determines that the calculated difference is less than the threshold value, the process proceeds to step S131 in Fig. 14; if the control unit 101 determines that the calculated difference is equal to or greater than the threshold value, the process proceeds to step S138.
[0236] In step S138, the control unit 101 reads the zoom value calculated in step S136 from the volatile memory 102, and transmits a zoom control command corresponding to the zoom value from the WS 100 to the sub-camera 400 via the communication unit 105. When the control unit 401 of the sub-camera 400 receives the zoom control command via the communication unit 405, it controls the PTZ driving unit 409 to achieve the received zoom value.
[0237] This makes it possible to bring the shooting angle of view of sub-camera 400 closer to the shooting angle of view of overhead camera 300 capturing the tracked subject, that is, to bring it closer to the state of Fig. 11(a) from the state of Fig. 11(c). Therefore, the difference in size of the tracked subject between the overhead image of overhead camera 300 and the sub-image of sub-camera 400 becomes smaller, and it becomes possible for multiple cameras to work together to track a specific subject even when the shooting positions and shooting directions of multiple cameras are placed apart, and part of the subject is outside the shooting angle of view or the subject is too small.
[0238] The trimming process performed prior to the zooming process can deal to some extent with cases where a part of a specific subject is outside the angle of view of the image captured by only one of the overhead camera 300 and the sub-camera 400 (usually the sub-camera). However, the inventors have found through their studies that, for example, if the overhead camera 300 captures a small image of the entire body of a specific subject and the sub-camera 400 captures the face of the same subject, the similarity of the feature information may decrease even if the trimming process is performed. Therefore, in this embodiment, if the subjects still do not match even after the trimming process (NO in step S121), zooming is performed, thereby increasing the possibility that the subjects will match if the same subject is being photographed. Also, by changing the zoom value only when the difference between the current zoom value of the sub-camera 400 and the zoom value of the sub-camera 400 that causes the size information hview of the tracked subject in the vertical field of view H of the overhead-view camera 300 to match the size information hview of the subject of the sub-camera 400 exceeds a threshold, it is possible to change the shooting angle of view of the sub-camera 400 only when the difference in subject size is so large that it is likely to affect the determination of the similarity of the feature information.
[0239] After transmitting the zoom control command, the control unit 101 ends the process of FIG. 15, and proceeds to step S131.
[0240] The processing from steps S131 to S133 is the same as the processing from steps S119 to S121 in FIG. 9(a).
[0241] According to the third embodiment described above, by controlling the photographing angle of view of the sub-camera 400 so that the subject does not deviate from the field of view, it is possible to acquire feature information of the entire subject.
[0242] [Embodiment 4] In the first to third embodiments, examples have been described in which either the WS 100 or the EB 200 controls the sub-camera 400. In the fourth embodiment, an example will be described in which the EB 200 is omitted and the WS 100 controls the sub-camera 400 based on the overhead image from the overhead camera 300 and the sub-image from the sub-camera 400.
[0243] In the fourth embodiment, the sub camera 400 is controlled using either the pan value / tilt value calculated based on the overhead image of the overhead camera 300 or the pan value / tilt value calculated based on the sub image of the sub camera 400.
[0244] The system configuration of the fourth embodiment is a configuration in which the EB 200 is omitted from the system configuration of Fig. 1, and differs from the first embodiment in that the sub-image of the sub-camera 400 is input to the WS 100. The operations other than those of the WS 100 are the same as those of the first embodiment.
[0245] In basic operation, the overhead camera 300 transmits an overhead image to the WS 100. The sub-camera 400 transmits a sub-image to the WS 100. The sub-camera 400 also has a PTZ function.
[0246] WS100 detects a subject from the overhead image of overhead camera 300 and the sub-image of sub-camera 400, and changes the imaging direction of sub-camera 400 to the direction of the tracked subject based on the subject recognition result. WS100 controls sub-camera 400 based on the subject recognition result of the overhead image of overhead camera 300 until the imaging direction of sub-camera 400 becomes the direction of the tracked subject. After the shooting direction of sub-camera 400 is aligned with the direction of the tracked subject, WS100 calculates feature information of the tracked subject from the overhead image of overhead-view camera 300, and calculates feature information of the subject from the sub-image of sub-camera 400. Then, based on this feature information, WS100 controls sub-camera 400. The feature information is information that can identify the same subject when the same subject is photographed by multiple cameras with different shooting positions and / or directions.
[0247] According to the fourth embodiment, the sub camera 400 can be controlled based on the object recognition result of either the overhead image of the overhead camera 300 or the sub image of the sub camera 400, and the object to be tracked can be tracked.
[0248] The hardware configuration of the WS 100, the overhead camera 300, and the sub-camera 400 is the same as that shown in FIG. 2 of the first embodiment.
[0249] First, the functional configuration of the WS 100 for realizing the control processing of this embodiment will be described with reference to FIG.
[0250] The functions of WS 100 are realized by hardware and / or software. If the functional units shown in Figure 16 are configured by hardware instead of software, it is sufficient to have a circuit configuration corresponding to each functional unit shown in Figure 16.
[0251] WS100 includes an image recognition unit 121, an object of interest determination unit 122, a tracking target determination unit 123, a control information generation unit 124, a feature information determination unit 125, a tracking state determination unit 126, an image recognition unit 127, and a tracking target determination unit 128. Software that realizes these functions is stored in non-volatile memory 103, and is loaded into volatile memory 102 and executed by control unit 101.
[0252] The functions of the image recognition unit 121, the target subject determination unit 122, the tracking target determination unit 123, and the feature information determination unit 125 are the same as those in FIG. 3 of the first embodiment.
[0253] The functions and basic operations of the WS 100 will be described with reference to FIGS.
[0254] The processing from step S501 to step S504 is the same as the processing from step S101 to S104 in FIG. 4(a) of the first embodiment.
[0255] In step S505, the control unit 101 transmits a shooting command to the sub-camera 400 via the communication unit 105, receives the captured sub-image from the sub-camera 400, stores it in the volatile memory 102, and proceeds to step S506.
[0256] In step S506, the control unit 101 executes the function of the image recognition unit 127 in FIG. 16, and the process proceeds to step S507.
[0257] The functions of the image recognition unit 127 can be achieved by replacing the control unit 201 with the control unit 101, the volatile memory 202 with the volatile memory 102, and the nonvolatile memory 203 with the nonvolatile memory 103 in the description of the image recognition unit 221 of the EB 200 in the first embodiment.
[0258] 16, compares the feature information calculated in steps S502 and S506, and updates the tracking state information STATE. Furthermore, the control unit 101 stores the tracking subject SELECT_ID and tracking state information STATE in the volatile memory 102, and the process proceeds to step S508.
[0259] The tracking state information STATE includes information on either "tracking using overhead image" or "tracking using sub-image." "Tracking using overhead image" indicates a state in which the tracking subject is being tracked by controlling the sub-camera 400 based on the subject recognition result of the overhead image of the overhead camera 300. "Tracking using sub-image" indicates a state in which the tracking subject is being tracked by controlling the sub-camera 400 based on the subject recognition result of the sub-image of the sub-camera 400. The processing of step S507 will be described in detail later.
[0260] The processing of steps S508 to S510 is executed by the function of the control information generating unit 124 in FIG.
[0261] In step S508, the control unit 101 reads the tracking state information STATE from the volatile memory 102, and determines whether “tracking using an overhead image” or “tracking using a sub-image” is in progress based on the tracking state information STATE. If the control unit 101 determines that “tracking using an overhead image” is in progress, the process proceeds to step S510; if the control unit 101 determines that “tracking using a sub-image” is in progress, the process proceeds to step S509.
[0262] In step S509, the control unit 101 calculates the pan value / tilt value of the sub camera 400 based on the subject recognition result of the sub image of the sub camera 400, and proceeds to step S511. The processing of step S509 can be performed by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the processing of the control information generation unit 223 in FIG.
[0263] In step S510, the control unit 101 calculates the pan value / tilt value of the sub-camera 400 based on the object recognition result of the overhead image of the overhead camera 300, and the process proceeds to step S511. The process of step S510 can be performed by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the process of the control information generation unit 223 in FIG.
[0264] In step S511, the control unit 101 executes the function of the control information generating unit 124 in FIG. 3, and the process proceeds to step S512.
[0265] The processes in steps S511 and S512 are similar to those in steps S108 and S109 in FIG. 4(a).
[0266] The above is the basic operation of the WS100.
[0267] Next, the control process of the WS 100 will be described with reference to FIG.
[0268] FIG. 18 shows the control process of the WS 100, and shows the detailed process of step S507 in FIG.
[0269] The process of step S520 is the same as that of step S110 in FIG. 9(a).
[0270] In step S521, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using an image from an overhead camera."
[0271] The tracking state determination unit 126 has a function of updating the tracking state information STATE stored in the volatile memory 102 .
[0272] In step S522, the control unit 101 reads the tracking state information STATE from the volatile memory 102, and determines whether “tracking using an overhead image” or “tracking using a sub-image” is in progress based on the tracking state information STATE. If the control unit 101 determines that “tracking using an overhead image” is in progress, the process proceeds to step S525; if the control unit 101 determines that “tracking using a sub-image” is in progress, the process proceeds to step S523.
[0273] The process of step S523 can be achieved by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the process of step S224 in FIG. 9(b).
[0274] In step S524, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using overhead image."
[0275] The processes in steps S525, S526a, and S527 are similar to those in steps S117, S118, and S119 in FIG. 9(a).
[0276] The processing of steps S528 to S530 can be achieved by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the processing of steps S211 to S214 in FIG. 9(b).
[0277] In step S531, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using sub-image", and ends the process.
[0278] According to the above-described fourth embodiment, WS 100 switches between the subject recognition results of the overhead image from overhead camera 300 and the sub-image from sub-camera 400 to control sub-camera 400. This eliminates the need for EB 200 in the first embodiment, and simplifies the system configuration while achieving the same effects as in the first embodiment.
[0279] [Embodiment 5] In the first to fourth embodiments, the overhead image is subjected to a trimming process to bring the feature information of the subject detected from the overhead image of the overhead camera 300 closer to the feature information of the subject detected from the sub-image of the sub-camera 400. In the fifth embodiment, an example will be described in which zoom control is performed so that part of the subject does not fall outside the shooting angle of view of the sub-camera 400.
[0280] The following description will focus on the differences from embodiments 1 to 4. First, the operation of the WS 100 will be described.
[0281] The control process of the fifth embodiment will be described with reference to FIG.
[0282] The processes in steps S110 to S117 and S119 to S122 in FIG. 19 are the same as the processes in steps S110 to S117 and S119 to S122 in FIG.
[0283] In step S150, the control unit 101 acquires subject information of the tracked subject included in the overhead image, subject information included in the sub-image, and shooting information from the sub-camera 400, thereby determining the zoom value of the sub-camera 400 and performing zoom control of the sub-camera 400.
[0284] 11, a method for calculating the zoom value of the sub camera 400 so that the size of the subject to be tracked in the overhead image and the size of the subject to be tracked in the sub image will be described in the zoom control in step S150 of Fig. 19. In this embodiment, an example will be described in which zoom control is performed based on the vertical size of the subject to be tracked.
[0285] In this embodiment, when the difference between the size of the tracked subject in the overhead image and the size of the tracked subject in the sub-image is large, the zoom value of the sub-camera 400 is controlled so that the size of the tracked subject in the overhead image and the size of the tracked subject in the sub-image become closer. This makes it less likely that the subject will be cut off, and therefore reduces the decrease in similarity of the feature information due to the subject being cut off.
[0286] In this embodiment, the zoom value of the sub camera 400 is calculated so that the size of the tracked subject in the sub image matches the size of the tracked subject in the overhead image. The calculated zoom value is then compared with the current zoom value of the sub camera 400, and if the difference between the calculated zoom value and the current zoom value of the sub camera 400 exceeds a threshold, the zoom value of the sub camera 400 is controlled so that the size of the tracked subject in the sub image approaches the size of the tracked subject in the overhead image. Note that "approaching" the size of the tracked subject in the sub image to the size of the tracked subject in the overhead image includes "matching." This reduces the possibility of erroneously determining that subjects that actually match are not matching.
[0287] In this embodiment, the zoom value of the sub-camera 400 is controlled so that the difference between the size information hview of the tracked subject in the vertical field of view H of the overhead camera 300 and the size information hview of the subject in the vertical field of view H of the sub-camera 400, calculated using the above formula 5, does not become too large.
[0288] If the difference between the current zoom value of sub camera 400, at which size information hview of the tracking subject in the vertical field of view H of overhead camera 300 matches size information hview of the subject in the vertical field of view H of sub camera 400, exceeds a threshold, the zoom value of sub camera 400 is controlled so that the difference is equal to or less than the threshold. On the other hand, if the difference between the current zoom value of sub camera 400 and the threshold is equal to or less than the threshold, the zoom value of sub camera 400 is not controlled. This makes it possible to control the shooting angle of view of sub camera 400 only when there is a large difference between the feature information of the tracking subject detected from the overhead image and the feature information of the subject detected from the sub-image, which affects the determination of similarity.
[0289] The zoom control process of the sub camera 400 in step S150 in FIG. 19 is as described with reference to FIG.
[0290] The zoom control of this embodiment makes it possible to control the shooting angle of view of the sub-camera 400 so that the difference between the size of the tracked subject in the overhead image and the size of the tracked subject in the sub-image does not become too large. Therefore, even if the shooting positions and shooting directions of multiple cameras are arranged apart and part of the subject is outside the shooting angle of view or the subject is too small, it is possible for multiple cameras to work together to track a specific subject.
[0291] Furthermore, after determining the similarity of the feature information of the subject, or when the similarity of the feature information of the subject is high, the zoom value of the sub-camera 400 may be returned to the zoom value before the change. In this case, when the zoom value of the sub-camera 400 is changed in step S138, the zoom value before the change is saved, and by transmitting the zoom value before the change to the sub-camera 400 before the processing of step S122, tracking can be started using the zoom value before the change.
[0292] In this case, the control unit 101 acquires the current zoom value of the sub-camera 400 in step S135 and stores it in the volatile memory 102, and before starting the processing of step S122, reads out the zoom value stored in step S135 from the volatile memory 102 and transmits it to the sub-camera 400. This makes it possible to start tracking after returning the imaging angle of view that was changed to increase the similarity of the feature information of the subject to the imaging angle of view before the change.
[0293] Furthermore, when the zoom value changed to increase the similarity of the feature information of the subject is returned to the zoom value before the change, the sub-camera 400 may notify the WS100 or the EB200 that control to the zoom value before the change is complete. In this case, when the control unit 401 of the sub-camera 400 receives notification from the PTZ driving unit 409 that the zoom operation is complete, the control unit 401 notifies the WS100 or the EB200 that the zoom operation is complete via the communication unit 405. This allows the image output to the WS100 or the EB200 to be quickly switched to the sub-image of the sub-camera 400 after zoom control of the sub-camera 400 is completed. Also, the image output from the WS100 or the EB200 to an external device can be quickly switched to the sub-image of the sub-camera 400.
[0294] Alternatively, instead of restoring the zoom value to the previous value, the WS100 or EB200 may store the previous zoom value and perform a cropping process to create a sub-image with the previous zoom value. In this case, the WS100 or EB200 acquires a sub-image from the sub-camera 400, calculates an image area according to the shooting angle of view corresponding to the previous zoom value, and performs cropping of the sub-image using the image processing units 307 and 407. This allows the image output to the WS100 or EB200 to be quickly switched to the sub-image of the sub-camera 400 after zoom control of the sub-camera 400 is completed. Furthermore, the image output from the WS100 or EB200 to an external device can be quickly switched to the sub-image of the sub-camera 400.
[0295] The trimming process may be performed by the sub-camera 400. In this case, the WS100 or EB200 instructs the sub-camera 400 to perform the trimming process. This allows the image output to the WS100 or EB200 to be quickly switched to the sub-image of the sub-camera 400 after zoom control of the sub-camera 400 is complete. Also, the image output from the WS100 or EB200 to an external device can be quickly switched to the sub-image of the sub-camera 400.
[0296] According to the fifth embodiment described above, the same subject can be recognized by the multiple cameras 300 and 400 with different shooting positions and shooting directions. Therefore, a specific subject can be tracked by appropriately switching between control of the sub-camera 400 by the WS 100 and control of the sub-camera 400 by the EB 200.
[0297] When the subject to be tracked does not exist in the sub-image of the sub-camera 400, the WS 100 controls the sub-camera 400, and when the subject to be tracked exists in the angle of view of the sub-camera 400, the control of the sub-camera 400 can be handed over from the WS 100 to the EB 200. Furthermore, when the subject to be tracked moves at high speed and is lost, or when the subject to be tracked is changed, the WS 100 controls the sub-camera 400, allowing tracking to continue.
[0298] In the fifth embodiment, an example has been described in which the WS100 and the EB200 switch whether or not to send pan / tilt values to the sub-camera 400, but the present invention is not limited to this example. For example, the WS100 and the EB200 may transmit pan / tilt values to the sub-camera 400 regardless of the tracking state, and the sub-camera 400 may control which device the pan / tilt values received from to perform the pan / tilt operation.
[0299] In this case, the processing of step S106 in Fig. 4(a) is omitted from the processing of WS100, and a process in which the control unit 101 transmits tracking state information STATE to the sub-camera 400 is added before the processing of step S107 in Fig. 4(a). The processing of EB200 is omitted from the processing of steps S205 and S206 in Fig. 4(b) and step S221 in Fig. 9(c).
[0300] If the tracking state information STATE received from the WS 100 is "tracking by the EB 200", the sub camera 400 performs control to perform panning / tilting operations in accordance with the control command received from the EB 200. If the tracking state information STATE received from the WS 100 is "tracking by the WS 100", the sub camera 400 performs control to perform panning / tilting operations in accordance with the control command received from the WS 100.
[0301] In addition, in the present embodiment, an example has been described in which the size of the tracked subject in the sub-image is controlled to approach the size of the tracked subject in the overhead image, but this is not limiting. For example, if overhead-view camera 300 includes an optical unit and a PTZ drive unit, the shooting angle of view of overhead-view camera 300 may be controlled to control the size of the tracked subject in the sub-image to approach the size of the tracked subject in the overhead image.
[0302] In addition, in the present embodiment, an example has been described in which the size of the tracked subject in the sub-image is controlled to approach the size of the tracked subject in the overhead image, but this is not limiting. For example, if the overhead-view camera 300 includes an optical unit and a PTZ drive unit, the size of the tracked subject in the overhead image may be controlled to approach the size of the tracked subject in the sub-image by controlling the shooting angle of the overhead-view camera 300.
[0303] Furthermore, in this embodiment, an example has been described in which WS 100 calculates the height hview of the subject by performing inference processing to recognize the subject based on the overhead image from overhead camera 300 and the sub-image from sub-camera 400, but this is not limiting. For example, overhead camera 300 or sub-camera 400 may perform inference processing to recognize the subject, and WS 100 may obtain the subject detection results from the inference processing from overhead camera 300 or sub-camera 400 via communication unit 105, and calculate the height hview of the subject.
[0304] Furthermore, in the present embodiment, an example has been described in which a threshold is set for the difference between the current zoom value and the zoom value of the sub-camera 400, so that the difference between the size of the tracked subject in the overhead-view image and the size of the tracked subject in the sub-image matches, but this is not limiting. Other configurations are possible as long as the difference in the height hview of the tracked subject within the vertical fields of view of the overhead-view camera 300 and the sub-camera 400 can be controlled to be small. For example, a threshold may be set for the difference between the size information hview of the tracked subject in the overhead-view image and the size information hview of the subject in the sub-camera 400, and the zoom value of the sub-camera 400 may be controlled when the difference in the size information hview of the subject exceeds this threshold.
[0305] Furthermore, in this embodiment, the zoom value of the sub camera 400 is controlled using the vertical size as the size of the tracking subject of the overhead camera 300 and the sub image, but the zoom value may also be controlled using the horizontal size.
[0306] [Embodiment 6] In the fifth embodiment, an example was described in which either the WS 100 or the EB 200 controls the sub-camera 400. In the sixth embodiment, similar to the fourth embodiment, the EB 200 is omitted, and an example is described in which the WS 100 controls the sub-camera 400 based on the overhead image from the overhead camera 300 and the sub-image from the sub-camera 400.
[0307] Next, the control process of the WS 100 will be described with reference to FIG.
[0308] FIG. 20 shows the control process of the WS 100, and shows the detailed process of step S507 in FIG.
[0309] The process in step S520 is the same as that in step S110 in FIG.
[0310] In step S521, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using an image from an overhead camera."
[0311] The tracking state determination unit 126 has a function of updating the tracking state information STATE stored in the volatile memory 102 .
[0312] In step S522, the control unit 101 reads the tracking state information STATE from the volatile memory 102, and determines whether “tracking using an overhead image” or “tracking using a sub-image” is in progress based on the tracking state information STATE. If the control unit 101 determines that “tracking using an overhead image” is in progress, the process proceeds to step S525; if the control unit 101 determines that “tracking using a sub-image” is in progress, the process proceeds to step S523.
[0313] The process of step S523 can be achieved by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the process of step S224 in FIG. 9(b).
[0314] In step S524, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using overhead image."
[0315] The processes in steps S525, S526b, and S527 are similar to those in steps S117, S150, and S119 in FIG.
[0316] The processing of steps S528 to S530 can be achieved by replacing the control unit 201 with the control unit 101 and the volatile memory 202 with the volatile memory 102 in the processing of steps S211 to S214 in FIG. 9(b).
[0317] In step S531, the control unit 101 executes the function of the tracking state determination unit 126 in FIG. 16 to change the tracking state information STATE to "tracking using sub-image", and ends the process.
[0318] According to the sixth embodiment described above, the WS 100 switches between the subject recognition results of the overhead image from the overhead camera 300 and the sub-image from the sub-camera 400 to control the sub-camera 400. This eliminates the need for the EB 200 in the fifth embodiment, and simplifies the system configuration while achieving the same effects as the fifth embodiment.
[0319] [Embodiment 7] In the first to sixth embodiments, an example of a system including an overhead camera 300 and a sub-camera 400 has been described.
[0320] In the seventh embodiment, an example of a system including a main camera 500 in addition to an overhead camera 300 and a sub-camera 400 will be described.
[0321] FIG. 21 is a system configuration diagram of the seventh embodiment.
[0322] The seventh embodiment differs from the first to sixth embodiments in that it includes a main camera 500 and determines the subject to be tracked by the sub-camera 400 based on the main image captured by the main camera 500. The following description will focus on the differences from the first to sixth embodiments.
[0323] In the seventh embodiment, the main camera 500 has a PTZ function. The target subject determination unit 122 of the WS 100 determines (estimates) the target subject of the main camera 500 from the shooting range of the main camera 500, and determines the tracking subject of the sub camera 400 based on the target subject of the main camera 500. The tracking subject of the sub camera 400 may be the same subject as the target subject of the main camera 500, or may be a different subject.
[0324] Next, an example of determining a subject to be tracked by the sub camera 400 based on a role set for the sub camera 400 will be described.
[0325] The role of the sub-camera 400 indicates the subject of interest for the main camera 500, the subject that the sub-camera 400 tracks in association with the zoom operation, and the control content of the zoom operation. The role of the sub-camera 400 can be set by the user via an operation unit provided on the WS100 or EB200.
[0326] Furthermore, if multiple sub-cameras are installed, any of the multiple sub-cameras can be set as the main camera, and the user may set the main camera settings using an operation unit provided on the WS100 or EB200. The role of the sub-camera 400 and the method for setting the main camera are not limited to the above methods. Any method is acceptable.
[0327] FIG. 22 shows an example of roles and contents that can be set for the sub-camera 400.
[0328] When the role (ROLE) is "main follow", the role (CAMERA_ROLE) of the sub camera 400 is to track the same subject as the subject that the main camera 500 is focusing on, and to perform zoom control in the same phase as the zoom operation of the main camera 500. A zoom control value of the sub camera 400 is calculated based on this role (CAMERA_ROLE). Here, the same phase in the zoom operation means that the zoom operations of the main camera 500 and the sub camera 400 are controlled in the same direction. For example, when the zoom control value of the main camera 500 is changed from the wide-angle side to the telephoto side, the zoom of the sub camera 400 is also changed from the wide-angle side to the telephoto side.
[0329] When the role (ROLE) is "main counter," the role (CAMERA_ROLE) of the sub camera 400 is to track the same subject as the subject that the main camera 500 is focusing on, and to perform zoom control in the opposite phase to the zoom operation of the main camera 500. The PTZ value of the sub camera 400 is calculated based on this role (CAMERA_ROLE). Here, the opposite phase in the zoom operation means that the zoom operations of the main camera 500 and the sub camera 400 are controlled in the opposite directions. For example, when the zoom control value of the main camera 500 is changed from the wide-angle side to the telephoto side, the zoom of the sub camera 400 is changed from the telephoto side to the wide-angle side.
[0330] When the role (ROLE) is "assist follow", the sub camera 400 tracks a subject other than the subject that the main camera 500 is focusing on, and performs zoom control in the same phase as the zoom operation of the main camera 500. Based on this role (CAMERA_ROLE), the zoom control value of the sub camera 400 is calculated.
[0331] When the role (ROLE) is "assist counter," the sub camera 400 tracks a subject other than the subject that the main camera 500 is focusing on, and performs zoom control in the opposite phase to the zoom operation of the main camera 500. Based on this role (CAMERA_ROLE), the zoom control value of the sub camera 400 is calculated. In the example of FIG. 22, "different from main (left side)" is exemplified as the control content of the tracking subject for "assist follow" and "assist counter," but there may also be "assist follow" and "assist counter" in which the tracking subject is "different from main (right side)."
[0332] Furthermore, when the tracking subject is "separate from the main subject," it may also serve to target a subject at a position other than left and right (up and down, or front and back).
[0333] If there are multiple sub-cameras, a role may be set for each sub-camera.
[0334] Furthermore, in the seventh embodiment, an example has been described in which the control content of the tracking subject and zoom is set as a role, but the control content of only the tracking subject may be set as a role, or other items may be added.
[0335] In addition, in the seventh embodiment, an example was described in which the tracking subject of the sub-camera 400 was set based on the main image of the main camera 500, and the seventh embodiment was combined with the first to fifth embodiments, but the seventh embodiment may also be combined with the sixth embodiment.
[0336] [Embodiment 8] In the eighth embodiment, an example will be described in which a dummy area is added to the sub-image so that it becomes a subject area corresponding to the overhead image, thereby bringing the feature information of the subject in the sub-image closer to the feature information of the subject in the overhead image. In this embodiment, the dummy area is called a "letter," and combining the dummy area with the subject area is called "adding a letter." Note that in the eighth embodiment, an all-black image is used as the dummy area, but this is not limiting.
[0337] The control process of the eighth embodiment will be described with reference to FIGS.
[0338] The processing of steps S110 to S122 in FIG. 23(a) except for steps S140 to S142 is the same as the processing of steps S110 to S122 in FIG. 9(a).
[0339] First, the letter adding process of the sub camera in step S140 of FIG. 23 will be described with reference to FIG.
[0340] In step S143, the control unit 101 reads the subject information from the volatile memory 102, similar to step S123, and the process proceeds to step S144.
[0341] In step S144, the control unit 101 acquires shooting information from the overhead camera 300 and the sub-camera 400 via the communication unit 105, and the process proceeds to step S145. In the sixth embodiment, the information acquired from the sub-camera 400 includes information on the subject area where the control unit 201 detects the subject by performing image recognition processing using an inference model for subject detection.
[0342] In step S145, the control unit 101 calculates the letter size based on the ratio of the subject area included in the overhead image to the subject area included in the sub-image within the shooting angle of view.
[0343] Here, the process of calculating the letter size will be described with reference to FIG.
[0344] FIG. 25(a) illustrates an example of a subject area in an overhead image, with dimensions of w0 (width) × h0 (height). FIGS. 25(b) and 25(c) illustrate example subject areas in sub-images, with dimensions of wc1 (width) × hc1 (height) and wc2 (width) × hc2 (height). In the eighth embodiment, because the subject's face is captured, letters must be added downward. Letters may be added in a direction corresponding to a specific part of the subject, such as when capturing a photograph of the subject's feet, where letters must be added upward. The direction in which letters are added may also be determined based on the position of the subject area. As shown in FIG. 25(c), when the subject area abuts the bottom edge of the sub-image, letters are added downward, and when the subject area abuts the right edge of the sub-image, letters are added right. The sizes hr and wr of the letters to be added can be calculated using Equation 10 below. In step S145, control unit 101 stores the letter size calculation result in volatile memory 102, and the process proceeds to step S141 in FIG. 23(a). (Formula 10) TIFF0007805396000010.tif25156 Alternatively, the degree of match r may be calculated using the above-mentioned formula 8, and the letter size may be calculated using the following formula 11. (Formula 11) TIFF0007805396000011.tif8161 In step S141, the control unit 101 executes the function of the feature information determination unit 125 in FIG. 3, similarly to step S119 in FIG. 9(a), and the process proceeds to step S142.
[0345] In step S141, control unit 101 transmits a tracking start command and feature information STAT[i] of the tracking subject to EB200 via communication unit 105, and proceeds to step S121. Here, control unit 101 reads the letter size calculated in step S145 from volatile memory 102. If the letter size is a positive number, control unit 101 transmits the letter size to EB200 via communication unit 105. If the letter size is 0 or less, a letter is not required and is not transmitted.
[0346] Next, the control of the EB 200 will be described with reference to FIG.
[0347] The eighth embodiment differs in that letter size is received and processed in the process of FIG. 23(b).
[0348] In step S230, control unit 201 determines whether or not it has received, via communication unit 205, a tracking start command, feature information STAT[i] of the tracking subject obtained from the overhead image of overhead camera 300, and letter size from WS 100. If control unit 201 has received the tracking start command and feature information STAT[i] of the tracking subject from WS 100, it proceeds to step S231; if not, it terminates the process. If control unit 201 has received letter size, it stores it in volatile memory 102, but if not, it does nothing and proceeds to step S231.
[0349] In step S231, the control unit 201 reads the volatile memory 102 to determine whether a letter-sized document was received in step S230, and if so, proceeds to step S231; if not, proceeds to step S233.
[0350] In step S232, the control unit 201 executes the function of the image recognition unit 221, but unlike step S202 in FIG. 4(b), it executes the letter assignment process described in FIG. 24 for the subject area in the sub-image. The control unit 201 determines the direction in which the subject area in the sub-image contacts the sub-image. For the determined direction, the control unit 201 calculates feature information STAT_SUB[m] for the subject area to which the letter has been assigned based on the letter size information received from the WS 100. The control unit 201 overwrites the volatile memory 102 with the newly calculated feature information STAT_SUB[m] and proceeds to step S233.
[0351] In steps S233 to S237, the control unit 201 performs the same processes as steps S211 to S214 in Fig. 9(b) by executing the function of the tracking target determination unit 222 in Fig. 3 and determining whether the feature information STAT[i] received from the WS 100 and the feature information STAT_SUB[m] obtained from the sub-image of the sub-camera 400 satisfy a predetermined condition.
[0352] According to the above-described eighth embodiment, by calculating feature information from the subject area to which letters are added based on the photographing information of the sub-camera 400, it is possible to obtain the same effects as those of the first to fifth embodiments.
[0353] [Embodiment 9] In the first to eighth embodiments, the height hview of the subject is calculated from the height hcam of the overhead camera 300 and the sub-camera 400, the vertical angle of view φ, the vertical field of view H, the distance L to the tracked subject, and the overall height hobj of the subject, as explained in Fig. 11. In the ninth embodiment, an example will be explained in which skeletal information of the subject is calculated instead of calculating the height hview of the subject.
[0354] FIG. 26 shows an example of skeletal information of a subject.
[0355] FIG. 26(a) illustrates an example of a subject in an overhead image. FIG. 26(b) illustrates an example of the result of executing a skeleton estimation process on the subject in FIG. 26(a). The skeleton estimation process is added as a function of the image recognition unit 121. The control unit 101 inputs the image read from the volatile memory 102 by the inference unit 104 into a trained model for skeleton estimation created by machine learning such as deep learning, and performs inference processing. The inference result is the coordinates of each body part in the image and a score indicating the likelihood of that coordinate, as shown in FIG. 26(c). For a nose, the coordinates are oxnose and oynose, and the score is osnose. It is assumed that the EB200 also has a similar function.
[0356] Similarly, Fig. 26(d) shows an example of a subject in a sub-image. Fig. 26(e) shows an example of the results of executing a skeleton estimation process on the subject in Fig. 26(d). Because the lower body of the subject in Fig. 26(d) is outside the field of view, the coordinates and score information shown in Fig. 26(f) only cover the waist.
[0357] In the first to eighth embodiments, the height hview of the subject is calculated, but in the ninth embodiment, skeletal information exemplified in Fig. 26(c) and Fig. 26(f) is calculated. That is, the height hview of the subject in Fig. 11(d) can be replaced with the difference in coordinates of each part of the skeletal information in Fig. 26.
[0358] For example, in Figure 26(c), let us assume that the difference between the y-coordinate of the nose (oynose) and the y-coordinate of the left ankle (oyLankle) is the largest. In the image shown in Figure 26(e), which only includes the waist, let us assume that the difference between the y-coordinate of the nose (cynose) and the y-coordinate of the left waist (cyLwaist) is the largest. In this case, the degree of match r obtained by Equation 8 can be calculated by the following Equation 12. (Formula 12) TIFF0007805396000012.tif12161 Figure 27(a) shows an example of the cropping amount T obtained from the above formula 12 and formula 9. By cropping the overhead image in Figure 27(a) using the skeleton information in Figures 26(c) and 26(f), an image of the same range as the sub-image in Figure 27(b) is obtained.
[0359] Similarly, for the letter size calculated in the eighth embodiment, the letter size hr shown in FIG. 27(c) can be calculated using the above formula 11.
[0360] In the ninth embodiment, the amount of trimming can be calculated by replacing the coordinate information of each part according to changes in the posture of the subject. However, in consideration of the case where skeleton estimation fails, it is desirable to also calculate the amount of trimming according to the first to fourth embodiments and check the consistency with the amount of trimming calculated according to the eighth and ninth embodiments. Furthermore, although the inference model used in this embodiment was for skeletal estimation of a person, similar processing may be performed for animals or objects by replacing it with a corresponding model.
[0361] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0362] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention.
[0363] The disclosure of this specification includes the following system, control device, and program. [Configuration 1] A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a first generating means for generating first feature information of the first region of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device second generation means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second imaging device so as to track the predetermined subject; A system comprising: [Configuration 2] The system described in configuration 1 is characterized in that the image processing means determines the amount of cropping for cropping the first region based on first size information indicating the overall size of the specified subject and second size information indicating the size of the subject within the shooting range of the second imaging device. [Configuration 3] The system according to configuration 2, wherein the image processing means acquires the second size information based on the first size information and position information and imaging range of the second imaging device. [Configuration 4] The system described in configuration 3, wherein the first size information is a height of the subject from a reference position, the second size information is a height of the subject in a shooting range of the second image capture device, and the shooting range is a vertical angle of view of the second image capture device. [Configuration 5] The system described in any one of configurations 2 to 4, wherein the image processing means determines the cropping amount based on a table that associates the ratio of the second size information to the first size information with the cropping amount. [Configuration 6] 5. The system according to any one of configurations 2 to 4, wherein the cutout amount is calculated from a ratio of the second size information to the first size information. [Configuration 7] the image processing means sets a plurality of cropping amounts, generates a plurality of pieces of first feature information for a plurality of first regions cropped by the plurality of cropping amounts, and transmits the generated first feature information to the second control device; The second control device 7. The system according to any one of configurations 1 to 6, wherein the plurality of first characteristic information received from the first control device is compared with the second characteristic information. [Configuration 8] the second control device compares the first characteristic information generated by the first control device with the second characteristic information generated by the second generation means; the first feature information and the second feature information are information that can identify the same subject when the same subject is photographed by a plurality of image capturing devices with different photographing directions, The system described in any one of configurations 1 to 7, characterized in that the first control device switches between a first state in which the first control device controls the second imaging device to track the specified subject based on the first feature information and a second state in which the second control device controls the second imaging device to track the specified subject based on the second feature information, based on a result of comparing the first feature information with the second feature information. [Configuration 9] the first control device transmits the first characteristic information to the second control device; the first control means switches to the second state when the first characteristic information and the second characteristic information satisfy a predetermined condition based on the result of the comparison in the second control device; 9. The system according to configuration 8, wherein if the first characteristic information and the second characteristic information do not satisfy the predetermined condition, the system switches to the first state. [Configuration 10] the predetermined condition is that the similarity between the first feature information and the second feature information is equal to or greater than a threshold; the second control device determines a similarity between the first feature information and the second feature information; The system according to configuration 9, wherein the result of the determination is notified to the first control device. [Configuration 11] The system described in configuration 9 or 10, wherein the first control means controls the second image capturing device so that the angle of view of the second image capturing device approaches the angle of view of the first image capturing device capturing the first image including the specified subject when the first feature information and the second feature information do not satisfy the specified condition. [Configuration 12] The system described in configuration 11, wherein the first control means controls the angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image is reduced. [Configuration 13] the second control means controls the second imaging device to track the predetermined subject when the predetermined subject is present within an imaging range of the second imaging device; If the predetermined subject is no longer present in the imaging range of the second imaging device, notify the first control device that tracking of the predetermined subject cannot be continued; 13. The system according to any one of configurations 8 to 12, wherein the first control means receives the notification and switches from the second state to the first state. [Configuration 14] 14. The system according to any one of configurations 8 to 13, wherein when the predetermined subject is changed, the first control means switches from the second state to the first state. [Configuration 15] 15. The system of claim 14, wherein when the specified subject is changed, the first control means switches from the first state to the second state if the first feature information and the second feature information satisfy a specified condition. [Configuration 16] The first control device a feature information determining means for determining the first feature information of the predetermined subject and transmitting the first feature information to the second control device; a first control information generating means for generating first control information for controlling the second control device so as to track the predetermined subject; The second control device and second control information generation means for generating second control information for controlling the second control device so as to track the predetermined subject. [Configuration 17] 17. The system of claim 16, wherein the control information includes at least one of a pan value and a tilt value. [Configuration 18] the first generation means generates the first feature information by performing inference processing using a trained model with the first image as an input; 18. The system according to any one of configurations 1 to 17, wherein the second generation means generates the second feature information by performing inference processing using a trained model with the second image as an input. [Configuration 19] the trained models include a first model for object detection and a second model for object identification; the first generation means performs inference processing using the first model with the first image as an input to generate first information indicating a position of a subject included in the first image; generating feature information of a subject included in the first image by performing inference processing using the second model with the first image and the first information as input; the second generation means generates second information indicating a position of a subject included in the second image by performing inference processing using the first model with the second image as an input; The system described in configuration 18, characterized in that it generates feature information of the subject contained in the second image by performing inference processing using the second model with the second image and the second information as input. [Configuration 20] The second model for identifying the subject is obtained by capturing images of a plurality of subjects from a plurality of different shooting directions. The captured images are used as learning data, and images of the same subject are compared to those with high similarity in feature information. The system according to configuration 19 is characterized in that the trained model is trained to achieve the following: [Configuration 21] A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; a first generating means for generating first feature information of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device second generation means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second control device so as to track the predetermined subject; and a first control means for controlling a photographing angle of view of the second imaging device in accordance with a size of the specified subject included in the first image when generating the first feature information to be transmitted to the second control device; [Configuration 22] The system described in configuration 21, wherein the first control means controls the angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image is reduced. [Configuration 23] The system described in Configuration 22, wherein the first control means calculates a shooting angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image becomes small, and controls the shooting angle of view of the second imaging device so that the difference between the calculated shooting angle of view and the current shooting angle of the second imaging device becomes equal to or less than a threshold value. [Configuration 24] 24. The system according to configuration 22 or 23, wherein the size of the subject is the height of the subject in the imaging angle of view of the imaging device, and the imaging angle of view is the vertical angle of view of the imaging device. [Configuration 25] 25. The system according to configuration 24, wherein the height of the subject is calculated from the height of the entire subject, the height position of the imaging device, and the vertical angle of view. [Configuration 26] 26. The system according to any one of configurations 21 to 25, wherein the subject included in the second image is the subject closest to the center of the angle of view of the second image. [Configuration 27] the second control device compares the first characteristic information generated by the first control device with the second characteristic information generated by the second generation means; the first feature information and the second feature information are information that can identify the same subject when the same subject is photographed by a plurality of image capturing devices with different photographing directions, the first control means stores the photographing angle of view of the second image capture device before changing the photographing angle of view, 27. The system according to any one of configurations 21 to 26, wherein after the comparison is performed, the angle of view of the second image capture device is returned to the angle of view before the change. [Configuration 28] The system described in Configuration 27, wherein the first control means returns the angle of view of the second imaging device to the angle of view before it was changed if the result of the comparison shows that the similarity between the first feature information and the second feature information is equal to or greater than a threshold. [Configuration 29] 29. The system according to any one of configurations 21 to 28, wherein the first control means controls the photographing angle of view by a zoom function. [Configuration 30] The system described in any one of configurations 27 to 29, characterized in that the second imaging device notifies the first control device and / or the second control device that the imaging angle of view of the second imaging device has returned to the imaging angle before it was changed. [Configuration 31] having an image processing means for cutting out a part of an image, the first control means stores the photographing angle of view of the second image capture device before the change; 29. The system according to claim 27 or 28, wherein after the comparison is performed, the image processing means creates the second image with the angle of view of the second imaging device before the change. [Configuration 32] The system described in any one of configurations 21 to 31, characterized in that the first control device switches between a first state in which the first control device controls the second imaging device to track the specified subject based on the first feature information and a second state in which the second control device controls the second imaging device to track the specified subject based on the second feature information, based on a result of comparing the first feature information with the second feature information. [Configuration 33] the first control device transmits the first characteristic information to the second control device; the first control means switches to the second state when the first characteristic information and the second characteristic information satisfy a predetermined condition based on the result of the comparison; 33. The system according to configuration 32, wherein if the first characteristic information and the second characteristic information do not satisfy the predetermined condition, the system switches to the first state. [Configuration 34] the predetermined condition is that the similarity between the first feature information and the second feature information is equal to or greater than a threshold; The system described in configuration 33, wherein the second control means determines the similarity between the first feature information and the second feature information and notifies the first control device of the result of the determination. [Configuration 35] the second control means controls the second imaging device to track the predetermined subject when the predetermined subject is present in a photographing angle of view of the second imaging device; If the predetermined subject is no longer present in the photographing angle of view of the second imaging device, notify the first control device that tracking of the predetermined subject cannot be continued; 33. The system according to configuration 32, wherein the first control means receives the notification and switches from the second state to the first state. [Configuration 36] 36. The system of any one of configurations 32 to 35, wherein the first control means switches from the second state to the first state when the predetermined subject is changed. [Configuration 37] 37. The system of claim 36, wherein when the specified subject is changed, the first control means switches from the first state to the second state if the first feature information and the second feature information satisfy a specified condition. [Configuration 38] The first control device a feature information determining means for determining the first feature information of the predetermined subject and transmitting the first feature information to the second control device; a first control information generating means for generating first control information for controlling the second control device so as to track the predetermined subject; The second control device and second control information generation means for generating second control information for controlling the second control device to track the predetermined subject. [Configuration 39] 39. The system of claim 38, wherein the control information includes at least one of a pan value and a tilt value. [Configuration 40] the first generation means generates the first feature information by performing inference processing using a trained model with the first image as an input; The system according to any one of configurations 21 to 39, wherein the second generation means generates the second feature information by performing inference processing using a trained model with the second image as input. [Configuration 41] the trained models include a first model for object detection and a second model for object identification; the first generation means performs inference processing using the first model with the first image as an input to generate first information indicating a position of a subject included in the first image; generating feature information of a subject included in the first image by performing inference processing using the second model with the first image and the first information as input; the second generation means generates second information indicating a position of a subject included in the second image by performing inference processing using the first model with the second image as an input; The system described in configuration 40, characterized in that it generates feature information of the subject contained in the second image by performing inference processing using the second model with the second image and the second information as input. [Configuration 42] The system described in configuration 41 is characterized in that the second model for identifying a subject is a trained model that has been trained using images of multiple subjects taken from multiple different shooting directions as training data so that the similarity of feature information between images of the same subject is high. [Configuration 43] A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a determining means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a generating means for generating first feature information of the first region of the predetermined subject; a control unit that controls the second imaging device so as to track the predetermined subject, The control device is characterized in that the control means transmits the first feature information to an external device that controls the second imaging device. [Configuration 44] A control device that controls a second imaging device so as to track a predetermined subject based on a second image captured by the second imaging device, the second image captured in a direction different from that of the first imaging device, a generating means for generating second feature information of a subject included in the second image; a determination means for determining the predetermined subject based on first feature information of the predetermined subject included in a first image captured by the first imaging device and acquired from an external device, and the second feature information; a control unit that controls the second imaging device so as to track the predetermined subject, The control device, wherein the first feature information is feature information of a first region of the specified subject cut out from the first image based on the shooting range of the second image. [Configuration 45] A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a first determination means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a first generating means for generating first feature information of the first region of the predetermined subject; second generation means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information; and a control unit that controls the second imaging device so as to track the predetermined subject based on the first feature information or the second feature information. [Configuration 46] A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a determining means for determining the predetermined subject from subjects included in the first image; a generating means for generating first feature information of the predetermined subject; a control unit that controls the second imaging device so as to track the predetermined subject, The control device is characterized in that, when generating the first feature information to be transmitted to an external device that controls the second imaging device, the control means controls the shooting angle of view of the second imaging device in accordance with the size of the specified subject included in the first image. [Configuration 47] A control device that controls a second imaging device so as to track a predetermined subject based on a second image captured by the second imaging device, the second image captured in a direction different from that of the first imaging device, a generating means for generating second feature information of a subject included in the second image; a determination means for determining the predetermined subject based on first feature information of the predetermined subject included in a first image captured by the first imaging device and acquired from an external device, and the second feature information; and a control unit that controls the second imaging device so as to track the predetermined subject. [Configuration 48] A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a first generating means for generating first feature information of the predetermined subject included in the first image; second generation means for generating second feature information of a subject included in the second image; a control unit that controls the second imaging device so as to track the predetermined subject based on the first feature information or the second feature information, The control device is characterized in that, when generating the second feature information from the second image, the control means controls the shooting angle of view of the second imaging device in accordance with the size of the specified subject included in the first image. [Configuration 49] A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; a calculation means for calculating an assigned area based on a ratio between the predetermined subject included in the first image and the subject included in the second image within a photographing angle of view; a first generating means for generating first feature information of a first region of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device second generation means for generating second feature information of a subject included in the second image; an image processing means for generating information in which the area of the subject included in the second image is expanded in accordance with the assigned area acquired from the first control device; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second imaging device so as to track the predetermined subject; A system comprising: [Configuration 50] The system described in configuration 49 is characterized in that the calculation means calculates a ratio indicating the second area relative to the first area and an assigned area based on first size information indicating the overall size of the specified subject and second size information indicating the size of the subject in the shooting range of the second imaging device. [Configuration 51] The system described in configuration 50, characterized in that the first size information is the height of the subject from a reference position, the second size information is the height of the subject in the shooting range of the second imaging device, and the shooting range is the vertical angle of view of the second imaging device. [Configuration 52] The system described in configuration 50, characterized in that the first size information is the width or height of the subject in the image captured by the first imaging device, and the second size information is the width or height of the subject in the image captured by the second imaging device. [Configuration 53] The first control device a first estimation means for estimating a bone structure of a subject in an image captured by the first imaging device; The second control device a second estimation means for estimating a bone structure of a subject in an image captured by the second imaging device; The system described in configuration 1, wherein the image processing means determines the amount of cropping based on the skeletal information obtained by the first estimating means and the skeletal information obtained by the second estimating means. [Configuration 54] The first control device a first estimation means for estimating a bone structure of a subject in an image captured by the first imaging device; The second control device a second estimation means for estimating a bone structure of a subject in an image captured by the second imaging device; The system described in any one of configurations 49 to 52, characterized in that the calculation means determines the assigned area based on skeletal information obtained by the first estimation means and skeletal information obtained by the second estimation means. [Configuration 55] A program for causing a computer to function as a control device according to any one of configurations 43 to 48. [Explanation of symbols]
[0364] 100...first control device (workstation / WS), 200...second control device (edge box / EB), 300...first imaging device (overhead camera), 400...second imaging device (sub-camera), 101, 201, 301, 401...control unit
Claims
1. A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a first generating means for generating first feature information of the first region of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device a second generating means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second imaging device so as to track the predetermined subject; A system comprising:
2. The system according to claim 1, characterized in that the image processing means determines the cropping amount for cropping the first region based on first size information indicating the overall size of the specified subject and second size information indicating the size of the subject within the shooting range of the second imaging device.
3. 3. The system according to claim 2, wherein the image processing means acquires the second size information based on the first size information and position information and imaging range of the second imaging device.
4. The system described in claim 3, characterized in that the first size information is the height of the subject from a reference position, the second size information is the height of the subject in the shooting range of the second imaging device, and the shooting range is the vertical angle of view of the second imaging device.
5. 3. The system according to claim 2, wherein the image processing means determines the cropping amount based on a table that associates the ratio of the second size information to the first size information with the cropping amount.
6. The system according to claim 2 , wherein the cutout amount is calculated from a ratio of the second size information to the first size information.
7. the image processing means sets a plurality of cropping amounts, generates a plurality of pieces of first feature information for a plurality of first regions cropped by the plurality of cropping amounts, and transmits the generated first feature information to the second control device; The second control device 2. The system according to claim 1, wherein the plurality of first characteristic information received from the first control device is compared with the second characteristic information.
8. the second control device compares the first characteristic information generated by the first control device with the second characteristic information generated by the second generation means; the first feature information and the second feature information are information that can identify the same subject when the same subject is photographed by a plurality of image capturing devices with different photographing directions, The system described in claim 1, characterized in that the first control device switches between a first state in which the first control device controls the second imaging device to track the specified subject based on the first feature information and a second state in which the second control device controls the second imaging device to track the specified subject based on the second feature information, based on a result of comparing the first feature information with the second feature information.
9. the first control device transmits the first characteristic information to the second control device; the first control means switches to the second state when the first characteristic information and the second characteristic information satisfy a predetermined condition based on the result of the comparison in the second control device; 9. The system according to claim 8, wherein the system switches to the first state when the first characteristic information and the second characteristic information do not satisfy the predetermined condition.
10. the predetermined condition is that the similarity between the first feature information and the second feature information is equal to or greater than a threshold; 10. The system according to claim 9, wherein the second control device determines a similarity between the first feature information and the second feature information and notifies the first control device of the result of the determination.
11. The system according to claim 9, characterized in that, when the first feature information and the second feature information do not satisfy the specified condition, the first control means controls the second image capturing device so that the angle of view of the second image capturing device approaches the angle of view of the first image capturing device that captures the first image including the specified subject.
12. The system according to claim 11, wherein the first control means controls the angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image is reduced.
13. the second control means controls the second imaging device to track the predetermined subject when the predetermined subject is present within a photographing range of the second imaging device; If the predetermined subject is no longer present in the imaging range of the second imaging device, notify the first control device that tracking of the predetermined subject cannot be continued; 9. The system according to claim 8, wherein the first control means receives the notification and switches from the second state to the first state.
14. 9. The system of claim 8, wherein said first control means switches from said second state to said first state when said predetermined subject is changed.
15. The system described in claim 14, wherein when the specified subject is changed, the first control means switches from the first state to the second state if the first feature information and the second feature information satisfy a specified condition.
16. The first control device a feature information determining means for determining the first feature information of the predetermined subject and transmitting the first feature information to the second control device; a first control information generating means for generating first control information for controlling the second control device so as to track the predetermined subject; The second control device 9. The system according to claim 8, further comprising: second control information generating means for generating second control information for controlling said second control device so as to track said predetermined subject.
17. 17. The system of claim 16, wherein the control information includes at least one of a pan value and a tilt value.
18. the first generation means generates the first feature information by performing an inference process using a trained model with the first image as an input; The system according to claim 1, characterized in that the second generation means generates the second feature information by performing inference processing using a trained model with the second image as input.
19. the trained model includes a first model for object detection and a second model for object identification; the first generation means performs inference processing using the first model with the first image as an input to generate first information indicating a position of a subject included in the first image; generating feature information of a subject included in the first image by performing an inference process using the second model with the first image and the first information as input; the second generation means performs inference processing using the first model with the second image as an input to generate second information indicating a position of a subject included in the second image; The system described in claim 18, characterized in that feature information of the subject contained in the second image is generated by performing inference processing using the second model with the second image and the second information as input.
20. The system described in claim 19, characterized in that the second model for subject identification is a trained model that has been trained using images of multiple subjects taken from multiple different shooting directions as training data so that the similarity of feature information between images of the same subject is high.
21. A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; a first generating means for generating first feature information of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device a second generating means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second control device so as to track the predetermined subject; and The system is characterized in that, when generating the first feature information to be transmitted to the second control device, the first control means controls the shooting angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image is reduced.
22. The system described in claim 21, characterized in that the first control means calculates a shooting angle of view of the second image capture device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image becomes small, and controls the shooting angle of view of the second image capture device so that the difference between the calculated shooting angle of view and the current shooting angle of the second image capture device is equal to or less than a threshold value.
23. 22. The system according to claim 21, wherein the size of the subject is a height of the subject in a photographing angle of view of the imaging device, and the photographing angle of view is a vertical angle of view of the imaging device.
24. 24. The system according to claim 23, wherein the height of the subject is calculated from the height of the entire subject, the height position of the imaging device, and the vertical angle of view.
25. 22. The system of claim 21, wherein the object included in the second image is the object closest to the center of the angle of view of the second image.
26. the second control device compares the first characteristic information generated by the first control device with the second characteristic information generated by the second generation means; the first feature information and the second feature information are information that can identify the same subject when the same subject is photographed by a plurality of image capturing devices with different photographing directions, the first control means stores the photographing angle of view of the second image capture device before the photographing angle of view is changed; 22. The system according to claim 21, wherein after the comparison is performed, the angle of view of the second image capture device is returned to the angle of view before the change.
27. The system described in claim 26, characterized in that, when the result of the comparison shows that the similarity between the first feature information and the second feature information is equal to or greater than a threshold, the first control means returns the angle of view of the second imaging device to the angle of view before it was changed.
28. 22. The system according to claim 21, wherein the first control means controls the photographing angle of view by a zoom function.
29. The system according to claim 26, characterized in that the second imaging device notifies the first control device and / or the second control device that the imaging angle of view of the second imaging device has returned to the imaging angle of view before it was changed.
30. having an image processing means for cutting out a part of an image, the first control means stores the photographing angle of view of the second image capture device before the change; 27. The system of claim 26, wherein after the comparison is performed, the image processing means creates the second image of the second image capture device at the angle of view before the change.
31. The system described in claim 21, characterized in that the first control device switches between a first state in which the first control device controls the second imaging device to track the specified subject based on the first feature information and a second state in which the second control device controls the second imaging device to track the specified subject based on the second feature information, based on a result of comparing the first feature information with the second feature information.
32. the first control device transmits the first characteristic information to the second control device; the first control means switches to the second state when the first characteristic information and the second characteristic information satisfy a predetermined condition based on the result of the comparison; 32. The system of claim 31, wherein the system switches to the first state if the first characteristic information and the second characteristic information do not satisfy the predetermined condition.
33. the predetermined condition is that the similarity between the first feature information and the second feature information is equal to or greater than a threshold; 33. The system according to claim 32, wherein the second control means determines a similarity between the first feature information and the second feature information, and notifies the first control device of the result of the determination.
34. the second control means controls the second imaging device to track the predetermined subject when the predetermined subject is present in a photographing angle of view of the second imaging device; If the predetermined subject is no longer present in the photographing angle of view of the second imaging device, notify the first control device that tracking of the predetermined subject cannot be continued; 32. The system of claim 31, wherein the first control means switches from the second state to the first state upon receiving the notification.
35. 32. The system of claim 31, wherein said first control means switches from said second state to said first state when said predetermined subject is changed.
36. The system described in claim 35, wherein when the specified subject is changed, the first control means switches from the first state to the second state if the first feature information and the second feature information satisfy a specified condition.
37. The first control device a feature information determining means for determining the first feature information of the predetermined subject and transmitting the first feature information to the second control device; a first control information generating means for generating first control information for controlling the second control device so as to track the predetermined subject; The second control device 32. The system according to claim 31, further comprising: second control information generating means for generating second control information for controlling said second control device so as to track said predetermined subject.
38. 38. The system of claim 37, wherein the control information includes at least one of a pan value and a tilt value.
39. the first generation means generates the first feature information by performing an inference process using a trained model with the first image as an input; The system according to claim 21, wherein the second generation means generates the second feature information by performing inference processing using a trained model with the second image as input.
40. the trained model includes a first model for object detection and a second model for object identification; the first generation means performs inference processing using the first model with the first image as an input to generate first information indicating a position of a subject included in the first image; generating feature information of a subject included in the first image by performing an inference process using the second model with the first image and the first information as input; the second generation means performs inference processing using the first model with the second image as an input to generate second information indicating a position of a subject included in the second image; The system described in claim 39, characterized in that feature information of the subject contained in the second image is generated by performing inference processing using the second model with the second image and the second information as input.
41. The system described in claim 40, characterized in that the second model for subject identification is a trained model that has been trained using images of multiple subjects taken from multiple different shooting directions as training data so that the similarity of feature information between images of the same subject is high.
42. A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a determining means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a generating means for generating first feature information of the first region of the predetermined subject; a control unit that controls the second imaging device so as to track the predetermined subject, The control device is characterized in that the control means transmits the first feature information to an external device that controls the second imaging device.
43. A control device that controls a second imaging device so as to track a predetermined subject based on a second image captured by the second imaging device, the second image captured in a direction different from that of the first imaging device, a generating means for generating second feature information of a subject included in the second image; a determination means for determining the predetermined subject based on first feature information of the predetermined subject included in a first image captured by the first imaging device and acquired from an external device, and the second feature information; a control unit that controls the second imaging device so as to track the predetermined subject, The control device, characterized in that the first feature information is feature information of a first region of the specified subject cut out from the first image based on the shooting range of the second image.
44. A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a first determination means for determining the predetermined subject from subjects included in the first image; an image processing means for cutting out a first region of the predetermined subject included in the first image based on a photographing range of the second image acquired from the second imaging device; a first generating means for generating first feature information of the first region of the predetermined subject; a second generating means for generating second feature information of a subject included in the second image; a second determination means for determining the predetermined subject based on the first feature information and the second feature information; and a control unit that controls the second imaging device so as to track the predetermined subject based on the first feature information or the second feature information.
45. A control device that controls a second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by a second imaging device having an imaging direction different from that of the first imaging device, a determining means for determining the predetermined subject from subjects included in the first image; a generating means for generating first feature information of the predetermined subject; a control unit that controls the second imaging device so as to track the predetermined subject, The control device is characterized in that, when generating the first feature information to be transmitted to an external device that controls the second imaging device, the control means controls the shooting angle of view of the second imaging device so that the difference between the size of the specified subject included in the first image and the size of the subject included in the second image is reduced.
46. A system including a first imaging device and a second imaging device having different imaging directions, and a first control device and a second control device that control the second imaging device to track a predetermined subject based on a first image captured by the first imaging device or a second image captured by the second imaging device, The first control device a first determination means for determining the predetermined subject from subjects included in the first image; a calculation means for calculating an assigned area based on a ratio between the predetermined subject included in the first image and the subject included in the second image within a photographing angle of view; a first generating means for generating first feature information of a first region of the predetermined subject; a first control means for controlling the second imaging device so as to track the predetermined subject; and The second control device a second generating means for generating second feature information of a subject included in the second image; an image processing means for generating information in which the area of the subject included in the second image is expanded in accordance with the assigned area acquired from the first control device; a second determination means for determining the predetermined subject based on the first feature information and the second feature information acquired from the first control device; a second control means for controlling the second imaging device so as to track the predetermined subject; A system comprising:
47. The system described in claim 46, characterized in that the calculation means calculates a ratio indicating the second area relative to the first area and an assigned area based on first size information indicating the overall size of the specified subject and second size information indicating the size of the subject in the shooting range of the second imaging device.
48. The system described in claim 47, characterized in that the first size information is the height of the subject from a reference position, the second size information is the height of the subject in a shooting range of the second image capture device, and the shooting range is the vertical angle of view of the second image capture device.
49. The system of claim 47, wherein the first size information is the width or height of the subject in the image captured by the first imaging device, and the second size information is the width or height of the subject in the image captured by the second imaging device.
50. The first control device a first estimation means for estimating a bone structure of a subject in an image captured by the first imaging device; The second control device a second estimation means for estimating a bone structure of a subject in an image captured by the second imaging device; 2. The system according to claim 1, wherein the image processing means determines the cutout amount based on the skeletal information obtained by the first estimating means and the skeletal information obtained by the second estimating means.
51. The first control device a first estimation means for estimating a bone structure of a subject in an image captured by the first imaging device; The second control device a second estimation means for estimating a bone structure of a subject in an image captured by the second imaging device; 50. The system according to claim 46, wherein the calculation means determines the assigned area based on the skeletal information obtained by the first estimation means and the skeletal information obtained by the second estimation means.
52. A program for causing a computer to function as the control device according to any one of claims 42 to 45.
Citation Information
Patent Citations
Method and device for automatically tracking intruder and image processor
JP2002290962A
Supervisory system
JP2003264822A
Imaging management system, imaging management apparatus, control method of them, and program
JP2015061239A
Control apparatus and control method
JP2015142181A
Systems and Methods for Measuring Scene Information While Capturing Images Using Array Cameras
US20160044257A1