Object state grasping device, object state grasping system, object state grasping method, and program
The object state grasping system effectively manages objects by linking and positioning data from multiple cameras with overlapping fields of view, improving tracking and visualization without additional hardware, addressing inefficiencies in existing methods.
Patent Information
- Application Number
- PCT/JP2025/003098
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-30
- Publication Date
- 2025-08-07
AI Technical Summary
Existing methods for managing objects using multiple cameras fail to organize the correspondence between objects captured in different image data, leading to inefficiencies in grasping the state of multiple objects.
An object state grasping system that utilizes multiple cameras with overlapping capture areas to detect two-dimensional coordinates, link object detection data, and estimate three-dimensional coordinates to accurately determine the state of objects without attaching devices.
Enables efficient and accurate management of objects by organizing correspondence between multiple camera captures, allowing for precise tracking and visualization of object states without additional hardware, reducing costs and enhancing user understanding.
Smart Images

Figure JP2025003098_07082025_PF_FP_ABST
Abstract
Description
Object state grasping device, object state grasping system, object state grasping method, and program
[0001] (Description of Related Applications) The present invention is based on the priority claim of Japanese Patent Application No. 2024-013437 (filed January 31, 2024), the entire contents of which are incorporated herein by reference. The present invention relates to an object state grasping device, an object state grasping system, an object state grasping method, and a program.
[0002] In the management of objects (e.g., luggage, animals, people, objects, equipment, and facilities)—such as inventory management in warehouses and work areas in the distribution industry, livestock management in the livestock industry, and people and objects management in offices, factories, construction sites, commercial facilities, and public spaces—data related to the objects is collected, recorded, and analyzed for the purpose of improving operations. This involves understanding the location history of the objects and their status, such as their movement paths (pathways) and movement volume. For example, in the livestock industry, tracking the location history of livestock and their movement paths (pathways) and movement volume is required to confirm their activity status and whether natural behaviors are being exhibited, with the aim of improving the quality of processed livestock products and understanding their health status as part of "animal welfare" initiatives. "Animal welfare" regards livestock as sentient beings and aims to reduce stress in a comfortable environment and build a happy relationship between humans and animals. Furthermore, because the number of objects to be monitored is often large, a method for monitoring the status of objects without attaching any device to them is needed. As a method for grasping the state of an object without attaching anything to the object, for example, there is a method for grasping (monitoring) the state of an object (e.g., an aquatic animal) using multiple camera images obtained from multiple cameras (imaging devices) that photograph (image) a specified observation space (e.g., an aquarium) from different directions (e.g., vertical and horizontal directions) (see, for example, Patent Document 1).
[0003] JP 2003-250382 A
[0004] The following analysis is provided by the present inventors.
[0005] In the method of Patent Document 1, when multiple objects are photographed with multiple cameras, the correspondence between the multiple objects captured in one image data and the multiple objects captured in another image data is not organized. Therefore, it is desirable to organize the correspondence between the multiple objects captured in each of the multiple image data even when multiple cameras are used, so that the state of the objects can be grasped. If the correspondence between the multiple objects is organized to grasp the state of the objects, the state of the objects can be grasped with higher accuracy, and therefore the objects can be managed easily and efficiently.
[0006] A main object of the present invention is to provide an object state grasping device, an object state grasping system, an object state grasping method, and a program that can contribute to easy and efficient management of objects.
[0007] The object state grasping device relating to a first viewpoint includes an object detection unit configured to detect the two-dimensional coordinates of all objects appearing in a plurality of image data captured by a plurality of cameras arranged in an observation area so that the capturing areas of at least two of the cameras are adjacent to or facing each other and overlap, based on the plurality of image data, and generate a plurality of object detection data; an object detection data linking unit configured to estimate a correspondence relationship between the plurality of object detection data for each of the objects and link the data to generate linked data; an object three-dimensional positioning unit configured to position the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data, and generate object three-dimensional coordinate data; and an object state estimation unit configured to estimate the state of the object based on the object three-dimensional coordinate data, and generate object state data.
[0008] The object state grasping system relating to the second viewpoint includes a plurality of cameras arranged so that the shooting areas of at least two of the cameras overlap in the observation target area, and an object state grasping device relating to the first viewpoint.
[0009] An object state grasping method according to a third perspective includes the steps of: an object state grasping device calculating, based on a plurality of image data captured by a plurality of cameras arranged in an observation area such that the capturing areas of at least two of the cameras are adjacent to or facing each other, the two-dimensional coordinates of all objects captured in the plurality of image data to generate a plurality of object detection data; estimating a correspondence between the plurality of object detection data for each of the objects and linking the data to generate linked data; measuring the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data to generate object three-dimensional coordinate data; and the object state grasping device estimating the state of the object based on the object three-dimensional coordinate data to generate object state data.
[0010] The program relating to the fourth viewpoint causes the object state grasping device to execute the following processes: a process of calculating the two-dimensional coordinates of all objects shown in multiple image data based on multiple image data captured by multiple cameras arranged in an observation area so that the capturing areas of at least two cameras that are adjacent to or facing each other overlap, to generate multiple object detection data; a process of estimating the correspondence between the multiple object detection data for each object and linking them to generate linked data; a process of measuring the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data, to generate object three-dimensional coordinate data; and a process of estimating the state of the object based on the object three-dimensional coordinate data to generate object state data.
[0011] The program can be recorded on a computer-readable storage medium. The storage medium can be a non-transitory medium such as a semiconductor memory, a hard disk, a magnetic recording medium, or an optical recording medium. The present disclosure can also be embodied as a computer program product. The program is input to a computer device via an input device or a communication interface from the outside, stored in a storage device, and drives a processor according to predetermined steps or processes. The processing results, including intermediate states as needed, can be displayed at each stage on a display device, or the computer device can communicate with the outside via the communication interface. For example, a computer device for this purpose typically includes a processor, a storage device, an input device, a communication interface, and, if necessary, a display device, all of which can be connected to each other via a bus.
[0012] The first to fourth aspects can contribute to the easy and efficient management of objects.
[0013] 1 is a block diagram schematically illustrating a first example of the configuration of an object state grasping system according to the present disclosure. FIG. 2 is a top-view image schematically illustrating an example of a camera's shooting area relative to an observation target area in the object state grasping system according to the present disclosure. FIG. 3 is an image schematically illustrating an example of the configuration of image data recorded in an image recording unit of an object state grasping device of the object state grasping system according to the present disclosure. FIG. 4 is an image schematically illustrating an example of the configuration of camera internal parameters held in a camera internal parameter holding unit of the object state grasping system of the present disclosure. FIG. 5 is an image schematically illustrating an example of the configuration of camera external parameters held in a camera external parameter holding unit of the object state grasping system of the present disclosure. FIG. 6 is an image schematically illustrating an example of the configuration of object detection data generated when an object is detected in an object detection unit of the object state grasping system according to the present disclosure. (A), (B), and (C) are schematic diagrams illustrating an example of an image when a correspondence relationship between object detection data is estimated in an object detection data linking unit of the object state grasping system of the present disclosure. 1 is a schematic diagram showing an example of an image when object detection data is linked in an object detection data linking unit of an object state grasping device of an object state grasping system according to the present disclosure. (A) and (B) are a top view image and a side view image showing an example of the positional relationship of a camera and an observed object with respect to a reference plane in an object state grasping system according to the present disclosure. (B) is an image showing an example of the configuration of object three-dimensional coordinate data generated when the three-dimensional coordinates of an object are measured in an object three-dimensional positioning unit of an object state grasping device of an object state grasping system according to the present disclosure. (C) is an image showing an example of the configuration of individual state data generated when the state of an individual is estimated in an individual state estimation unit of an object state estimation unit of an object state grasping device of an object state grasping system according to the present disclosure. (D) is an image showing an example of the configuration of swarm state data generated when the state of a swarm is estimated in a swarm state estimation unit of an object state estimation unit of an object state grasping device of an object state grasping system according to the present disclosure.1 is an image diagram schematically showing examples of (A) a linear flow line, (B) a curved flow line, and (C) a broken line flow line when the flow line of an object is visualized in a state visualization unit of the object state grasping device of the object state grasping system according to the present disclosure. FIG. 2 is a flowchart schematically showing an example of the operation of the object state grasping device of the object state grasping system according to the present disclosure. FIG. 3 is a flowchart schematically showing an example of details of the object detection data linking operation of the object state grasping device in the object state grasping system according to the present disclosure. FIG. 4 is a block diagram schematically showing a second example of the configuration of the object state grasping system according to the present disclosure. FIG. 5 is a block diagram schematically showing an example of the configuration of an object state grasping device according to the present disclosure. FIG. 6 is a block diagram schematically showing a configuration of hardware resources.
[0014] The following description of the embodiments will be made with reference to the drawings. Note that, where reference numerals are used in this application, they are intended solely to facilitate understanding and are not intended to limit the present invention to the illustrated embodiments. Furthermore, the following embodiments are merely exemplary and do not limit the present invention. Furthermore, connecting lines between blocks in the drawings, etc., referred to in the following description, include both bidirectional and unidirectional lines. Unidirectional arrows are used to schematically indicate the flow of the main signal (data) and do not exclude bidirectionality. Furthermore, although not explicitly shown, input and output ports exist at the input and output ends of each connecting line in the circuit diagrams, block diagrams, internal configuration diagrams, connection diagrams, etc., shown in this disclosure. The same applies to input / output interfaces. A program is executed via a computer device, which includes, for example, a processor, a storage device, an input device, a communication interface, and, if necessary, a display device. The computer device is configured to communicate with internal or external devices (including computers) via the communication interface, whether wired or wireless.
[0015] [Form 1] An object state grasping system according to Form 1 will be described with reference to the drawings. FIG. 1 is a block diagram schematically illustrating a first example of the configuration of an object state grasping system according to the present disclosure. FIG. 2 is a top-view image schematically illustrating an example of a camera's shooting area relative to an observation target area in the object state grasping system according to the present disclosure. FIG. 3 is an image schematically illustrating an example of the configuration of image data recorded in an image recording unit of an object state grasping device of the object state grasping system according to the present disclosure. FIG. 4 is an image schematically illustrating an example of the configuration of camera internal parameters held in a camera internal parameter holding unit of the object state grasping system of the object state grasping system according to the present disclosure. FIG. 5 is an image schematically illustrating an example of the configuration of camera external parameters held in a camera external parameter holding unit of the object state grasping system of the object state grasping system according to the present disclosure. FIG. 6 is an image schematically illustrating an example of the configuration of object detection data generated when an object is detected by an object detection unit of an object state grasping device of the object state grasping system according to the present disclosure. 7(A), 7(B), and 7(C) are schematic diagrams showing an example of an image when the object detection data correspondence is estimated in the object detection data linking unit of the object state grasping device of the object state grasping system according to the present disclosure. FIG. 8 is a schematic diagram showing an example of an image when object detection data is linked in the object detection data linking unit of the object state grasping device of the object state grasping system according to the present disclosure. FIGS. 9(A) and 9(B) are top and side bird's-eye views schematically showing an example of the positional relationship of the camera and the observed object relative to a reference plane in the object state grasping system according to the present disclosure. FIG. 10 is an image schematically showing an example of the configuration of object three-dimensional coordinate data generated when the three-dimensional coordinates of an object are measured in the object three-dimensional positioning unit of the object state grasping device of the object state grasping system according to the present disclosure. FIG. 11 is an image schematically showing an example of the configuration of individual state data generated when the state of an individual is estimated in the individual state estimation unit of the object state estimation unit of the object state grasping device of the object state grasping system according to the present disclosure.Fig. 12 is an image that schematically shows an example of the configuration of swarm state data generated when the state of a swarm is estimated by the swarm state estimation unit of the object state estimation unit of the object state assessment device in the object state assessment system according to the present disclosure. Fig. 13 is an image that schematically shows examples of (A) a linear flow line, (B) a curved flow line, and (C) a broken line flow line when the flow line of an object is visualized by the state visualization unit of the object state assessment device in the object state assessment system according to the present disclosure.
[0016] The object state grasping system 1 is a system for grasping the state of an object (5A to 5H in FIG. 2) as an observation target object present in an observation target area (3 in FIG. 2) (see FIG. 1).
[0017] Here, virtual coordinates used within the object state grasping system 1 are set for the observation target area 3. Furthermore, a virtual reference plane 4 that is parallel to the XY plane (assumed to be a horizontal plane) is set for the observation target area 3. The reference plane 4 can be, for example, a virtual horizontal ground. If the objects 5A-5H are livestock animals, the observation target area 3 can be an area where the animals are allowed to roam freely by being allowed to graze, or an area where they can move freely within a fenced area. The observation target area 3 is photographed by the cameras 10A-10N so that at least a portion of the photographing areas 11A-11N of at least two cameras 10A-10N that are adjacent to or facing each other overlap.
[0018] Furthermore, the objects 5A to 5H may be any objects that can move or change freely within a certain space and have some state. "Changing freely within a certain space" means that the objects can change their posture, behavior, and shape without being constrained by the space. The objects 5A to 5H may be, for example, moving objects such as people, robots, vehicles, drones, animals, livestock (cows, pigs, sheep, chickens, etc.), or non-moving objects such as goods, luggage, transported goods, structures, equipment, facilities, raw materials, and feed that are not moving but can be moved or changed by the action of a moving object.
[0019] Furthermore, understanding the state of an object means that the user can know when, where, and what happened (what state it was) at the moment the object was in a certain state. This includes the user being able to know the distribution of events, active events, passive events, and environmental information. It also includes the user being able to know the continuity and trends of events by aggregating and statistically processing when, where, and what happened (what state it was) over a certain time span. It also includes the user being able to sense when an abnormality is occurring.
[0020] The cameras 10A to 10N are devices that capture images of objects 5A to 5H present in the observation area 3 (see FIG. 1). Each time an image is captured, the cameras 10A to 10N transmit the captured image data to the object state grasping device 20. The cameras 10A to 10N are arranged adjacent to or facing each other so that two cameras share at least a portion of their capture range. The cameras 10A to 10N are arranged in pairs (camera 10A and camera 10B, camera 10A and camera 10C, camera 10A and camera 10D, etc.) that are adjacent to or facing each other. The cameras 10A to 10N may have the same specifications or different specifications. Any number of cameras 10A to 10N can capture multiple objects 5A to 5H (especially moving objects), and the cameras 10A to 10N can be flexibly arranged according to the shape and size of the observation area 3. Each camera 10A to 10N is assigned a unique camera ID (identifier). Camera internal parameters are set inside the cameras 10A to 10N. The camera internal parameters are parameters such as ISO (International Organization for Standardization) sensitivity, aperture value, shutter speed, focal length, angle of view, and aspect ratio that are set inside the cameras 10A to 10N. The cameras 10A to 10N can provide various data of the camera internal parameters set therein to a camera internal parameter storage unit 22 of the object state grasping device 20.
[0021] The object state grasping device 20 is a device that grasps the state of an object (5A to 5H in FIG. 2) as an observation target object located in an observation target area (3 in FIG. 2) (see FIG. 1). By executing a predetermined program, the object state grasping device 20 can virtually be configured to include an image recording unit 21, a camera internal parameter storage unit 22, a camera external parameter storage unit 23, an object detection unit 24, an object detection data linking unit 25, an object three-dimensional positioning unit 26, an object state estimation unit 27, a state visualization unit 28, an information storage unit 29, and an information provision unit 30.
[0022] The image recording unit 21 is a functional unit that acquires and records image data (see FIG. 3) (see FIG. 1). The image data is data that associates a timestamp, a camera ID, and an image. The timestamp is, for example, the time when each of the cameras 10A to 10N captured the image data. Alternatively, it may be the time when the image recording unit 21 acquired the image data.
[0023] The camera internal parameter storage unit 22 is a functional unit (see FIG. 1) that stores camera internal parameter data (see FIG. 4). The camera internal parameter data is data that associates a camera ID with camera internal parameters. Various types of camera internal parameter data can be acquired from the cameras 10A to 10N.
[0024] The camera external parameter storage unit 23 is a functional unit that stores camera external parameter data (see FIG. 5) (see FIG. 1). The camera external parameter data is data that associates a camera ID with camera external parameters. The camera external parameters are parameters such as the coordinates, orientation, tilt, and torsion of the installation position that are set outside the cameras 10A to 10N. Various types of camera external parameter data can be set in advance by the user, or can be acquired from the cameras 10A to 10N if the cameras 10A to 10N are equipped with orientation sensors, angle sensors, torsion sensors, etc.
[0025] The object detection unit 24 is a functional unit that detects the locations (two-dimensional object coordinates) of objects (5A to 5H in FIG. 2) (see FIG. 1). The object detection unit 24 detects characteristic parts of the objects 5A to 5H (e.g., the highest point, head, back, etc. of livestock animals) based on image data (see FIG. 3). Based on the detection results, the object detection unit 24 calculates two-dimensional object coordinates of the locations of the detected objects 5A to 5H using internal camera parameter data (see FIG. 4) and external camera parameter data (see FIG. 5). The two-dimensional object coordinates are coordinates on the reference plane 4 and are the two-dimensional coordinates of the positions of the objects 5A to 5H when viewed from the cameras 10A to 10N and projected onto the reference plane 4 (see FIGS. 9A and 9B). The object detection unit 24 generates object detection data (see FIG. 6) for each object, which associates the calculated two-dimensional object coordinates of the locations of the objects 5A to 5H, a timestamp, and a camera ID. The object detection unit 24 generates object detection data each time the image recording unit 21 acquires image data. The camera ID is acquired from the image data that was the basis for generating the two-dimensional coordinates of the object. The timestamp is the timestamp of the image data.
[0026] The object detection data linking unit 25 is a functional unit that estimates and links the correspondence between multiple object detection data (see FIG. 6) for each of the objects 5A to 5H (see FIG. 1). When linking the object detection data, the object detection data linking unit 25 estimates the correspondence between multiple object detection data related to the same object based on the generated object detection data (see FIG. 6) and previously generated object detection data, using internal camera parameters (see FIG. 4) and external camera parameters (see FIG. 5). The previously generated object detection data has the same camera ID and is the object detection data with the timestamp immediately preceding the generated object detection data. In estimating the correspondence, the object detection data linking unit 25 generates all combinations of object detection data pairs based on the generated object detection data, calculates the probability of a match between pairs of object detection data in the generated combinations, and selects the combination of object detection data with the highest probability. In linking the object detection data, the object detection data linking unit 25 links the selected combinations of object detection data to generate linked data.
[0027] 7A and 7B, when object detection data A-1 to A-6 are generated from image data from camera 10A and object detection data B-1 to B-6 are generated from image data from camera 10B, object detection data linking unit 25 generates all paired combinations of object detection data A-1 to A-6 and object detection data B-1 to B-6. For example, as shown in FIG. 7C, for object detection data A-1, six paired combinations are generated with object detection data B-1 to B-6. The same applies to object detection data A-2 to A-6. Combinations of object detection data are generated not only between a specific pair of cameras (cameras 10A and 10B in FIG. 2), but also between other pairs of cameras (cameras 10A and 10C, cameras 10A and 10D, cameras 10B and 10C, cameras 10B and 10D, cameras 10C and 10D, cameras 10C and 10E, cameras 10C and 10F, cameras 10D and 10E, cameras 10D and 10F, cameras 10E and 10F, ... in FIG. 2).
[0028] In calculating the probability of a match between object detection data, the object detection data linking unit 25 calculates the probability that a combination of object detection data relating to objects 5A to 5H will match, using factors such as the difference between the two-dimensional coordinates of the objects in the object detection data, the closest distance of the vectors from cameras 10A to 10N to the objects (5A to 5H in Figure 2), the difference in the movement speeds of objects 5A to 5H, the difference in the movement direction (angle) of objects 5A to 5H per unit time, the difference in distance on the coordinates from other detected individuals in the vicinity, the difference in orientation on the coordinates from other detected individuals in the vicinity, and the difference in probability distribution as a result of the previous calculation. When calculating the probability, the probability of matching between object detection data increases as the difference between the two-dimensional coordinates of the objects decreases, the probability of matching between object detection data increases as the closest distance of the vectors from cameras 10A to 10N to objects 5A to 5H decreases, the probability of matching between object detection data increases as the difference in the movement speed of objects 5A to 5H decreases, the probability of matching between object detection data increases as the difference in movement direction (angle) per unit time decreases, the probability of matching between object detection data increases as the distance on coordinates from other detected individuals in the vicinity increases, the probability of matching between object detection data increases as the difference in orientation on coordinates from other detected individuals in the vicinity increases, and the probability of matching between object detection data increases as the difference in probability distribution as a result of the previous calculation decreases.
[0029] Based on the calculated probability, the object detection data linking unit 25 links the combinations of object detection data relating to objects 5A to 5H with the highest probability. In the example of FIG. 7C, object detection data A-1 and object detection data B-1, which have the highest probability of matching between object detection data (0.96), are linked. The number of object detection data linked may not be limited to two, but may be three or more. For example, as in the case of object detection data relating to object 5C in FIG. 8, four object detection data A-3, B-3, C-1, and D-1 may be linked. Note that object detection data C-1 is object detection data for object 5C generated from image data from camera 10C, and object detection data D-1 is object detection data for object 5C generated from image data from camera 10D. The same applies to the other object detection data C-2 to C-3, D-2 to D-3, E1 to E4, and F1 to F4. When linking object detection data, the object detection data linking unit 25 assigns an object identifier (e.g., 5A to 5H) to each group of links, thereby enabling combinations of object detection data for the same object to be linked.
[0030] The three-dimensional object positioning unit 26 is a functional unit that measures the three-dimensional coordinates of objects (5A to 5H in FIG. 2) (see FIG. 1). For each identifier (object 5A to 5H), the three-dimensional object positioning unit 26 calculates the three-dimensional object coordinates of the locations of the objects 5A to 5H using camera external parameter data (see FIG. 5) based on two arbitrary or predetermined sets of object detection data among the object detection data linked in the linking data. The two object detection data may be, for example, the two object detection data corresponding to the top two camera IDs, prioritized in advance. The three-dimensional object coordinates can be calculated by triangulation using, for example, the two-dimensional object coordinates in the two object detection data and the coordinates of the cameras 10A to 10N in the camera external parameter data (see FIG. 9B). Alternatively, the three-dimensional object coordinates may be calculated using, but are not limited to, the inter-camera distance, focal length, and parallax of the multiple cameras 10A to 10N. The height and body height of the objects 5A to 5H can be determined from the three-dimensional object coordinates. The object 3D positioning unit 26 generates object 3D coordinate data (see FIG. 10) for each object, which associates the calculated object 3D coordinates, a timestamp, and a camera ID (a combination of the camera IDs of the two object detection data used for the calculation). This makes it possible to obtain unique, unique object 3D coordinate data for each of the objects 5A to 5H.
[0031] The object state estimation unit 27 is a functional unit that estimates the state of the objects (5A to 5H in FIG. 2) (see FIG. 1). The object state estimation unit 27 estimates the state of the objects 5A to 5H based on the object 3D coordinate data (see FIG. 10), its history, characteristic area information (e.g., feeding area, sandbox, etc.) set for a predetermined area in the observation target area 3, etc. The object state estimation unit 27 generates object state data (individual state data in FIG. 11 and group state data in FIG. 12) related to the estimated state of the objects 5A to 5H. The object state estimation unit 27 includes an individual state estimation unit 27a and a group state estimation unit 27b.
[0032] The individual state estimation unit 27a is a functional unit that estimates the individual states of the objects 5A to 5H (see FIG. 1). The individual state estimation unit 27a estimates the individual states of the objects 5A to 5H based on the object three-dimensional coordinate data (see FIG. 10), its history, characteristic area information (e.g., feeding area, sandbox, etc.) set for a predetermined area in the observation target area 3, and the like. The individual state estimation unit 27a generates individual state data (see FIG. 11) for each individual (object) that associates the estimated individual state of the objects 5A to 5H, the object three-dimensional coordinates, a timestamp, and a camera ID (a combination of the camera IDs of two object detection data). Here, the state of an individual can be, for example, the position of each individual, movement line (a chronological order of the positions of each individual, and how they moved), behavior, actions (e.g., eating / not eating, lying down / standing, moving / not moving, unusual behavior that does not fall into any of the above categories, etc.), events that occur between individuals (e.g., fighting, attacking, being chased away from the herd, etc.), the continuation of the above state (e.g., whether it is a momentary event or a continuing event, the frequency of the event, etc.), etc.
[0033] The group state estimation unit 27b is a functional unit that estimates the state of a group of multiple objects 5A-5H (see FIG. 1). In this embodiment, for example, if multiple objects are present and their positions remain within a predetermined threshold range for a certain period of time, the group is considered to be a group. The group state estimation unit 27b estimates the group state based on the individual state data (see FIG. 11). The group state estimation unit 27b generates group state data (see FIG. 12) for each group, associating the estimated group state, object 3D coordinates, a timestamp, and a camera ID (a combination of camera IDs from two object detection data). Here, the group state may be, for example, the number (number) of individual objects 5A-5H in a certain area, the number (number) of individual objects 5A-5H in a certain state, the number of occurrences of a certain event, the distribution and density of the objects 5A-5H, the distribution of high-density locations, the distribution of low-density locations, the distribution of areas in a certain state (such as the number of occurrences), and the location where a behavior is occurring.
[0034] The state visualization unit 28 is a functional unit that visualizes the states of objects (5A to 5H in FIG. 2) (see FIG. 1). The state visualization unit 28 visualizes the states (states of individuals and states of groups) of the objects 5A to 5H based on object state data (individual state data in FIG. 11 and group state data in FIG. 12). In this visualization, the estimated results of the positions, movement paths, and states of the objects 5A to 5H are visualized in a manner that is easy for the user to see. Examples of visualization methods include plotting on a map, representing using colors or icons, creating various graphs, and creating lists. For example, methods for visualizing the movement paths of objects include connecting the movement paths between positions 91 and 92 (measured position, estimated position) with a straight line as shown in FIG. 13(A), connecting with a curve (e.g., a Bezier curve, a spline curve, etc.) as shown in FIG. 13(B), or connecting with a broken line as shown in FIG. 13(C). The state visualization unit 28 generates visualized information (state visualization information).
[0035] The information storage unit 29 is a functional unit that stores information (see FIG. 1 ). The information storage unit 29 stores object detection data (see FIG. 6 ), object three-dimensional coordinate data (see FIG. 10 ), individual state data (see FIG. 11 ), group state data (see FIG. 12 ), and state visualization information (see, for example, FIG. 13 ).
[0036] The information providing unit 30 is a functional unit that provides (transmits) information in the information holding unit 29 to the user terminal 40 in response to a request from the user terminal 40 (see FIG. 1).
[0037] The user terminal 40 is a terminal used by a user (see FIG. 1 ). There may be a plurality of user terminals 40 in the object state grasping system 1. The user terminal 40 is communicably connected to the object state grasping device 20 via a network 50. The user terminal 40 can transmit data input by a user's operation to the object state grasping device 20. The user terminal 40 can output (display, output audio, etc.) data from the object state grasping device 20.
[0038] The network 50 is an information communication network that communicatively connects the cameras 10A to 10N, the object state grasping device 20, and the user terminal 40 (see FIG. 1). As the network 50, for example, a communication network such as a PAN (Personal Area Network), a LAN (Local Area Network), a MAN (Metropolitan Area Network), a WAN (Wide Area Network), or a GAN (Global Area Network) can be used.
[0039] Next, the operation of the object state grasping device 20 in the object state grasping system 1 according to the first embodiment will be described with reference to the drawings. Fig. 14 is a flowchart schematically illustrating an example of the operation of the object state grasping device 20 in the object state grasping system 1 according to the present disclosure. Please refer to Fig. 1 for the configurations of the object state grasping system 1 and the object state grasping device 20. Here, it is assumed that the camera internal parameter data (see Fig. 4) and the camera external parameter data (see Fig. 5) are already stored.
[0040] First, the image recording unit 21 of the object state grasping device 20 acquires and records image data (FIG. 3) of the photographed areas 11A to 11N from the cameras 10A to 10N (step A1).
[0041] Next, the object detection unit 24 of the object state grasping device 20 detects characteristic portions of the positions (object two-dimensional coordinates) of the objects (5A to 5H in FIG. 2) in the acquired image data. Based on the detection results, the object detection unit 24 calculates the object two-dimensional coordinates of the positions of the detected objects 5A to 5H using the camera internal parameter data (see FIG. 4) and the camera external parameter data (see FIG. 5). Then, object detection data (see FIG. 6) is generated that associates the calculated object two-dimensional coordinates of the positions of the objects 5A to 5H, a timestamp, and a camera ID (step A2).
[0042] Next, the object detection data linking unit 25 of the object state grasping device 20 estimates correspondence relationships between the plurality of object detection data (see FIG. 6) for each of the objects 5A to 5H, using the camera internal parameters (see FIG. 4) and the camera external parameters (see FIG. 5), based on the generated object detection data (see FIG. 6) and the previously generated object detection data. Then, the object detection data is linked based on the estimated correspondence relationships to generate linked data (step A3).
[0043] Next, the three-dimensional object positioning unit 26 of the object state grasping device 20 calculates, for each of the objects 5A to 5H, the three-dimensional object coordinates of the positions of the objects 5A to 5H, using the external camera parameter data (see FIG. 5), based on a combination of any two or predetermined two of the object detection data linked in the generated linking data. The three-dimensional object coordinate data (see FIG. 10) is generated, associating the calculated three-dimensional object coordinates, a timestamp, and a camera ID (the combination of the camera IDs of the two object detection data) (step A4).
[0044] Next, the object state estimation unit 27 of the object state grasping device 20 estimates the states of the objects 5A to 5H based on the generated object 3D coordinate data (see FIG. 10), its history, characteristic area information (e.g., feeding area, sandbox, etc.) set in a predetermined area in the observation target area 3, etc. Then, it generates object state data (individual state data in FIG. 11, group state data in FIG. 12) related to the estimated states of the objects 5A to 5H (step A5).
[0045] Next, the state visualization unit 28 of the object state grasping device 20 visualizes the states (states of individuals, states of groups) of the objects 5A to 5H based on the object state data (individual state data in FIG. 11, group state data in FIG. 12), and generates visualized information (state visualization information) (step A6).
[0046] Next, the information storage unit 29 of the object state grasping device 20 stores the state visualization information (see, for example, FIG. 13) (step A7), and then ends and returns to the start.
[0047] Next, details of the object detection data linking operation (step A3 in FIG. 14 ) of the object state determination device 20 in the object state determination system 1 according to form 1 will be described with reference to the drawings. FIG. 15 is a flowchart schematically showing an example of details of the object detection data linking operation of the object state determination device 20 in the object state determination system 1 according to the present disclosure. Please refer to FIG. 1 for the configurations of the object state determination system 1 and the object state determination device 20.
[0048] After step A2 in FIG. 14, the object detection data linking unit 25 of the object state grasping device 20 generates all combinations of object detection data pairs based on the generated plurality of object detection data (step B1).
[0049] Next, the object detection data linking unit 25 of the object state grasping device 20 calculates the probability of a match between a pair of object detection data in the generated combination (step B2).
[0050] Next, the object detection data linking unit 25 of the object state grasping device 20 links the combinations of object detection data with the highest calculated probabilities to generate linked data (step B3), and then proceeds to step A4 in Figure 14.
[0051] According to the first aspect, object detection data relating to objects 5A-5H in an observation target area 3 is generated based on multiple image data from multiple cameras 10A-10N. Then, the correspondence between the multiple object detection data for each object is estimated and linked. The linking results are used to calculate the three-dimensional coordinates of the objects 5A-5H, and the state of the objects is estimated. Therefore, even when multiple cameras 10A-10N are used, the correspondence between the multiple objects captured in each of the multiple image data can be organized, contributing to understanding the state of the objects. This allows the state of the objects to be understood with greater accuracy, making it possible to easily and efficiently manage the objects.
[0052] Furthermore, according to the first aspect, the plurality of cameras 10A to 10N can capture the objects 5A to 5H in a wide observation target area 3. Therefore, the cameras 10A to 10N can be flexibly arranged in accordance with the shape and size of the observation target area 3.
[0053] Furthermore, according to form 1, the object state estimation unit 27 can estimate the states (including positions, movement lines, etc.) of the objects 5A to 5H from the three-dimensional coordinates, heights, and their history of the objects 5A to 5H. Therefore, since the correspondences of the objects 5A to 5H in the multiple cameras are linked, the number of objects can be determined without overlap between the cameras 10A to 10N, and the estimated states of the objects 5A to 5H can accurately determine the number of objects 5A to 5H in a certain state.
[0054] Furthermore, according to form 1, the state visualization unit 28 generates state visualization information that makes the estimated results of the states of the estimated objects 5A to 5H easier for the user to see, and provides this to the user terminal 40 via the information providing unit 30, thereby allowing the information to be presented to the user.
[0055] Furthermore, according to form 1, camera internal parameter data (see FIG. 4) and camera external parameter data (see FIG. 5) are set for each of the multiple cameras 10A to 10N, so there is no need for the cameras 10A to 10N to have the same specifications.
[0056] Furthermore, according to the first aspect, the states of the objects 5A to 5H can be grasped without attaching wireless communication devices to the objects 5A to 5H. This eliminates the cost of wireless communication devices and the hassle of attaching wireless communication devices and dealing with malfunctions, dead batteries, etc. Furthermore, the objects 5A to 5H are more comfortable because they do not need to wear wireless communication devices. For these reasons, the object state grasping system 1 of the first aspect can be used to manage the states of a large number of objects 5A to 5H or groups of these objects.
[0057] Furthermore, according to the first aspect, the state of the objects 5A to 5H can be grasped without attaching markers such as two-dimensional barcodes to the objects 5A to 5H, which eliminates the need to deal with stains, tears, peeling, etc. of the markers (cleaning, replacing, etc.), and can be used to manage the state of a large number of objects 5A to 5H or a group of them.
[0058] Furthermore, according to the first aspect, since the objects 5A to 5H are photographed from different directions by the two cameras 10A to 10N, the height positions of the objects 5A to 5H can be accurately determined. In managing the states of the objects 5A to 5H, it is possible to avoid being affected by the size of the objects 5A to 5H.
[0059] Furthermore, according to the first embodiment, the observation target area 3 is photographed so that the photographing areas 11A-11N of at least two cameras 10A-10N at least partially overlap. The correspondence between the object detection data relating to the detected objects 5A-5H is estimated and linked using the photographed image data in this manner, so that the objects 5A-5H are not duplicated or missed. Therefore, according to the first embodiment, the total number of objects 5A-5H can be accurately grasped, and the status of each object 5A-5H can be thoroughly grasped.
[0060] Furthermore, according to the first aspect, the observation target area 3 is photographed by at least two cameras 10A to 10N so that the photographing areas 11A to 11N of the cameras 10A to 10N at least partially overlap, and the same individual is tracked using a combination of the cameras 10A to 10N. Therefore, even if there are multiple individuals that are very similar in shape and size, it is possible to track the same individual.
[0061] [Embodiment 2] An object state grasping system 1 according to embodiment 2 will be described with reference to the drawings. Fig. 16 is a block diagram schematically showing a second example of the configuration of the object state grasping system 1 according to the present disclosure.
[0062] A second embodiment is a modification of the first embodiment, in which a camera extrinsic parameter calculation unit 31 is added to the object state grasping device 20 of the first embodiment. The camera extrinsic parameter calculation unit 31 calculates the camera extrinsic parameters (e.g., coordinates, orientation, tilt, torsion, etc. of the installation positions set outside the cameras 10A to 10N) of the corresponding cameras 10A to 10N based on image data acquired from the cameras 10A to 10N, and automatically sets the camera extrinsic parameters (stored in the camera extrinsic parameter storage unit 23). The camera extrinsic parameter calculation unit 31 can calculate the camera extrinsic parameters based on the form (position, size, range, length, spacing, shape, etc.) of characteristic locations in the observation target area 3 captured in the image data. The characteristic locations are located at fixed positions in the observation target area 3, and their coordinates are set in advance. Not only one characteristic location but also multiple characteristic locations may be used. Natural objects (e.g., trees, mountains, ponds, etc.), artificial objects (e.g., fences, facilities, etc.), markers (e.g., poles, dial plates, barcodes, etc.), etc., in the observation target area 3 can be used as characteristic locations. The other configurations and operations are the same as those of the first embodiment.
[0063] According to the second aspect, similar to the first aspect, even when multiple cameras 10A-10N are used, the correspondence between multiple objects captured in each of the multiple image data can be organized, thereby contributing to grasping the state of the objects. Furthermore, since the system has a function for automatically setting the external camera parameters of the cameras 10A-10N, the external camera parameters of the cameras 10A-10N can be automatically set even if the cameras 10A-10N are replaced with cameras 10A-10N having different specifications from the cameras before replacement. This reduces the effort required for maintaining the object state grasping system 1.
[0064] [Form 3] The object state grasping device 20 according to Form 3 will be described with reference to the drawings. Fig. 17 is a block diagram schematically showing an example of the configuration of the object state grasping device 20 according to the present disclosure.
[0065] The object state grasping device 20 is a device that grasps the state of an object in the observation target area 3. The object state grasping device 20 includes an object detection unit 24, an object detection data linking unit 25, an object three-dimensional positioning unit 26, and an object state estimation unit 27.
[0066] The object detection unit 24 is configured to detect all objects captured in the multiple image data and generate their two-dimensional coordinates as multiple object detection data. The multiple image data are captured by multiple cameras 10A-10N arranged in the observation target area 3. The multiple cameras 10A-10N are arranged in the observation target area 3 so that at least a portion of the capture areas of at least two cameras (e.g., cameras 10A and 10B) adjacent to or facing each other overlap. The object detection data linking unit 25 is configured to estimate a correspondence relationship between the multiple object detection data for each object, link the data, and generate linked data. The object three-dimensional positioning unit 26 is configured to locate the three-dimensional coordinates of the object based on any or predetermined two object detection data among the object detection data linked by the linked data, thereby generating object three-dimensional coordinate data. The object state estimation unit 27 is configured to estimate the state of the object based on the object three-dimensional coordinate data and generate object state data.
[0067] According to the third aspect, object detection data relating to objects in the observation target area 3 is generated based on a plurality of image data from the plurality of cameras 10A-10N. Then, the correspondence between the plurality of object detection data for each object is estimated and linked, and then the three-dimensional coordinates of the object are calculated to estimate the state of the object. Therefore, even when a plurality of cameras 10A-10N are used, the correspondence between the plurality of objects captured in each of the plurality of image data can be organized, thereby contributing to understanding the state of the objects.
[0068] In the object state grasping device 20 according to the first to third aspects, each of the cameras 10A-10N acquires image data multiple times at a predetermined timing. The object detection unit 24 detects characteristic portions of the objects 5A-5H each time each of the cameras 10A-10N acquires image data. The object detection unit 24 then calculates two-dimensional object coordinates of the positions of the detected objects 5A-5H from the detection results using the camera internal parameter data (see FIG. 4) and the camera external parameter data (see FIG. 5). The object detection unit 24 then generates and accumulates object detection data (see FIG. 6) for each object, which associates the calculated two-dimensional object coordinates of the positions of the objects 5A-5H, a timestamp, and a camera ID.
[0069] The timing at which each of the cameras 10A to 10N acquires image data may be changed depending on the observation target, or may be changed according to the specifications of the object state grasping device 20. The timing may also be set in advance by the user.
[0070] Each of the cameras 10A to 10N may synchronize and capture images multiple times at the same timing. In this case, the timing of capturing images may be instructed, for example, from the object state grasping device 20. In this case, the object detection data linking unit 25 links the object detection data based on the object detection data generated from image data having the same timestamp among the image data captured by each of the cameras 10A to 10N.
[0071] Each of the cameras 10A-10N may acquire images asynchronously multiple times. In this case, the object detection data linking unit 25 may link the correspondence of the object detection data based on the object detection data generated from the latest image data of each of the cameras 10A-10N. In this case, it is desirable that the changes in the positions and states of the objects 5A-5H are so slight that the correspondence of the multiple objects can be linked, and that the image data be within an allowable time range.
[0072] Furthermore, in the object state grasping device 20 according to the first to third aspects, the individual state estimation unit 27a may estimate the state of an individual by combining a time series of position information and characteristic area information with respect to behaviors such as eating / not eating, lying down / standing, moving / not moving, etc. Specifically, in the case of eating, if the individual is located in a feeding area and has remained there without moving for a certain period of time or more, it is estimated that the individual is eating.
[0073] Furthermore, unusual behaviors that do not fall under the above categories may be estimated based on differences between statistical movement line data of an individual's location information and animal husbandry know-how. Specifically, with regard to differences between statistical movement line data of location information, for example, if the distance traveled by an individual on the day is clearly shorter than the distance traveled the previous day, the possibility of injury or illness is estimated. Note that animal husbandry know-how is information derived by producers, veterinarians, etc. from experience and statistics about the correspondence between the location and movement line of an individual and its condition. This animal husbandry know-how may be used as input for estimating the individual's condition.
[0074] Furthermore, for events that occur between individuals (for example, fights, etc.), in addition to the same estimation method as for unusual behavior that does not fall into these categories, estimation can also be made by obtaining positional information of predetermined characteristic parts of the individuals (here, the position of the head, the position of the tail, etc.) at regular intervals using a camera or the like, and observing changes over time.
[0075] Furthermore, a prediction model may be trained using training data that associates time-series changes in location with the state (behavior) of the individual, and the state of the individual may be estimated using the prediction model.
[0076] The group state estimation unit 27b may estimate the group state by, for example, aggregating information on individuals estimated by each of the above methods.
[0077] The object state grasping device 20 according to the first to third aspects can be configured by so-called hardware resources (information processing device, computer), and can use a device having the configuration shown in Fig. 18. For example, the hardware resource 100 includes a processor 101, a memory 102, a network interface 103, etc., which are interconnected by an internal bus 104.
[0078] 18 is not intended to limit the hardware configuration of the hardware resource 100. The hardware resource 100 may include hardware (e.g., an input / output interface) that is not shown. Furthermore, the number of units, such as the processor 101, included in the device is not intended to be limited to the example shown in FIG. 18, and for example, multiple processors 101 may be included in the hardware resource 100. The processor 101 may be, for example, a central processing unit (CPU), a microprocessor unit (MPU), a graphics processing unit (GPU), or the like.
[0079] The memory 102 may be, for example, a random access memory (RAM), a read only memory (ROM), a hard disk drive (HDD), or a solid state drive (SSD).
[0080] The network interface 103 may be, for example, a LAN (Local Area Network) card, a network adapter, a network interface card, or the like.
[0081] The functions of the hardware resource 100 are realized by the processing modules described above. The processing modules are realized, for example, by the processor 101 executing a program stored in the memory 102. The programs can be updated by downloading them over a network or by using a storage medium that stores the programs. Furthermore, the processing modules may be realized by semiconductor chips. In other words, it is sufficient that the functions performed by the processing modules can be realized by executing software on some kind of hardware.
[0082] Some or all of the above aspects may be described as, but are not limited to, the following supplementary notes.
[0083] [Supplementary Note 1] An object state grasping device comprising: an object detection unit configured to calculate two-dimensional coordinates of all objects appearing in a plurality of image data captured by a plurality of cameras arranged in an observation target area so that the capturing areas of at least two of the cameras adjacent to or facing each other overlap, and generate a plurality of object detection data; an object detection data linking unit configured to estimate a correspondence between the plurality of object detection data for each of the objects and link the data to generate linked data; an object three-dimensional positioning unit configured to position the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data, and generate object three-dimensional coordinate data; and an object state estimation unit configured to estimate a state of the object based on the object three-dimensional coordinate data, and generate object state data. [Supplementary Note 2] The object state grasping device according to Supplementary Note 1, further comprising: a state visualization unit configured to visualize the state of the object based on the object state data, and generate state visualization information. [Supplementary Note 3] The object state grasping device according to Supplementary Note 1 or 2, wherein the object detection unit is configured to detect characteristic parts of the object based on the image data, calculate the two-dimensional coordinates of the object based on the detection results, and generate the object detection data. [Supplementary Note 4] The object detection data linking unit is configured to generate all combinations of pairs of object detection data based on at least the plurality of object detection data related to the at least two cameras whose shooting areas overlap, calculate a match probability between the pairs of object detection data in the generated combinations, and link the combinations of object detection data with the highest match probability to generate the linked data. [Supplementary Note 5] The object state grasping device according to Supplementary Note 4, wherein the object detection data linking unit is configured to calculate the probability that the combinations of object detection data related to the object match, using a difference between the two-dimensional coordinates of the object in the pair of object detection data, a closest distance of vectors from the camera to the object, and a difference in moving speeds of the objects, when calculating the probability.[Supplementary Note 6] The object state grasping device according to any one of Supplementary Notes 1 to 5, wherein the object state estimation unit comprises: an individual state estimation unit configured to estimate a state of an individual of the object based on the object three-dimensional coordinate data and generate individual state data as the object state data; and a group state estimation unit configured to estimate a state of a group of the individual objects based on the individual state data and generate group state data as the object state data. [Supplementary Note 7] The object state grasping device according to Supplementary Note 3, wherein the object detection unit is configured to calculate the two-dimensional coordinates of the object using the camera internal parameter data and the camera external parameter data based on the detection result. [Supplementary Note 8] The object state grasping device according to Supplementary Note 7, further comprising: a camera external parameter calculation unit configured to calculate the camera external parameter data of the camera corresponding to the image data acquired from the camera, and have the calculated camera external parameter data stored in the camera external parameter storage unit. [Supplementary Note 9] The object state grasping device according to Supplementary Note 8, wherein the camera extrinsic parameter calculation unit is configured to calculate the camera extrinsic parameters based on the form of a characteristic location in the observation target area captured in the image data. [Supplementary Note 10] The object state grasping device according to any one of Supplements 1 to 9, comprising: an information storage unit configured to store the object detection data and information; and an information providing unit configured to provide the information stored in the information storage unit to a user terminal in response to a request from the user terminal. [Supplementary Note 11] An object state grasping system comprising: a plurality of cameras arranged in an observation target area so that the shooting areas of at least two cameras overlap; and the object state grasping device according to any one of Supplements 1 to 10.[Supplementary Note 12] An object state grasping method comprising: a step in which an object state grasping device calculates two-dimensional coordinates of all objects shown in a plurality of image data based on the plurality of image data taken by a plurality of cameras arranged in an observation target area such that the shooting areas of at least two of the cameras adjacent to or facing each other overlap, thereby generating a plurality of object detection data; a step in which a correspondence relationship between the plurality of object detection data for each of the objects is estimated and linked to generate linked data; a step in which the object is positioned three-dimensionally based on any or predetermined two of the object detection data linked by the linked data, thereby generating object three-dimensional coordinate data; and a step in which the object state grasping device estimates the state of the object based on the object three-dimensional coordinate data, thereby generating object state data. [Supplementary Note 13] A program that causes an object state grasping device to execute the following processes: based on a plurality of image data captured by a plurality of cameras arranged in an observation area such that the capturing areas of at least two of the cameras are adjacent to or facing each other and overlap, a process of calculating two-dimensional coordinates of all objects captured in the plurality of image data to generate a plurality of object detection data; a process of estimating a correspondence relationship between the plurality of object detection data for each of the objects and linking the data to generate linked data; a process of measuring the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data to generate object three-dimensional coordinate data; and a process of estimating the state of the object based on the object three-dimensional coordinate data to generate object state data.
[0084] Note that Supplements 12 and 13 relating to the object state grasping method and program can be expanded in the same manner as Supplements 2 to 10.
[0085] The disclosures of the above-mentioned patent documents are incorporated herein by reference and may be used as the basis or part of the present invention, if necessary. Modifications and adjustments of the embodiments and examples are possible within the scope of the entire disclosure of the present invention (including the claims and drawings), and further based on the basic technical concepts thereof. Furthermore, various combinations and selections (or non-selections, if necessary) of the various disclosed elements (including each element of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible within the scope of the entire disclosure of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible by a person skilled in the art in accordance with the entire disclosure and technical concepts, including the claims and drawings. Furthermore, with regard to the numerical values and numerical ranges described in this application, any intermediate values, lower values, and smaller ranges are deemed to be included, even if not explicitly stated. Furthermore, the disclosures of the above-cited documents, when used in part or in whole in combination with the disclosures herein as part of the disclosure of the present invention, are also deemed to be included in (belong to) the disclosures of this application, in accordance with the spirit of the present invention.
[0086] 1 Object state grasping system 3 Observation target area 4 Reference plane 5A to 5H Object 10A to 10N Camera 11A to 11N Photographing area 20 Object state grasping device 21 Image recording unit 22 Camera internal parameter storage unit 23 Camera external parameter storage unit 24 Object detection unit 25 Object detection data linking unit 26 Object 3D positioning unit 27 Object state estimation unit 27a Individual state estimation unit 27b Group state estimation unit 28 State visualization unit 29 Information storage unit 30 Information provision unit 31 Camera external parameter calculation unit 40 User terminal 50 Network 91, 92 Position 100 Hardware resources 101 Processor 102 Memory 103 Network interface 104 Internal bus
Claims
1. An object state grasping device comprising: an object detection unit configured to calculate the two-dimensional coordinates of all objects appearing in multiple image data captured by at least two cameras positioned adjacent to or facing each other in an observation area so that the capturing areas of the cameras overlap, and generate multiple object detection data; an object detection data linking unit configured to estimate a correspondence between the multiple object detection data for each object and link the data to generate linked data; a three-dimensional object positioning unit configured to measure the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data, and generate object three-dimensional coordinate data; and an object state estimation unit configured to estimate the state of the object based on the three-dimensional object coordinate data, and generate object state data.
2. The object state grasping device according to claim 1, further comprising a state visualization unit configured to visualize the state of the object based on the object state data and generate state visualization information.
3. The object state grasping device according to claim 1 or 2, wherein the object detection unit is configured to detect characteristic parts of the object based on the image data, calculate the two-dimensional coordinates of the object based on the detection results, and generate the object detection data.
4. An object state grasping device as described in any one of claims 1 to 3, wherein the object detection data linking unit is configured to generate all combinations of object detection data pairs based on the object detection data from the at least two cameras whose shooting areas overlap, calculate the probability of a match between the pair of object detection data in the generated combinations, link the combination of object detection data with the highest probability, and generate the linked data.
5. An object state grasping device as described in any one of claims 1 to 4, wherein the object state estimation unit comprises: an individual state estimation unit configured to estimate the state of an individual object based on the object three-dimensional coordinate data and generate individual state data as the object state data; and a group state estimation unit configured to estimate the state of a group of the individual objects based on the individual state data and generate group state data as the object state data.
6. An object state grasping device as described in claim 3, comprising: a camera internal parameter storage unit configured to store camera internal parameter data set inside the camera; and a camera external parameter storage unit configured to store camera external parameter data set outside the camera, wherein the object detection unit is configured to calculate the two-dimensional coordinates of the object using the camera internal parameter data and the camera external parameter data based on the detection result.
7. The object state grasping device according to claim 6, further comprising a camera external parameter calculation unit configured to calculate the camera external parameter data of the corresponding camera based on the image data acquired from the camera and store the calculated data in the camera external parameter storage unit.
8. An object state grasping system comprising: a plurality of cameras arranged so that the photographing areas of at least two cameras overlap in an observation target area; and an object state grasping device according to any one of claims 1 to 7.
9. A method for grasping an object state, comprising: a step in which an object state grasping device calculates, based on a plurality of image data taken by a plurality of cameras arranged in an observation area so that the shooting areas of at least two of the cameras adjacent to or facing each other, the two cameras overlap, and generates a plurality of object detection data; a step in which a correspondence relationship between the plurality of object detection data for each of the objects is estimated and linked to generate linked data; a step in which, based on any or predetermined two of the object detection data linked by the linked data, the three-dimensional coordinates of the object are measured and object three-dimensional coordinate data is generated; and a step in which the object state grasping device estimates the state of the object based on the object three-dimensional coordinate data and generates object state data.
10. A program that causes an object state grasping device to execute the following processes: a process of calculating the two-dimensional coordinates of all objects captured in multiple image data taken by multiple cameras in an observation area, the cameras being arranged so that the capture areas of at least two cameras adjacent to or facing each other overlap, and generating multiple object detection data; a process of estimating the correspondence between the multiple object detection data for each object and linking them to generate linked data; a process of measuring the three-dimensional coordinates of the object based on any or predetermined two of the object detection data linked by the linked data, and generating object three-dimensional coordinate data; and a process of estimating the state of the object based on the object three-dimensional coordinate data and generating object state data.
Citation Information
Patent Citations
Camera calibration device and method, and vehicle
JP2008193188A
Object identification method
JP2016057998A
Identical person detection method and identical person detection system
JP2016162306A
Method and device for monitoring a monitoring region
US20140327780A1
People flow estimation device, display control device, people flow estimation method, and recording medium
WO2018025831A1