Position detection system and position detection method
Patent Information
- Application Number
- JP2025538926
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-04
- Filing Date
- 2023-08-04
- Publication Date
- 2026-03-05
- Estimated Expiration
- 2043-08-04
AI Technical Summary
Existing methods for detecting the position of an object using a monocular camera face challenges due to variations in object appearance affecting detection accuracy, especially when internal camera parameters are unknown, and require additional installation and known information, limiting applicability and accuracy.
A position detection system that utilizes a monocular camera and a measuring device to calculate a transformation matrix based on positional relationships at multiple points, converting image coordinates to actual object positions, allowing for high-accuracy detection with minimal scene and object restrictions.
Enables accurate position detection using a general-purpose monocular camera with an easy-to-use method applicable to various scenes and objects, without the need for extensive installation or additional devices.
Smart Images

Figure 00000011_0000 
Figure 00000011_0001 
Figure 00000011_0002
Description
[Technical Field]
[0001] The present invention relates to a position detection system that detects the position of an object based on an image captured by a camera that captures an area of interest. [Background technology]
[0002] In visual information processing using a monocular camera, various pieces of information possessed by a photographed object can be estimated by using AI (Artificial Intelligence) based on machine learning. Examples of information that can be estimated by AI include the position and type of an object in a photographed image, and the distance from the camera to the object. The task of estimating the position and type of an object in a photographed image is generally called object detection, and is disclosed in Non-Patent Documents 1 and 2, etc. The task of estimating the distance from the camera to an object is generally called depth estimation, and is disclosed in Non-Patent Document 3, etc.
[0003] As AI advances through deep learning, the types of information processing that can be performed are increasing and accuracy is also dramatically improving. However, detecting a two-dimensional position from a bird's-eye view from an image captured with a horizontal angle of view using a monocular camera is a difficult task. While it is theoretically difficult to identify the position of a target object based on a single image alone, this can be achieved by utilizing prerequisites and known information in addition to the image information.
[0004] For example, Patent Document 1 discloses a method for detecting the position of a moving object on a floor by drawing a dot pattern on the floor and associating this dot pattern with coordinate values on the floor in advance. Patent Document 2 discloses a method for probabilistically calculating the position of a person in three-dimensional space by detecting the top of the person's head in an image and additionally using the probability distribution of the person's height as known information. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-102585 [Patent Document 2] Japanese Patent Application Laid-Open No. 2004-302700 [Non-patent literature]
[0006] [Non-Patent Document 1] Joseph Redmon, Santosh Divvala, Ross Girshick, Ali Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” June 8, 2015, [online], https: / / arxiv.org / abs / 1506.02640. [Non-patent document 2] Mingxing Tan, Ruoming Pang, Quoc V. Le, “EfficientDet: Scalable and Efficient Object Detection,” November 20, 2019, [online], https: / / arxiv.org / abs / 1911.09070. [Non-patent document 3] Ashutosh Saxena, Sung Chung, Andrew Ng, “Learning Depth from Single Monocular Images,” Advances in Neural Information Processing Systems 18, 2005. Summary of the Invention [Problem to be solved by the invention]
[0007] When detecting the position of an object using images captured by a monocular camera, a challenge arises: even if the relative positions of the camera and the object are constant, the appearance of the object will vary depending on the camera's internal parameters. When using a method that directly utilizes the pixel distance in the captured image, differences in the appearance of the object will affect the detection results, so the internal parameters must be fixed to ensure accurate detection. Therefore, when applying the above detection method to a camera with unknown internal parameters, detection accuracy cannot be guaranteed. Furthermore, when estimating the position of an object using AI, the generalization performance of the AI may reduce estimation errors due to differences in the appearance of the object, but this remains one of the factors that reduce estimation accuracy. These issues pose practical constraints and obstacles, such as limiting the number of cameras that can be used and requiring additional information and effort when using a new camera.
[0008] When using devices other than monocular cameras, the price, size, availability, etc. of the device become practical issues. For example, when using sensors such as stereo cameras or LiDAR (Light Detection and Ranging) instead of monocular cameras, the price of the device is often more expensive than a monocular camera. Furthermore, monocular cameras are often installed in familiar devices such as smartphones, making them easy to obtain, and detection methods using monocular cameras can be said to have a wide range of applications.
[0009] When using a device other than an optical sensor instead of a monocular camera, practical issues arise, such as the detectable range, detection accuracy, and applicable environments. Typical sensor methods and systems include wireless ones, such as Bluetooth (registered trademark), GPS (Global Positioning System), RFID (Radio Frequency Identification), and UWB (Ultra Wide Band). All of these methods have in common the need for both the object to be detected and devices such as tags and signal receivers. For example, when detecting the location of a person, the person must always carry a tag or signal receiver, making it impossible to detect the locations of an unspecified number of people.
[0010] When using a device in addition to a monocular camera, a typical example is a method as described in Patent Document 1, in which devices serving as positional references or guides are installed throughout the entire target area or placed at regular intervals. Such methods have issues, such as the need for extensive installation work at each application location and the difficulty of installing the device throughout the entire target area depending on the application location. Furthermore, when using known information that is not dependent on a specific device, as described in Patent Document 2, additional work is required when expanding the scope of application to unknown targets. Furthermore, depending on the type and diversity of targets, it may be difficult to meet prerequisites, making application difficult, posing issues in terms of the scope of applicability.
[0011] The present invention has been made in consideration of the above-described conventional circumstances, and aims to detect the position of an object with high accuracy using a general-purpose and widely used monocular camera, in a method that is easy to use and has few restrictions on the scenes and objects to which it can be applied. [Means for solving the problem]
[0012] In order to achieve the above object, a position detection system according to one aspect of the present invention is configured as follows: That is, the position detection system includes a camera that captures an image of a target area, and a position detection device that detects the position of an object based on the image captured by the camera, and further includes a measuring device that acquires the positional relationship between the camera and the object, wherein the position detection device, in a preparation stage, performs a process of calculating a transformation matrix for converting the coordinates of the object in the image captured by the camera into the actual position of the object, based on the positional relationships between the camera and the object at multiple points acquired using the measuring device, and in an operation stage, performs a process of calculating the coordinates of the object in the image captured by the camera and converting them into the actual position of the object using the transformation matrix.
[0013] Here, in the above position detection system, the camera is installed so that its horizontal axis is parallel or approximately parallel to the plane of the target area, and the position detection device can perform a process in the preparation stage to calculate the transformation matrix based on the positional relationship between the camera and the target object obtained using the measuring instrument at two points within the target area, and the positional relationship at two other points obtained by flipping this left and right.
[0014] In the above position detection system, the measuring device may be a wireless sensor that performs positioning wirelessly.
[0015] In the position detection system, the camera may be installed facing sideways or diagonally downward.
[0016] Furthermore, in the above-mentioned position detection system, during the operational phase, the position detection device acquires a rectangle surrounding the object contained in the image captured by the camera, and if the ratio of the height of the rectangle to the width of the rectangle is smaller than a predetermined threshold, it complements the rectangle downward so that the ratio of the height of the rectangle to the width of the rectangle becomes a predetermined ratio, and performs a process of calculating the coordinates of the object in the image captured by the camera based on the complemented rectangle.
[0017] A position detection method according to another aspect of the present invention is configured as follows: That is, a position detection method for detecting the position of an object based on an image captured by a camera capturing an image of a target area, the method comprising: in a preparation stage, using a measuring device for acquiring the positional relationship between the camera and the object, acquiring the positional relationship between the camera and the object at a plurality of points; calculating a transformation matrix for converting coordinates of the object in the image captured by the camera into an actual position of the object based on the positional relationship between the camera and the object at the plurality of points; and in an operation stage, calculating the coordinates of the object in the image captured by the camera, and converting them into the actual position of the object using the transformation matrix. [Effects of the Invention]
[0018] According to the present invention, the position of an object can be detected with high accuracy using a general-purpose and widely used monocular camera, with an easy-to-use method that has few restrictions on the scene or object to which it can be applied. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a diagram illustrating an example of the configuration of a position detection system according to an embodiment of the present invention. [Figure 2] FIG. 10 is a diagram showing an example of a camera installed in a room, viewed from the side of the camera. [Figure 3A] FIG. 10 is a diagram showing an example in which the camera is installed without tilting relative to the optical axis, as viewed from the front of the camera. [Figure 3B] FIG. 10 is a diagram showing an example in which a camera is installed at an angle relative to the optical axis, as viewed from the front of the camera. [Figure 4] FIG. 10 is a diagram showing an example of an image of a target area for position detection captured by a camera. [Figure 5] FIG. 10 is a diagram showing an example of a target area for position detection viewed from directly above. [Figure 6] 2 is a diagram showing an overview of processing steps of a position detection method executed by the position detection system of FIG. 1. [Figure 7]FIG. 10 is a diagram showing an example of a person detection result based on an image captured by a camera. [Figure 8] FIG. 10 is a diagram illustrating a method for determining points A' and B' in an image captured by a camera. [Figure 9] FIG. 10 is a diagram illustrating a method for determining points A' and B' when the target area is viewed from directly above. [Figure 10] 10A and 10B are diagrams illustrating an example of complementation when a part of an object to be detected for position detection is hidden by an obstruction. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the following description is an example, and the present invention is not limited thereto. Here, the description will be given taking position detection of a person as an example, but the subject of the present invention is not limited to a person, and the present invention may also be applied to subjects other than a person. Note that, as a condition for applying the present invention, it is desirable that the subject always be in contact with the ground in the space in which position detection is performed, and that the location of the subject is on a flat surface without any steps.
[0021] 1 shows an example of the configuration of a position detection system according to one embodiment of the present invention. The position detection system shown in the figure includes a camera 10 that captures an image of an area to be detected, a measuring instrument 20 that acquires the positional relationship between the camera 10 and an object, and a position detection device 30 that detects the position of the object based on an image captured by the camera 10. In this example, a monocular camera is used as the camera 10.
[0022] The position detection device 30 can be realized by a computer equipped with hardware such as a CPU (Central Processing Unit) and a memory. The position detection device 30 may also include other processors such as a DSP (Digital Signal Processor), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit). The position detection device 30 may be configured as a single computer or as a plurality of computers operating in cooperation with each other. The position detection device 30 is configured to realize each function according to the present invention, for example, by loading a predetermined program into memory and executing the program using a processor such as a CPU.
[0023] FIG. 2 shows an example of a camera 10 installed in a room, viewed from the side. The environment in which the present invention is implemented is not limited to a room, and the camera 10 may also be implemented outdoors where there are no walls or ceilings. The camera 10 is installed facing sideways (i.e., facing horizontally) or diagonally downward (i.e., facing downward at an angle). The camera 10 may be attached to a wall as shown in FIG. 2, or may be attached to the ceiling, or may be installed using equipment such as a tripod.
[0024] It is preferable to install camera 10 so that the person whose position is to be detected is within the angle of view from the top of the head to the feet. In Figure 2, the area between the two dashed lines indicates the shooting range of camera 10, and camera 10 is installed so that it can optimally shoot people standing at points A and B.
[0025] 3A and 3B show the camera 10 as viewed from the front in the installation example of the camera 10 shown in FIG. 2. As mentioned above, the camera 10 may be tilted in the depression angle direction. However, it is desirable that the camera 10 be installed so that its horizontal axis is parallel or approximately parallel to the plane of the target area for position detection. In other words, as shown in FIG. 2A, it is preferable that the camera 10 be installed perpendicular or approximately perpendicular to the floor surface of the room, without being tilted relative to its optical axis.
[0026] Fig. 4 shows an example of an image of the target area for position detection captured by camera 10 in the imaging environment shown in Fig. 2. Hereinafter, the representation of positions on the image captured by camera 10 as shown in Fig. 4 will be referred to as the camera coordinate system.
[0027] Figure 5 shows an example of a bird's-eye view of the target area for position detection in the shooting environment shown in Figure 2. Hereinafter, the representation of positions from a bird's-eye view such as that in Figure 5 will be referred to as a bird's-eye coordinate system.
[0028] Here, the camera coordinate system and the overhead coordinate system represent the same space to be photographed from different viewpoints, and the position on the floor in one of these can be considered as the position on the floor in the other obtained by perspective projection transformation. Therefore, with regard to the camera coordinate system and the overhead coordinate system, the coordinate (x', y') on the floor in one coordinate system that corresponds to any coordinate (x, y) on the floor in the other coordinate system can be calculated using the following formula (1).
[0029]
number
[0030] In equation (1), c is a constant term, and M is a third-order square matrix called a perspective projection transformation matrix. The constant term c is not particularly important in the calculation, and once the perspective projection transformation matrix M is determined, the position in the overhead coordinate system corresponding to any point on the floor in the camera coordinate system can be uniquely determined.
[0031] The perspective projection transformation matrix M for a certain perspective projection transformation is uniquely determined once the correspondence between the coordinates of four points before transformation and the coordinates of four points after transformation is clear. Equation (1) and a method for deriving the perspective projection transformation matrix M are publicly known. For example, the perspective projection transformation matrix M can be easily calculated by running OpenCV (https: / / github.com / opencv / opencv), an open source software for image processing, on a computer operating as the position detection device 30.
[0032] Figure 6 shows an overview of the processing steps of the position detection method executed by the position detection system of this example. As shown in Figure 6, the position detection method of this example is roughly divided into two stages of implementation procedures. In the first stage, the preparation stage, a perspective projection transformation matrix M that represents the correspondence between the camera coordinate system and the overhead coordinate system is calculated. In the second stage, the operation stage, the perspective projection transformation matrix M is used to calculate the position of the person in the overhead coordinate system from the results of person detection in the camera coordinate system based on the captured image.
[0033] The preparatory stage processing must be performed each time the type and settings of the camera 10 (for example, zoom magnification settings, etc., related to how the captured image is captured), or the installation position, height, or orientation of the camera 10 changes. In other cases, once the preparatory stage processing is performed, it does not need to be performed again. In the preparatory stage, to calculate the perspective projection transformation matrix M, the coordinates of four different points are identified in the camera coordinate system and the bird's-eye coordinate system. In this example, for simplicity of implementation, the coordinates of only two of the four points (referred to as points A and B) are measured, and the coordinates of the remaining two points (referred to as points A' and B') are identified based on the coordinate values of points A and B.
[0034] As shown in Figure 6, the detailed procedure of the preparation stage includes a step (S01) of detecting and locating a person at point A, a step (S02) of detecting and locating a person at point B, a step (S03) of calculating the coordinates of point A' from the coordinates of point A, a step (S04) of calculating the coordinates of point B' from the coordinates of point B, and a step (S05) of calculating a perspective projection transformation matrix M from the coordinates of points A, B, A', and B'.
[0035] In step S01, the coordinates of point A in the camera coordinate system are calculated by capturing an image of a person standing upright at point A using camera 10 and transmitting the captured image to position detection device 30. The position detection device 30 analyzes the captured image using a pre-prepared object detection AI to detect the person. The detection method and type of object detection AI employed here can be arbitrarily selected from known methods as long as they can identify the coordinates of the feet of the detection target in the captured image. For example, YOLO (You Only Look Once, see Non-Patent Document 1), EfficientDet (see Non-Patent Document 2), etc. can be used to detect people.
[0036] Fig. 7 shows an example in which the person detection result is superimposed on the image captured by camera 10. In Fig. 7, the rectangles displayed surrounding each person at point A and point B indicate the coordinates of the person estimated by person detection. In this example, the coordinates of the midpoint of the base of the rectangle showing the person detection result are set to the coordinates of the target person's feet.
[0037] In step S01, as an example of a method for determining the coordinate of point A in the bird's-eye coordinate system, the relative position of point A from a reference point may be measured using a tool such as a tape measure as measuring instrument 20. As another example of a method for determining the coordinate of point A in the bird's-eye coordinate system, the position of a person present at point A may be measured using a positioning sensor such as a wireless sensor as measuring instrument 20. Note that in step S01, any point in the shooting environment may be set as point A. However, taking into consideration the procedures of steps S03 and S04 described below, it is preferable to set point A at a position closer to the right or left edge of the captured image rather than at a position near the center of the captured image.
[0038] In step S02, the coordinates of point B can be obtained in the same manner as the coordinates of point A were obtained in step S01. However, with regard to which point in the shooting environment is set as point B, it is preferable to set a point that is further back or closer to point A and sufficiently far away from point A as point B. Steps S01 and S02 may be performed in the order shown in FIG. 5 by having one person move from point A to point B. Furthermore, if possible, steps S01 and S02 may be performed simultaneously by placing two people at points A and B.
[0039] In step S03, position detection device 30 determines the coordinates of point A' based on the coordinates of point A determined in step S01. With regard to the camera coordinate system in step S03, as shown in Fig. 8, a vertical line that divides the image captured by camera 10 into left and right halves is set as the axis, and point A' is the coordinate that is line-symmetric to point A with respect to this axis. Also, with regard to the overhead coordinate system in step S03, as shown in Fig. 9, a line drawn from camera 10 in the direction in front of the camera is set as the axis, and point A' is the coordinate that is line-symmetric to point A with respect to this axis. In other words, the coordinates of point A' are determined by flipping the coordinates of point A left and right.
[0040] In step S04, the position detection device 30 determines the coordinates of point B' based on the coordinates of point B determined in step S02. The coordinates of point B' can be determined in the same manner as the coordinates of point A' determined in step S03. Steps S03 and S04 may be performed simultaneously, or in reverse order, if possible.
[0041] In step S05, the position detection device 30 calculates a perspective projection transformation matrix M that represents such a perspective projection transformation, by defining the coordinates of points A, B, A', and B' in the camera coordinate system as the coordinates of the four points before transformation, and defining the coordinates of points A, B, A', and B' in the overhead coordinate system as the coordinates of the four points after transformation.
[0042] As shown in Fig. 6, the detailed procedure of the operation stage includes a step (S06) of performing person detection and a step (S07) of performing perspective projection transformation on the person detection result using the perspective projection transformation matrix M. In addition, as conditional branching related to the repetition of the processing in the operation stage, there are a step (S08) of determining whether to continue the processing and a step (S09) of determining whether there is a change in the camera conditions.
[0043] In step S06, the position detection device 30 performs person detection by analyzing the image captured by the camera 10 using an object detection AI, and calculates the coordinates (x, y) of the target person's feet in the camera coordinate system. The person detection method here may be the same as in step S01.
[0044] In step S07, the position detection device 30 calculates the coordinates (x', y') of the target person in the overhead coordinate system based on the perspective projection transformation matrix M obtained in step S05 and the coordinates (x, y) of the target person's feet in the camera coordinate system obtained in step S06, in accordance with the above equation (1).
[0045] The procedure of steps S01 to S07 described above allows the position of a target person to be detected from an image captured by camera 10, which is a monocular camera. The accuracy of position detection depends on the accuracy of position estimation in the camera coordinate system by the object detection AI in steps S01, S02, and S06, the accuracy of measurement in the bird's-eye coordinate system in steps S01 and S02, and the size of the area surrounded by the four points A, B, A', and B'. When the four points A, B, A', and B' are sufficiently far apart and the area surrounded by the four points is large, position detection according to the present invention functions well. However, when the four points are close together and the area surrounded by the four points is small, it is expected that the error in position detection will increase with increasing distance from the area surrounded by the four points.
[0046] The procedure of steps S06 and S07 allows the position of a target person to be detected from a single still image captured by camera 10, but position detection can be repeated multiple times by passing through the conditional branching of steps S08 and S09. In step S08, a decision to continue or terminate the process can be made arbitrarily based on the number of times or time, or by the person performing the decision each time. In step S09, it is determined whether the conditions of camera 10, such as the type and settings of camera 10 (e.g., zoom magnification settings, etc., related to how the captured image appears), and the position, height, and orientation of camera 10, have changed since the previous preparation stage. If the conditions of camera 10 have changed, the process returns to step S01 and starts again from the preparation stage. If the conditions of camera 10 have not changed, the process returns to step S06 and starts again from the operation stage.
[0047] Several examples of the position detection system according to the present invention will be given below, but the present invention is not limited to these.
[0048] (First Example) In the first embodiment, in steps S01 and S02, a UWB sensor array is installed at the same position as camera 10 or at a position sufficiently close thereto (for example, directly below the camera), and people standing at points A and B carry UWB sensors to measure the person's position (the coordinates of points A and B) in a bird's-eye coordinate system. In this example, a UWB sensor array and a UWB sensor are used as measuring device 20, and the former sensor array is considered to be the anchor (wireless base station) and the latter sensor is considered to be the tag (wireless slave station), making it possible to measure the relative position of the tag as seen from the anchor. By using a sensor array as the anchor UWB sensor, it is possible to utilize the difference in the arrival times of signals, known as TDoA (Time Difference of Arrival), enabling relatively high-accuracy positioning.
[0049] The wireless sensor method and wireless sensor system used can be selected from any of UWB, Bluetooth, GPS, RFID, etc. Furthermore, the positioning method can be selected from any of methods using RSSI (Received Signal Strength Indicator), etc., in addition to TDoA. In other words, it is possible to flexibly select the wireless sensor method, wireless sensor system, and positioning method according to requirements such as the scale and characteristics of the environment in which this system is operated, the required position detection accuracy, and the cost required to prepare the wireless sensors.
[0050] The wireless sensor described above is used only in steps S01 and S02 and is not required in other steps. Therefore, the person whose location is to be detected does not need to carry a wireless sensor. Furthermore, even when implementing the present invention in multiple areas using multiple monocular cameras, the wireless sensor can be reused, so it is possible to implement the present invention by preparing only one pair of anchor and tag.
[0051] (Second Example) In the explanation so far, steps S03 and S04 are performed to determine the coordinates of points A' and B'. This applies to the case where the camera 10 is installed perpendicular or approximately perpendicular to the plane of the target area (e.g., the floor of a room) without being tilted with respect to the optical axis, as shown in FIG. 3A. In the second embodiment, instead of steps S03 and S04, the coordinates of points A' and B' are determined using a method similar to steps S01 and S02. This makes it possible to implement the present invention even when the camera 10 is installed tilted with respect to the optical axis, as shown in FIG. 3B.
[0052] (Third Example) In the third embodiment, a video camera is used as the camera 10, so that position detection is performed in real time based on live video or video, not just a single still image. In this case, steps S06, S07, S08, and S09 are performed for each frame of the video. According to the third embodiment, it is possible to detect the movement trajectory of the target person based on the video captured by the camera 10.
[0053] (Fourth Example) In practical application of the present invention, it is expected that multiple objects whose positions are to be detected exist within the same area. In this case, if a portion of an object located further back relative to the camera 10 is hidden by an object located closer to the camera, the position of the object located further back cannot be properly detected. In particular, when capturing an image with a depressed angle of view as shown in FIG. 2, the upper part of the object whose position is to be detected may be clearly visible, but the lower part may be hidden by an object located closer to the camera. If the lower part of the object whose position is to be detected is hidden, and the feet (the part that is in contact with the ground) cannot be captured, there is a concern that the position error in the depth direction will be significant. To address this issue, in cases where the apparent width-to-height ratio of the object is expected to be relatively constant, in the fourth embodiment, when the ratio of the height to the width of the detection result is smaller than a threshold, the lower part of the detection result is interpolated so that the width-to-height ratio of the detection result becomes a predetermined ratio.
[0054] Figure 10 shows an example of complementation when a part of the object of position detection is hidden by an obstruction. In Figure 10, two detection results, detection result 1 and detection result 2, are shown by solid lines as the results of analyzing the object of position detection from the image captured by camera 10 using object detection AI. In other words, two rectangles are displayed so as to surround the two objects detected from the image captured by camera 10, respectively.
[0055] Here, in detection result 1, the rectangle surrounding the object is vertically long and the ratio of width to height of the rectangle is normal. On the other hand, in detection result 2, the rectangle surrounding the object is nearly square, and looks unnatural for a standing person. Therefore, the position detection device 30 acquires detection result 2' by extending the bottom side of detection result 2, and performs position detection using detection result 2' instead of detection result 2. That is, in the operational stage, if a rectangle indicating the range of the object in the image captured by camera 10 is acquired and the ratio of height to width is smaller than a threshold, the rectangle is interpolated downward so that the ratio of height to width becomes a predetermined ratio, and then the coordinates of the object are calculated.
[0056] As described above, the position detection system of this example includes a camera 10 that captures an image of a target area, a measuring instrument 20 that acquires the positional relationship between the camera 10 and an object, and a position detection device 30 that detects the position of the object based on the image captured by the camera 10. In a preparation stage, the position detection device 30 performs a process of calculating a perspective projection transformation matrix M for converting the coordinates of the object in the image captured by the camera 10 to the actual position of the object based on the positional relationship between the camera 10 and the object at multiple points acquired using the measuring instrument 20, and in an operation stage, it calculates the coordinates of the object in the image captured by the camera 10 and converts them to the actual position of the object using the perspective projection transformation matrix M. This makes it possible to detect the position of an object with high accuracy using a general-purpose and widely-used monocular camera, with an easy-to-use method that has few restrictions on the scenes and objects to which it can be applied.
[0057] In the above description, the camera 10 has been installed facing sideways or diagonally downward. However, if an AI model that detects objects from images with a top-down angle of view is available, the camera 10 may be installed facing downward or approximately downward. In this case, however, unless the camera 10 is installed at a high position, a sufficient shooting range cannot be secured, and the range of position detection will be narrowed. For this reason, in an environment where the camera 10 cannot be installed at a high position, it is preferable to install the camera 10 facing sideways or diagonally downward. Furthermore, since there is a concern that the accuracy of position detection of objects located at the back will decrease in images captured when the camera 10 is installed sideways, it is more preferable to install the camera 10 facing diagonally downward.
[0058] Although the embodiments of the present invention have been described above, these embodiments are merely illustrative and do not limit the technical scope of the present invention. The present invention can take on various other embodiments, and various modifications such as omissions and substitutions can be made without departing from the spirit of the present invention. These embodiments and modifications thereof are included in the scope and spirit of the invention described in this specification, etc., and are included in the invention described in the claims and their equivalents.
[0059] Furthermore, the present invention can be provided not only as devices such as those described above or as systems composed of these devices, but also in the form of methods executed by these devices, programs for realizing the functions of these devices using a processor, storage media on which such programs are stored in a computer-readable format, etc. [Industrial Applicability]
[0060] The present invention can be used in a position detection system that detects the position of an object based on an image captured by a camera that captures an area of interest. [Explanation of symbols]
[0061] 10: camera, 20: measuring instrument, 30: position detection device
Claims
1. A position detection system including a camera that photographs a target area and a position detection device that detects the position of a target object based on an image photographed by the camera, further comprising a measuring device for acquiring a positional relationship between the camera and the object; the camera is positioned so that a horizontal axis of the camera is parallel or substantially parallel to a plane of the target area; The position detection device In a preparation stage, a process is performed to calculate a transformation matrix for converting the coordinates of the object in the image captured by the camera into the actual position of the object, based on the positional relationship between the camera and the object at two points in the target area obtained using the measuring device and the positional relationship between two other points obtained by flipping the positional relationship between the camera and the object. A position detection system characterized in that, in an operational stage, the coordinates of the object in the image captured by the camera are calculated and converted into the actual position of the object using the transformation matrix.
2. A position detection system comprising a camera that photographs a target area and a position detection device that detects the position of an object based on an image photographed by the camera, further comprising a measuring device for acquiring a positional relationship between the camera and the object; The position detection device In a preparation stage, a process is performed to calculate a transformation matrix for transforming the coordinates of the object in the image captured by the camera into an actual position of the object, based on the positional relationship between the camera and the object at a plurality of points acquired using the measuring instrument; A position detection system characterized in that, in an operational stage, a rectangle surrounding the object contained in the image captured by the camera is obtained, and if the ratio of the height of the rectangle to the width of the rectangle is smaller than a predetermined threshold, the rectangle is interpolated downward so that the ratio of the height of the rectangle to the width of the rectangle becomes a predetermined ratio, the coordinates of the object in the image captured by the camera are calculated based on the interpolated rectangle, and a process is performed to convert the coordinates to the actual position of the object using the transformation matrix.
3. 3. The position detection system according to claim 1, A position detection system characterized in that the measuring device is a wireless sensor that performs positioning wirelessly.
4. 3. The position detection system according to claim 1, A position detection system characterized in that the camera is installed facing sideways or diagonally downward.
5. A position detection method for detecting the position of an object based on an image captured by a camera that captures an area of interest, comprising: the camera is positioned so that a horizontal axis of the camera is parallel or substantially parallel to a plane of the target area; In a preparation stage, a measuring device for acquiring the positional relationship between the camera and the object is used to acquire the positional relationship between the camera and the object at two points within the target area, and a process is performed to calculate a transformation matrix for converting the coordinates of the object in the image captured by the camera into the actual position of the object, based on the positional relationship between the camera and the object at two points within the target area and the positional relationship between another two points obtained by flipping this left and right; A position detection method characterized in that, in an operation stage, the coordinates of the object in the image captured by the camera are calculated and converted to the actual position of the object using the transformation matrix.
6. A position detection method for detecting the position of an object based on an image captured by a camera that captures an area of interest, comprising: In a preparation stage, a measuring device for acquiring a positional relationship between the camera and the object is used to acquire the positional relationship between the camera and the object at a plurality of points, and a process is performed to calculate a transformation matrix for converting the coordinates of the object in the image captured by the camera into an actual position of the object based on the positional relationship between the camera and the object at the plurality of points; A position detection method characterized in that, in an operational stage, a rectangle surrounding the object contained in the image captured by the camera is obtained, and if the ratio of the height of the rectangle to the width of the rectangle is smaller than a predetermined threshold, the rectangle is interpolated downward so that the ratio of the height of the rectangle to the width of the rectangle becomes a predetermined ratio, the coordinates of the object in the image captured by the camera are calculated based on the interpolated rectangle, and a process is performed to convert the coordinates to the actual position of the object using the transformation matrix.