Three-dimensional position identifying apparatus

The three-dimensional position specifying device enhances accuracy by using a two-dimensional image sensor, range image sensor, and ring buffer storage to increase point clouds and reliability information, addressing the low precision of conventional LiDAR systems.

JP2025130185APending Publication Date: 2025-09-08BRAINS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024027189
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2025-09-08

Smart Images

  • Figure 2025130185000001_ABST
    Figure 2025130185000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of a three-dimensional position identified by a depth-image sensor.SOLUTION: A three-dimensional position identifying apparatus includes: a two-dimensional image sensor; a depth-image sensor; object recognition means which recognizes an object from a two-dimensional image; point cloud extraction means which extracts a point cloud included in a region of the recognized object; three-dimensional position identifying means which identifies a three-dimensional position using data corresponding to the extracted point cloud; and three-dimensional position information output means which outputs three-dimensional position information of the identified object. The three-dimensional position identifying apparatus is configured to sequentially write the data corresponding to the extracted point cloud, to a ring buffer that can store data for a predetermined period, in association with object identification information uniquely assigned to each of recognized objects.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a three-dimensional position specifying device for specifying the three-dimensional position of an object. [Background technology]

[0002] Conventionally, in a surveillance system that uses a network camera to monitor a specific space such as an intersection, a railroad crossing, or inside a factory, it has been disclosed that a LiDAR, which is a distance image sensor, is used to identify the position of an object present in the specific space (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-124496 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the LiDAR, which is a distance image sensor disclosed in Cited Document 1, emits laser light and determines the distance of reflection points (point clouds) using the reflection of the laser light. Therefore, it is necessary to scan a specific area by rapidly and sequentially changing the direction of emission of the laser light. Scanning the entire specific area takes a certain amount of time, and the reflection points (point clouds) whose distances are determined are coarser than two-dimensional images. Therefore, depending on the scanning conditions, the number of reflection points (point clouds) included within the range of the object may become extremely small. If the number of reflection points (point clouds) included within the range of the object is extremely small, there is a problem that the accuracy of the three-dimensional position of the object determined by the distances of these reflection points (point clouds) becomes very low.

[0005] The present invention has been made in light of these problems, and aims to provide a three-dimensional position identification device that can improve the accuracy of three-dimensional positions identified using a range image sensor. [Means for solving the problem]

[0006] The three-dimensional position specifying device of claim 1 a two-dimensional image sensor capable of capturing a two-dimensional image of a specific space; a range image sensor capable of generating a three-dimensional point cloud range image of the specific space including a plurality of point clouds whose distances are specified; an object recognition means for recognizing an object from the two-dimensional image captured by the two-dimensional image sensor; a point cloud extraction means for extracting a point cloud included in the area of ​​the object recognized by the object recognition means from the point cloud included in the three-dimensional point cloud range image generated by the range image sensor; a three-dimensional position specifying means for specifying a three-dimensional position of the object using data corresponding to the point cloud extracted by the point cloud extracting means; a three-dimensional position information output means for outputting three-dimensional position information of the object identified by the three-dimensional position identifying means; A three-dimensional position specifying device comprising: sequentially writing data corresponding to the point cloud extracted by the point cloud extraction means as a point cloud included in the area of ​​each object, in association with object identification information uniquely assigned to each object recognized by the object recognition means, into a ring buffer capable of retaining data for a predetermined period of time; It is characterized by the following. According to this feature, the point cloud data contained in the area of ​​each object extracted over a specified period is written sequentially into the ring buffer, allowing more point clouds to be stored in the ring buffer, thereby improving the accuracy of the identified three-dimensional positions, and by using a ring buffer that does not require data erasure, it is possible to prevent the processing load required to improve this accuracy from increasing.

[0007] The three-dimensional position specifying device of claim 2 is the three-dimensional position specifying device according to claim 1, The object recognition means is capable of identifying the type of object, data corresponding to the point cloud extracted by the point cloud extraction means is written sequentially into ring buffers having different storage capacities corresponding to the types of objects recognized by the object recognition means; It is characterized by the following. This feature makes it possible to prevent the storage capacity of the ring buffer from becoming unnecessarily large.

[0008] The three-dimensional position specifying device of claim 3 is the three-dimensional position specifying device according to claim 1 or 2, an occlusion determination means for determining the occurrence of occlusion in the specific space, when the occurrence of occlusion is determined by the occlusion determination means, the three-dimensional position specifying means specifies the three-dimensional position using data corresponding to the point cloud at the time of the occurrence as well as data corresponding to the point cloud stored in the ring buffer during a period before the occurrence. It is characterized by the following. This feature makes it possible to prevent the accuracy of the three-dimensional position from being significantly reduced due to occlusion.

[0009] The three-dimensional position specifying device of claim 4 is the three-dimensional position specifying device according to claim 1 or 2, a reliability information generating means for generating reliability information of the position identified by the three-dimensional position identifying means based on the storage status of the ring buffer; The three-dimensional position information output means outputs the three-dimensional position information to which the reliability information generated by the reliability information generation means has been added. It is characterized by the following. According to this feature, in a system using a plurality of three-dimensional position specifying devices, it is possible to specify three-dimensional position information with higher accuracy by comparing reliability information.

[0010] The three-dimensional position specifying device of claim 5 is the three-dimensional position specifying device according to claim 1, the object is a wireless communication terminal capable of wireless communication, the three-dimensional position information output means outputs the three-dimensional position information of the wireless communication terminal to a radio wave propagation control means that controls radio wave propagation in the specific space. It is characterized by the following. According to this feature, it is possible to output highly accurate three-dimensional position information to the radio wave propagation control means, thereby enabling highly accurate control of radio wave propagation in a specific space.

[0011] Furthermore, the present invention may have only the invention-specific matters set forth in the claims of the present invention, or may have the invention-specific matters set forth in the claims of the present invention as well as configurations other than the invention-specific matters. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a plan view showing a room as an example of a specific space to which a three-dimensional position specifying device according to an embodiment of the present invention is applied; [Figure 2] 1 is a block diagram showing the configuration of a radio wave propagation control system having a three-dimensional position specifying device according to an embodiment of the present invention; [Figure 3] 1 is a block diagram showing the configuration of a three-dimensional position specifying device according to an embodiment; [Figure 4] FIG. 2 is an explanatory diagram showing the configuration of a RAM in the three-dimensional position specifying device of the embodiment. [Figure 5] FIG. 10 is a diagram illustrating the relationship between a ring buffer and an object type in the three-dimensional position specifying device according to the embodiment. [Figure 6] FIG. 2 is a flowchart showing a process flow in the three-dimensional position specifying device of the embodiment. [Figure 7] FIG. 4 is a flowchart showing the processing contents of image and point cloud data integration processing in the three-dimensional position specifying device of the embodiment. [Figure 8] FIG. 1 is an explanatory diagram showing an example of the state of a point cloud obtained by LiDAR used in the three-dimensional position specifying device of the embodiment. [Figure 9]FIG. 2 is an explanatory diagram showing the relationship between a point cloud obtained by LiDAR used in the three-dimensional position specifying device of the embodiment and an object. DETAILED DESCRIPTION OF THE INVENTION

[0013] The following describes a mode for implementing the three-dimensional positioning device of the present invention based on an embodiment. Note that in the following embodiment, an embodiment in which the three-dimensional positioning device is applied to a radio wave propagation control system is illustrated, but the present invention is not limited to this, and it goes without saying that the present invention can be applied to systems other than radio wave propagation control systems. [Example]

[0014] FIG. 1 is a diagram showing a room R, which is a specific space to which a radio wave propagation control system including a three-dimensional positioning device of the present invention is applied, and FIG. 2 is a block diagram showing the configuration of a radio wave propagation control system including a three-dimensional positioning device of the present invention.

[0015] As shown in FIG. 1, a room R, which is a specific space in this embodiment, is rectangular and is provided with a door for entering the room R.

[0016] As shown in Figure 1, room R is equipped with desks D1 and D2 on which wireless LAN (Wi-Fi) terminals such as laptops (note PCs) and smartphones, which are sources of radio waves, can be placed, allowing employees to work or rest at desks D1 and D2.

[0017] In the room R, chairs and the like (not shown) are placed, and office equipment such as a printer is also placed in predetermined positions.

[0018] In room R, as shown in FIG. 1, wireless LAN (Wi-Fi) access point terminals (AP) 11 and 12 are placed on the wall surfaces at diagonally opposite corners of room R.

[0019] The access point terminals (AP) 11, 12, like publicly known wireless LAN (Wi-Fi) access point terminals, are capable of two-way wireless communication with wireless LAN terminals using radio waves in the 2.4 GHz band or 5 GHz band, and in particular have the function of transmitting data on the session status (communication status) with the wireless LAN terminal to the propagation path control device 3 described below, and the function of three-dimensional beamforming, which concentrates radio waves of a specified frequency (channel) toward a three-dimensional position specified by instructions from the propagation path control device 3.

[0020] For these three-dimensional beamforming techniques, known methods can be used, such as three-dimensional beamforming techniques using planar antenna arrays, which are being considered for use in 5G mobile phones.

[0021] In addition, the access point terminals AP1 and AP2 are connected to the external network, the Internet, via LAN cables and a switching hub with router functionality, and wireless LAN (Wi-Fi) terminals present inside room R can connect to the Internet via the access point terminals AP1 and AP2 and the switching hub.

[0022] As shown in Figure 1, on the wall near the ceiling in the corner opposite the door that serves as the entrance to Room R, A sensor unit 1 (see Figure 2) having an RGB camera (two-dimensional image sensor) 11 capable of capturing two-dimensional images (also called RGB images) and a laser scanner (LiDAR) 12, which is an active distance sensor that can emit laser light and measure the distance to the point where the emitted laser light is reflected, is installed so as to be able to overlook the interior of room R.

[0023] In addition, on the wall near the sensor unit 1, there are provided a processing device 2 having a box-shaped housing connected to the sensor unit 1, and a propagation path control device 3 for controlling the propagation path of radio waves within the room R to improve the data communication environment, and the processing device 2 and the propagation path control device 3 are connected to enable two-way data communication.

[0024] The three-dimensional position specifying device of the present invention is composed of a sensor unit 1 and a processing device 2.

[0025] As hardware, the processing device 2 has relatively excellent processing capabilities and is highly reliable and less prone to malfunctions, and can suitably be a known industrial PC (personal computer) having a LAN communication interface, etc., and these industrial PCs are used with various programs installed to provide the functions required by the processing device 2, but the present invention is not limited to this, and the processing device 2 can be any computer as long as it is capable of providing each of the functions of the processing device 2 described below.

[0026] Specifically, as shown in FIG. 3, the processing device 2 of this embodiment includes a central processing unit (CPU) 21 that performs various processes such as object detection and three-dimensional position identification of the detected object through calculations, an image processing unit (GPU) 28 that mainly performs calculations related to a deep learning neural network (DNN), a random access memory (RAM) 22 that is used as a work memory or the like in the processing by the central processing unit (CPU) 21 and the image processing unit (GPU) 28, and a memory for storing various processes executed by the central processing unit (CPU) 21 and the image processing unit (GPU) 28. It is a relatively small computer connected to a flash memory (F-MEM) 23 that stores processing programs for the system and various parameters related to initial settings, a real-time clock (RTC) 26 that outputs time information and generates system time, a laser scanner interface (I / F) 24 that inputs and outputs various data to and from the laser scanner (LiDAR) 12, a camera interface (I / F) 25 that inputs and outputs various data to and from the RGB camera 11, and a communication interface (I / F) 27 for data communication with the propagation path control device 3.

[0027] The laser scanner interface (I / F) 24 has a function of receiving point cloud data (also called range image data) output from the laser scanner (LiDAR) 12 and transferring it to the central processing unit (CPU) 21 via the data bus 20. The camera interface (I / F) 25 has a function of receiving two-dimensional image data output from the RGB camera 11 and transferring it to the central processing unit (CPU) 21 via the data bus 20.

[0028] As the laser scanner interface (I / F) 24 and the camera interface (I / F) 25, specifically, a well-known USB interface (I / F) used as an interface (I / F) with peripheral devices in a computer can be suitably used.

[0029] The RGB camera 11 is capable of capturing two-dimensional images at a frame rate of 30 FPS (30 frames per second), which is the normal frame rate of video, and outputting color image data, and can be a known CCD camera capable of capturing video.

[0030] Furthermore, a known laser scanner (LiDAR) can be used as the laser scanner (LiDAR) 12, but it is preferable to use one that is capable of relatively high-speed scanning.

[0031] Furthermore, a propagation path control device 3 is disposed near the processing device 2 and is communicatively connected to the processing device 2. Access point terminals AP1 and AP2 are connected to the propagation path control device 3 via a wired LAN, and the propagation path control device 3 is capable of executing three-dimensional beamforming control of these access point terminals AP1 and AP2, thereby actively controlling radio wave propagation in room R, which is a specific space.

[0032] As with the processing device 2, the propagation path control device 3 is a highly reliable piece of hardware that has relatively excellent processing capabilities and is less susceptible to breakdowns, and can be suitably configured using a computer such as a known industrial PC that has a LAN communication interface, etc.

[0033] The propagation path control device 3 stores module programs for providing various functions as shown in Fig. 2, as well as various data used by these module programs. Specifically, the module programs that can be run on the operation system (OS) of the propagation path control device 3 include a propagation path calculation module program that calculates the propagation paths between the access point terminals AP1 and AP2 and the wireless LAN terminals in room R, and an AP control module program that controls the direction of three-dimensional beamforming optimized for the propagation paths calculated by the propagation path calculation module program. In Fig. 2, thick rectangular dashed lines indicate module programs, and thin cylindrical dashed lines indicate data.

[0034] In addition, the data used by these module programs includes a communication status database (DB) that stores session information including the three-dimensional position coordinates of the access point terminals AP1 and AP2 within room R, the local IP addresses and MAC addresses of the wireless LAN terminals within room R with which each access point terminal AP1 and AP2 is communicating, the radio frequency (channel) used for transmission and reception, and communication level information, etc.; a terminal database that sequentially updates the three-dimensional position of an object that is a wireless LAN terminal within room R whose position is identified by the processing device 2, and information on the radio frequency emitted from the wireless LAN terminal; and terminal correspondence table data that stores the correspondence between the terminal ID used in the terminal DB and the MAC address that serves as terminal identification information in the communication status database.

[0035] The processing device 2 also stores module programs for providing various functions as shown in FIG. 2, as well as various data used by these module programs.

[0036] As shown in Figure 2, the module programs that can be run on the operating system (OS) of the processing device 2 include an image processing module program for performing corrections, etc. on image data sent from the RGB camera 11 by calculations using image processing data consisting of determinants, etc.; a deep learning (Deep Neural Network; DNN) module program that performs deep learning (DNN) using training data to identify the area and type of object; an object detection module program that uses the trained deep learning (DNN) to detect objects present in room R from image data (captured image) captured by the RGB camera 11; and a detected object position identification module program that identifies the three-dimensional position of objects present in room R from point cloud data of the laser scanner (LiDAR) 12 corresponding to the area of ​​the object detected by the object detection module program.

[0037] The data stored in advance includes image processing data, which is made up of various data used by the image processing module program to correct images.

[0038] Next, the data stored in the RAM 22 of the processing device 2 will be briefly described with reference to Fig. 4. As shown in Fig. 4, the RAM 22 stores a processing program booted from the flash memory (F-MEM) 23, and also mainly stores two-dimensional image data with a system time attached, LiDAR point cloud data with a system time attached, and an object detection list in which objects extracted from the two-dimensional image data are registered. The RAM 22 also has a frame ring buffer configured to store a predetermined number of frames (e.g., five frames) of two-dimensional image data and LiDAR point cloud data set by initial setting based on parameters. The two-dimensional image data and LiDAR point cloud data are sequentially overwritten and stored in these frame ring buffers, thereby updating the data without the need to erase old two-dimensional image data or LiDAR point cloud data.

[0039] In addition, RAM 22 is also set with individual ring buffers for objects in which point cloud data existing within the area of ​​the object extracted from the two-dimensional image data is accumulated and stored for a predetermined accumulation time, and new point cloud data is sequentially overwritten and stored in these individual ring buffers for objects, so that the point cloud data is accumulated and stored over the accumulation time without the need to perform a process to erase old point cloud data.

[0040] As these individual ring buffers for objects, multiple types of ring buffers with different accumulation times are defined and set. Specifically, as shown in the example in Figure 5, a different accumulation time is defined for each type of object identified by object detection. For example, for a desk which is very unlikely to move, 600 seconds is defined, for a chair, 300 seconds is defined, for a printer, 60 seconds is defined, for a laptop, 10 seconds is defined, and for a person who is likely to move, 1 second is defined.

[0041] In this way, for an object whose three-dimensional position is unlikely to change when it moves (moves), such as a desk, the accuracy of the identified three-dimensional position can be improved by increasing the amount of point cloud data used to identify the three-dimensional position by making point cloud data from a long period available since the object's position is unlikely to change.On the other hand, for an object whose three-dimensional position is likely to change when it moves (moves), such as a person, the object's position is likely to change, so if point cloud data from a long period is used, even point cloud data from before the movement will be used to identify the three-dimensional position, which could actually reduce the accuracy of the three-dimensional position.Therefore, it is sufficient to define a short accumulation time.Note that the accumulation time shown in Figure 5 is just an example and can be determined appropriately depending on the situation of the specific space (indoors, outdoors, etc.).

[0042] Furthermore, instead of having only one ring buffer set for each type of ring buffer corresponding to these types, multiple ring buffers are set for each type, such as, for example, in the case of a desk, there is a ring buffer of type desk in which point cloud data for desk D1 is stored, and a ring buffer of type desk in which point cloud data for desk D2 is stored.

[0043] In other words, as will be described later, when a new object is detected in object detection, the individual ring buffer for objects is individually assigned to each object ID assigned to the object, and point cloud data is accumulated and stored over an accumulation time in association with the object ID. However, the present invention is not limited to this, and it is also possible to provide only one of these ring buffers for each type of object, and when two desks, desk D1 and desk D2, are detected as objects, the data at that time may be stored as the data for that time, along with the object ID assigned to desk D1, and the data for the point cloud located within the area of ​​desk D2 may also be stored as the data for that time.

[0044] In this embodiment, an example is given in which the ring buffer is defined by the accumulation time, but the present invention is not limited to this. In the case of an embodiment in which a ring buffer is individually assigned to each object, as in this embodiment, the ring buffer may be defined by a storage capacity that is at least capable of storing point cloud data of the accumulation time.

[0045] Next, the processes executed by the processing device 2 will be described with reference to Figures 6 and 7. Note that some of the processes described below are executed by the CPU 28 alone, and some are executed by the CPU 28 and the GPU 28 working together.

[0046] The processing device 2 first acquires two-dimensional image data from the RGB camera 11 and adds the system time as time information (step S1). As described above, the two-dimensional image data to which the system time has been added is stored in the ring buffer for two-dimensional image data.

[0047] In addition, the point cloud data is acquired from the laser scanner (LiDAR) 12 and the system time is added as time information (step S2). Note that the point cloud data with the system time added is stored in the ring buffer for point cloud data, as described above.

[0048] Next, an object detection process using instance segmentation is performed on the two-dimensional image data acquired in step S1 to extract all objects contained in the two-dimensional image, identify the type of each extracted object, and create a detection region of interest (ROI) and mask for each object, which are then registered in an object detection list in association with the detection number assigned to each object (step S3).

[0049] Instance segmentation is a deep learning neural network (DNN) object detection method that performs real-time object detection on a pixel-by-pixel basis in two-dimensional images. For example, YOLACT (Daniel Bolya Chong Zhou Fanyi Xiao Yong Jae Lee. YOLACT: Real-time Instance Segmentation. In ICCV, 2019) can be suitably used.

[0050] In addition, in this embodiment, an example is given of a form using YOLACT using a deep learning neural network (DNN) because it can perform object detection (extraction) in two-dimensional images at high speed and with high accuracy, but the present invention is not limited to this. If the movement speed of the object to be measured is relatively slow and the imaging period of the two-dimensional image is relatively long, and object detection (extraction) can be performed with sufficient accuracy using an object detection method that does not use a deep learning neural network (DNN) by spending a sufficiently long time on the object detection (extraction) process, a form that does not use a deep learning neural network (DNN) may be used, or even if a deep learning neural network (DNN) is used, a method other than the above-mentioned YOLACT may be used. The method of object detection (extraction) in these two-dimensional images can be selected appropriately based on the time available for object detection (extraction) processing based on the imaging period corresponding to the movement speed of the object whose speed is to be measured, and the required recognition accuracy.

[0051] In the deep learning neural network (DNN) of this embodiment, images of various laptops, printers, tablets, smartphones, etc. are input as training data and trained to enable highly accurate extraction of objects such as laptops, printers, tablets, and smartphones, which are wireless terminals. This allows the deep learning neural network (DNN) to be trained, thereby improving the accuracy of extraction of laptops, printers, tablets, and smartphones in particular.

[0052] Next, the image and point cloud data integration process shown in Fig. 7 is executed. In the image and point cloud data integration process of this embodiment, as shown in Fig. 7, a movement detection process is executed to determine whether or not the object detected in step S3 has moved (step S401).

[0053] In the movement detection process of this embodiment, the two-dimensional image acquired in step S1 at that time and stored in the two-dimensional image frame ring buffer is compared with the two-dimensional image acquired immediately before and stored in the two-dimensional image frame ring buffer (the two-dimensional image of the previous frame), and the presence or absence of movement of the detected object (instance) is determined from the difference between the two two-dimensional images.

[0054] It should be noted that the two-dimensional image of the previous frame is used to determine whether or not there is movement, but the present invention is not limited to this. As mentioned above, the frame ring buffer for two-dimensional images stores five frames of two-dimensional images including the current data, so the presence or absence of movement may be determined with higher accuracy by also referring to the two-dimensional image of the frame immediately before the current one.

[0055] Then, an identical object determination process is executed to determine whether the detected objects are the same (step S402). In this identical object determination process, for an object (instance) whose movement has not been detected in the movement detection process, it is determined whether the object is the same based on whether the detection regions (ROI) registered in the object detection list match. On the other hand, for an object (instance) whose movement has been detected in the movement detection process, it is determined based on the type registered in the object detection list, the degree of match in mask shape, the degree of match in detection region (ROI) size, etc.

[0056] Then, it is determined whether or not a non-identical object (instance), that is, a new object such as a person who has newly entered the room R or a smartphone carried by that person, has been detected (step S403).

[0057] In the determination, if there are no non-identical objects (instances) (step S403; N), the process proceeds to step S406. In this case, for the objects determined to be the same object, the same object ID as the object ID already assigned to the object (instance) determined to be the same continues to be assigned.

[0058] On the other hand, if a non-identical object (instance) exists (step S403; Y), a new object ID is issued and assigned to the non-identical object, and the object ID is registered as data for the object in the object detection list (step S404).

[0059] Next, in response to the detection of a new object, an individual ring buffer for the object having a storage time corresponding to the type of the non-identical object for which an object ID has been issued is defined and allocated (step S405).

[0060] In this way, an individual ring buffer for objects is assigned to objects that are newly detected in the object detection process and for which a new object ID is issued, so that the individual ring buffer for objects contains ring buffers corresponding to all objects (instances) detected in the object detection process.

[0061] Then, by referring to the LiDAR point cloud data acquired in step S2 and stored in the LiDAR point cloud frame ring buffer, all point clouds within the detected objects are extracted for each object (step S406).

[0062] To extract the point clouds within these objects, for example, in the case of the laptop computer shown in Figure 9(a), the detection region (ROI) of the object stored in the object detection list is the region of the laptop computer shown in Figure 9(a), so it is sufficient to extract the point clouds that exist within this detection region (ROI).

[0063] Then, the extracted LiDAR point cloud data is additionally written by overwriting into the individual object ring buffer assigned to the corresponding object (step S407).

[0064] By doing this, the point clouds present within the detection region (ROI) of all detected objects are sequentially added and stored in the individual object ring buffer. As a result, as shown in Figure 9(b), the number of point clouds extracted within the region of each object increases, and the point cloud filling rate (the ratio of the number of point clouds to pixels) increases.

[0065] Returning to FIG. 6, after the image and point cloud data integration process is performed, the point cloud data stored in the individual object ring buffer for each detected object is read out, and a three-dimensional position calculation process is performed to calculate the three-dimensional position of each object based on the distance, etc., determined from the point cloud data (step S5).

[0066] The three-dimensional position of an object is calculated as a reference position of the object, and this reference position may be, for example, an average position of three-dimensional positions identified from the point cloud, or may be the center of gravity, etc. Furthermore, these reference positions may be different for each type of object.

[0067] Next, an output information process is executed to output three-dimensional position information including the three-dimensional positions of the individual objects calculated in the three-dimensional position calculation process and reliability information of the three-dimensional positions to the propagation path control device 3 (step S6).

[0068] Note that the three-dimensional position information of only objects whose type corresponds to the radio wave terminal is output to the propagation path control device 3, rather than the three-dimensional position information of all objects. However, for example, when an obstacle is detected on the propagation path and control is performed such as switching from access point terminal AP1 to AP2, the three-dimensional position information of all objects may be output to the propagation path control device 3 so that the propagation path control device 3 can detect the presence of an obstacle on the propagation path.

[0069] In addition, the reliability information may be calculated using a predetermined formula with parameters such as the point cloud filling rate based on the number of point clouds stored in the individual ring buffer for objects at the time the three-dimensional position information is calculated, the distance from the laser scanner (LiDAR) 12, and the object type (whether it is a type that is likely to move or not).

[0070] As described above, according to the three-dimensional position identification device of this embodiment, point cloud data contained in the area of ​​each object is extracted and sequentially written into a ring buffer and accumulated therein, so that by increasing the number of point cloud data used to identify the three-dimensional position, the accuracy of the three-dimensional position identified by this point cloud data can be improved.

[0071] In other words, as shown in Figure 8, LiDAR point cloud data is obtained by scanning by sequentially changing the radiation direction of the emitted laser light, so the number of point clouds is significantly smaller than the pixels in the two-dimensional image, and the distance image based on the point cloud is coarse.In addition, the number of point clouds contained in the detection area of ​​an object may be significantly smaller, for example, if the object is located in an area of ​​room R where the laser scan density is low.However, even in such cases, by using an individual ring buffer for objects, the point cloud filling rate (the ratio of the number of point clouds to pixels), which is the number of point clouds extracted within an area in the two-dimensional image of each object, can be increased, thereby improving the accuracy of the three-dimensional position of the object identified by this point cloud data.

[0072] Furthermore, as shown in Figure 9(c), when occlusion occurs due to a person operating a laptop computer, the person's hand is detected as the area of ​​the person, which narrows the area of ​​the laptop computer. If the three-dimensional position were determined using only the point cloud data contained in this narrowed area of ​​the laptop computer, the position would be inaccurate. However, according to the present invention, the data before the occlusion occurred, which is stored in the individual ring buffer for objects that serves as a storage means, is also used to determine the three-dimensional position, thereby preventing a decrease in the accuracy of the three-dimensional position due to the occurrence of occlusion.

[0073] Furthermore, according to the three-dimensional position determination device of this embodiment, a ring buffer is used that does not require data erasure, so there is no need to perform a process to erase old data, which prevents the processing load from increasing due to the accumulation of point cloud data to improve accuracy.

[0074] Although the embodiments of the present invention have been described above with reference to the drawings, the specific configuration is not limited to these embodiments, and the present invention also includes modifications and additions that do not deviate from the gist of the present invention.

[0075] For example, in the above embodiment, an example is given in which an individual ring buffer for objects is used as storage means, but the present invention is not limited to this, and these storage means may be ordinary buffers other than ring buffers.

[0076] Furthermore, in the above embodiment, the image data and point cloud data acquired from the RGB camera 11 and the laser scanner (LiDAR) 12 are also stored in the frame ring buffer, thereby eliminating the processing load of data deletion, but the present invention is not limited to this, and the image data and point cloud data may also be stored in a normal buffer.

[0077] Furthermore, in the above embodiment, an example is given of a form in which ring buffers with different accumulation times are defined for each type of object, thereby preventing the storage capacity of the ring buffer (storage capacity of RAM) from becoming unnecessarily large, but the present invention is not limited to this, and these ring buffers may only be ring buffers for the type with the longest accumulation time.

[0078] Furthermore, in the above embodiment, the occurrence of occlusion is not actively determined, but the present invention is not limited to this. The occurrence of occlusion as shown in FIG. 9(c) may be actively determined (occlusion determination means) in the same object determination process (step S402), for example, and if the determination shows that no occlusion has occurred, only the point cloud data acquired at that time may be used, whereas if occlusion has occurred, the point cloud data accumulated and stored in the individual object ring buffer may be used in addition to the point cloud data acquired at that time. Alternatively, if no occlusion has occurred, point cloud data of a first period that is shorter than the point cloud data accumulated and stored in the individual object ring buffer may be used, and if occlusion has occurred, point cloud data of a second period that is longer than the first period may be used to identify the three-dimensional position.

[0079] Furthermore, in the above embodiment, the configuration has only one sensor unit 1, but the present invention is not limited to this. For example, a plurality of sensor units 1 may be connected to a processing device, and for each sensor unit 1, the three-dimensional position and reliability information of each object may be calculated, and the reliability information may be compared, and three-dimensional position information including the three-dimensional position of the sensor unit 1 with the highest reliability may be output.

[0080] Furthermore, in the above embodiment, the reliability information does not include information on the presence or absence of occlusion, but the reliability information may include information on the presence or absence of occlusion.

[0081] Furthermore, in the above embodiment, the specific space is illustrated as an indoor room R, but the present invention is not limited to this, and these specific spaces may be outdoors, such as parking lots or roads, and the object to be managed may be a vehicle that does not emit radio waves rather than a wireless terminal. [Explanation of symbols]

[0082] 1 sensor unit 2 Processing equipment 3 Propagation path control device 11 RGB camera (two-dimensional image sensor) 12 Laser scanner (LiDAR) R Room (specific space) AP1 Access Point AP2 Access Point

Claims

1. a two-dimensional image sensor capable of capturing a two-dimensional image of a specific space; a range image sensor capable of generating a three-dimensional point cloud range image of the specific space including a plurality of point clouds whose distances are specified; an object recognition means for recognizing an object from the two-dimensional image captured by the two-dimensional image sensor; a point cloud extraction means for extracting a point cloud included in the area of ​​the object recognized by the object recognition means from the point cloud included in the three-dimensional point cloud range image generated by the range image sensor; a three-dimensional position specifying means for specifying a three-dimensional position of the object using data corresponding to the point cloud extracted by the point cloud extracting means; a three-dimensional position information output means for outputting three-dimensional position information of the object identified by the three-dimensional position identifying means; A three-dimensional position specifying device comprising: sequentially writing data corresponding to the point cloud extracted by the point cloud extraction means as a point cloud included in the area of ​​each object, in association with object identification information uniquely assigned to each object recognized by the object recognition means, into a ring buffer capable of retaining data for a predetermined period of time; A three-dimensional position specifying device characterized by:

2. The object recognition means is capable of identifying the type of object, sequentially writing data corresponding to the point cloud extracted by the point cloud extraction means into ring buffers having different storage capacities corresponding to the types of objects recognized by the object recognition means; 2. The three-dimensional position specifying device according to claim 1.

3. an occlusion determination means for determining the occurrence of occlusion in the specific space, when the occurrence of occlusion is determined by the occlusion determination means, the three-dimensional position specifying means specifies the three-dimensional position using data corresponding to the point cloud at the time of the occurrence as well as data corresponding to the point cloud stored in the ring buffer during a period before the occurrence.

3. The three-dimensional position specifying device according to claim 1 or 2.

4. a reliability information generating means for generating reliability information of the position identified by the three-dimensional position identifying means based on the storage status of the ring buffer; The three-dimensional position information output means outputs the three-dimensional position information to which the reliability information generated by the reliability information generation means has been added.

3. The three-dimensional position specifying device according to claim 1 or 2.

5. the object is a wireless communication terminal capable of wireless communication, the three-dimensional position information output means outputs the three-dimensional position information of the wireless communication terminal to a radio wave propagation control means that controls radio wave propagation in the specific space.

2. The three-dimensional position specifying device according to claim 1.

Citation Information

Patent Citations

  • Light detection and ranging (LIDAR) device

    JP2021124496A