Object dimension marking device
The device uses a monocular camera and radar to calculate and assign absolute dimensions to three-dimensional depth-estimated images, addressing the lack of absolute dimensions in existing methods and enabling efficient, precise dimension assignment in confined spaces.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2026-04-06
AI Technical Summary
Existing three-dimensional depth-estimated images obtained by monocular cameras lack absolute dimensions, necessitating time-consuming methods like using gauge blocks for calibration, and current methods are inefficient for assigning absolute dimensions in confined spaces.
A device and method using a monocular camera, radar, and three-dimensional imaging to calculate and assign absolute dimensions to depth-estimated images, employing alignment techniques and radar-based imaging to create an occupied grid map for accurate dimension assignment.
Enables easy and accurate assignment of absolute dimensions to three-dimensional depth-estimated images without the need for physical gauge blocks, allowing high-precision alignment and dimension assignment in confined spaces.
Smart Images

Figure 0007840532000001 
Figure 0007840532000002 
Figure 0007840532000003
Abstract
Description
Technical Field
[0001] The present invention relates to an object dimension value - attaching device and an object dimension value - attaching method for attaching absolute - value dimensions to a three - dimensional depth - estimated image calculated from information obtained by a monocular camera.
Background Art
[0002] Currently, many inspection operations in narrow areas are carried out, such as plant equipment inspection, building maintenance, bridge support inspection, and ceiling leak inspection. In narrow areas, for example, in addition to direct visual inspection, fiber scopes and video scopes are used. If a video scope with a binocular camera is used, a three - dimensional image can be obtained, but an image cannot be obtained unless the object is approached to about 4 cm, so it is not suitable when it is desired to view the whole object from a certain distance.
[0003] On the other hand, if a monocular camera is used, an image can be obtained even from a position about 1 m away from the object. Also, for the image obtained by the monocular camera, for example, by using AI (artificial intelligence) such as Visual - SLAM or CNNs, FCNs, etc., a three - dimensional depth - estimated image of the object can be obtained. However, the three - dimensional depth - estimated image obtained by the monocular camera has a problem that although it has relative - value dimensions, the absolute dimensions are unknown. Therefore, it was necessary to attach absolute - value dimensions to the three - dimensional depth - estimated image.
[0004] In Patent Document 1, for an input image, the depth of each pixel of the input image is estimated based on a depth - estimation model for estimating depth that has been pre - learned, and the input of dimension information of a part in the input image whose dimension in the real space is known is received. Using the depth of each pixel corresponding to the part of the estimated dimension information and the dimension information, the depth of each pixel of the estimated input image is corrected. However, in Patent Document 1, since a part whose dimension in the real space is known is required, for example, it is necessary to place a block gauge, etc., take a picture together, and calibrate, which is a problem of being time - consuming. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2018-132477 [Overview of the project] [Problems that the invention aims to solve]
[0006] This invention was made based on the above problems, and aims to provide an object dimension value assignment device and object dimension value assignment method that can easily assign absolute dimensions to three-dimensional depth estimation images. [Means for solving the problem]
[0007] The object dimension value assignment device of the present invention comprises a monocular camera for photographing an object, a three-dimensional depth estimation image calculation means for calculating a three-dimensional depth estimation image of the object from information obtained by the monocular camera, a radar for irradiating the object with radio waves and measuring the reflected waves, a three-dimensional imaging means for creating a three-dimensional imaging image of the object based on information obtained by the radar, and an absolute dimension value assignment means for assigning absolute dimensions to the three-dimensional depth estimation image based on the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means and the three-dimensional imaging image obtained by the three-dimensional imaging means.
[0008] The object dimension value assignment method of the present invention includes a three-dimensional depth estimation image calculation procedure for calculating a three-dimensional depth estimation image of an object from information obtained by photographing the object with a monocular camera; a three-dimensional imaging procedure for irradiating the object with radio waves using radar, measuring the reflected waves, and creating a three-dimensional imaging image of the object based on the obtained information; and an absolute dimension value assignment procedure for assigning absolute dimensions to the three-dimensional depth estimation image based on the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation procedure and the three-dimensional imaging image obtained by the three-dimensional imaging procedure. [Effects of the Invention]
[0009] According to the present invention, absolute dimensions are assigned to a three-dimensional depth estimation image using a three-dimensional imaging image. This eliminates the need to place gauge blocks or other objects that indicate dimensions in real space, and allows for easy assignment of dimensions through image processing or the like.
[0010] Furthermore, by calculating the edges of the planar image viewed from above or below in the Y-axis direction, or the edges of the XZ planar image at a predetermined position, as comparison shapes for the three-dimensional depth estimation image and the three-dimensional imaging image, and then calculating the similarity of each comparison shape to perform alignment, alignment can be easily performed and values can be assigned with high accuracy.
[0011] Furthermore, by creating an object occupancy grid map as a three-dimensional imaging image, the similarity between the three-dimensional depth estimation image and the three-dimensional imaging image can be easily calculated.
[0012] Furthermore, for unclassified pixels located between two discrete points that have been classified based on a three-dimensional imaging image, classifying their pixel position in the Z direction in the three-dimensional depth estimation image based on the dimensional difference between the two discrete points allows for accurate classification of absolute dimensions that exceed the resolution of distance obtained from radar.
[0013] Furthermore, by calculating the sound source intensity level from the information obtained by the microphone system, and combining the obtained sound source intensity level with the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means and the three-dimensional imaging image obtained by the three-dimensional imaging means, the location of the sound source can be identified and the distance to the sound source can be calculated, thereby allowing the acoustic power level to be calculated. Therefore, an index of the generated sound that is independent of the measurement location can be obtained from the obtained acoustic power level.
[0014] Furthermore, the device includes a moving member for moving an object to be moved, a support member that supports the moving member and is provided with a through-hole in the shape of a V or curved shape for guiding the movement of the moving member, a rotating member that is rotatably disposed relative to the support member and holds the moving member so that it can move in the longitudinal direction and has a through-hole that extends in the longitudinal direction of the moving member, a guide rod that is disposed relative to the moving member and passes through the through-hole and the guide hole, a mounting base that is rotatably disposed relative to the moving member and is provided with a projection that can move along a straight cut, and a connecting member that connects the projection and the guide rod, wherein the guide hole is positioned such that when the guide rod is positioned in the center, the center is located closer to the base end in the longitudinal direction of the moving member compared to both ends, so that by moving the guide rod along the guide hole, the rotating member rotates, and consequently the moving member rotates and moves toward the front end in the longitudinal direction, thereby moving the object to be moved. Furthermore, a projection connected to the guide rod moves along the notch, causing the mounting platform to rotate and adjust the orientation of the object being moved. This allows the object to be moved in parallel, making it easy to assign a price. [Brief explanation of the drawing]
[0015] [Figure 1] This is a block diagram showing the functional configuration of an object dimensioning device according to the first embodiment of the present invention. [Figure 2] This figure conceptually represents the three-dimensional depth estimation image of an object calculated by the three-dimensional depth estimation image calculation means shown in Figure 1. [Figure 3] Figure 1 is a conceptual diagram representing the state in which radio waves are irradiated onto an object from the radar shown. [Figure 4] Figure 1 shows the range of information from the radar used in the three-dimensional imaging means. [Figure 5] This diagram explains the occupied grid based on information obtained from the radar shown in Figure 1. [Figure 6] This figure shows the occupied grid map created by the three-dimensional imaging means shown in Figure 1. [Figure 7]This is a diagram for explaining the deviation correction when moving in the vertical direction in the deviation correction means shown in FIG. 1. [Figure 8] This is a diagram for explaining the deviation correction when moving in the horizontal direction in the deviation correction means shown in FIG. 1. [Figure 9] This is a diagram for explaining the ridge line of the planar image viewed from above in the Y-axis direction calculated from the three-dimensional depth estimation image in the alignment means shown in FIG. 1. [Figure 10] This is a diagram for explaining the ridge line of the planar image viewed from above in the Y-axis direction calculated from the three-dimensional imaging image in the alignment means shown in FIG. 1. [Figure 11] This is a diagram for explaining the value assignment for unassigned pixels existing between discrete points in the supplementary means shown in FIG. 1. [Figure 12] This is a diagram showing an example of the hardware configuration of the three-dimensional depth estimation image calculation means, the three-dimensional imaging means, and the absolute value dimension value assignment means. [Figure 13] This is a flowchart showing the procedure in the object dimension value assignment method. [Figure 14] This is a block diagram showing the functional configuration of the object dimension value assignment device according to the second embodiment of the present invention. [Figure 15] This is a diagram conceptually showing the sound source intensity level calculated by the sound source intensity level calculation means shown in FIG. 14 superimposed in front of the three-dimensional depth estimation image. [Figure 16] This is a diagram showing the configuration of the parallel movement jig of the object dimension value assignment device according to the third embodiment of the present invention. [Figure 17] This is a diagram showing a partial decomposition of the parallel movement jig shown in FIG. 16.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, embodiments of the present invention will be described in detail.
[0017] (First Embodiment) Figure 1 is a block diagram showing the functional configuration of an object dimensioning device 10 according to a first embodiment of the present invention. This object dimensioning device 10 includes a monocular camera 11 for photographing an object M, a three-dimensional depth estimation image calculation means 12 for calculating a three-dimensional depth estimation image of the object M from information obtained by the monocular camera 11, a radar 13 for irradiating the same object M photographed by the monocular camera 11 with radio waves and measuring the reflected waves, and a three-dimensional imaging means 14 for creating a three-dimensional imaging image of the object M based on information obtained by the radar 13.
[0018] The three-dimensional depth estimation image calculation means 12 is configured, for example, by a computer, and is configured to calculate a three-dimensional depth estimation image of the object M from information obtained by the monocular camera 11, for example, from images captured by the monocular camera, by executing a program. The three-dimensional depth estimation image of the object M can be calculated using known depth estimation techniques. For example, it can be calculated using Visual-SLAM (Simultaneous Localization and Mapping), or AI (artificial intelligence) such as CNNs (Convolutional Neural Networks) or FCNs (Fully Convolutional Networks).
[0019] Figure 2 shows a conceptual example of a three-dimensional depth-estimated image of an object M calculated by the three-dimensional depth-estimated image calculation means 12. Note that Figure 2 is an image where the object M is placed on the back wall W, and the object M is represented in gray. The obtained three-dimensional depth-estimated image is an image with relative values, without absolute dimensions.
[0020] The radar 13 emits radio waves toward the object M and measures the reflected waves to detect the position information (distance and angle) of the object M, as well as its relative velocity to the object M. A millimeter-wave radar using millimeter waves in the 24 GHz to 140 GHz wavelength range is preferred for the radar 13. This is because a wider bandwidth allows for measurements with higher distance resolution. It is preferable to position the radar 13 in the same direction as the monocular camera 11 relative to the object M. This facilitates alignment between the three-dimensional depth estimation image calculated from the information obtained by the monocular camera 11 and the three-dimensional imaging image created from the information obtained by the radar 13 in the absolute dimension value setting means 15 described later. Note that even if the radar 13 is not positioned in the same direction as the monocular camera 11, and the three-dimensional estimation image from the monocular camera 11 and the three-dimensional imaging image from the radar 13 are obtained separately, alignment of the two images is still possible.
[0021] Figure 3 conceptually illustrates the state in which radio waves are irradiated from radar 13 onto object M. In Figure 3, object M is represented in gray. The radio waves from radar 13 spread out in a fan shape and irradiate object M, and the shape of the irradiated area is elliptical. That is, the radio waves spread in a flattened cone shape in two orthogonal directions, with a wider spread in one direction and a narrower spread in the other. Figure 3 shows horizontal polarization. Depending on the irradiation location and object M, vertical polarization may also be used.
[0022] Preferably, the radar 13 uses a two-dimensional radar that can obtain azimuth distance, and detects while moving it vertically and horizontally, or while rotating it vertically and horizontally around its central axis. This is because moving or rotating it allows the radar 13 to use fewer antennas, thus enabling miniaturization and lower costs. The movement or rotation of the radar 13 may be done mechanically or manually. Manual movement is preferable because it allows for miniaturization of the device and application to confined spaces. Furthermore, moving it manually is preferable because it allows for measurement of areas that would be obscured by shadows during rotation. For example, if the radar 13 and monocular camera 11 are fixed to a support rod (not shown), and a mechanism is provided to slide the support rod vertically and horizontally, it is preferable that the radar 13 can be moved mechanically and still be applied to confined spaces.
[0023] The three-dimensional imaging means 14 is configured, for example, by a computer, and is configured to create a three-dimensional imaging image of the object M based on information obtained by the radar 13 by executing a program. The three-dimensional imaging means 14 preferably uses information from the radar 13 in a range narrower than 0.25 rad from the center (i.e., a range of 0.5 rad from the center). Figure 4 shows the range of information from the radar 13 used in the three-dimensional imaging means 14. Since radio waves spread out in a fan shape from the radar 13, they are in polar coordinates and require conversion to Cartesian coordinates. However, if the range is narrower than 0.25 rad from the center, it can be processed by considering it as a Cartesian coordinate value without coordinate conversion. In the direction where the radio waves spread out widely, information in a range of 120° is obtained, so information in a range narrower than 0.25 rad from the center is extracted and used.
[0024] Furthermore, it is preferable that the three-dimensional imaging means 14 includes, for example, an occupied grid map creation means 14a that creates an occupied grid map based on information obtained by the radar 13 as a three-dimensional imaging image. The occupied grid map divides the three-dimensional space into a plurality of voxels of a rectangular parallelepiped or cube, and displays the voxels in which the object exists as an occupied grid N. The information obtained from the radar 13 can be considered as an occupied grid N consisting of a collection of rectangular parallelepipeds or cubes, where the distance resolution, azimuth resolution, and the width of the elevation / depression angle (antenna half-power angle) are the three sides emanating from one vertex, as shown in Figure 5. Therefore, a three-dimensional occupied grid map can be obtained by moving the radar 13 in the vertical and horizontal directions, or by rotating it in the vertical and horizontal directions based on the central axis, and connecting the occupied grid N. Figure 6 shows an example of an occupied grid map created by the three-dimensional imaging means 14. Figure 6(A) shows the occupied grid N together with grid lines, and the occupied grid N is shown in gray. Figure 6(B) shows only the occupied grid N.
[0025] Furthermore, when the radar 13 is moved by hand, vertical and horizontal displacement (fluctuation) occurs, so it is preferable to have displacement correction means 14b to correct the displacement. The displacement correction means 14b can be configured to perform displacement correction in the following way, for example.
[0026] Figure 7 illustrates the shift correction when the radar 13 is moved vertically. In Figure 7, the dashed line is the virtual vertical line of the Y axis, and the black stars are the scatter points of the radar 13. Figure 7 shows the case when the radar 13 is moved vertically from top to bottom. For shifts in the front-to-back direction (i.e., the distance direction), for example, if the distance to the same target changes after the radar 13 is moved, the entire system is corrected so that the same target is at the same distance at t-1 (seconds). Similarly, for shifts in the lateral direction (i.e., the azimuth direction), for example, if the azimuth position of the same target changes after the radar 13 is moved, the entire system is corrected so that the same target is at the same azimuth position at t-1 (seconds). In shift correction, the same target is defined as a target obtained at approximately the same distance and azimuth as at t-1 (seconds). t is the time of current sampling (measurement). Multiple targets are corrected for shift according to their distance.
[0027] Furthermore, it is preferable to consider the elevation resolution (full width at half maximum) of the radar 13 (defined as the reference region). Due to the elevation resolution of the radar 13, a new point cannot be obtained unless the radar 13 is moved by more than half width. In the case of manual operation, it is assumed that the distance coordinates will not be obtained on the straight line of the Y axis (referred to as the virtual vertical line) due to fluctuations. However, since the azimuth angle of distant targets does not change easily, the deviation from the virtual vertical line is small, while the deviation from the virtual vertical line is large for nearby targets because the azimuth angle changes easily. Therefore, correction in the Y-axis direction is performed based on the degree of positional deviation from the virtual vertical line for distant and nearby targets.
[0028] Figure 8 illustrates the correction of misalignment when the radar 13 is moved horizontally. In Figure 8, the dashed line is the virtual horizontal line of the X-axis, and the black stars are the scatter points of the radar 13. Figure 8 shows the case when the radar 13 is moved horizontally from right to left across the paper. When moving horizontally, it is preferable to correct and connect the occupied grids using known matching techniques, such as the ICP method, NDT (Normal Distribution Transform) method, likelihood function, cosine similarity, or Pose Graph Optimization (PGO), so that the occupied grids match. This connects parts that appear to be the same target (parts with a strong correlation in the reference region or measurement window). The measurement window refers to the cell range that is the target of matching using known matching techniques. In this case, if the distances are different, it is preferable to adjust the entire distance to t-1 (seconds). As with vertical movement, the horizontal movement distance can be obtained from the Doppler amount.
[0029] Furthermore, the correction of displacement when moving in the lateral direction may be configured to be performed by other methods. For example, using a similar approach to camera attitude estimation such as Visual-SLAM, even if the radar 13 is not necessarily moving in parallel, the object is a stationary, identical rigid body, so the position and attitude of the radar 13 may be estimated (ego-motion, self-localization) based on the reflection intensity. Alternatively, the monocular camera 11 may be moved horizontally, vertically, and forward and backward using pose graph optimization (PGO), and the coordinates of the radar 13 may be calculated based on the coordinates of the monocular camera 11 at that time.
[0030] The object dimension value assignment device 10 further includes an absolute dimension value assignment means 15 that assigns absolute dimensions to the three-dimensional depth estimation image based on the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means 12 and the three-dimensional imaging image obtained by the three-dimensional imaging means 14. The absolute dimension value assignment means 15 is configured, for example, by a computer and is configured to assign absolute dimensions to the three-dimensional depth estimation image by executing a program.
[0031] Preferably, the absolute dimension value assignment means 15 includes, for example, an alignment means 15a that calculates the similarity between a three-dimensional depth estimation image and a three-dimensional imaging image and performs alignment, and a value assignment means 15b that assigns absolute dimension values to the three-dimensional depth estimation image aligned by the alignment means 15a based on the three-dimensional imaging image.
[0032] Preferably, the alignment means 15a is configured to calculate, for example, the edges of a planar image viewed from above or below in the Y-axis direction, or the edges of an XZ planar image at a predetermined position, as comparison shapes for the three-dimensional depth estimation image and the three-dimensional imaging image, and to perform alignment by calculating the similarity of each comparison shape. This is because the similarity can be easily calculated by extracting feature parts as comparison shapes. A planar image viewed from above or below in the Y-axis direction means a planar image of a so-called plan view viewed from above or below in the Y-axis direction, and an XZ planar image at a predetermined position means a planar image of a cross-section at a predetermined position on the Y-axis.
[0033] Figure 9 shows a planar image of the three-dimensional depth estimation image viewed from above along the Y-axis. In Figure 9, the object M is represented with a textured surface, and its edges are represented by thick dashed lines. Figure 10 shows a planar image of the three-dimensional imaging image viewed from above along the Y-axis. Figure 10(A) shows the occupied grid N together with the grid lines, with the occupied grid N represented in gray. Figure 10(B) shows the occupied grid N extracted and represented together with its edges, with the occupied grid N represented with a textured surface, and its edges represented by thick dashed lines.
[0034] In the three-dimensional depth estimation image, the ridges are, for example, formed by sequentially connecting the boundary lines of objects M in the X-axis direction on the side of the monocular camera 11. For example, if there are multiple objects M and they overlap as seen from the monocular camera 11, it is preferable to combine each object and draw a ridge line as one. If the multiple objects M are located far apart, lines may be drawn between each object M to connect them, or no connecting lines may be drawn. Figure 9 shows the case where lines are drawn between objects M located far apart. In the three-dimensional imaging image, the ridges are, for example, formed by sequentially connecting the boundary lines of each occupied grid N in the X-axis direction on the side of the radar 13. For example, if there are multiple occupied grids N and each occupied grid N is adjacent to another, it is preferable to combine each occupied grid N into one and draw a ridge line. If the occupied grids N overlap with a gap in the Z direction as viewed from the radar 13, a ridge line is drawn to the occupied grid N closest to the radar 13. If multiple occupied grids N are adjacent in the X direction and separated in the Z direction, lines may be drawn between each occupied grid N to connect them, or no lines may be drawn to connect them. Figure 10 shows the case where lines are drawn between occupied grids N that are separated in the Z direction.
[0035] The alignment means 15a is configured to calculate the similarity of each edge of the three-dimensional depth estimation image and the three-dimensional imaging image shown in Figures 9 and 10, and to perform alignment. The similarity can be calculated using known matching techniques, such as the ICP method, NDT method, likelihood function, cosine similarity, or Pose Graph Optimization (PGO). The alignment means 15a is configured to determine that the images match when the similarity is higher than a predetermined value, and to align the three-dimensional depth estimation image and the three-dimensional imaging image.
[0036] The value assignment means 15b is configured to assign values to the three-dimensional depth estimation image, for example, to the absolute dimensions of the three-dimensional imaging image, for locations where the three-dimensional depth estimation image and the three-dimensional imaging image coincide due to alignment.
[0037] The absolute dimension valuation means 15 preferably also includes a supplementary means 15c for valuing unvalued pixels that exist between two discrete points valued based on a three-dimensional imaging image in a three-dimensional depth estimation image. Valuation of a three-dimensional depth estimation image cannot be obtained beyond the distance resolution of the radar 13, and points valued by the three-dimensional imaging image become discrete points with respect to the distance resolution of the radar 13. Therefore, there may be pixels with unvalued Z values between two discrete points that have been valued.
[0038] Figure 11 illustrates the assignment of values to unassigned pixels located between discrete points in the acquisition means 15c. Figure 11 shows an object M with an inclined surface. Since the radar 13 often cannot obtain reflections and measure distances on inclined surfaces, the inclined surface portion is empty in the occupied grid map. That is, in Figure 11(A), occupied grids N can be obtained for the surfaces M1 and M2 shown as textured, and therefore distances can be obtained, but occupied grids N cannot be obtained for the inclined surface portion, and therefore distances cannot be obtained. Assigning values to pixels with unassigned Z values can be done based on the value of a single point assigned based on the three-dimensional imaging image, but it is preferable to calculate them based on the dimensional difference between two discrete points assigned based on the three-dimensional imaging image. This is because, due to distortion of the camera lens, the pixel values in the Z direction may not be linear, so accurate assignment can be achieved by using the distance difference between two points. In other words, it is preferable that the supplementation means 15c is configured to assign values to the pixel positions in the Z direction in the three-dimensional depth estimation image for unassigned pixels existing between two discrete points, based on the dimensional difference between the two discrete points for which distances have been obtained.
[0039] Let's explain this in more detail using Figure 11. For example, if the distance (Z value) of surface M1, shown as a textured surface by radar 13, is 80 cm, and the distance (Z value) of surface M2, also shown as a textured surface, is 100 cm, then the dimensional difference between the two discrete points is 20 cm. For the portion of the inclined surface where distance measurement was not possible, relative values regarding the positional relationship with surfaces M1 and M2 can be obtained from the three-dimensional depth estimation image, as indicated by the arrows in Figure 11(B). Therefore, the pixel position in the Z direction can be assigned according to these relative values.
[0040] Figure 12 shows an example of the hardware configuration of the three-dimensional depth estimation image calculation means 12, the three-dimensional imaging means 14, and the absolute value dimension value setting means 15. The three-dimensional depth estimation image calculation means 12, the three-dimensional imaging means 14, and the absolute value dimension value setting means 15 include, for example, a CPU (Center Processing Unit) 16A, a ROM (Read Only Memory) 16B, a RAM (Random Access Memory) 16C, an HDD (Hard Disk Drive) 16D, and an operation interface (operation I / F) 16E. The CPU 16A executes various processes according to various programs recorded in the ROM 16B or various programs loaded from the HDD 16D to the RAM 16C. The RAM 16C also appropriately stores data necessary for the CPU 16A to execute various processes. The HDD 16D stores various data.
[0041] The object dimensioning device 10 can, for example, assign dimensions to an object M as follows. Figure 13 is a flowchart showing the procedure for assigning dimensions to an object M.
[0042] First, for example, the object M is photographed using a monocular camera 11, and the three-dimensional depth estimation image calculation means 12 calculates a three-dimensional depth estimation image of the object M from the information obtained by the monocular camera 11 (Step S110; three-dimensional depth estimation image calculation procedure).
[0043] Furthermore, for example, the radar 13 is used to irradiate the same object M with radio waves and measure the reflected waves, and a three-dimensional imaging image is created from the obtained information by the three-dimensional imaging means 14 (step S120; three-dimensional imaging procedure). In this case, it is preferable to measure while moving the radar 13 in the vertical and horizontal directions, or while rotating it in the vertical and horizontal directions based on the central axis. In creating the three-dimensional imaging image, for example, it is preferable to create an occupied grid map as the three-dimensional imaging image by the occupied grid map creation means 14a. Also, for example, when moving the radar 13 by hand, it is preferable to correct the displacement by the displacement correction means 14b.
[0044] Next, for example, the absolute dimension value assignment means 15 assigns absolute dimensions to the three-dimensional depth estimation image based on the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation procedure (step S110) and the three-dimensional imaging image obtained by the three-dimensional imaging procedure (step S120) (step S130; absolute dimension value assignment procedure).
[0045] In the absolute dimension value assignment procedure (step S130), first, for example, the alignment means 15a calculates the edges of the planar image viewed from above or below in the Y-axis direction, or the edges of the XZ planar image at a predetermined position, as comparison shapes for the three-dimensional depth estimation image and the three-dimensional imaging image, respectively (step S131; comparison shape calculation procedure). Next, for example, the alignment means 15a calculates the similarity of the calculated comparison shapes and aligns the three-dimensional depth estimation image and the three-dimensional imaging image (step S132; alignment procedure).
[0046] Next, for example, the absolute dimensions are assigned based on the three-dimensional imaging image aligned to the three-dimensional depth estimation image by the assignment procedure 15b (step S133; assignment procedure). Then, for example, it is preferable to assign values to unassigned pixels in the three-dimensional depth estimation image that exist between two discrete points assigned based on the three-dimensional imaging image by the supplementation means 15c (step S134; supplementary procedure). Specifically, for example, based on the dimensional difference between two discrete points for which distances have been obtained, values are assigned to the pixel positions in the Z direction in the three-dimensional depth estimation image for unassigned pixels that exist between the two discrete points.
[0047] As described above, according to this embodiment, absolute dimensions are assigned to the three-dimensional depth estimation image using a three-dimensional imaging image. Therefore, there is no need to place block gauges or other objects that indicate dimensions in real space, and the dimensions can be easily assigned by image processing or the like.
[0048] Furthermore, by calculating the edges of the planar image viewed from above or below in the Y-axis direction, or the edges of the XZ planar image at a predetermined position, as comparison shapes for the three-dimensional depth estimation image and the three-dimensional imaging image, and then calculating the similarity of each comparison shape to perform alignment, alignment can be easily performed and values can be assigned with high accuracy.
[0049] Furthermore, by creating an occupied grid map of the object M as a three-dimensional imaging image, the similarity between the three-dimensional depth estimation image and the three-dimensional imaging image can be easily calculated.
[0050] In addition, for unclassified pixels that exist between two discrete points classified based on a three-dimensional imaging image, if the pixel position in the Z direction in the three-dimensional depth estimation image is classified based on the dimensional difference between the two discrete points, then the absolute dimensions exceeding the resolution of the distance obtained from the radar 13 can be accurately classified.
[0051] (Second Embodiment) Figure 14 is a block diagram showing the functional configuration of an object dimension marking device 20 according to a second embodiment of the present invention. This object dimension marking device 20 has the same configuration as the object dimension marking device 10 of the first embodiment, except that, in addition to the components described in the first embodiment, it includes a microphone system 21 for detecting sound from an object M, a sound source intensity level calculation means 22 for calculating the sound source intensity level from the information obtained by the microphone system 21, and a sound power level calculation means 23 for calculating the sound power level of the sound source at a desired position based on the sound source intensity level obtained by the sound source intensity level calculation means 22. Therefore, the same reference numerals are used for components that are the same as in the first embodiment, and their detailed descriptions are omitted.
[0052] The microphone system 21 integrates multiple microphones for detecting sound from an object M. The multiple microphones are arranged and integrated in any shape, such as a spherical, planar, hemispherical, or spiral shape. The microphone system 21 is preferably arranged in conjunction with, for example, a monocular camera 11 or a radar 13.
[0053] For example, by integrating the microphone system 21 and the monocular camera 11, when moving the monocular camera 11 to photograph an object M in order to calculate a three-dimensional depth estimation image, the sound source intensity level can be calculated using the sound received by the microphone system 21 at any position. By storing the camera coordinates of the calculation position, the sound source intensity level can be superimposed on that coordinate position when the three-dimensional depth estimation image is displayed. The camera coordinates can be calculated using existing methods such as SVO (Fast Semi-Direct Monocular Visual Odometry) for pose estimation processing of the monocular camera 11.
[0054] Furthermore, for example, by integrating the microphone system 21 and the radar 13, when creating a three-dimensional imaging image, the sound source intensity level can be calculated at the location of the radar 13's occupied grid map, and the radar coordinates can be stored as the sound source intensity level calculation location. This allows the sound source intensity level to be superimposed on the location of the occupied grid map when it is displayed.
[0055] The sound source intensity level calculation means 22 calculates the sound source intensity level in the direction of sound direction and elevation angle without moving the microphone system 21. The sound source intensity level is the intensity distribution of the combined sound pressure level of multiple microphones, calculated by processing the sound received by multiple microphones. Figure 15 shows a conceptual example of the sound source intensity level calculated by the sound source intensity level calculation means 22. In Figure 15, the sound source intensity level is superimposed on the three-dimensional depth estimation image of the object M calculated by the three-dimensional depth estimation image calculation means 12. In Figure 15, the sound source intensity level is shown with finer textures depending on the intensity, and the three-dimensional depth estimation image of the object M is shown with gray. Note that Figure 15 shows the case where the microphone system 21 and the monocular camera 11 are integrated and the sound source intensity level is superimposed on the camera coordinate position where the sound source intensity level was calculated. The sound source intensity level may also be calculated using multiple camera locations or multiple grid map locations occupied by the radar 13.
[0056] The acoustic power level calculation means 23 is configured, for example, by a computer, and is configured to calculate the acoustic power level of a sound source at a desired location based on the sound source intensity level obtained by the sound source intensity level calculation means 22, the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means 12, and the three-dimensional imaging image obtained by the three-dimensional imaging means 14, by executing a program. Specifically, for example, by superimposing the calculated sound source intensity level with the three-dimensional depth estimation image or the three-dimensional imaging image, the position of the sound source in the three-dimensional depth estimation image, for which the absolute value dimension value value means 15 has assigned an absolute value dimension, can be identified, the distance between the microphone system 21 and the sound source can be calculated, and the acoustic power level can be calculated based on acoustic theory from the determined distance to the sound source and the sound pressure level of that sound source.
[0057] This allows the acoustic power level of the sound obtained from the object M to be acquired, providing an index of the generated sound that is independent of the measurement position, and which can be used for comparative evaluation of noise, etc. Furthermore, to determine whether the obtained sound is abnormal or not, for example, the Mahalanobis-Taguchi method or a Variational Autoencoder (VAE) can be used. In addition, when the acoustic power level calculation means 23 calculates the sound source as a point source, line source, or surface source, the length of the side may be calculated from the absolute value dimensioning means 15 for line sources or surface sources.
[0058] In this object dimensioning device 20, the dimensions of an object M can be assigned in the same manner as in the first embodiment, and the acoustic power level of a sound source emanating from the object M can also be calculated. Specifically, for example, first, the microphone system 21 detects sound from the object M, and the sound source intensity level calculation means 22 calculates the sound source intensity level from the information obtained by the microphone system 21. Next, for example, the acoustic power level calculation means 23 superimposes the sound source intensity level with a three-dimensional depth estimation image or a three-dimensional imaging image to identify the position of the sound source in the three-dimensional depth estimation image on which absolute dimensions have been assigned, calculates the distance between the microphone system 21 and the sound source, and calculates the acoustic power level from the obtained distance to the sound source and the sound pressure level of that sound source.
[0059] Thus, according to this embodiment, in addition to the effects described in the first embodiment, the sound power level of the sound source can be calculated, and an indicator of the generated sound that is independent of the measurement location can be obtained from the obtained sound power level.
[0060] (Third embodiment) The object dimension marking device according to this embodiment has the same configuration as the first or second embodiment, except that, in addition to the components described in the first or second embodiment, it is equipped with a translation jig 30 for translating at least one of the monocular camera 11 and radar 13. Therefore, the same reference numerals are used for components identical to those in the first embodiment, and their detailed descriptions are omitted.
[0061] Figure 16 shows the configuration of the translation jig 30. Figure 17 shows a disassembled view of a part of the translation jig 30. The translation jig 30 is used to translate at least one of the monocular camera 11 and the radar 13 as the object to be moved. Using the translation jig 30, the monocular camera 11 and the radar 13 can be easily moved, and the pricing work can be easily performed. In addition, the monocular camera 11 or the radar 13 may have a microphone system 21 integrated into it, as described in the second embodiment.
[0062] The parallel movement jig 30 comprises a moving member 31 for moving an object to be moved, a support member 32 for supporting the moving member 31, a rotating member 33 rotatably disposed with respect to the support member 32 and for movably holding the moving member 31, a rod-shaped guide rod 34 disposed with respect to the moving member 31 and for guiding the movement of the moving member 31, a mounting base 35 rotatably disposed with respect to the moving member 31 and for placing the object to be moved, and a connecting member 36 connecting the mounting base 35 and the guide rod 34.
[0063] The movable member 31 has, for example, a longitudinal shape extending in one direction, and supports the object to be moved at its longitudinal end. The support member 32 is, for example, a flat plate, and the movable member 31 is positioned on one side of the support member 32 and configured to move along that side of the support member 32. The support member 32 is positioned, for example, so that one side is horizontal. The support member 32 is provided with a through-hole V-shaped or curved guide hole 32A for guiding the movement of the movable member 31 in accordance with the movement of the movable member 31. A guide rod 34 passes through the guide hole 32A. The guide hole 32A is positioned such that when the guide rod 34 is positioned in the center of the guide hole 32A, the center is located closer to the base end in the longitudinal direction of the movable member 31 compared to both ends. This is because when the guide rod 34 is moved along the guide hole 32A, the movable member 31 rotates and moves toward the longitudinal end, causing the end of the movable member 31 to move linearly. Furthermore, it is preferable that the support member 32 is provided with a notched plate in the guide hole 32A that can catch on to indicate movement in increments of dx = λ / 4. λ is the wavelength of the radar 13 used. It is preferable that the notched plate be configured to be replaceable depending on the frequency.
[0064] The rotating member 33 changes the extending direction of the movable member 31 by rotating relative to the support member 32. The rotating member 33 is positioned, for example, between the movable member 31 and the support member 32, and holds the base end of the movable member 31 in a longitudinally movable manner. Specifically, for example, U-shaped recesses 33A are formed on both sides of the rotating member 33 in the width direction, surrounding the sides of the movable member 31, and these recesses 33A hold the movable member 31 in a longitudinally movable manner. The rotating member 33 is also provided with a through hole 33B that extends in the longitudinal direction of the movable member 31. A guide rod 34 passes through the through hole 33B. The guide rod 34 is positioned on the base end side in the longitudinal direction relative to the movable member 31, and protrudes from the movable member 31 through the through hole 33B and the guide hole 32A, onto the support member 32 on the side opposite to the movable member 31. In other words, by moving the guide rod 34 along the guide hole 32A, the guide rod 34 moves along the through hole 33B of the rotating member 33, causing the rotating member 33 to rotate. Consequently, the extending direction of the moving member 31 changes, and the moving member 31 moves toward the end in the longitudinal direction.
[0065] The mounting base 35 is, for example, positioned on the longitudinal end side of the movable member 31. The mounting base 35 has, for example, a linear notch 35A on the side of the guide rod 34. The notch 35A may or may not be a through-hole. The notch 35A is provided with a projection 35B that can move along the notch 35A. The projection 35B is connected to the guide rod 34 and a connecting member 36 such as a non-extendable wire. As a result, when the guide rod 34 is moved along the guide hole 32A, the end side of the movable member 31 moves linearly, and the projection 35B moves along the notch 35A, causing the mounting base 35 to rotate and allowing the object to be moved placed on the mounting base 35 to move in parallel.
[0066] The parallel movement jig 30 preferably also includes a vertical movement means 37 for moving the support member 32 in the vertical direction. The vertical movement means 37 is configured, for example, to move the rack 37B in the vertical direction by rotating an adjustment screw 37A. The adjustment screw 37A has a scale indicating movement of dy = λ / 4, and it is preferable that the scale is configured to be interchangeable according to frequency, where λ is the wavelength of the radar 13 used.
[0067] The parallel movement jig 30 may further include an extension means 38 that connects the support member 32 and the vertical movement means 37, extending the distance between the mounting base 35 and the vertical movement means 37. This is because the monocular camera 11 and radar 13 can be positioned closer to the object M when the object M is in a narrow space. The extension means 38 includes, for example, an extension support member 38A that supports the support member 32 and connects to the vertical movement means 37, an operating lever 38B connected to a guide rod 34, and a V-shaped or curved operating hole 38C cut into the extension support member 38A corresponding to the guide hole 32A. The extension support member 38A is preferably hollow cylindrical, for example, because it makes it easier to move the operating lever 38B left and right. The operating hole 38C may or may not be a through hole. The extension means 38 may also be provided with an operating hole 38C and a notched plate that can be used to indicate movement in increments of dx = λ / 4. λ is the wavelength of the radar 13 used. In this case, it is preferable that the notch plate be configured to be replaceable depending on the frequency.
[0068] Furthermore, the parallel movement jig 30 may be operated not only manually but also electrically.
[0069] The parallel movement jig 30 can be operated as follows. For example, when the guide rod 34 is moved left and right along the guide hole 32A, the guide rod 34 moves along the through hole 33B of the rotating member 33, causing the rotating member 33 to rotate. Consequently, the extension direction of the moving member 31 changes, and the moving member 31 moves toward the longitudinal end. As a result, the end of the moving member 31 moves linearly left and right. Also, as the guide rod 34 moves, the projection 35B moves along the notch 35A, and the mounting base 35 rotates accordingly. Even if the extension direction of the moving member 31 changes, the orientation of the mounting base 35 remains constant, and the object to be moved placed on the mounting base 35 moves in parallel. Furthermore, when the adjustment screw 37A of the vertical movement means 37 is rotated, the support member 32 moves vertically, and the object to be moved placed on the mounting base 35 moves vertically.
[0070] Thus, according to this embodiment, in addition to the effects described in the first and second embodiments, the addition of a translation jig 30 allows for easy translation of at least one of the monocular camera 11 and radar 13 by placing it on the mounting base 35, and thus enables easy price setting.
[0071] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments and can be modified in various ways. For example, in the above embodiment, when assigning absolute dimensions to the three-dimensional depth estimation image using the absolute dimension value assignment means 15, the case in which the similarity between the three-dimensional depth estimation image and the three-dimensional imaging image is calculated and alignment is performed was described. However, alignment may also be performed by superimposing the three-dimensional depth estimation image and the three-dimensional imaging image. In that case, scaling is performed to match the size of the three-dimensional depth estimation image and the three-dimensional imaging image. Furthermore, the resolution may be increased by smoothing the occupied grid map, or the three-dimensional depth estimation image may be processed into blocks.
[0072] Furthermore, although each component has been described in detail in the above embodiment, the specific structure and shape of each component may differ. Moreover, it is not necessary to have all of the above-mentioned components, and other components may also be included. [Explanation of Symbols]
[0073] 10, 20... Object dimension value setting device, 11... Monocular camera, 12... Three-dimensional depth estimation image calculation means, 13... Radar, 14... Three-dimensional imaging means, 13a... Occupancy grid map creation means, 14b... Shift correction means, 15... Absolute value dimension value setting means, 15a... Alignment means, 15b... Value setting means, 15c... Supplementation means, 16A... CPU, 16B... ROM, 16C... RAM, 16D... HDD, 16E... Operation interface, M... Object, N... Occupancy grid, 21 ...microphone system, 22...sound source intensity level calculation means, 23...acoustic power level calculation means, 30...parallel movement jig, 31...moving member, 32...support member, 32A...guide hole, 33...rotating member, 33A...recess, 33B...through hole, 34...guide rod, 35...mounting base, 35A...notch, 35B...projection, 36...connecting member, 37...vertical movement means, 37A...adjustment screw, 37B...rack, 38...extension means, 38A...extension support member, 38B...operating lever, 38C...operating hole 38C
Claims
1. A monocular camera for photographing objects, A three-dimensional depth estimation image calculation means calculates a three-dimensional depth estimation image of the object from information obtained by the monocular camera, A radar that irradiates the aforementioned object with radio waves and measures the reflected waves, A three-dimensional imaging means that creates a three-dimensional imaging image of the object based on information obtained by the radar, The system includes an absolute dimension value assignment means for assigning absolute dimensions to the three-dimensional depth estimation image based on the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means and the three-dimensional imaging image obtained by the three-dimensional imaging means, The absolute value dimensioning means is, A alignment means that calculates the similarity between the three-dimensional depth estimation image and the three-dimensional imaging image and aligns them, The system includes a value assignment means for assigning absolute dimensions to the three-dimensional depth estimation image, which has been aligned by the alignment means, based on the three-dimensional imaging image. A device for determining the dimensions of an object, characterized by the following features.
2. The object dimensioning device according to claim 1, characterized in that the alignment means calculates the edges of the planar image viewed from above or below in the Y-axis direction, or the edges of the X-Z planar image at a predetermined position, as comparison shapes for the three-dimensional depth estimation image and the three-dimensional imaging image, respectively, and aligns the objects by calculating the similarity of each comparison shape.
3. The object dimensioning device according to claim 1 or 2, characterized in that the three-dimensional imaging means has an occupied grid map creation means that, as the three-dimensional imaging image, creates an occupied grid map in which the voxels in which the object exists are displayed as an occupied grid in a three-dimensional space divided into a plurality of voxels based on information obtained by the radar.
4. The object dimension valuation device according to any one of claims 1 to 3, characterized in that the absolute dimension valuation means has supplementary means for valuing the pixel position in the Z direction in the three-dimensional depth estimation image for unvalued pixels that exist between two discrete points valued based on the three-dimensional imaging image, based on the dimensional difference between the two discrete points.
5. A microphone system comprising multiple microphones integrated to detect sound from the aforementioned object, A sound source intensity level calculation means that calculates the sound source intensity level from the information obtained by the microphone system, A sound power level calculation means calculates the sound power level of a sound source at a desired location based on the sound source intensity level obtained by the sound source intensity level calculation means, the three-dimensional depth estimation image obtained by the three-dimensional depth estimation image calculation means, and the three-dimensional imaging image obtained by the three-dimensional imaging means. An object dimension marking device according to any one of claims 1 to 4, characterized by comprising:
6. A moving member having an elongated shape, for moving at least one of the monocular camera and the radar as the object to be moved, A support member that supports the movable member and is provided with a through-hole that is V-shaped or curved in shape for guiding the movement of the movable member, A rotating member is rotatably disposed with respect to the support member, holds the longitudinal base end of the movable member so as to be movable in the longitudinal direction, and has a through hole extending in the longitudinal direction of the movable member, thereby changing the direction of extension of the movable member by rotating with respect to the support member, A rod-shaped guide rod is provided on the longitudinal base end side of the movable member, and extends from the movable member through the through hole and the guide hole, protruding from the support member on the side opposite to the movable member, A mounting base is rotatably disposed on the longitudinal end of the moving member, on which the object to be moved is placed, and which has a linear notch and a projection that is movable along this notch, The system includes a connecting member that connects the projection and the guide rod, When the guide rod is positioned in the center, the central part of the guide hole is located closer to the base end in the longitudinal direction of the moving member compared to both ends. The object dimension marking device according to any one of claims 1 to 5.
Citation Information
Patent Citations
JP1974042106A
Three-dimensional radar device
JP2000111635A
Three-dimensional data-processing device
JP2001012922A
Depth estimation device, dimension estimation device, depth estimation method, dimension estimation method, and program
JP2018132477A
Body contour information analysis system
JP2018163031A