Object extraction system, object extraction method and object extraction program

JP2024141898A5Active Publication Date: 2025-06-20HITACHI SOFTWARE ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023053768
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-06-20
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Existing systems struggle to accurately extract the plane of reinforcing bars in construction sites due to misalignment between the camera optical axis and the normal direction of the plane, leading to errors in measurement, especially when the camera cannot be positioned perpendicular to the reinforcing bars.

Method used

The system divides an image into regions, counts pixel numbers, identifies local maxima, extracts pixels within a predetermined distance range, calculates planes for each region, groups these regions based on plane relationships, and selects the appropriate group to determine the extraction target plane accurately.

Benefits of technology

Enables accurate determination of the plane representing the extraction target even when the camera is not perpendicular to the reinforcing bars, ensuring stable and user-independent extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To accurately determine a plane that represents an extraction object.SOLUTION: An object extraction system that extracts information on an object from an image, includes: an arithmetic operation unit which executes arithmetic processing; and a storage unit which can be accessed by the arithmetic operation unit. The arithmetic operation unit acquires a distance image including distance information measured for each pixel, divides the distance image into multiple regions, counts the number of pixels for each distance in the divided area, extracts the pixels within a prescribed distance range from the maximum value of the shortest distance from an imaging point among the maximum values of the counted number of pixels, calculates the first plane representing the extracted pixels for each divided area, groups the divided areas using the relation between the first planes calculated for each area, selects one of the groups generated by the grouping, and calculates the second plane representing the extracted pixels in the area of the selected group.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an object extraction system that extracts a plane on which an object appears in a range image. [Background technology]

[0002] In the past, there have been problems at construction sites where the number, diameter, and spacing of rebars were different from the design, leading to stricter inspection of rebars and a greater burden on on-site workers. To reduce this burden, three-dimensional data of rebars is acquired using cameras and other sensors. The point cloud represented by the acquired three-dimensional coordinates contains information other than the rebar to be measured, so it is necessary to extract the point cloud representing the target rebar.

[0003] The following prior art is included as background technology in this technical field. Patent Document 1 (JP Patent Publication 2015-1146 A) describes a reinforcing bar inspection support device that detects images of multiple adjacent nodes of a reinforcing bar to be inspected from an image obtained by photographing the reinforcing bar to be inspected with a photographing device, derives the distance between adjacent nodes based on the images of the multiple detected nodes, and identifies the diameter of the reinforcing bar to be inspected by reading the diameter of the reinforcing bar corresponding to the derived distance from a secondary storage unit that previously stores the distance between the adjacent nodes and the corresponding diameter of the reinforcing bar in association with each other for multiple predetermined types of diameter of reinforcing bars. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2015-1146 A Summary of the Invention [Problem to be solved by the invention]

[0005] In the above-mentioned known technology, in order to extract point cloud data on the plane of the rebar position, a histogram of the distance to the object and the number of pixels is used, and the point cloud data is rotated three-dimensionally so that the camera optical axis and the normal direction of the plane are parallel. In other words, in order to extract the plane of the rebar position using the histogram, it is assumed that the camera optical axis and the plane are perpendicular. However, it is difficult to accurately position the camera so that the optical axis is perpendicular to the plane of the rebar position, and errors occur due to a misalignment between the optical axis and the normal of the plane. In addition, depending on the on-site environment, it may not be possible to position the camera so that the optical axis is perpendicular to the plane of the rebar position.

[0006] Therefore, what is needed is a system that can accurately determine a plane that represents an extracted object without the need to position the camera perpendicular to the extracted object. [Means for solving the problem]

[0007] A representative example of the invention disclosed in the present application is as follows: That is, an object extraction system for extracting information about an object from an image includes a calculation unit that executes calculation processing and a storage unit accessible by the calculation unit, and the calculation unit acquires a distance image including distance information measured for each pixel, divides the distance image into a plurality of regions, counts the number of pixels for each distance in the divided regions, extracts pixels within a predetermined distance range from the maximum value of the counted number of pixels at the shortest distance from the shooting point, calculates a first plane representing the extracted pixels for each divided region, groups the divided regions using the relationship between the first planes calculated for each region, selects one of the groups generated by the grouping, and calculates a second plane representing the extracted pixels in the region of the selected group. Effect of the Invention

[0008] According to one aspect of the present invention, a plane representing an extraction target can be accurately determined. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]

[0009] [Figure 1] 1 is a block diagram showing a configuration of an object extraction system according to an embodiment of the present invention. [Diagram 2] 3 is a diagram showing an example of the data structure of a distance image handled in the present embodiment; FIG. [Diagram 3] 1 is a flowchart of a conventional object extraction process. [Figure 4] FIG. 13 is a diagram showing the designation of three points on a rebar in a conventional image. [Diagram 5] FIG. 1 is a top view of a conventional reinforcing bar. [Figure 6] FIG. 13 is a diagram showing an extracted area estimated to be a conventional reinforcing bar. [Figure 7] 11 is a flowchart of an object extraction process according to the present embodiment. [Figure 8] FIG. 2 is a diagram showing the definition of a region on an image in this embodiment. [Figure 9] FIG. 13 is a diagram showing an example of a histogram of the number of pixels and distance within a region in the embodiment. [Figure 10] FIG. 2 is a diagram showing an example of grouped regions in the present embodiment. [Figure 11] FIG. 11 is a top view of the extracted coordinates obtained in this embodiment. [Figure 12] FIG. 13 is a diagram showing an example of a station conveyance device allocation number setting screen in the present embodiment. [Figure 13] FIG. 13 is a diagram showing an extracted region estimated to be a reinforcing bar in this embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] An embodiment to which the present invention is applied will be described below.

[0011] FIG. 1 is a block diagram showing the configuration of an object extraction system 100 according to this embodiment.

[0012] The object extraction system 100 of this embodiment is configured by a computer having a processor (CPU) 1, a memory 2, an auxiliary storage device 3, and a communication interface 4. The object extraction system 100 may also have an input interface 5 and an output interface 8.

[0013] The processor 1 is a calculation device that executes programs stored in the memory 2. The processor 1 executes various programs to realize the functions of the functional units of the object extraction system 100. Note that some of the processes performed by the processor 1 by executing the programs may be executed by other calculation devices (for example, hardware such as ASIC and FPGA).

[0014] The memory 2 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores unchanging programs (e.g., BIOS) and the like. The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 1 and data used when the programs are executed.

[0015] The auxiliary storage device 3 is, for example, a large-capacity non-volatile storage device such as a magnetic storage device (HDD) or a flash memory (SSD). The auxiliary storage device 3 also stores data used by the processor 1 when executing a program, and the program executed by the processor 1. That is, the program is read from the auxiliary storage device 3, loaded into the memory 2, and executed by the processor 1 to realize each function of the object extraction system 100.

[0016] The communication interface 4 is a network interface device that controls communications with other devices in accordance with a predetermined protocol.

[0017] The input interface 5 is an interface to which input devices such as a keyboard 6 and a mouse 7 are connected and which receives input from an operator. The output interface 8 is an interface to which output devices such as a display device 9 and a printer (not shown) are connected and which outputs the results of program execution in a format that can be viewed by a user. Note that a user terminal connected to the object extraction system 100 via a network may provide the input device and the output device. In this case, the object extraction system 100 may have a function of a web server, and the user terminal may access the object extraction system 100 using a predetermined protocol (e.g., http).

[0018] The program executed by the processor 1 is provided to the object extraction system 100 via a removable medium (such as a CD-ROM or a flash memory) or a network, and is stored in a non-volatile auxiliary storage device 3, which is a non-transitory storage medium. For this reason, the object extraction system 100 may have an interface for reading data from the removable medium.

[0019] The object extraction system 100 is a computer system configured on one physical computer, or on multiple logically or physically configured computers, and may operate on a virtual computer constructed on multiple physical computer resources. For example, multiple programs that realize the functions of the object extraction system 100 may each operate on a separate physical or logical computer, or multiple programs may be combined to operate on a single physical or logical computer.

[0020] FIG. 2 is a diagram showing an example of the data structure of an image 401 including distance information handled in this embodiment.

[0021] Range image data has a data structure that contains four pieces of information per pixel: Red, Green, Blue, and Depth. Range image data can be acquired using an imaging device that can acquire distance data for each pixel, such as a TOF camera, LiDAR, or stereo camera. Range image data may also be generated by combining image data and distance data of the same resolution acquired by multiple devices (for example, a camera that outputs RGB pixel data and a laser rangefinder that outputs point cloud distance data).

[0022] Note that in Figure 2, a color image having three colors of RGB brightness data for each pixel has been described, but it is also possible to use RAW data having one color of brightness data for each pixel, or a monochrome image having black and white brightness data for each pixel, as long as it has distance data for each pixel.

[0023] Next, a conventional object extraction procedure will be described with reference to Figures 3 to 6. In the plane estimation method in this conventional object extraction procedure, an operator specifies three points on the object.

[0024] FIG. 3 is a flowchart of a conventional object extraction process using plane estimation.

[0025] The object extraction system 100 receives the specification of three points on an object to be extracted (e.g., a reinforcing bar) in an image showing the object to be extracted, and stores the coordinates of the specified pixels on the image (301). The specification of the three points on the image is performed by an operator specifying three points 403 on the reinforcing bars 501 and 502 in the image 401, as shown in FIG.

[0026] Next, the object extraction system 100 converts the coordinates of the three specified points 403 on the image into three-dimensional coordinates using the camera position and shooting data (302).

[0027] Next, the object extraction system 100 calculates the normal vector of a plane representing the three coordinates in the three-dimensional space (303).

[0028] Next, the object extraction system 100 determines a region b having a depth of a predetermined value Δ in the normal direction based on the plane including the three coordinates (304). For example, as shown in Fig. 5, a rectangular parallelepiped region b507 is determined within a range of a predetermined value Δ506 in front and behind a plane 504 including the specified point 403 (a plane defined by a normal vector 505 calculated from the specified three points 403).

[0029] Next, the object extraction system 100 extracts pixels included in the region b (305). For example, as shown in Fig. 6, by extracting pixels located in the region b 507, a region estimated to be reinforcing bars 501 and 502 can be extracted.

[0030] 7 is a flowchart of the object extraction process executed by the object extraction system 100 of this embodiment. The object extraction process shown in FIG.

[0031] First, the object extraction system 100 divides an area of ​​the image 401 from which an object is to be extracted into regions 803 of a predetermined number or size. The regions 803 are rectangular areas smaller than a predetermined size, and it is preferable to equally divide the area from which an object is to be extracted. Then, the number of pixels in the divided regions 803 is counted for each distance, a local maximum value of the number of pixels is searched for, the local maximum value with the smallest distance is identified, and pixels included within a predetermined distance around the local maximum value are extracted (701). As shown in FIG. 2, the image 401 includes brightness information and distance information of each RGB color for each pixel. Also, in the image 401, the camera does not need to face the grid of the reinforcing bars 501 and 502 directly.

[0032] For example, as shown in FIG. 8, when an operator specifies an area in the image 401 in which the reinforcing bars 501, 502 to be extracted are shown, the object extraction system 100 divides the specified area into regions 803 of a specified size. Also, when the operator specifies a point in the image 401 in which the reinforcing bars 501, 502 to be extracted are shown, a region 803 of a specified size including that point may be defined, and adjacent regions 803 on the top, bottom, left and right may be set to the entire image 401. The region 803 may be set to the entire image 401, but it does not have to be set to an area of ​​a specified width on the top, bottom, left and right of the image 401 (the outer edge of the image). This is because the outer edge of the image may show objects other than the object to be extracted (for example, left and right walls, the ceiling above, and the floor below), and it is desirable to avoid detecting these objects.

[0033] Then, the number of pixels is counted for each distance in a certain region 803. Usually, an image is taken with the object to be extracted in the foreground, and the number of pixels of the object to be extracted is greater than the number of pixels of other objects. Therefore, as shown in FIG. 9, when the number of pixels is represented in a histogram with the vertical axis representing the number of pixels and the horizontal axis representing the distance, the pixels of the reinforcing bars 501 and 502, which are the objects to be extracted, are present at a peak at a close distance on the histogram. For this reason, pixels within a predetermined distance 903 before and after the maximum value 901 with the smallest distance are extracted. When the object to be extracted is a reinforcing bar, the predetermined distance 903 may be set to the maximum diameter of the reinforcing bar. Note that the maximum values ​​902 other than the maximum value 901 with the smallest distance are those of objects present in the background of the object to be extracted.

[0034] Next, the object extraction system 100 converts the coordinates of the extracted pixels on the image into three-dimensional coordinates using the camera position and shooting data (702).

[0035] Next, the object extraction system 100 calculates the normal vector of a plane representing the extracted pixel in a three-dimensional space, and calculates the center of gravity of the three-dimensional coordinates of the extracted pixel in the region 803 (703). For example, it is preferable to calculate the normal vector of the plane that minimizes the sum of the distances from the extracted pixel. Alternatively, the center of gravity (i.e., the average coordinate value) of the extracted pixels in the region 803 may be calculated by the RANSAC method, vectors from the calculated center of gravity to each pixel may be calculated, and the normal vector of the plane may be calculated using the eigenvectors of these pixel groups. Then, a plane having the calculated normal vector and passing through the center of gravity of the pixel group becomes the plane representing the three-dimensional coordinates of the pixel.

[0036] Next, the object extraction system 100 groups adjacent regions 803 using the calculated relationship between the plane and the point cloud between the regions 803 (704). For example, if the average value of the distance between the plane representing a certain first region 803 and the center of gravity of the pixel group extracted in the adjacent second region 803 and the distance between the plane representing the second region 803 and the center of gravity of the pixel group extracted in the first region 803 is smaller than a predetermined threshold, the first region 803 and the second region 803 may be grouped into one group. Note that instead of the average value of the two distances, one distance (for example, the distance between the plane representing the already grouped first region 803 and the center of gravity of the pixel group extracted in the adjacent second region 803) may be used. The predetermined threshold may be, for example, twice the maximum diameter of the reinforcing bar.

[0037] Other methods may be used for grouping the regions 803. For example, if the angle between the normal vectors calculated for two adjacent regions 803 is smaller than a predetermined threshold, the two regions 803 may be grouped into one group. Also, if the distance between the centers of gravity of the pixel groups extracted in the two adjacent regions 803 is smaller than a predetermined threshold, the two regions 803 may be grouped into one group.

[0038] A single index may be calculated using a plurality of the methods described above to comprehensively determine whether or not to group.

[0039] In this way, by grouping the regions 803 according to their relationships, it is possible to accurately determine the plane that represents the position where the object exists. In addition, it is possible to exclude regions 803 that do not include reinforcing bars within the region, or regions 803 that mainly show reinforcing bars behind them.

[0040] Next, the object extraction system 100 selects one of the grouped region groups (705). Since the object to be extracted is usually photographed in the foreground, the region 803 with the smallest distance may be selected, and when an object in the background is to be extracted, the region 803 with the second closest distance may be selected. For example, as shown in Fig. 10, the regions 803 including the reinforcing bars 501 and 502 are grouped into the same group shown in gray, and the other regions 803 are excluded.

[0041] Next, the object extraction system 100 calculates (706) a normal vector of a plane that represents the coordinates of the extracted pixels in the selected region 803. For example, as shown in Fig. 11, the plane is calculated from the coordinates 1104 of the extracted pixels in the selected region 803. For example, similarly to step 703, the normal vector may be calculated by the RANSAC method. The position of the object to be extracted can be identified by the plane calculated in step 706.

[0042] Next, the object extraction system 100 determines an area b having a depth of a predetermined value Δ in the normal direction based on the plane including the three coordinates (707). For example, as shown in Fig. 12, a rectangular parallelepiped area b 1207 is determined within a range of a predetermined value Δ 1206 in front and behind the plane 1203 (normal vector 1205) calculated from the coordinates of the pixel 1204 in step 706.

[0043] Next, the object extraction system 100 extracts pixels included in the region b (708). For example, as shown in Fig. 13, by extracting pixels located in the region b 1207, a region estimated to be the reinforcing bars 501 and 502 can be extracted.

[0044] As described above, according to the embodiment of the present invention, even if the camera is not directly facing the plane consisting of the extraction object (e.g., a rebar grid), the plane representing the extraction object can be accurately determined. In addition, since pixels within a predetermined range are extracted from the plane consisting of the extraction object, the object can be extracted stably regardless of the user, and dependency on individual users can be eliminated.

[0045] The present invention is not limited to the above-described embodiments, and includes various modified examples and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the configurations described. Furthermore, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, a part of the configuration of each embodiment may be added, deleted, or replaced with another configuration.

[0046] In addition, each of the above-mentioned configurations, functions, processing units, processing means, etc. may be realized in hardware, for example by designing some or all of them as an integrated circuit, or may be realized in software by a processor interpreting and executing a program that realizes each function.

[0047] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC card, an SD card, or a DVD.

[0048] In addition, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In reality, it can be considered that almost all components are connected to each other. [Explanation of symbols]

[0049] 1...processor, 2...memory, 3...auxiliary storage device, 4...communication interface, 5...input interface, 6...keyboard, 7...mouse, 8...output interface, 9...display device, 100...object extraction system, 401...image, 403...specified point, 501, 502...rebar, 803...region, 901, 902...local maximum,

Claims

1. An object extraction system for extracting information of an object from an image, comprising: an arithmetic unit that executes arithmetic processing, and a storage unit accessible by the arithmetic unit; The arithmetic unit: obtains a distance image including distance information measured for each pixel; divides the distance image into a plurality of regions; counts the number of pixels for each distance in the divided regions; extracts pixels within a predetermined distance range from the maximum value of the shortest distance from the shooting point among the maximum values of the counted number of pixels; calculates a first plane representing the extracted pixels for each of the divided regions; groups the divided regions using the relationship between the first planes calculated for each region; selects one of the groups generated by the grouping, and calculates a second plane representing the extracted pixels within the region of the selected group. An object extraction system characterized by that.

2. The object extraction system according to claim 1, wherein: the arithmetic unit selects a group with the shortest distance from the shooting point among the groups generated by the grouping. An object extraction system characterized by that.

3. The object extraction system according to claim 1, wherein: the arithmetic unit extracts pixels at a predetermined distance from the calculated second plane as pixels of the object. An object extraction system characterized by that.

4. The object extraction system according to claim 1, wherein: when the distance between the center of gravity of the pixel group extracted in the second region adjacent to the plane representing the first region is smaller than a predetermined threshold value, the arithmetic unit groups the first region and the second region into one group. An object extraction system characterized by that.

5. The object extraction system according to claim 1, wherein: The calculation unit groups two adjacent regions into one group if the angle between the normal vectors calculated in the two adjacent regions is smaller than a predetermined threshold value. The object extraction system is characterized by this.

6. The object extraction system according to claim 1, The calculation unit groups two adjacent regions into one group if the distance between the centroids of the pixel groups extracted in the two adjacent regions is smaller than a predetermined threshold value. The object extraction system is characterized by this.

7. An object extraction method in which an object extraction system extracts information on an object from an image, The object extraction system is configured by a computer having a calculation unit that executes arithmetic processing and a storage unit accessible to the calculation unit, The object extraction method is as follows. The calculation unit acquires a distance image including distance information measured for each pixel, The calculation unit divides the distance image into a plurality of regions, The calculation unit counts the number of pixels for each distance in the divided regions, The calculation unit extracts pixels within a predetermined distance range from the maximum value of the shortest distance from the shooting point among the maximum values of the counted number of pixels, The calculation unit calculates a first plane representing the extracted pixels for each of the divided regions, The calculation unit groups the divided regions using the relationship between the first planes calculated for each region, The calculation unit selects one of the groups generated by the grouping and calculates a second plane representing the extracted pixels within the region of the selected group. The object extraction method is characterized by this.

8. An object extraction program in which an object extraction system extracts information on an object from an image, The object extraction system is configured by a computer having an arithmetic unit that executes arithmetic processing and a storage unit accessible to the arithmetic unit. The object extraction program includes a procedure for acquiring a distance image including distance information measured for each pixel, a procedure for dividing the distance image into a plurality of regions, a procedure for counting the number of pixels for each distance in the divided regions, a procedure for extracting pixels within a predetermined distance range from the maximum value of the shortest distance from the shooting point among the maximum values of the counted number of pixels, a procedure for calculating a first plane representing the extracted pixels for each of the divided regions, a procedure for grouping the divided regions using the relationship between the first planes calculated for each region, and a procedure for calculating a second plane representing the extracted pixels within the region of the selected group by selecting one group generated by the grouping, and causing the arithmetic unit to execute the procedures. An object extraction program characterized by this.