Method, device and computer program for detecting 3D objects from 2D images

By detecting the direction candidates of volume blocks based on projection geometry in the 3D coordinate system, the problem of the object being partially hidden or cut off when detecting 3D objects from 2D images is solved, and higher detection accuracy is achieved.

CN111145139BActive Publication Date: 2025-05-16SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910998360.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-01
Filing Date
2019-10-21
Publication Date
2025-05-16
Estimated Expiration
2039-10-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect 3D objects from 2D images, especially when the objects are partially hidden or cut off in the 2D image.

Method used

Through an iterative search method based on projection geometry, the direction candidates of volume blocks are searched in a 3D coordinate system, and the volume blocks are detected based on the search results.

Benefits of technology

Even if the object is partially hidden or cut off in the 2D image, this method can accurately detect the volume blocks of the 3D object, improving the detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111145139B_ABST
    Figure CN111145139B_ABST
Patent Text Reader

Abstract

The present application provides a method, device and computer program for detecting a 3D object from a 2D image. The method comprises: receiving a 2D image including an object; acquiring an object detection region from the 2D image; iteratively searching for candidates for directions of a volume block of the object in the 2D image in a 3D coordinate system based on the object detection region; and detecting the volume block from the 3D coordinate system based on the result of the iterative search.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] Korean Patent Application No. 10-2018-0133044, entitled “Method and Apparatus for Detecting 3D Objects from 2D Images”, filed on November 1, 2018 in the Korean Intellectual Property Office, is hereby incorporated by reference in its entirety. Technical Field

[0003] Various embodiments relate to methods and apparatus for detecting 3D objects from 2D images. Background Art

[0004] Object detection techniques are used to detect areas containing objects from an image. For example, object detection techniques can be used to detect a two-dimensional (2D) bounding box around an object from a 2D image. The 2D bounding box can be defined by the position and size of the 2D bounding box in the image. Object detection techniques can be performed by neural network-based image processing. In addition, a three-dimensional (3D) bounding box refers to a volume block surrounding an object in a 3D coordinate system, and can be defined, for example, by the position, size, and orientation of the 3D bounding box in the 3D coordinate system. Applications that require a 3D bounding box can include, for example, driving applications. Summary of the invention

[0005] According to an aspect of an embodiment, a method for detecting a 3D object from a 2D image is provided, the method comprising: receiving a 2D image including an object; acquiring an object detection area from the 2D image; iteratively searching for candidates for a direction of a volume block including the object in a 3D coordinate system based on the object detection area; and detecting the volume block from the 3D coordinate system based on a result of the search.

[0006] According to another aspect of the embodiment, a device for detecting a 3D object from a 2D image is provided, the device comprising: a memory configured to store a 2D image including an object; and at least one processor configured to obtain an object detection area from the 2D image, iteratively search for candidates for a direction of a volume block including the object in a 3D coordinate system based on the detection area, and detect the volume block from the 3D coordinate system based on a result of the search. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Various features will become apparent to those skilled in the art from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0008] Figure 1 A diagram showing an object detection method according to an embodiment;

[0009] Figure 2 An operation flow chart showing an object detection method according to an embodiment;

[0010] Figure 3 A detailed flow chart showing the operation of the object detection method according to an embodiment;

[0011] Figure 4 A diagram showing directions according to an embodiment;

[0012] Figure 5 A diagram showing candidates for directions of volumes according to an embodiment;

[0013] Fig. 6A A diagram showing a method of determining a position of a volume according to an embodiment;

[0014] Figure 6B A diagram illustrating a correspondence between a 2D bounding box and a 3D bounding box according to an embodiment;

[0015] Fig. 7A and Figure 7B A diagram showing a method of calculating the position of a volume block according to an embodiment;

[0016] Figure 8 A diagram showing a method of iteratively determining candidates for directions of volume blocks according to an embodiment;

[0017] Fig. 9 A flowchart showing the operation of a method of detecting a 3D object from a 2D image according to an embodiment; and

[0018] Fig.10 A block diagram of an object detection apparatus according to an embodiment is shown. DETAILED DESCRIPTION

[0019] The specific structural or functional descriptions presented in this specification are example descriptions for describing embodiments according to the technical concept, and the embodiments can be implemented in various other forms without being limited to the forms described in this specification.

[0020] Although the terms "first" and "second" are used to describe various elements, these terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element.

[0021] It should be understood that when an element is referred to as being "coupled" or "connected" to another element, the element may be directly coupled or connected to the other element, or any other element may be present between the two elements. Conversely, it should be understood that when an element is referred to as being "directly coupled" or "directly connected" to another element, there are no intervening elements between the two elements. Expressions used to describe the relationship between elements, such as "on", "directly on", "between", "directly between", "adjacent" or "directly adjacent" should be interpreted in the same manner.

[0022] Unless otherwise specified, terms in the singular may include plural forms. In this specification, it should be understood that terms such as "include", "have" or "comprises" are intended to be present in an attribute, fixed number, step, process, element, component or combination thereof, but are not intended to exclude one or more other attributes, fixed numbers, steps, processes, elements, components or combinations thereof.

[0023] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art. Terms such as those defined in general dictionaries may be interpreted as having the same meaning as the contextual meaning in the relevant field, and unless explicitly defined herein, should not be interpreted as having an ideal or overly formal meaning.

[0024] Hereinafter, various embodiments will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same elements throughout.

[0025] Figure 1 is a diagram illustrating an object detection method according to an embodiment. Figure 1 , shows a 2D image and a 3D coordinate system according to an embodiment. The 2D image contains an object.

[0026] For example, to detect an object from a 2D image, the object region can be detected by inputting the 2D image into a learning neural network. The detected region of the object can be a 2D bounding box surrounding the object in the 2D image.

[0027] For example, the 2D image may be an image in which a vehicle is traveling. Assume that a first frame 110, a second frame 120, a third frame 130, and a fourth frame 140 are input over time. In each frame, another vehicle traveling in an adjacent lane may be detected as an object. A first bounding box 115 may be detected in the first frame 110, a second bounding box 125 may be detected in the second frame 120, a third bounding box 135 may be detected in the third frame 130, and a fourth bounding box 145 may be detected in the fourth frame 140.

[0028] The 2D bounding box can be rectangular and can be defined in various ways. For example, the 2D bounding box can be defined using the coordinates of the four corner points. Alternatively, the 2D bounding box can be defined by a position-size combination. The position can be represented by the coordinates of the corner or center point, and the size can be represented by the width or height.

[0029] According to an embodiment, a volume block containing an object in a 3D coordinate system may be detected based on a detection area of ​​the object in a 2D image. The 3D coordinate system may be a world coordinate system. The volume block may be a 3D bounding box surrounding the object in the 3D coordinate system. For example, a 3D bounding box 150 corresponding to the first frame 110, the second frame 120, the third frame 130, and the fourth frame 140 may be detected from the 3D coordinate system. When the object moves over time, the 3D bounding box 150 also moves in the 3D coordinate system.

[0030] The 3D bounding box can be a cuboid and can be defined in various ways. For example, the coordinates of eight corner points can be used to define the 3D bounding box. Alternatively, a combination of position and size can be used to define the 3D bounding box. The position can be represented by the coordinates of the corner points on the bottom surface or the coordinates of the center point on the bottom surface, and the size can be represented by the width, length, or height. The direction can be represented by a direction vector of a line perpendicular to the surface. The direction vector can correspond to the degree (e.g., yaw, pitch, roll) of the 3D bounding box from the three axes (e.g., x-axis, y-axis, z-axis) of the 3D coordinate system. The direction can also be called an orientation.

[0031] Since a 2D image does not include depth information in the z-axis direction, projective geometry may be used to detect the 3D bounding box 150 from the 2D bounding box. Projective geometry relates to properties of geometric objects in a 2D image that may not change when they undergo a projective transformation.

[0032] According to one embodiment, in a 2D image, an object may be partially hidden by another object, or may be partially cut off along the boundary of the 2D image. In this case, the detection area of ​​the object may not completely include the object. For example, referring to the first frame 110, the left side of other vehicles traveling in the adjacent lane is cut off along the boundary of the image, so the first bounding box 115 does not include the portion not displayed in the image, that is, the first bounding box 115 does not include the cut-off portion of the other vehicle.

[0033] If the region of the object detected from the 2D image does not completely include the object, the volume block of the object may not be accurately detected in the 3D coordinate system. However, according to an embodiment, even if the object is partially hidden or cut out in the 2D image, the object will be represented as a complete object in the 3D coordinate system. The following embodiment describes such a technology: by iteratively searching for candidates for the direction of the volume block in the 3D coordinate system based on projective geometry, the volume block in the 3D coordinate system is accurately detected, even if the object is partially hidden or cut out in the 2D image.

[0034] Figure 2 is a flowchart illustrating an operation of an object detection method according to an embodiment.

[0035] refer to Figure 2 , the object detection method of this embodiment includes: an operation 210 of receiving a 2D image including an object, an operation 220 of obtaining an object detection region from the received 2D image, an operation 230 of iteratively searching for volume block direction candidates in a 3D coordinate system based on the object detection region obtained in operation 220, and an operation 240 of detecting a volume block in the 3D coordinate system based on a result of the iterative search in operation 230. At this time, at least a portion of the object included in the 2D image received in operation 210 may be hidden by another object, or may be partially cut off along a boundary of the received 2D image.

[0036] In detail, in operation 210, the received 2D image may be an image captured with a camera.

[0037] In operation 220, the detection area detected from the received 2D image may be a 2D bounding box. The detection area may be detected from a 2D bounding box as described above (eg, a 2D bounding box defined using coordinates of four corner points or by a position-size combination).

[0038] In operation 220, not only the detection area is detected, but also the direction of the object (hereinafter referred to as the local direction) can be acquired from the received 2D image. For example, a neural network can be used, and the neural network can receive the 2D image and output the local direction as the direction of the object in the 2D image.

[0039] The direction of an object in a 2D image (hereinafter referred to as a local direction) can be converted into the direction of the object in a 3D coordinate system (hereinafter referred to as a global direction) based on projective geometry. Figure 4 , a ray direction 410 is defined from the camera to the center of the object in the 2D image. The camera can be aligned with an axis 430 (e.g., the x-axis) of the 3D coordinate system. The ray angle θ can be used ray(ie, an angle between the ray direction 410 and the axis 430 of less than 180°) represents the ray direction 410. In a 2D image, the local angle θ between the ray direction 410 and the direction 420 of the object can be used. L (For example, Figure 4 The local direction is represented by the angle greater than 180° between the direction 410 of the light and the direction 420 of the vehicle. Alternatively, the global angle θ between the direction 435 parallel to the axis 430 and the direction 420 of the object may be used. G to represent the global direction in the 3D coordinate system. ray and the local angle θ L Add to calculate the global angle θ G .

[0040] If the object detection region in the 2D image completely contains the object, then the object can be detected by moving the camera in the direction toward the center of the object detection region (e.g. Figure 4 The direction from the camera to the center of the vehicle) determines the ray angle θ ray However, if at least a portion of the object is hidden or cut away, the ray angle θ determined by the direction from the camera toward the center of the object detection area ray May be inaccurate.

[0041] As described below, in operation 220, the size of the volume block in the 3D coordinate system may be further acquired. The volume block may be a cuboid, and the size of the volume block may include dimensions of width, length, and height. For example, a learning neural network may identify a category or type of an object to be detected. The neural network may output the size of the volume block according to the identified category or type. For ease of description, it has been described that the size of the volume block is determined after the category or type of the object is identified. However, in some embodiments, the neural network may be an end-to-end neural network that receives a 2D image and directly outputs the size of the volume block.

[0042] At least some neural networks may be used in operation 220. For example, a first neural network configured to determine and output a detection region, a second neural network configured to determine and output a local direction, a third neural network configured to determine and output a volume size, etc. may be combined into a single neural network.

[0043] According to an embodiment, in operation 230 of the object detection method, a volume block direction candidate corresponding to the direction of the volume block is iteratively searched. Therefore, even if at least a portion of the object is hidden or cut off, the volume block can be accurately detected or reconstructed. Figure 3 Operation 230 is described in detail.

[0044] In operation 240, after the volume direction is detected based on the iterative search results in operation 230, the detected volume area of ​​the object containing the 2D image can be defined by its position, size and orientation in the 3D coordinate system. As described above, the method of defining the volume can be modified in various ways, for example, it can be modified to a method of defining the volume using the coordinates of eight corner points.

[0045] In the following, reference will be made to Figure 3 The operation 230 of iteratively searching for volume direction candidates is described in detail. Figure 3 is a detailed flow chart of operation 230 according to an embodiment.

[0046] refer to Figure 3 , a volume block 350 is detected from a 2D image 310. The 2D image 310 is Figure 2 The image captured by the camera in operation 210 and includes an object that is at least partially hidden or cut away.

[0047] First, when a 2D image 310 is received (in operation 210), as described above with reference to Figure 2 As described in operation 220, an object detection area 317 and an object local direction 319 are obtained based on the 2D image 310. In addition, as described above, a size 315 of a volume block of the detection area based on the 2D image received in operation 210 may also be obtained. The above description may be applied herein, and therefore, the detailed description will not be repeated.

[0048] Next, candidates 320 of global directions that can orient the volume 350 in the 3D coordinate system are generated. The initial global direction candidates can be generated based on a predetermined search range and resolution. Note that "candidate" refers to possibilities, so candidates for directions refer to possible directions that are checked to determine the actual direction of the object in the 3D coordinate system.

[0049] Figure 5 Eight global direction candidates 510, 520, 530, 540, 550, 560, 570, and 580 generated in the xz plane of the 3D coordinate system are shown. In this case, the search range may be from -π to π in the xz plane, and the resolution may be π / 4. According to an embodiment, since the shape of the volume block 350 is symmetric in the direction of the volume block 350, the search range may be set to be from 0 to π.

[0050] According to an embodiment, a global direction determined by a ray direction from a camera toward a detection region (e.g., a 2D bounding box) and an object local direction 319 may also be included in the initial global direction candidate. When a ray direction pointing to a detection region of an object that is at least partially hidden or cut away is used, an inaccurate global direction may be calculated, but the inaccurate global direction may be used as a starting point for searching for global direction candidates.

[0051] A volume block position candidate 330 may be estimated in a 3D coordinate system based on the global direction candidate 320 and the volume block size 315. Fig. 6A , based on projective geometry, shows the correspondence between the pixels of a 2D image and the 3D coordinates. Fig. 6A The relationship shown in projects the 3D coordinates onto the pixels of the 2D image.

[0052] exist Fig. 6A , the coordinates (x, y, z) correspond to the 3D coordinates in the 3D coordinate system. Assuming that the center point 610 of the lower surface of the volume block is located at the origin of the 3D coordinate system, the 3D coordinates of the eight corners of the volume block can be represented by the size of the volume block. For example, if the width of the volume block is w, the length of the volume block is l, and the height of the volume block is h, the 3D coordinates of the four corners of the lower surface of the volume block can be (-w / 2, 0, -1 / 2), (w / 2, 0, -1 / 2), (w / 2, 0, 1 / 2), (-w / 2, 0, 1 / 2), and the 3D coordinates of the four corners of the upper surface of the volume block can be (-w / 2, -h, -1 / 2), (w / 2, -h, -1 / 2), (w / 2, -h, 1 / 2), (-w / 2, -h, 1 / 2).

[0053] In addition, reference Fig. 6A , T is the motion matrix associated with the position of the volume block. This can be solved by Fig. 6A The position of the volume block is estimated by T obtained from the equation of . R is a rotation matrix, which can be determined by the global direction of the volume block. Candidates corresponding to the global direction candidates of the position of the volume block can be estimated. K represents the intrinsic parameters of the camera, and S represents the scale factor. In addition, (x_img, y_img) represents the pixel coordinates of the 2D image.

[0054] Figure 6BThe correspondence between the first feature points included in the object detection area 710 of the 2D image and the second feature points included in the object volume block 720 in the 3D coordinate system is shown. Some of the first feature points and some of the second feature points may match each other. For example, a pixel (x_img_min, y_img_min) in the detection area 710 may match the 3D coordinates (x_min, y_min, z_min) of the object volume block 720. According to an embodiment, when determining the pixels to be matched, pixels corresponding to hidden or cut-off portions may be excluded.

[0055] Reference again Fig. 6A , if the coordinates of the matched pixel are set to (x_img, y_img) and the matched 3D coordinates are set to (x, y, z), an equation is obtained where the position of the volume block is unknown. Here, the 3D coordinates may be 3D coordinates that can be represented by the size of the volume block when assuming that the volume block is located at the origin of the 3D coordinate system.

[0056] Fig. 7A and 7B A method for calculating the position of a volume block by matching 2D image pixels with 3D coordinates is shown. Fig. 7A In the left 2D image of , when viewed from the inside of the vehicle, the x-coordinate and y-coordinate of the pixel corresponding to the upper right end of the front side of the vehicle are x min and min In addition, the y coordinate of the pixel corresponding to the lower left end of the front side of the vehicle when viewed from the inside of the vehicle is y max , and the x coordinate of the pixel corresponding to the lower left end of the rear side of the vehicle is x max .

[0057] like Fig. 7A The 3D coordinate system on the right shows that the pixels match the corners of the volume, so we can use Figure 7B The relationship shown is used to calculate T x , T y and T z . Coordinates (T x ,T y ,T z ) may be a 3D coordinate corresponding to the location of the volume block (eg, the center of the lower surface of the volume block).

[0058] When there are three or more pairs of pixels and 3D coordinates that match each other, the position of the volume block in the 3D coordinate system can be explicitly calculated. According to an embodiment, even when there are two pairs of pixels and 3D coordinates that match each other, the position of the volume block can be determined by further considering the inclination of the ground plane in the 2D image. For example, the objects in the driving image may be adjacent vehicles, and it can be assumed that the vehicles are traveling parallel to the ground plane in the driving image. Therefore, if the inclination of the volume block is set to be the same as the inclination of the ground plane, the position of the volume block can be explicitly determined even if there are only two pairs of pixels and 3D coordinates that match each other. Alternatively, even when there are three or more pairs of pixels and 3D coordinates that match each other, the position of the volume block can be determined by further considering the inclination of the ground plane.

[0059] Reference again Figure 3 , a candidate 330 of a position of a volume block corresponding to the global direction candidate 320 may be estimated. One of the volume block position candidates 330 may be selected as a position candidate 335. For example, after estimating the volume block position candidate 330, the volume block corresponding to the volume block position candidate 330 may be projected onto a 2D image, and then one of the volume block position candidates 300 may be selected as the position candidate 335 by comparing the object detection region 317 in the 2D image with the projection region to which the volume block position candidate 330 is projected.

[0060] For example, in a 3D coordinate system, a volume block can be defined based on position, size, and orientation, and a position candidate can be determined for a global direction candidate, respectively. Given the size of the volume block, a volume block candidate corresponding to the position candidate can be determined. A projection area can be obtained by projecting the volume block candidate onto a 2D image. A projection area that has the maximum overlap with the detection area 317 or a projection area having a value equal to or greater than a predetermined critical value can be selected, and a position candidate 335 corresponding to the selected projection area can be selected.

[0061] When a position candidate 335 is selected, the light direction is calculated using the projection image corresponding to the position candidate 335. The global direction 340 can be determined by adding the light direction and the local direction 319, as previously described with reference to Figure 4 As described.

[0062] Once the global direction 340 is determined in the current iteration, the next global direction candidate may be generated in the next iteration based on the global direction 340. For example, a global direction candidate having a search range smaller than that in the previous iteration but having a higher resolution than that in the previous iteration may be generated based on the global direction 340.

[0063] refer to Figure 8, it can be assumed that direction 810 is determined as a global direction in a previous iteration. In this case, global direction candidates 820, 830, 840, 850, 860, and 870 can be generated at a resolution of π / 8 within a search range of 0 to π. According to an embodiment, since the shape of the volume block is symmetric in the direction of the volume block, the search range can be set to be from π / 4 to 3π / 4, which is smaller than the search range of the previous iteration.

[0064] After the final global direction is determined by iteratively searching for global direction candidates, the volume block 350 can be detected or reconstructed based on the final global direction. After the final global direction is determined, since the size of the volume block is given, Fig. 6A The final position of the volume block 350 is estimated by the relationship shown in . Once the final global orientation, final position and final size of the volume block 350 are determined, the volume block 350 can be detected or reconstructed.

[0065] Fig. 9 is a flowchart illustrating a method of detecting a 3D object from a 2D image according to an embodiment.

[0066] refer to Fig. 9 In operation 1, global orientation candidates are selected by quantizing the range of [-π,π] (910). The global orientation candidates correspond to the global direction candidates respectively.

[0067] In operation 2, the center position of the 3D box is calculated by using the size of the 3D box given as input and applying the projection geometry to the global orientation candidate selected in operation 1 (920). The 3D box corresponds to the 3D volume block, and the center position of the 3D box corresponds to the position of the 3D volume block. According to an embodiment, in operation 2, the center position of the 3D box may be calculated by further considering the inclination of the ground plane.

[0068] In operation 3, the ray angle is calculated using the optimal value in the center position of the 3D frame calculated in operation 2 and corresponding to the global orientation candidate (930). The optimal value may be determined based on an overlap area between the 2D detection area and the projected image of the 3D frame corresponding to the global orientation candidate.

[0069] In operation 4, the global orientation is calculated (940) by adding together the local orientation given as input and the ray angle calculated in operation 3. The local orientation corresponds to the local direction, and the global orientation corresponds to the global direction.

[0070] In operation 5, a global orientation candidate close to the global orientation calculated in operation 4 is selected (950).

[0071] In operation 6, the center position of the 3D box is calculated by using the size of the 3D box given as input and applying projection geometry to the global orientation candidate selected in operation 5 (960). According to an embodiment, in operation 6, the center position of the 3D box may be calculated by further considering the inclination of the ground plane.

[0072] In operation 7, the ray angle is calculated using the optimal value in the center position of the 3D box calculated in operation 6 and corresponding to the global orientation candidate (970). In operation 8, the final global orientation is calculated by adding the local orientation given as input and the ray angle calculated in operation 7 (980).

[0073] In operation 9, a final center position of the 3D box is calculated (990) by using the size of the 3D box given as input and applying projection geometry to the final global orientation calculated in operation 8. According to an embodiment, in operation 9, the final center position of the 3D box may be calculated by further considering the inclination of the ground plane.

[0074] Fig.10 is a block diagram showing an electronic system according to an embodiment. Fig.10 , the electronic system includes at least one processor 1020 and a memory 1010. The electronic system may further include a sensor 1030. The processor 1020, the memory 1010, and the sensor 1030 may communicate with each other via a bus.

[0075] The processor 1020 may execute the above reference Figure 1 and Figure 2 At least one of the methods described. The memory 1010 may store images captured using the sensor 1030. The memory 1010 may be a volatile memory or a non-volatile memory. The processor 1020 may execute a program and may control an electronic system. Program codes executable on the processor 1020 may be stored in the memory 1010.

[0076] The electronic system may be connected to an external device (eg, a personal computer or a network) through an input / output device, and may exchange data with the external device. The electronic system may include various electronic systems, for example, a server device or a client device.

[0077] The above-mentioned embodiments can be realized with hardware elements, software elements and / or a combination of hardware elements and software elements. For example, the devices, methods and elements described in the above-mentioned embodiments can be realized with at least one general or special-purpose computer (such as, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor or any other device capable of executing instructions and responding to instructions). The processing device can execute an operating system (OS) and at least one software application running on the operating system. In addition, the processing device can access, store, manipulate, process and generate data in response to the execution of software. For ease of understanding, the situation of using a single processing device can be described. However, those of ordinary skill in the art will recognize that the processing device can include multiple processing elements and / or multiple types of processing elements. For example, the processing device can include multiple processors, or a processor and a controller. Other processing configurations such as parallel processors are also possible.

[0078] Software may include a computer program, code, instructions, or a combination of at least one of them. In addition, the processing device may be configured to operate in a desired manner and may be instructed independently or collectively. Software and / or data may be permanently or temporarily embodied in a specific machine, component, physical device, virtual device, computer storage medium or device, or propagate signal waves so that instructions or data are interpreted by or provided to the processing device. Software may be distributed on network-coupled computer systems and may be stored and executed in a distributed manner. Software and data may be stored in at least one computer-readable recording medium.

[0079] The method of the embodiment can be implemented as a program instruction that can be executed on various computers and then stored in a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., individually or in combination. The program instructions stored in the medium may be program instructions designed and configured according to the embodiment or known in the computer software industry. The computer-readable recording medium may include hardware specially configured to store program instructions and execute program instructions, and examples of hardware include magnetic media, such as hard disks, floppy disks, and tapes; optical media, such as CD-ROMs and DVDs; magneto-optical media, such as floppy disks; and ROMs, RAMs, and flash memories. Examples of program instructions may include machine code generated by a compiler and high-level language codes executable on a computer using an interpreter. The above-mentioned hardware device may be configured to operate via one or more software modules to perform operations according to the embodiment, and vice versa.

[0080] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and interpreted in a general and descriptive sense only and not for purposes of limitation. Unless otherwise expressly indicated, in some cases, as would be apparent to one of ordinary skill in the art at the time of filing this application, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the spirit and scope of the invention as set forth in the appended claims.

Claims

1. A method for detecting a 3D object from a 2D image, the method comprising: receiving a 2D image including an object; Acquire an object detection region from the 2D image; Iteratively searching for candidates for directions of a volume block of the object in the 2D image in a 3D coordinate system based on the object detection area; as well as detecting the volume block from the 3D coordinate system based on a result of the iterative search, Wherein, the iterative search includes: generating a candidate for the orientation of the volume block in the 3D coordinate system; estimating a candidate for the position of the volume block in the 3D coordinate system based on the generated candidate for the direction of the volume block and the size of the volume block; selecting a candidate for the position of the volume block from among the estimated candidates for the position of the volume block based on the object detection region and a projection region of the volume block corresponding to the estimated candidates for the position of the volume block; Based on the selected candidate for the position of the one volume block and the orientation of the object in the 2D image, the orientation of the volume block is determined in the 3D coordinate system.

2. The method of claim 1, wherein: In the 2D image, at least a portion of the object is hidden by another object or cut off along a boundary of the 2D image.

3. The method of claim 1, wherein: Generating candidates for the direction of the volume block includes generating candidates for the direction of the volume block based on directions of the volume block determined in previous iterations.

4. The method of claim 1, wherein: Generating the candidate for the direction of the volume block includes generating the candidate for the direction of the volume block based on a search range smaller than a search range of a previous iteration and a resolution higher than a resolution of the previous iteration.

5. The method according to claim 1, wherein: Candidates for generating the direction of the volume block include at least one of the following: generating a plurality of candidates for the direction of the volume block based on a predetermined search range and a predetermined resolution; as well as A candidate for the direction of the volume block is generated corresponding to the direction of the object in the 2D image and the direction of light directed to the center point of the object detection area.

6. The method of claim 1, wherein: Estimating the candidates for the position of the volume block includes determining candidates for the position of the volume block corresponding to the candidates for the direction of the volume block and the size of the volume block based on a correspondence between feature points of the object detection area and feature points of the volume block.

7. The method of claim 6, wherein: Estimating candidates for the position of the volume block further includes excluding feature points corresponding to a cut-away portion or a hidden portion of the object from the feature points of the object detection area.

8. The method of claim 6, wherein: Determining the candidate for the position of the volume block includes determining the candidate for the position of the volume block corresponding to the candidate for the direction of the volume block and the size of the volume block by further considering the inclination of the ground plane of the 2D image.

9. The method of claim 1, wherein: Selecting a candidate for the position of a volume block from the estimated candidates for the position of the volume block comprises: calculating a size of an overlapping area between the projection area and the object detection area; selecting a projection area from the projection areas based on a size of the overlapping area; A candidate for the position of the volume block corresponding to the selected projection region is selected.

10. The method of claim 1, wherein: Detecting the volume block includes determining a position of the volume block in the 3D coordinate system based on a size of the volume block obtained from the 2D image and a direction of the volume block obtained from a result of the search.

11. The method of claim 10, wherein: Determining the position of the volume block includes determining the position of the volume block corresponding to the direction of the volume block and the size of the volume block based on a correspondence between feature points of the object detection area and feature points of the volume block.

12. The method of claim 1, wherein: Acquiring the object detection area includes acquiring the object detection area including the object, a direction of the object in the 2D image, and a size of the volume block in the 3D coordinate system by using a neural network that recognizes the 2D image.

13. A computer program product comprising a computer program / instructions, wherein: The computer program / instructions implement the steps of the method according to claim 1 when executed by a processor.

14. A device for detecting a 3D object from a 2D image, the device comprising: A memory for storing a 2D image including an object; as well as at least one processor, configured to obtain an object detection region from the 2D image, iteratively search for candidates for a direction of a volume block including the object in a 3D coordinate system based on the object detection region, and detect the volume block from the 3D coordinate system based on a result of the search, Wherein, for the iterative search, the at least one processor: generating a candidate for the orientation of the volume block in the 3D coordinate system; estimating a candidate for a position of the volume block in the 3D coordinate system based on the candidate for the direction of the volume block and the size of the volume block; selecting a candidate for the position of the volume block from among the candidates for the position of the volume block based on the object detection region and a projection region of the volume block corresponding to the candidate for the position of the volume block; and Based on the selected candidate for the position of the one volume block and the orientation of the object in the 2D image, the orientation of the volume block is determined in the 3D coordinate system.

15. The apparatus of claim 14, wherein: In the 2D image, at least a portion of the object is hidden by another object or cut off along a boundary of the 2D image.

16. The apparatus of claim 14, wherein: The at least one processor generates a candidate for an orientation of the volume based on an orientation of the volume determined in a previous iteration.

17. The apparatus of claim 14, wherein: The at least one processor generates candidates for the direction of the volume based on a search range that is smaller than a search range of a previous iteration and a resolution that is higher than a resolution of the previous iteration.

18. The apparatus of claim 14, wherein: The at least one processor generates a plurality of candidates for the direction of the volume block based on a preset search range and a preset resolution.

19. The apparatus of claim 14, wherein: The at least one processor generates a candidate for the direction of the volume corresponding to the direction of the object in the 2D image and a direction of light rays directed to a center point of the object detection area.

20. The apparatus of claim 14, wherein: The at least one processor determines a candidate for a position of the volume block corresponding to a candidate for a direction of the volume block and a size of the volume block based on a correspondence relationship between feature points of the object detection area and feature points of the volume block.

21. The apparatus of claim 20, wherein: The at least one processor excludes feature points corresponding to a cut-away portion or a hidden portion of the object from the feature points of the object detection area.

22. The apparatus of claim 20, wherein: The at least one processor determines a candidate for the position of the volume block corresponding to the candidate for the direction of the volume block and the size of the volume block by further considering the inclination of the ground plane of the 2D image.

23. The apparatus of claim 14, wherein: The at least one processor selects a candidate for the position of the volume block from the candidates for the position of the volume block in the following manner: calculating a size of an overlapping area between the projection area and the object detection area; selecting a projection area from the projection areas based on a size of the overlapping area; A candidate for the position of the volume block corresponding to the selected projection region is selected.

Citation Information

Patent Citations

  • 3D printer extruder

    KR1020180133044A

  • Systems and methods for extracting information about objects from scene information

    US20170220887A1

  • Object tracking techniques

    US9563955B1