A method for 3D scanning of real objects
The method facilitates user-friendly 3D scanning of large objects using standard cameras by guiding users to align a virtual 3D box and capture frames from predefined poses, addressing the inefficiencies of existing methods and ensuring precise reconstruction.
Patent Information
- Application Number
- JP2021117287
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-22
- Filing Date
- 2021-07-15
- Publication Date
- 2026-03-02
- Estimated Expiration
- 2041-07-15
AI Technical Summary
Existing 3D scanning methods are tedious, require specialized hardware, and yield uncertain results, especially when scanning large objects with standard cameras.
A method using a camera-equipped handheld device with augmented reality to guide users through placing a virtual 3D box around the object, aligning it with the object's symmetry, and capturing frames from predefined tile poses to ensure accurate 3D reconstruction.
Enables user-friendly, high-quality 3D scanning of large objects using standard cameras without the need for specialized hardware, ensuring precise and efficient reconstruction.
Smart Images

Figure 0007822138000002 
Figure 0007822138000003 
Figure 0007822138000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of 3D content creation, more particularly to a method for 3D scanning of real objects (hereinafter referred to as real objects). Hereinafter, 3D scanning refers to the process of capturing the shape and appearance of a real object to create a virtual 3D representation. This process is also referred to as 3D digitization and 3D capture. 3D scanning of real objects is applied in fields such as computer-aided design (CAD) and 3D printing. [Background technology]
[0002] For decades, photogrammetry has been proposed as a way to scan objects (ranging in size from everyday objects to parts of the Earth's surface, such as mountain waves). Photogrammetry is based on the principle of automatically recognizing and correlating corresponding points in images acquired from different viewpoints. To do this, the user moves around the object to be scanned and takes as many photographs as possible. The user does not know exactly how many photographs need to be taken, how much time they need to spend acquiring them, or from what positions the photographs should be taken. Once the images are acquired, they are sorted. Some of them are deemed useful—with enough detail, adequate lighting, and sufficient overlap—while others are ignored. Some are blurry, have too much overlap, or not enough overlap, or were taken in poor lighting conditions. The user then lets the algorithm create a 3D model entirely without further user control. Therefore, this task of 3D scanning can be tedious and frustrating for users, as the quality of the results is uncertain.
[0003] 3D scanning applications such as [1] utilize augmented reality (AR) technology to scan objects in a user-friendly way. AR technology allows virtual objects to be overlaid on a real scene captured by a camera. The user's pose in the real world is updated by the AR system, and the pose of the virtual object is updated accordingly, so that the virtual object appears anchored in the real world.
[0004] In the aforementioned 3D scanning application, the user prints a checkerboard on paper to serve as a tracking marker, places the object in the center of the checkerboard, and opens the application. The tracking marker allows the algorithm to solve more quickly, allowing results to be displayed immediately after capture. The user places the object on the tracking marker and launches the application. The object is then displayed surrounded by a dome composed of hundreds of segments. The user rotates around the object (or rotates the object on a rotating support) until all segments are clear. Because the application must print the checkerboard, the user must carry a marker or have a printer available, which is very inconvenient. Furthermore, the dome must fit the dimensions of the tracking marker (the height of the dome is the same as the length of the tracking marker). This is rarely feasible using a standard office printer for anything other than A3 format.
[0005] Other 3D scanning applications available on the market, such as "Scandy Pro 3DScanner" or "itSeez3D," make the 3D scanning process very easy and do not require printed media. These applications require the use of a specific camera called an RGB-D camera, which stands for "Red / Green / Blue-Depth." An RGB-D camera simultaneously provides a color image and a depth map (based on known patterns, such as facial patterns) to characterize the distance of objects in the image. However, RGB-D cameras are not available on all smartphones and tablets, and are only available on some of the latest models. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] "Qlone 3D Scanner" (https: / / www.qlone.pro / ) developed by EyeCue Vision Tech [Non-patent document 2] "3D Bounding Box Estimation Using Deep Learning and Geometry," Arsalan Mousavian et al., CVPR 2017 [Non-patent document 3] "Mesh R-CNN," Georgia Gkioxari et al., 2019 IEEE / CVF International Conference on Computer Vision ICCV [Non-patent document 4] "Shape Theory through Spatial Sculpture" (Kiriakos N. Kutulakos et al., International Journal of Computer Vision 38(3),199-218,2000) [Non-Patent Document 5] "Autofocusing of Diatoms in Bright-Field Microscopy: A Comparative Study," Pech-Pacheco et al., Proceedings of the 15th International Conference on Pattern Recognition. ICPR-2000 Summary of the Invention [Problem to be solved by the invention]
[0007] Therefore, there is a need to provide a method for 3D scanning of real objects that is user-friendly, compatible with standard cameras, and capable of scanning large objects. [Means for solving the problem]
[0008] The object of the invention is a computer-implemented method for 3D scanning of a real object using a camera having a 3D position, comprising the following steps: a) Receive an image of the real object from the camera. b) An image of a real object surrounded by a virtual 3D box is displayed on the screen in an augmented reality view, and a virtual structure consisting of a set of planar tiles is overlaid on the real object and fixed to the virtual 3D box, each tile corresponding to a given pose of the camera. c) Detect when the camera is pointed at the tile. d) Obtain a frame of the virtual 3D box from the camera and verify the tile with it. This frame is the projection of the virtual 3D box on the image.
[0009] Steps a) to d) are repeated for different 3D positions of the camera until a sufficient number of tiles have been verified to scan the real object. e) Perform a 3D reconstruction algorithm on all captured frames.
[0010] Additional features and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a flow chart of a method according to the present invention; [Figure 2] A close-up of a real object surrounded by a virtual 3D box. [Figure 3] FIG. 1 is a perspective view illustrating an example of a virtual structure. [Figure 4] FIG. 1 is a top view illustrating an example of a virtual structure. [Figure 5] FIG. 10 is a side view showing an example of a virtual structure. [Figure 6] FIG. 10 is an illustration showing a screenshot of an augmented reality view, along with a view in which a user is capturing frames. [Figure 7] FIG. 1 is a block diagram of a computer system suitable for carrying out methods in accordance with the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0012] FIG. 1 shows the main steps of the method of the present invention.
[0013] Throughout the process, the user holds the camera of a handheld device compatible with Augmented View. For example, the handheld device may be a smartphone or a tablet equipped with a screen and a camera for displaying the acquired images on the screen. Thanks to the invented method, the user is guided from one pose of the camera CAM to another by following the visual instructions of the wizard, as described below.
[0014] The real environment is provided with an object OBJ to be scanned. In the example shown, the object to be scanned is a teddy bear placed on a table. In a first step a), an image IM of the real object OBJ is received from the camera CAM. The pose of the camera CAM relative to the pose of the object OBJ can be obtained by calibrating and registering the camera CAM, i.e., by determining the intrinsic and extrinsic parameters of the camera. Knowing these parameters, it is possible to calculate the transformation (rotation / translation) of 3D points in the 3D scene to the corresponding pixels in the captured image.
[0015] In a second step b) of the method, an image IM of the real object OBJ is displayed on the screen in real time in an augmented reality view. Because the camera is calibrated, the object may be surrounded by a virtual 3D box VBO (rectangular prism) that surrounds the object OBJ. As the user moves around the object OBJ, the orientation of the virtual 3D box VBO changes accordingly. The virtual 3D box VBO is overlaid on the real object OBJ in the augmented view enabled by the handheld device. The virtual 3D box VBO can be positioned and scaled in the augmented reality view so that the symmetry of the virtual 3D box VBO matches the underlying symmetry of the object OBJ. However, this condition is not essential for the method to work.
[0016] Accurately placing a virtual 3D box VBO in a 3D scene can be challenging, especially when considering typical user input consisting of 2D interactions on the surface of a tablet, or imprecise in-air hand gestures in the case of devices such as augmented reality head-mounted display smart glasses.
[0017] Furthermore, it should be noted that in most cases, the object to be scanned is placed on a flat surface. Therefore, in a preferred embodiment, step a) includes detecting in the image IM a horizontal plane PLA on which the real object OBJ is located and prompting the user to place the virtual 3D box VBO on the plane PLA. By using the constraint of the detected plane, the user is assisted in placing the virtual 3D box VBO in the 3D scene.
[0018] The planes may be detected, tracked, and displayed, for example (and without limitation), by using the ARKit software development kit running on an iOS device and dedicated to augmented reality applications.
[0019] The user is asked to select a point on the visible part of the plane that is close to the object to be scanned OBJ. A virtual 3D box VBO is then created near the user-defined point, which can be modified to fit the object to be scanned.
[0020] You can adjust the virtual 3D box VBO using intuitive gestures. If you're using a smartphone, you can translate it by simply touching the virtual 3D box VBO and moving your finger on the screen. To facilitate easy manipulation, movement is limited by the previously detected plane. Therefore, the virtual 3D box VBO can be easily dragged on a surface and overlaid on the object to be scanned.
[0021] A virtual grid, shown in Figure 2, can be placed on the plane the box is placed on to further assist the user. Such visual feedback helps anchor the virtual object over the real world by providing scale and perspective. It also helps determine whether the plane has been accurately detected by checking the alignment of the grid with the real plane.
[0022] The virtual 3D box VBO can also be rotated using the rotation handles, which can be represented as arcs near the virtual 3D box VBO, specifically below it. The user drags the rotation handle to one side or the other of the arc. This is especially useful when the shape of the object OBJ to be scanned is close to a box. In this case, you need to align the edges of the real shape and the virtual box as much as possible to get a better fit and match the potential symmetry plane.
[0023] The virtual 3D box VBO can also be resized. When using a smartphone screen, one possible implementation of this feature is to tap and hold on the face to modify it. The face changes color, becomes detached from the rest of the cube, and can be moved along its normals. When you release the face, the box is rebuilt to take into account the face's new position, and the box grows or shrinks depending on the user's gestures.
[0024] By moving slightly around the object, the user can iteratively adjust the box as described above, ensuring that the object to be captured fits as closely inside as possible while still leaving room to account for inevitable imprecision in tracking.
[0025] If an object isn't positioned on a plane, a virtual box can still be created and adjusted, but the process can be lengthy. Indeed, without a physical plane to provide a useful constraint, the user must move around the object and adjust every face individually. For example, for a lampshade hanging from the ceiling, the user must roughly place a virtual box over the lampshade, then move it so that all parts of the lampshade fit within the box, adjusting every face in all x, y, and z directions.
[0026] Instead of manually providing a virtual 3D box VBO that surrounds objects in an extended view, you can provide a virtual 3D box VBO automatically.
[0027] In fact, many machine learning object detection algorithms (e.g., classification) can provide bounding rectangles for all objects detected in an image. The user can select the object to reconstruct (by touch, voice, gaze, or simply pointing the camera). The virtual 3D box VBO can be centered at the intersection between the center of the bounding rectangle's base and a plane. Two of the three dimensions of the virtual 3D box VBO are known. The user only needs to adjust one dimension (depth) and can possibly rotate the box to improve alignment. Rotation can also be replaced by moving around the object and relying on machine learning to iteratively adjust the virtual 3D box VBO.
[0028] Recent advances in machine learning can go a step further by providing 3D bounding boxes (as disclosed in the article [2]) or even approximations of the shape of objects if they belong to a known class (as disclosed in the article [3]).
[0029] Alternatively, the virtual 3D box VBO can be provided automatically using computer vision. Some augmented reality libraries are able to detect 3D points that correspond to notable features in the real environment, resulting in a more or less dense 3D point cloud. If there are enough points, the virtual 3D box VBO is calculated automatically when the user points to the object to capture.
[0030] During step b) of the invented method, a virtual structure VST, created from a set of planar tiles TIL and fixed to the virtual 3D box VBO, is also displayed in the augmented reality view, as shown, for example, in FIG. 6. The tiles are planar and preferably rectangular. The arrangement of tiles TIL is arranged around the virtual 3D box VBO, with each tile TIL corresponding to a predefined pose of the camera CAM. Thus, by pointing the camera at each tile while positioning the entire virtual 3D box VBO within the camera's field of view, the user takes exactly one valid image to ensure fast reconstruction of the object with good 3D reconstruction quality.
[0031] In a third step c), it is detected that the tile TIL is pointed at by the camera CAM. Naturally, the user will tend to point incorrectly at the tile (not exactly perpendicular to the tile), so the user may move the camera CAM slightly before pointing accurately at the tile TIL.
[0032] If a tile TIL is specified, a frame of the virtual 3D box VBO is acquired and the tile TIL is enabled (step d). The frame is the projection of the virtual 3D box VBO onto the image IM. Depending on the reconstruction algorithm, only the projection of the virtual 3D box VBO is processed, rather than the entire 3D scene displayed in the image. Thus, the maximum volume to be scanned is precisely defined, and the reconstruction algorithm can focus on details within that volume instead of trying to reconstruct the entire scene. The virtual 3D box VBO is projected onto the captured frame. This can be verified by projecting each corner point of the bounding box onto the camera frame and checking that the resulting coordinates of the projected points fit within the known frame size. This ensures that the model to be scanned is fully visible in the captured frame and is not accidentally cropped due to poor user aim.
[0033] To acquire a frame and detect whether the camera is correctly positioned and oriented relative to the object, a ray is cast from the camera. When the ray collides with an unverified tile, the angle between the ray and the normal direction of the unverified tile is calculated, and if the angle is less than a predetermined value, the tile is validated. We found that a predetermined value between 15 and 25 degrees is appropriate. In particular, a predetermined value of approximately 18 degrees, corresponding to the cosine of 0.95, is particularly suitable because it provides sufficient accuracy without being difficult to reach. Whether the ray direction is close to the tile's normal direction can also be measured by the dot product between the two directions, which must be higher than a threshold. Pointing the tile in a direction close to the normal direction is ideal for the latter 3D reconstruction because it ensures that the captured frame represents a perspective close to that defined by the tile's pose.
[0034] In a preferred embodiment, the light beam is cast periodically at predetermined time intervals. Therefore, the user does not need to press any buttons or touch the screen of the handheld device during the scanning process, and the user can hold the camera with both hands, improving the stability of the captured frames. The periodicity of the light beam casting can be embodied by displaying a timer, such as a progress bar, on the screen.
[0035] A visual indicator, such as a target, may appear in the center of the screen corresponding to the camera's optical axis, as shown in Figure 6. If the visual indicator is within a tile, a timer is displayed.
[0036] Next, steps a) to d) are repeated for different 3D positions of the camera CAM until a sufficient number of tiles have been verified to scan the real object OBJ. Then, in step e), a 3D reconstruction algorithm is performed on all captured frames. In a preferred embodiment, the reconstruction algorithm is a 3D carving algorithm, and the positions of the tiles of the virtual structure VST are determined accordingly. The carving algorithm rarely requires very specific poses. The principles of carving are explained in the article "Carving 3D Models: A Virtual Structure of Virtual Structures VST" [4]. Each frame is used to "carve" the 3D model.
[0037] In a preferred embodiment, the virtual structure VST is arranged as a hemisphere (or dome) covering the top of the virtual 3D box VBO, as shown in Figures 3, 4, and 5. Because the tiles are planar, the virtual structure VST does not fit exactly on the hemisphere. Figure 3 shows an example virtual structure in perspective view, Figure 4 shows a top view of the virtual structure, and Figure 5 shows a side view.
[0038] Looking at the top view of the virtual structure VST (Figure 4), the tiles include a first subset of tiles ST1 that form a cross with its center coinciding with the topmost point TMP of the virtual structure VST. The cross is aligned with the sides of the virtual 3D box VBO. That is, the top-center tile is parallel to the top surface of the virtual 3D box VBO, and the direction of the cross matches the direction of the virtual 3D box VBO. The cross-shaped arrangement of the tiles allows the user to take photos of the object according to the perspectives typically used in computer-aided design: front, back, top, left, and right perspectives.
[0039] The second subset of tiles ST2 is positioned around the periphery of the virtual structure VST at a height midway between the bottom point BOT and the top point TMP of the virtual structure VST. The tiles of the second subset of tiles ST2 may be larger than the tiles of the first subset of tiles ST1. Therefore, a limited number of tiles of the second subset of tiles ST2 are required to quickly obtain a complete set of frames. Four tiles of the second subset of tiles ST2 are aligned with the cross of the first subset of tiles ST1 so as to be aligned with the face of the virtual 3D box VBO.
[0040] Finally, a third subset of tiles ST3 is placed at the bottom BOT of the periphery of the virtual structure VST. The tiles of the third subset of tiles ST3 may be larger than the tiles of the second subset of tiles ST1 for the same reasons as above. The four tiles of the third subset of tiles ST3 are aligned with the crosses of the first subset of tiles ST1 so as to be aligned with the faces of the virtual 3D box VBO.
[0041] In a preferred embodiment of the present invention, the first, second, and third subsets of tiles (from highest to lowest along the z-axis) have 5 (including the top tile), 8, and 16 tiles, respectively, for a total of 29 tiles in the dome. Thus, the user does not need to take as many photographs without sacrificing the quality of the reconstruction.
[0042] As can be seen in Figure 3, the first subset of tiles ST1, the second subset of tiles ST2, and the third subset of tiles ST3 can be spaced apart from one another by a constant polar angle θ. The constant polar angle θ is measured between the centers of the tiles aligned along the z-axis. The third subset of tiles ST3 is elevated relative to the center CTR of the virtual 3D box VBO to facilitate frame acquisition by the user.
[0043] TIFF0007822138000001.tif35150
[0044] As can be seen in Figure 4, tiles in the same subset are spaced apart from each other by a constant azimuth angle. The constant azimuth angles within the first, second, and third subsets of tiles are π / 2, π / 4, and π / 8, respectively.
[0045] As can be seen in Figure 5, the radius r of the virtual structure VST, which corresponds to the radius of the circle underlying the dome, is set to the maximum dimension of the virtual 3D box VBO. x , d y , d z Considering this, r=max(d x , d y , d z ) The resulting virtual 3D box VBO is adjusted to the size of the object and the radius of the dome circle, allowing even large objects to be scanned.
[0046] The virtual structure VST patterns disclosed in Figures 3 to 5 allow scanning and viewing the object from all sides while obtaining good reconstruction results. Note that the bottom tiles may be slightly tilted rather than vertical, since they are above the plane on which the object is placed, forcing the user to assume an uncomfortable pose.
[0047] If a pure photogrammetry algorithm were chosen (instead of a carving algorithm), the dome would appear denser and more regular, even without such alignment constraints.
[0048] As an alternative to the aforementioned hemispherical shape of the virtual structure VST, the virtual structure may be arranged as a pattern corresponding to the category of the object OBJ. The category of the object is determined during step a): the user provides information about the object category of the 3D object from a set of categories of objects (e.g., vehicles or human faces) or from categories of objects automatically detected by pattern recognition.
[0049] Thus, a virtual structure VST that conforms to the shape of an object takes into account the complexity of the object. Tiles are arranged to capture as much of the object's detail as possible. These tiles can be of different sizes and granularity and can be displayed in several steps: a first set of tiles can be presented to capture the object's coarse features, ideal for sculpting the object's overall shape (e.g., a vehicle), and a second set of tiles with a finer pattern can be used to capture the finer surface details of the object's subcomponents (e.g., an exterior rearview mirror).
[0050] Whatever the type of virtual structure VST (dome or pattern corresponding to the object category), the user rotates the object OBJ and the tiles are examined one by one, as shown in Figure 6, where the right corner image shows the user is around the object and capturing frames using the camera CAM (of a handheld device such as a tablet or smartphone), and the center image shows what the user is looking at on the screen of the handheld device.
[0051] Temporary visual feedback on the pointed tile indicates that the corresponding frame is being retrieved. The temporary feedback may be a predefined color or the tile flashing. Once the tile is retrieved, the temporary visual feedback is converted to persistent visual feedback, for example, of a different color. In Figure 6, the tile containing the teddy bear's head has already been retrieved on the screen, but the user is attempting to retrieve the tile containing the teddy bear's belly on the screen. By providing the user with visual feedback, the complexity of the inputs required for the reconstruction algorithm is hidden from the user and instead presented as an entertaining game of "tile hunt."
[0052] Advantageously, having the user take as clear a photograph as possible, avoiding blur, ensures that the image will be usable by the reconstruction algorithm.
[0053] To this end, according to a first variant, a check is performed in advance (before frame acquisition): during the frame acquisition step, the amplitude of the camera CAM movement and / or the speed of the handheld device CAM movement are measured. If the amplitude and / or speed of the movement exceed a predetermined value, frame capture is prohibited. Thus, no frame is acquired. Stability can be calculated using internal sensors of the handheld device, such as a gyroscope, accelerometer, or compass. Stability can also be calculated by verifying that the point at the intersection of the camera direction and the tile being aimed at is within a predetermined position range during tile validation.
[0054] According to a second variant, a check is made a posteriori (after frame acquisition): once the frame is acquired, the sharpness of the captured frame is measured and the frame is validated only if it is sufficiently sharp. For the evaluation of the frame sharpness, the Laplacian deviation algorithm can be used. The algorithm is described in the article "Image Processing with 3D Imaging," by Chris Bauer, "Image Processing with 3D Imaging," and "Image Processing with 3D Imaging," by Chris Bauer, ... and "Image Processing with 3D Imaging."
[0055] These conditions allow a high level of validity in the automatically captured frames, in the sense that the localization, as well as the quality of the frames themselves, allows for a high degree of utilization by subsequent reconstruction algorithms: all captured frames are used with confidence and none are discarded.
[0056] In some cases, not all sides of the object OBJ are visible or even accessible to the user. This is especially true when the object is too large for the user to see all its sides (especially the top view), or when the object to be scanned is placed against an obstacle (such as a wall) and cannot be moved. In such cases, it should be possible to handle partial acquisition.
[0057] In particular, in step a), it is determined whether the object OBJ has symmetry. To this end, the user can provide instructions to the system regarding the shape of the object to be scanned, such as a plane, a plane, an axis of symmetry, or an axis of rotation. Since adding such instructions to the object is not easy using only the image projected on the screen, the user may interact with the virtual 3D box VBO to indicate that the object OBJ has symmetry after appropriately adjusting the virtual 3D box VBO to fit the object to be captured. If the object is general (e.g., a chair, a bottle, etc.), another instruction provided by the user could be the class of the object to be scanned. Next, a 3D reconstruction algorithm is performed based on the verified tiles and the symmetry of the object.
[0058] Therefore, returning to FIG. 1, if the object has symmetry, it is not necessary to examine all tiles in order to perform the 3D reconstruction algorithm and reconstruct it.
[0059] The methods of the present invention can be implemented by a suitably programmed general-purpose computer or virtual reality system, optionally including a computer network, storing a suitable program in non-volatile form on a computer-readable medium such as a hard disk, solid state disk or CD-ROM, and using its microprocessor(s) and memory to execute said program.
[0060] A computer CPT suitable for carrying out the method according to an exemplary embodiment of the present invention will now be described with reference to Fig. 7. In Fig. 7, the handheld device HHD comprises a central processing unit CPU which carries out the above-mentioned method steps while executing an executable program, i.e. a set of computer-readable instructions, stored in a memory device such as RAM M1 or ROM M2, or stored remotely. The created 3D content is stored in one of the memory devices or stored remotely.
[0061] The claimed invention is not limited by the form of the computer-readable medium on which the computer-readable instructions and / or data structures of the inventive processes are stored. For example, the instructions and files may be stored on a CD, DVD, flash memory, RAM, ROM, PROM, EPROM, EEPROM, hard disk, or other information processing device with which the computer communicates, such as a server or computer. The programs and files may be stored on the same memory device or on different memory devices.
[0062] Furthermore, a computer program suitable for carrying out the methods of the present invention may be provided as a utility application, a background daemon, or a component of an operating system, or a combination thereof, running in conjunction with a central processing unit CPU and an operating system such as Microsoft VISTA, Microsoft Windows 10, UNIX, Solaris, LINUX, Apple MAC-OS, and other systems known to those skilled in the art.
[0063] The central processing unit CPU can be a Xenon processor from Intel Corporation of America, an Opteron processor from AMD Corporation of America, or other processor types such as a Freescale ColdFire, IMX, or ARM processor from Freescale Corporation of America. Alternatively, the CPU can be a processor such as a Core2Duo from Intel Corporation of America, or can be implemented using FPGA, ASIC, PLD, or discrete logic circuitry, as will be appreciated by those skilled in the art. Furthermore, the central processing unit can be implemented as multiple processors working in cooperation to execute the computer-readable instructions of the processes of the present invention described above.
[0064] The virtual reality system of FIG. 7 also includes a network interface NI, such as an IntelEthernetPRO network interface card from Intel Corporation of the United States, for interfacing with a network such as a local area network (LAN), a wide area network (WAN), or the Internet.
[0065] The handheld device HHD includes a camera CAM for acquiring images of the environment, a screen DY for displaying the acquired images to the user, and a sensor SNS for calculating the 3D orientation and 3D position of the handheld device HHD.
[0066] Any method steps described herein should be understood as representing a module, segment, or portion of code that comprises one or more executable instructions for implementing a particular logical function or step in processing, and alternative implementations are included within the scope of the exemplary embodiments of the present invention.
Claims
1. A computer-implemented method for 3D scanning a real object (OBJ) using a camera (CAM) having a 3D position, comprising: a) receiving an image (IM) of said real object (OBJ) from said camera (CAM); b) displaying on a screen in an augmented reality view said image (IM) of said real object (OBJ) enclosed in a virtual 3D box (VBO), wherein a virtual structure (VST) consisting of a set of planar tiles (TIL) is superimposed on said real object (OBJ) and fixed to said virtual 3D box (VBO), each tile (TIL) corresponding to a given pose of said camera (CAM); c) detecting that the camera (CAM) is pointed at a tile (TIL); d) obtaining a frame of the virtual 3D box (VBO) from the camera (CAM) showing the virtual 3D box (VBO) projected onto the image (IM), thereby validating the tile (TIL); Steps a) to d) are repeated for different 3D positions of the camera (CAM) until a sufficient number of tiles have been verified to scan the real object (OBJ); e) performing a 3D reconstruction algorithm on all frames captured by said camera (CAM); Including, the virtual structure (VST) is arranged as a hemisphere covering the virtual 3D box (VBO), and the radius of the virtual structure (VST) is set to the maximum dimension of the virtual 3D box (VBO); The set of tiles (TIL) including the top (TMP) of the virtual structure (VST) to the bottom (BOT) of the virtual structure (VST) is a first subset (ST1) of tiles forming a cross with its center coinciding with the topmost point (TMP) of said virtual structure (VST), said cross being aligned with the sides of said virtual 3D box (VBO); a second subset of tiles (ST2) partially aligned with the cross; a third subset of tiles (ST3), partially aligned with the cross; Including, the first (ST1), second (ST2), and third (ST3) subsets of tiles are spaced apart from one another by a constant polar angle (θ); the tiles (TIL) are spaced apart from each other by a constant azimuth angle (θ1, θ2, θ3) within the same subset, a third subset (ST3) of tiles is elevated relative to the center (CTR) of the virtual 3D box (VBO); 10. A computer-implemented method comprising:
2. 2. The method of claim 1, wherein step a) comprises detecting, in an image (IM), a horizontal plane (PLA) on which the real object (OBJ) is located and prompting a user to place the virtual 3D box (VBO) on the horizontal plane (PLA).
3. the first (ST1), the second (ST2) and the third (ST3) subsets of tiles comprise 5, 8 and 16 tiles, respectively; said constant polar angle (θ) is 2π / 15; the constant azimuthal angles within the first (ST1), second (ST2) and third (ST3) subsets of tiles are π / 2, π / 4 and π / 8, respectively; - the third subset (ST3) of tiles is elevated at an angle of π / 10 relative to the center (CTR) of the virtual 3D box (VBO), 2. The method of claim 1 .
4. A computer-implemented method for 3D scanning a real object (OBJ) using a camera (CAM) having a 3D position, comprising: a) receiving an image (IM) of said real object (OBJ) from said camera (CAM); b) displaying on a screen in an augmented reality view said image (IM) of said real object (OBJ) enclosed in a virtual 3D box (VBO), wherein a virtual structure (VST) consisting of a set of planar tiles (TIL) is superimposed on said real object (OBJ) and fixed to said virtual 3D box (VBO), each tile (TIL) corresponding to a given pose of said camera (CAM); c) detecting that the camera (CAM) is pointed at a tile (TIL); d) obtaining a frame of the virtual 3D box (VBO) from the camera (CAM) showing the virtual 3D box (VBO) projected onto the image (IM), thereby validating the tile (TIL); Steps a) to d) are repeated for different 3D positions of the camera (CAM) until a sufficient number of tiles have been verified to scan the real object (OBJ); e) performing a 3D reconstruction algorithm on all frames captured by said camera (CAM); Including, Step a) includes detecting or receiving information about a category of the real object (OBJ) from a set of object categories, and the virtual structure is arranged in a pattern corresponding to the category of the object.
10. A computer-implemented method comprising:
5. 5. The method of claim 4, wherein first the entire real object (OBJ) is 3D scanned by using the pattern, and then sub-components of the real object (OBJ) are 3D scanned by using finer patterns as virtual structures corresponding to the sub-components.
6. 6. The method of claim 1, wherein step d) comprises displaying temporary visual feedback on the screen on the indicated tile to indicate that a corresponding frame of the virtual 3D box (VBO) has been acquired, and wherein once the frame of the virtual 3D box (VBO) has been acquired, the temporary visual feedback is converted into persistent visual feedback.
7. A computer-implemented method for 3D scanning a real object (OBJ) using a camera (CAM) having a 3D position, comprising: a) receiving an image (IM) of said real object (OBJ) from said camera (CAM); b) displaying on a screen in an augmented reality view said image (IM) of said real object (OBJ) enclosed in a virtual 3D box (VBO), wherein a virtual structure (VST) consisting of a set of planar tiles (TIL) is superimposed on said real object (OBJ) and fixed to said virtual 3D box (VBO), each tile (TIL) corresponding to a given pose of said camera (CAM); c) detecting that the camera (CAM) is pointed at a tile (TIL); d) obtaining a frame of the virtual 3D box (VBO) from the camera (CAM) showing the virtual 3D box (VBO) projected onto the image (IM), thereby validating the tile (TIL); Steps a) to d) are repeated for different 3D positions of the camera (CAM) until a sufficient number of tiles have been verified to scan the real object (OBJ); e) performing a 3D reconstruction algorithm on all frames captured by said camera (CAM); Including, Step d) is - projecting a ray from the 3D position of said camera in the direction of the optical axis of said camera (CAM); - when detecting a collision between said ray and an unverified tile, calculating the angle between said ray and the normal direction of said unverified tile; - validating said unvalidated tile if said angle is less than a predetermined value; 10. A computer-implemented method comprising:
8. 8. The method of claim 7, wherein the predetermined value is comprised between 15 and 25 degrees.
9. 9. The method according to claim 7, wherein the light beam is applied periodically at predetermined time intervals.
10. Step d) is measuring the amplitude of the movement of said camera (CAM) and / or the speed of the movement of said camera (CAM); - inhibiting the capture of frames by said camera (CAM) if the amplitude and / or speed of said movement exceeds a predetermined value; 10. The method according to claim 1, comprising:
11. Step d) is measuring the sharpness of the frames captured by said camera (CAM); - verifying the frames captured by said camera (CAM) if they are sufficiently clear; 11. The method according to claim 1, comprising:
12. Step a) comprises: a substep of determining whether the real object (OBJ) has symmetrical parts, in which case only tiles corresponding to parts of the real object (OBJ) that differ from each other are examined; the 3D reconstruction algorithm is performed based on the frames captured by the camera (CAM) and on the symmetry of the real object; 12. The method according to any one of claims 1 to 11.
13. 13. The method of claim 1, wherein the 3D reconstruction algorithm comprises a space carving algorithm.
14. A computer program comprising instructions enabling a computer to carry out the method according to any one of claims 1 to 13.
15. 15. A non-transitory computer-readable data storage medium having recorded thereon the computer program of claim 14.
16. A computer system comprising a processor (P) coupled to a memory (M1, M2), a screen (DY) and a camera (CAM), characterized in that said memory stores computer-executable instructions that cause the computer system to perform the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and apparatus for representing a physical scene
JP2016534461A
Enhanced depth map images for mobile devices
JP2019534515A
Caching and updating of dense 3D reconstruction data
US20190197786A1
Laser scanning system, laser scanning method, moving laser scanning system, and program
WO2017130770A1