Progressive 3D Point Cloud Segmentation of Objects and Background from a Tracking Session
By using fuzzy segmentation and tracking in augmented reality, the technique automatically classifies 3D point clouds into object and background elements, addressing the limitations of existing methods and achieving efficient and reliable segmentation.
Patent Information
- Application Number
- JP2021568784
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-21
- Filing Date
- 2020-05-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-05-06
AI Technical Summary
Existing methods for segmenting 3D point clouds into object and background elements are either manual and time-consuming or rely on heuristic methods that fail when objects do not conform to assumed sizes and distances from backgrounds.
The technique employs fuzzy segmentation and tracking in augmented reality applications to automatically distinguish between object and background points in a 3D point cloud, using probability-based classification and continuous learning from tracking sessions.
This approach enables efficient, automatic, and reliable object-background segmentation without manual intervention, improving object detection and tracking across various environments and continuously refining results over time.
Smart Images

Figure 0007685956000001 
Figure 0007685956000002 
Figure 0007685956000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for generating augmented reality content by automating point cloud cleanup using fuzzy segmentation of points as corresponding to an object or background.
Background Art
[0002] A 3D point cloud is a standard format for storing 3D models that can contain thousands of points, each having its position and color. Usually, a point cloud is created during 3D reconstruction (a computational pipeline that calculates the 3D geometry of a model of an object or scene based on its video). Typically, when reconstructing a 3D object from a video, the resulting point cloud includes not only the object but also some background elements such as the table, wall, floor on which the object is placed, and other objects seen in the video. For many applications of 3D models, it is important to have a clean point cloud that lists only the points belonging to the object. Therefore, the task of cleaning and removing segments within the model that do not belong to the object is very important. This task is sometimes referred to as "segmentation into 'object' and 'background'".
[0003] One approach for segmentation is the manual removal of background points using an interactive application that enables labeling background regions and deleting them. Tools such as Blender and MeshLab are examples of this approach, but this requires manual and time investment and is therefore not scalable. A second approach is to use some heuristic method, for example, setting a threshold on the distance from the camera or calculating a compact blob. These methods do not work well when the object does not conform to the assumptions about its size and distance from the background elements.
[0004] Therefore, there is a need for a technique that provides an automatic and reliable performance of the point cloud object-background segmentation task.
Summary of the Invention
[0005] Embodiments of the present system and method can provide techniques for the automatic and reliable performance of point cloud segmentation tasks. Embodiments can provide the ability to gradually learn object-background segmentation from tracking sessions in augmented reality (AR) applications. The 3D point cloud can be used for tracking in AR applications by matching points from the cloud to regions within the live video. 3D points with many matches are more likely to be part of an object, while 3D points with few matches are more likely to be part of the background. Advantages of this approach include that no manual work is required to perform the segmentation, and the results can be continuously improved over time as the object is tracked in multiple environments.
[0006] For example, in one embodiment, a method of generating augmented reality content may be implemented in a computer comprising a processor, a memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor. The method includes, in a computer system, tracking a 3D model of an object in a scene within a video stream as a camera moves around within the scene, where tracking includes determining, for each point in each frame of the video stream, whether the point corresponds to the object or the background. Segmenting, in the computer system, each of a plurality of points into either a point corresponding to the object or a point corresponding to the background based on the probability that the point corresponds to the object and the probability that the point corresponds to the background, where tracking identifies a point as an inlier if, for each point, there are more cases where the probability that the point corresponds to the object is higher, and tracking identifies a point as an outlier if, for each point, there are more cases where the probability that the point corresponds to the background is higher. Generating, in the computer system, augmented reality content based on the segmented 3D model of the object.
[0007] In an embodiment, tracking may include estimating the pose of a camera with respect to a 3D model of an object using a segmented 3D model of the object. The object 3D model may first be segmented using default settings of probabilities per point, and for each frame, tracking updates the probability per point based on whether it was determined that the point corresponds to the object or to the background. The default setting of the probability per point may be.5, and when tracking matches a point to a pixel in the video frame, the probability may be increased, and when tracking does not match a point to any of the pixels in the video frame, the probability may be decreased. The 3D model of the object in the scene may include a 3D point cloud including a plurality of points including at least some points corresponding to the object and at least some points corresponding to the background of the scene. The video stream may be obtained from a camera moving around in the scene.
[0008] In one embodiment, a system for generating augmented reality content may include a processor, a memory accessible by the processor, and computer program instructions stored in the memory, the computer program instructions being executable by the processor to track a 3D model of an object in a scene in a video stream as the camera moves around in the scene, tracking including determining, for each point per frame, whether the point corresponds to the object or the point corresponds to the background, segmenting each of the plurality of points into either a point corresponding to the object or a point corresponding to the background based on the probability that the point corresponds to the object and the probability that the point corresponds to the background, such that when, for each point, it is more often the case that the probability that the point corresponds to the object is higher, tracking determines that the point corresponds to the object, and when, for each point, it is more often the case that the probability that the point corresponds to the background is higher, tracking determines that the point corresponds to the background, generating augmented reality content based on the segmented 3D model of the object.
[0009] In one embodiment, a computer program product for generating augmented reality content may comprise a non-transitory computer-readable storage having program instructions incorporated therein, the program instructions being executable by a computer to cause the computer to perform a method, the method comprising tracking a 3D model of an object within a scene in a video stream as a camera moves around within the scene, the tracking including determining, for each point in each frame, whether the point corresponds to the object or the background, segmenting each of the plurality of points into either a point corresponding to the object or a point corresponding to the background based on a probability that the point corresponds to the object and a probability that the point corresponds to the background, such that, for each point, if the probability that the point corresponds to the object is higher more often, the tracking determines that the point corresponds to the object, and for each point, if the probability that the point corresponds to the background is higher more often, the tracking determines that the point corresponds to the background, and generating augmented reality content based on the segmented 3D model of the object.
[0010] The details of the present invention can be most deeply understood by referring to the accompanying drawings with respect to both its structure and operation. In the drawings, like reference numerals and symbols refer to like elements.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0012] Embodiments of the present system and method can provide techniques for automatically and reliably performing a point cloud object-background segmentation task. Embodiments can provide the ability to perform automatic segmentation of a 3D point cloud into object and background segments by gradually learning object-background segmentation from a tracking session in an augmented reality (AR) application. In an AR application, a 3D model of one or more objects to be displayed within a scene in a video stream is generated and can be displayed on the scene or on an object within the scene. To appropriately display such AR content, the position of the camera with respect to the scene and with respect to the objects within the scene on which the AR content is to be displayed is determined and can be tracked as the camera moves. The model of an object within the scene can be generated by generating a 3D point cloud of the object using 3D reconstruction into an image or video stream. The 3D point cloud can be used for tracking in an AR application by matching points from the cloud to regions in a live video.
[0013] Typically, a generated 3D model that includes a 3D point cloud cannot be used as is because the point cloud may include points that are located on the background of the scene rather than on the object to be tracked. To remove the background or to segment the points into points that are part of the object to be tracked and points that are not, a form of "cleaning" of the image or video stream may be applied. Embodiments of the present system and method can automatically perform such cleaning by segmenting the points into points that are part of the object to be tracked and points that are not. Such automatic segmentation can result in improved performance.
[0014] For example, embodiments provide a method for progressive segmentation of a 3D point cloud into an object and a background. Embodiments can progressively associate points with object and background classes using a tracking session. Embodiments may use fuzzy segmentation rather than binary segmentation, and the segmentation may be considered a "side effect" of the tracking process. 3D points that have many matches to pixels in the video are considered inlier points for tracking the object over time and are more likely to be part of the object, while 3D points that have few matches to pixels in the video are more likely to be part of the background. The advantages of this approach can include enabling improved object detection and tracking when the object is repositioned in various environments, further including that no manual effort is required to perform the segmentation, and the results can be continuously improved over time as the object is tracked in multiple environments.
[0015] FIG. 1 shows an exemplary system 100 in which embodiments of the present system and method may be implemented. As shown in the example of FIG. 1, system 100 may include a platform 101 that is or may include a camera for capturing video images of a 3D object 104 while the camera moves through a plurality of viewpoints 102A-H to form a video stream 103. The platform 101 may move around following a path or trajectory 108. The video stream 103 may be transmitted to an image processing system 106 where processes included in embodiments of the present system and method may be implemented. For example, the image processing system 106 may perform 3D reconstruction by generating a 3D point cloud of a scene or an object within the scene or both, according to embodiments of the present system and method. FIG. 3 shows an example of such a model 300 of an object. After the 3D point cloud of the scene including the model 300 is generated, the points within the 3D point cloud may be segmented into points 302 within the object represented by the model 300 and points 304 belonging to the background.
[0016] FIG. 2 shows an embodiment of a process 200 of the operation of embodiments of the present system and method. Process 200 may start at 202 where a 3D point cloud may be generated. For example, a video stream 103 may be received and a 3D point cloud for a model of an object may be reconstructed from the video stream 103 having n points. Each point may be assigned a probability of being a part of an object within the scene in the video stream 103. In an embodiment of fuzzy segmentation, each point i within the 3D point cloud may be assigned a probability of belonging to the object, represented as i At this time, 1 - p i is the probability of belonging to the background. The probability p i that point i is a part of the object may be initialized to, for example, 50%. The probabilities of all the points within the point cloud may be formed into a vector. At 204, the point cloud and the probability vector are represented as P = [p 0 , p 1 ... p n , where p iSegmentation can be performed by considering all points having a threshold value T as belonging to the object.
[0017] In 206, in order to be used for tracking in an augmented reality (AR) session of the point cloud, several descriptors of the visual appearance of each point from various viewpoints can be assigned. The descriptors can be calculated from the video used to reconstruct the point cloud. In 208, a model of the object can be tracked using the 3D point cloud in the AR session. As shown in FIG. 4, for each frame 402, the tracker 404 can estimate the pose 406 of the camera with respect to the model 410 using the segmentation information 408, and for the points in the AR video stream 103, matching can be calculated from the point descriptors. For example, the estimation can be based on matching visual descriptors such as scale-invariant feature transform (SIFT) that can associate 3D points with image points. Segmentation can be used to select which 3D points should be considered. For example, the tracking can be performed using only points having p i >T. Or, for example, the tracking can be performed based on sampled points and their probabilities. Since the tracker can output the estimated camera pose, the result of the tracking can be the object position in the video stream 103. The matches that support the estimated pose may be considered inliers for the pose, while the other matches may be considered outliers.
[0018] In 210, the 3D points belonging to the inlier matches may have their p i values increased, while the points belonging to the outlier matches may have their p i values decreased. In 212, over time, the p i values are processed such that points belonging to the object can have high p i values, while background points have low p iIt will be adjusted so that it can have a value. In 214, in order to improve the segmentation, the determined segmentation can be updated. In an embodiment, 3D-2D matching statistics, which can be internal data within the tracker, can be used to improve the segmentation. For example, 3D points belonging to an object segment may generally be considered to be "correctly" matched, while points belonging to the background may be considered to be "badly" matched. As shown in FIG. 5, in an embodiment, for each point, the model tracker 502 can output its matching error 504 and whether it contributed to the pose estimation (inlier-outlier). That is, the model tracker 502 can indicate which points were actually used for tracking (inliers) and which points were not used (outliers). The segmentation update 506 can update the current segmentation 508 by increasing the p i value for inlier points and decreasing the p i value for outlier points, which can be used to update the segmentation information 510, which can then be used by, for example, the model tracker 502. Therefore, information about which points were actually used for tracking and which points were not used can be used to update the segmentation, which can then be used to improve the tracking.
[0019] FIG. 6 shows an exemplary block diagram of a computer system 600 in which the processes described herein may be implemented. The computer system 600 may be implemented using one or more programmed general purpose computer systems such as an embedded processor, a system-on-chip, a personal computer, a workstation, a server system, and a minicomputer or mainframe computer, or within a distributed networked computing environment. The computer system 600 may include one or more processors (CPUs) 602A - 602N, an input / output circuitry 604, a network adapter 606, and a memory 608. The CPUs 602A - 602N execute program instructions to implement the functions of the present communication system and method. Typically, the CPUs 602A - 602N are one or more microprocessors such as INTEL CORE(R) processors. FIG. 6 shows one embodiment in which the computer system 600 is implemented as a single multi-processor computer system where multiple processors 602A - 602N share system resources such as the memory 608, the input / output circuitry 604, and the network adapter 606. However, the present communication system and method also include embodiments in which the computer system 600 is implemented as multiple networked computer systems, which may be single-processor computer systems, multi-processor computer systems, or a combination thereof.
[0020] The input / output circuit mechanism 604 provides the ability to input data into the computer system 600 or output data therefrom. For example, the input / output circuit mechanism may include input devices such as a keyboard, mouse, touch pad, track ball, scanner, analog-to-digital converter, etc., output devices such as a video adapter, monitor, printer, etc., and input / output devices such as a modem, etc. The network adapter 606 interfaces the device 600 with the network 610. The network 610 can be any public or private LAN or WAN, including but not limited to the Internet.
[0021] Memory 608 stores program instructions executed by CPU 602 to perform the functions of computer system 600, as well as data used and processed. Memory 608 can be, for example, electronic memory devices such as random-access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc., and can also include electromechanical memories such as magnetic disk drives, tape drives, optical disk drives, etc., which can use an integrated drive electronics (IDE) interface, or a variation or extension thereof such as enhanced IDE (EIDE) or ultra-direct memory access (UDMA), or a small computer system interface (SCSI)-based interface, or a variation or extension thereof such as first SCSI, wide SCSI, first and wide SCSI, etc., or Serial Advanced Technology Attachment (SATA), or a variation or extension thereof, or a fiber channel-arbitrated loop (FC-AL) interface.
[0022] The content of the memory 608 may vary according to the functions programmed for the computer system 600 to execute. In the example shown in FIG. 6, exemplary memory contents representing routines and data for the embodiments of the processes described above are shown. However, those skilled in the art will recognize that these routines, along with the memory contents associated with those routines, need not be included on a single system or device, but rather may be distributed among multiple systems or devices based on well-known engineering considerations. The present communication system and method may include any such configuration.
[0023] In the example shown in FIG. 6, the memory 608 may include a point cloud generation routine 612, a segmentation determination routine 614, a tracking routine 616, a segmentation update 618, video stream image data 620, and an operating system 622. The point cloud routine 612 may include software routines for generating a 3d point cloud for a model of an object that can assign a probability that each point in the point cloud is part of an object, can assign descriptors of the visual appearance from various viewpoints, and can be reconstructed from the video stream image data 620. The segmentation determination routine 614 may include software routines for determining and updating the probability that each point is part of an object based on information received from the tracking routine 616. The tracking routine 616 may include software routines for tracking the position of an object within the video stream image data 620, outputting an estimated camera pose, and outputting information indicating whether each point was used to track the object, and this information may be used by the segmentation determination routine 614 for determining and updating the probability that each point is part of an object. The operating system 634 may provide overall system functionality.
[0024] As shown in FIG. 6, the present communication system and method may include implementation on a system or group of systems that provide multi-processor, multi-tasking, multi-process, or multi-threaded computing or a combination thereof, as well as implementation on a system that provides only single-processor, single-threaded computing. Multi-processor computing includes performing computing using more than one processor. Multi-tasking computing includes performing computing using more than one operating system task. A task is an operating system concept that refers to a combination of a program to be executed and accounting information used by the operating system. Whenever a program is executed, the operating system necessarily creates a new task for it. A task is like an envelope for a program in that it identifies the program using a task number and attaches other accounting information to it. Many operating systems, including Linux, UNIX(R), OS / 2(R), and Windows(R), have the ability to execute many tasks simultaneously and are called multi-tasking operating systems. Multi-tasking is the ability of an operating system to execute more than one executable file simultaneously. Each executable file is executing within its own address space. That is, the executable files cannot share any of their memory. This has the advantage that no program can harm the execution of any other program running on the system. However, programs can only exchange information other than through the operating system (or by reading files stored on the file system). Since the terms task and process are often used interchangeably, multi-process computing is similar to multi-tasking computing. However, some operating systems distinguish between the two.
[0025] The present invention can be a system, a method, a computer program product, or a combination thereof at any possible technical detail integration level. The computer program product can include a computer-readable storage medium (or a group of media) having computer-readable program instructions for causing a processor to implement aspects of the present invention. The computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device.
[0026] The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, punched cards, mechanically encoded devices such as raised structures within a groove in which instructions are recorded, and any suitable combination of the foregoing. The computer-readable storage medium should not be construed as being a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through an electrical wire.
[0027] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0028] Computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or any combination of object-oriented programming languages such as Smalltalk(R), C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or an external computer connection may be made (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit mechanism including a programmable logic circuit mechanism, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit mechanism to perform aspects of the present invention.
[0029] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0030] These computer-readable program instructions may be provided to the processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, or other devices to function in a particular manner, such that the storage medium containing instructions comprises an article of manufacture including instructions for implementing the function / act specified in the block or blocks of the flowchart and / or block diagram.
[0031] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the block or blocks of the flowchart and / or block diagram.
[0032] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the drawings. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of a block diagram or flowchart diagram, or both, and combinations of blocks in a block diagram or flowchart diagram, or both, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.
[0033] Although specific embodiments of the present invention have been described, it will be understood by those skilled in the art that there are other embodiments that are equivalent to the described embodiments. Therefore, it should be understood that the present invention is not limited by the specific exemplary embodiments shown, but only by the scope of the appended claims.
Claims
1. A method for segmenting a 3D point cloud into segments of an object and a background, comprising: a processor tracking an object in a scene within a video stream as the camera moves around in the scene; progressively segmenting and updating the segmentation of each of a plurality of points included in a frame within the video stream into either a point corresponding to the object or a point corresponding to the background, based on a probability that the point corresponds to the object and a probability that the point corresponds to the background, wherein when the probability that the point corresponds to the object is higher, the point is identified as an inlier point for tracking the object over time, and when the probability that the point corresponds to the background is higher, the point is identified as an outlier point not used for tracking, and updating the segmentation; generating augmented reality content by assigning a descriptor of a visual appearance to a 3D model of the object obtained by tracking the inlier points over time; A method for performing the above steps.
2. The method according to claim 1, further comprising estimating a pose of the camera with respect to the 3D model of the object using the 3D model of the object.
3. The method according to claim 1, wherein the probability for each point is updated based on whether the point is determined to correspond to the object or to the background.
4. The method according to claim 3, wherein a default setting of the probability for each point is 0.5, and the probability is increased when the point is determined to correspond to the object, and the probability is decreased when the point is determined to correspond to the background.
5. A system comprising a processor configured to execute the method according to any one of claims 1 to 4.
6. A program for causing a processor to execute the method according to any one of claims 1 to 4.
7. A storage medium storing the program according to claim 6.
Citation Information
Patent Citations
Mobile object extraction device, method, and program
JP2017016333A
Moving Object Segmentation Using Depth Images
US20120195471A1
Apparatus and method for foreground object segmentation
US20150269739A1
System and method for detecting, tracking, and classifiying objects
US20160210512A1