Three-dimensional object portion segmentation using machine learning models
Through the combination of visual language pre-training model and the three-dimensional fusion engine, two-dimensional bounding boxes and semantic marking super points are generated, which solves the accuracy and efficiency of partial segmentation of three-dimensional objects and realizes efficient three-dimensional partial segmentation.
Patent Information
- Application Number
- CN202380073537.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-01
- Filing Date
- 2023-09-13
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively identify and segment key parts of three-dimensional objects, especially in the absence of training data, and there are challenges in semantic segmentation and depth estimation in two-dimensional prediction tasks.
The visual language pre-training model is used to combine the three-dimensional fusion engine, and the super points are merged to generate a three-dimensional point cloud to achieve three-dimensional partial segmentation.
It improves the accuracy and efficiency of partial segmentation of three-dimensional objects, and can achieve high-precision partial segmentation under a small amount of labeled data, which is suitable for a variety of application scenarios.
Smart Images

Figure CN120303696A_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure generally relate to object segmentation. For example, aspects of the present disclosure relate to systems and techniques for performing three-dimensional part segmentation (e.g., zero-shot and / or few-shot three-dimensional part segmentation) by applying three-dimensional fusion of three-dimensional data and two-dimensional data output from a machine learning model or system (e.g., a vision-language pre-trained model). Background Art
[0002] Three-dimensional object part segmentation may include using a three-dimensional representation of an object to identify and segment different parts of the object. For example, parts of a chair may include a backrest, armrests, a seat, chair legs, etc. In some cases, a system may be interested only in certain key parts (e.g., a handle of a cabinet, a button on an appliance, etc.). However, the system may not be trained or designed to identify or segment such key parts. Summary of the Invention
[0003] A simplified summary of one or more aspects related to the present disclosure is presented below. Accordingly, the following summary should not be considered an exhaustive overview of all contemplated aspects, nor should it be considered to identify key or critical elements of all contemplated aspects or to delineate the scope associated with any particular aspect. Thus, the sole purpose of the following summary is to present certain concepts related to one or more aspects related to the mechanisms disclosed herein in a simplified form prior to the detailed description presented below.
[0004] Systems and techniques for implementing zero-shot and few-shot three-dimensional object part segmentation using a vision-language pre-trained model or a similar model are disclosed. According to at least one example, an apparatus for performing part segmentation is provided. The apparatus includes at least one memory and at least one processor, the processor being coupled to the at least one memory and configured to: generate one or more two-dimensional images of an object from a three-dimensional capture of the object; receive data identifying parts of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies parts of the object based on a vision-language pre-trained model and the data; perform part segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; perform semantic labeling on each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merge at least one subgroup of superpoints associated with parts of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and perform three-dimensional part segmentation of the parts of the object based on the three-dimensional point cloud.
[0005] In another example, a method for performing partial segmentation is provided. The method includes: generating one or more two-dimensional images of an object from a three-dimensional capture of the object; receiving data identifying a part of the object; processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; performing partial segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merging at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and performing three-dimensional partial segmentation of the part of the object based on the three-dimensional point cloud.
[0006] In another example, a non-transitory computer-readable medium is provided, having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to generate one or more two-dimensional images of an object from a three-dimensional capture of the object; receive data identifying a part of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; perform partial segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merging at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and performing three-dimensional partial segmentation of the part of the object based on the three-dimensional point cloud.
[0007] In another example, an apparatus for performing partial segmentation is provided. The apparatus includes: means for generating one or more two-dimensional images of an object from a three-dimensional capture of the object; means for receiving data identifying a part of the object; means for processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; means for performing partial segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; means for semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; means for merging at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and means for performing three-dimensional partial segmentation of the part of the object based on the three-dimensional point cloud.
[0008] In another example, an apparatus for performing part segmentation includes at least one memory and at least one processor, the at least one processor being coupled to the at least one memory and configured to: receive a three-dimensional image of an object; receive, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and perform three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0009] In another example, a method for performing part segmentation includes: receiving a three-dimensional image of an object; receiving, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and performing three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0010] In another example, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receive a three-dimensional image of an object; receive, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and perform three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0011] In another example, an apparatus for performing part segmentation includes: means for receiving a three-dimensional image of an object; and means for receiving, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and means for performing three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0012] In some aspects, one or more of the devices described herein are and / or include the following and / or are part of the following: an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device, a wireless communication device, a camera, a personal computer, a laptop computer, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one camera or multiple cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-described device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyroscopic testers, one or more accelerometers, any combination thereof, and / or other sensors).
[0013] The features and technical advantages of examples in accordance with the present disclosure have been outlined rather broadly above so that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for accomplishing the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein, both as to their organization and method of operation, as well as associated advantages, will be better understood from the following description when considered in conjunction with the accompanying drawings. Each of the drawings provided is for the purpose of illustration and description and is not a definition of the limits of the claims.
[0014] While aspects are described herein by way of illustration of some examples, those skilled in the art will understand that such aspects can be implemented in many different arrangements and scenarios. The techniques described herein can be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects can be embodied via an integrated chip or other non-module-component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / shopping devices, medical devices, and / or artificial intelligence devices). Aspects can be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and / or system-level components. Devices incorporating the described aspects and features can include additional components and features for implementing and practicing the claimed and described aspects. The aspects described herein are intended to be practiced in a variety of devices, components, systems, distributed arrangements, and / or end-user devices of various sizes, shapes, and configurations.
[0015] Based on the drawings and the detailed description, other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0016] The foregoing and other features and aspects will become more apparent when reference is made to the following specification, claims, and appended drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings are presented to assist in describing the various aspects of the present disclosure, and the drawings are provided for illustration only and not to limit the aspects.
[0018] Figure 1 Illustrates the usage of robotic manipulation where different parts of an object can be accurately identified to improve the object, according to some examples;
[0019] Figure 2 Illustrates a seat and a bounded region identifying a part on the seat, according to some examples;
[0020] Figure 3A Is a diagram illustrating an example system for performing three-dimensional object part segmentation, according to some examples;
[0021] Figure 3B Is a diagram illustrating an example system for performing three-dimensional object part segmentation using a cabinet as an input, according to some examples;
[0022] Figure 4Illustrate a seat according to some embodiments, where different parts of a specific portion are highlighted.
[0023] Figure 5 Illustrate various bounding boxes identifying seat portions according to some examples;
[0024] Figure 6 Is a diagram illustrating an aggregation process associated with three-dimensional object part segmentation according to some examples;
[0025] Figure 7 Is a diagram illustrating a different part segmentation process between the disclosed principles and traditional methods according to some examples;
[0026] Figure 8 Is a diagram illustrating a hint tuning result associated with the object segmentation process disclosed herein according to some examples;
[0027] Figure 9 Is a flowchart illustrating an example of a process for performing three-dimensional object part segmentation according to some examples;
[0028] Figure 10 Is a flowchart illustrating another example of a process for performing three-dimensional object part segmentation according to some examples; and
[0029] Figure 11 Is a block diagram illustrating an example of a computing system according to some examples. Detailed Description
[0030] For illustrative purposes, certain aspects of the present disclosure are provided below. Alternative aspects can be designed without departing from the scope of the present disclosure. Additionally, well-known elements of the present disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the present disclosure. Some aspects described herein can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation to provide a thorough understanding of the aspects of the present application. However, it is evident that the various aspects can be practiced without these specific details. The drawings and the description are not intended to be restrictive.
[0031] The following description only provides example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the example aspects will provide those skilled in the art with a description that can be used to implement the example aspects. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the scope of the present application as set forth in the appended claims.
[0032] Object segmentation can be difficult in some cases, such as when trying to identify a certain part or certain parts of an object. For example, there may be a lack of labeled data for training a machine learning model (e.g., a neural network model) to perform three-dimensional part segmentation, such as in terms of the quantity of available data and the categories of parts. The lack of such data poses challenges to learning-based methods (e.g., supervised learning or training of a neural network model). For example, standard supervised training only allows a machine learning model (e.g., a neural network designed to process 3D point cloud input) to identify object parts observed during training (e.g., those object parts labeled in the training image dataset). In an illustrative example, the training dataset of an image may only include twenty object part categories. A model trained to perform object segmentation using such a dataset (e.g., a conventional 3D neural network) cannot identify or segment new parts not included in the dataset. Even for parts included in the training dataset, due to the various object and / or part appearances and the lack of training data, the recognition performance may not be satisfactory during testing / deployment.
[0033] In addition, tasks involving two-dimensional (2D) prediction, such as semantic segmentation and depth estimation, may be problematic. A monocular network consumes a single image and outputs a prediction for that single image. In some scenarios, multiple posed images with overlapping views (e.g., measured camera positions and / or orientations for each image) may be obtained, and several of these images can be collectively processed by a machine learning model or system. However, existing monocular networks cannot utilize such posed images because they can only operate on a single image.
[0034] Part segmentation information can be greatly beneficial to various systems, such as robotic manipulation systems (e.g., for scene navigation, object manipulation, etc.), vehicle systems (e.g., for autonomous or semi-autonomous driving, safety warning systems, etc.), and so on. For example, the segmentation information can allow or help a robot identify where to apply force / action to move an object, such as moving a seat by pushing it in a specific direction, opening a cabinet door by grasping a handle, etc. Figure 1Illustrate the usage of robot manipulation where the manipulation of an object can be improved when different parts of the object can be accurately identified. As shown in the figure, the robot 100 includes the robot 102 shown in two views, which manipulates an object such as a drawer in the cabinet 104a or the cabinet 104b; two views of the robot 106, which manipulates an object such as the container 108 (shown in two views); two views of the robot 110, which manipulates an object such as a door in the cabinet 112a or the cabinet 112b; and two views of the robot 114, which manipulates an object such as the seat 116a or the cabinet drawer 116b. Each of the objects 104, 108, 112a, 112b, and 116 is different and has different parts, such as a handle or a lid or an armrest or a chair leg. Depending on the task of the robot, it is important for the robot to correctly identify the respective parts so that they can be manipulated by the robot.
[0035] A solution for improved part segmentation is needed, which can utilize various types of models, even models with the limitations described above. The systems and techniques described herein can utilize pre-trained machine learning models (e.g., models for visual language pre-training (VLP)) to perform three-dimensional (3D) part segmentation (e.g., zero-shot and / or few-shot 3D part segmentation). For example, machine learning models (e.g., VLP models) can be widely pre-trained on two-dimensional (2D) image-text data. In one example, given an image, the VLP model can identify one or more objects and one or more parts of the one or more objects based on the provided data or text that can identify one or more parts. In some cases, the systems and techniques can be applied to any 2D visual prediction task, such as semantic segmentation and depth estimation. Visual prediction tasks can be a component of many applications or systems, such as extended reality (XR), vehicle systems (e.g., autonomous or semi-autonomous driving, safety systems, etc.), camera image / video processing, robots (e.g., as Figure 1 shown) and / or other applications or systems.
[0036] One example usage where the systems and techniques described herein can be applied is human body part segmentation. The ability to identify different parts of the human body (e.g., arms, legs, head, etc.) is useful for many applications or tasks (e.g., extended reality (XR), vehicle systems, medical applications, etc.). In an illustrative example, 3D part segmentation can be used to segment different parts of a person depicted in one or more images, and virtual clothing / attachments can be placed on the person's parts. In another illustrative example, 3D part segmentation can be performed to segment different parts of a vehicle depicted in one or more images, which can allow a robotic system to identify different parts of the vehicle and facilitate assembly and repair tasks of the vehicle in the physical world. Other usage cases can include construction, furniture manufacturing and manipulation, equipment manufacturing and manipulation, precision spraying and coating, etc.
[0037] The pre-trained model can include any type of machine learning model. Non-limiting illustrative examples of open-source models include CLIP (Contrastive Language-Image Pretraining), GLIP (General Language-Image Pretraining). The output of the pre-trained model (e.g., VLP model) is two-dimensional and can include bounding regions, such as bounding boxes or a bounding zone or region having another shape. For example, the output of the pre-trained model is not three-dimensional. Figure 2 Illustrate the seat 200 and the bounded region 202 (shown as a bounding box) that identifies a part (the back of the seat) on the identification seat 201. Of course, other parts of the seat can also be identified by the bounding region 202, such as the seat, chair legs, rollers, and so on. The system and technique can include a 3D fusion engine that performs 3D part segmentation based on the 2D output from the model.
[0038] Figure 3A FIG. is a diagram illustrating an example of a system 300 for performing 3D object part segmentation. Given a 3D capture 302 of an object, the system 300 can segment the requested part without having to train a machine learning model. In some cases, the system 300 can use a pre-trained machine learning model 306, such as a VLP model or other types of machine learning models. In certain aspects, a small amount of labeled data can be used to improve the model (e.g., by using a small amount of labeled data, such as using supervised learning, to further train the pre-trained model to fine-tune the model).
[0039] The system 300 can receive a 3D capture 302 of an object as input. The process can include rendering multiple views from the 3D capture Figure 2The D-image 304. The system 300 can also receive partial data 308 (e.g., such as text) associated with parts such as the "backrest" or "legs" of a seat. In some cases, the system 300 can receive other data, such as data identifying a specific part (e.g., image, audio data, video data, or other types of data). The segmentation can be considered "0-shot", which means that no data of any kind of label is given in the target task domain. For example, if the task is to identify a handle, there is no labeled data for the "handle" part. The system 300 can generate multi-view Figure 2 from the 3D capture 302 of the object to obtain the D-image 304. Using the multi-view Figure 2 D-image 304 and the partial data 308, the machine learning model 306 (e.g., a VLP model) can determine or generate one or more 2D bounding boxes 314 that identify the parts. The one or more 2D bounding boxes 314 are provided to the 3D fusion engine 312 that also receives the 3D capture 302 of the object. The 3D fusion engine 312 can output a 3D point cloud 310 of the object with partial segmentation.
[0040] To generate the 3D point cloud 310 with partial segmentation, the 3D fusion engine 312 can perform one or more operations. For example, the 3D fusion engine 312 can perform over-segmentation. To perform over-segmentation, the 3D fusion engine 312 can segment the 3D capture 302 (which can be assumed to be a point cloud) into "superpoints" (e.g., sub-parts) based on the corresponding 3D normals at each point in the point cloud of the 3D capture 302. Figure 4 is a view illustrating the 3D capture of the seat 400. For example, Figure 4 the object parts of the seat 400 (e.g., the backrest 404, the seat 406, the legs 408, etc.) can be segmented into multiple sub-parts during over-segmentation. In an illustrative example, the leg 408 can be segmented into a first sub-part 410 and a second sub-part 412. The seat 406 or any other part of the seat 400 can also be segmented into sub-parts.
[0041] The 3D fusion engine 312 can also perform semantic labeling of each superpoint (or in other cases less than all superpoints) based on the 2D bounding boxes 314 output from the machine learning model 306. For each part (p) category i and each superpoint (sp) j, a score can be calculated using the image k, such as using the following equation:
[0042]
[0043] The equation uses "bbox" to represent the bounding box. The fraction indicates the amount by which the corresponding superpoint is included in one or more bounding boxes associated with the corresponding part category in one or more 2D images (e.g., how much the superpoint is covered by the bounding box of a certain category across all views), where the one or more 2D images include parts associated with the corresponding part category. The 3D fusion engine 312 can normalize the fraction [i]. For each superpoint j, the 3D fusion engine 312 can label it as argmax i (score[i,j]).
[0044] The 3D fusion engine 312 can also merge superpoints belonging to the same part. For example, if two superpoints have the same semantic label and if the two superpoints are adjacent to each other in the 3D capture 302, the 3D fusion engine 312 can merge the two superpoints. In another case, the 3D fusion engine 312 can merge two superpoints where, for the views in which the two superpoints are visible (e.g., images from different views), given each bounding box, both superpoints are either within the bounding box or not within the bounding box.
[0045] Figure 3B is an example system 330 that illustrates an example system for performing 3D object part segmentation using the 3D capture 302 of the cabinet 334 (e.g., the point cloud rendering of an object generated using data from the camera 336 or other sensors) as input. The zero-shot method provided by the system 332 can include receiving the 3D capture 302 as input and generating 2D images 338 that are provided to a machine learning model 342 (e.g., a VLP model). Partial text 344 (e.g., which can be a text prompt or other data) can also be provided as input, as previously described.
[0046] In some cases, an inter-view consistency process 340 for providing inter-view consistency (also referred to as inter-view aggregation or feature aggregation) can be performed, as Figure 6As further shown. For example, the inter-view consistency processing 340 can provide multi-view feature aggregation in three dimensions to enhance a single-view machine learning model (e.g., a neural network) for 2D prediction tasks. The method can use non-learning aggregation schemes, such as feature averaging / merging, which can be used directly at inference time on any existing monocular network for 2D prediction tasks. The inter-view consistency processing 340 can be applied to any layer of the network. The method is a general technique for improving the performance of any 2D dense prediction network. In some cases, short sequences of images can be used as input data (e.g., short videos, image bursts, etc.). The camera pose can also be obtained via one or more sensors (e.g., inertial measurement unit (IMU), accelerometer, gyroscope, etc.), a visual-inertial odometry (VIO) system, and / or other devices or systems. Therefore, the proposed scheme can be widely applied to different scenarios and tasks.
[0047] In some aspects, the system can utilize a learning-based aggregation scheme, e.g., a small neural network that can be trained together with the monocular network. When testing the method, performance improvements have been shown. For the 2D detection task, via this method, the mAP50 has increased from 0.68 to 0.74. "mAP50" is the average of average precisions, and its bounding box matching threshold is the mean intersection over union (mIoU) = 50%.
[0048] Although numerically evaluated for 2D detection tasks, the proposed scheme can be used for any 2D dense prediction task, such as segmentation, depth estimation, denoising, etc.
[0049] The output 346 of the machine learning model 342 (e.g., a VLP model) can include various images with detected boxes, and the detected boxes are provided to the 3D fusion engine 312 that fuses 2D bounding boxes into 3D segmentation. The output is a 3D point cloud 348 with partial segmentation. A tuning process 349 can be further implemented to output a 3D shape with ground truth segmentation 350. This tuning process 349 can be referred to as a "few-shot" or "few-shot prompt tuning" method, where an input is provided for tuning.
[0050] The methods disclosed herein utilize the machine learning models 306 / 342 to achieve zero-shot 3D partial segmentation. Additional innovations include using multiple views of the object of interest to improve accuracy and providing an efficient tuning process 349. The scheme can consume a very small amount of labeled 2D data to significantly improve accuracy (e.g., using few-shot 3D partial segmentation).
[0051] The performance of system 330 can be measured by the mean intersection over union (mIoU) in terms of segmentation accuracy. The IoU value is defined as the intersection of the predicted ground truth (GT) segmentation mask divided by the union of the predicted GT segmentation mask. In a test, semantic segmentation mIoU was performed on 81 chair shapes with 1 shot taken, and the accuracy of different parts ranged from 63% to over 88%.
[0052] Figure 5 Illustrative identification of various bounding boxes of seat portion 500 and shows how multi-view feature aggregation can occur. The seat is shown in various views, and the machine learning model 306 / 342 may predict good results in some views while predicting poor results in other views. Seat 502 shows one bounding box 504 for the entire seat and another bounding box 506 for the seat base. In Figure 5 each bounding box of the bounding boxes, it shows how well each corresponding bounding box captures the corresponding score or confidence of a particular part. For example, the score or confidence of bounding box 504 is 0.60, indicating how well it captures the seat base. The score or confidence can be the corresponding score / confidence of the corresponding part category among multiple part categories. This score or confidence can indicate the amount of corresponding superpoints included in one or more bounding boxes associated with the corresponding part category in a two-dimensional image including the part associated with the corresponding part category. Bounding box 506 has a score or confidence of 0.52. Seat 508 shows one bounding box 512 for the side view of the seat and another bounding box 510 for the seat base. Seat 514 shows one bounding box 516 for the seat from the rear view. Seat 518 shows one bounding box 520 for the seat in a side view and another bounding box 522 for the seat base. Seat 524 has one bounding box 526 for the top view, and seat 528 shows one bounding box 530 for the side view. The input 3D capture 302 to the disclosed system is a 3D shape, and there is a correspondence between the multi-point cloud rendering and the 2D output of the machine learning model 306 / 342.
[0053] Figure 6FIG. 600 illustrates an aggregation process associated with 3D object part segmentation. Given n feature maps (e.g., 50x50) from n views, the process can generate n fused feature maps. For each corresponding 3D point (e.g., see the points identified on seat views 602, 604, 606, 608, 610), the process can average the corresponding 2D pixel features across different views. The 3D pixel feature refers to the feature associated with the 2D pixel position on the image. For each 2D pixel, there can be multiple 3D points, and the process can average the corresponding point features. The point feature refers to the feature associated with the 3D point. For example, it can be the averaged 2D features from different views of the 3D point. Since for each 2D pixel, there may be multiple associated 3D points, the method averages the point features from the 3D points to obtain the feature of the 2D pixel. Optional methods also include ignoring boundary pixels. Graph 612 shows the aggregation of data on the 3D points.
[0054] Figure 7 FIG. 700 illustrates a different part segmentation process between using the disclosed principles and traditional methods. A method without fusion 702 can result in example views as shown in the figure, such as seat 706 with bounding boxes 710 and 708 associated with the armrest. Seat 712 can have a bounding box 716 associated with the seat and another bounding box 713. Seat 718 can have a bounding box 720 and another bounding box 722, each associated with the backrest. Seat 724 can have a bounding box 726 associated with the backrest and another bounding box 728. As Figure 7 shown, bounding box 728 appears to be more focused on the armrest of the seat rather than the backrest of the seat. Seat 730 can have a bounding box 732 for the chair leg. The various seats and bounding boxes without fusion 702 illustrate how multiple bounding boxes can be generated, which may lead to confusion regarding the specific part to which the corresponding bounding box should be associated.
[0055] With the fusion process 704, the results can be better. Seat 734 has a bounding box 736 for the armrest. Seat 738 has a bounding box 740 for the seat. Seat 742 shows a bounding box 744 for the backrest. Seat 746 shows a bounding box 748 for the backrest. Seat 750 shows a bounding box 752 for one leg, another bounding box 754 for another leg, and a third bounding box 756 for another leg. When using fusion, the improvement in the accuracy of the bounding box associated with a specific part is evident.
[0056] Figure 8 FIG. 800 illustrates the tuning results associated with the object segmentation process disclosed herein. Figure 8 The tuning results of Figure 3BResults of the tuning process 349) described. A set of images illustrate object segmentation before tuning 802. Teapot 806 includes a mouth bounding box 808 that covers the entire image. A probability or success value of 0.55 for the mouth of the teapot appropriately covered by the bounding box 808 is shown. Teapot 810 includes a bounding box 812 for the handle and another box 814 for the entire teapot. Pliers 824 include a handle bounding box 816 for the entire set of pliers, and pliers 834 also include a handle bounding box 836 for the handle and another smaller and more concentrated handle bounding box 837. After only one shape adjustment 804, the system can learn the text mapping and can generalize it to other instances. Teapot 826 is shown in a more accurate spout bounding box 818 that covers the spout of the teapot rather than the entire teapot, with a score of 0.89. Teapot 820 has a spout bounding box 822 that is more accurate, with a score of 0.94. Pliers 828 have a first bounding box 830 and a second bounding box 832, each surrounding a handle with scores of 0.74 and 0.71. Pliers 838 include a handle bounding box 840 that more accurately surrounds the two handles of pliers 838 (score of 0.70) and another box 842 that surrounds one handle of pliers 838 (score of 0.71).
[0057] Figure 9 is a flowchart example illustrating a process 900 for performing three-dimensional object part segmentation. Operations of process 900 can be implemented as software components executed and run on one or more processors (e.g., Figure 11 processor 1110 and / or other processors).
[0058] At block 902, process 900 can include generating one or more two-dimensional images of an object from a three-dimensional capture of the object (e.g., Figure 3A and / or Figure 3B 3D capture 302). The three-dimensional capture can include an initial three-dimensional point cloud. In one aspect, process 900 can include obtaining a three-dimensional capture of the object.
[0059] At block 904, process 900 can include receiving data identifying a part of the object (e.g., Figure 3A part text 308). The data can be in any form. One example form is text identifying a part of the object. The three-dimensional capture can be a three-dimensional point cloud, and process 900 can include generating a modified three-dimensional point cloud based on three-dimensional part segmentation of at least one part of the object (e.g., Figure 3A and / or Figure 3B 3D point cloud 310 / 334).
[0060] At block 906, process 900 may include processing one or more two-dimensional images of an object to generate at least one two-dimensional bounding box (e.g., Figure 3A and / or Figure 3B 2D bounding boxes 314 / 346), where the at least one two-dimensional bounding box identifies parts of the object based on a vision-language pre-trained model and data. For example, the at least one two-dimensional bounding box may identify at least one part of the object based on the training of a machine learning model (e.g., Figure 3A and / or Figure 3B models 306 / 342), as shown in, for example, Figure 3B . In one illustrative example, the machine learning model is a vision-language pre-trained (VLP) model.
[0061] In some cases, the machine learning model may be pre-trained on two-dimensional images and text data (e.g., Figure 3A and / or Figure 3B partial data or text 308 / 344), as described herein. For example, the text / data may identify at least one part of the object. In one illustrative example, process 900 may include receiving data or text identifying an interesting part of an object configured with an object. Process 900 may process one or more two-dimensional images of the object based on the received data or text to generate at least one two-dimensional bounding box associated with at least one part of the object.
[0062] At block 908, process 900 may include performing partial segmentation of a three-dimensional capture of an object to generate a plurality of superpoints associated with the object. The plurality of superpoints may include sub-parts of parts of the object based on the three-dimensional normal at each point of the plurality of superpoints. In one aspect, based on the respective three-dimensional normals associated with each point in the plurality of superpoints, the plurality of superpoints may be associated with one or more sub-parts of the object. Based on performing three-dimensional partial segmentation of at least one part of the object, process 900 may output a three-dimensional representation (e.g., 3D point cloud) of the object with partial segmentation.
[0063] At block 910, process 900 may include semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; in one aspect, as part of process 900, semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box may include generating a corresponding score for a corresponding part category in a plurality of part categories and a corresponding superpoint associated with the corresponding part category. In one example, the corresponding score may indicate the amount by which the corresponding superpoint is included in one or more bounding boxes associated with the corresponding part category in at least one of the one or more two-dimensional images, where the at least one two-dimensional image includes a part associated with the corresponding part category.
[0064] At block 912, process 900 may include merging at least one subgroup of hyperpoints associated with a portion of the object from multiple hyperpoints based on hyperpoints of multiple semantic tags to generate a three-dimensional point cloud. In one aspect, as part of process 900, merging at least one subgroup of hyperpoints associated with a portion of the object may include merging two hyperpoints based on at least one of the following: two hyperpoints having the same semantic tag, a first hyperpoint of the two hyperpoints being adjacent to a second hyperpoint of the two hyperpoints, or whether the two hyperpoints are included in corresponding bounding boxes of a two-dimensional image in one or more two-dimensional images including the two hyperpoints.
[0065] At block 914, process 900 may include performing three-dimensional part segmentation on a portion of the object based on the three-dimensional point cloud. In one aspect, the three-dimensional point cloud may be generated based on an initial three-dimensional point cloud. In one aspect, performing three-dimensional part segmentation of a portion of the object further includes performing multi-view feature aggregation. In another aspect, performing multi-view feature aggregation as part of process 900 further includes averaging corresponding two-dimensional pixel features on one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
[0066] In some aspects, process 900 for performing multi-view feature aggregation further includes averaging corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in one or more two-dimensional images of the object.
[0067] In another aspect, an apparatus for performing part segmentation includes at least one memory; and at least one processor coupled to the at least one memory and configured to: generate one or more two-dimensional images of an object from a three-dimensional capture of the object; receive data identifying a portion of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the portion of the object based on a vision-language pre-training model and the data; perform part segmentation of the three-dimensional capture of the object to generate a plurality of hyperpoints associated with the object; perform semantic tagging on each hyperpoint of the plurality of hyperpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically tagged hyperpoints; merge at least one subgroup of hyperpoints associated with the portion of the object from the plurality of hyperpoints based on the plurality of semantically tagged hyperpoints to generate a three-dimensional point cloud; and perform three-dimensional part segmentation on the portion of the object based on the three-dimensional point cloud.
[0068] Figure 10 is another flowchart illustrating an example of process 1000 for performing three-dimensional object part segmentation. Process 1000 may be performed by any device or group of devices. Operations of process 1000 may be implemented as software components executed and run on one or more processors (e.g., Figure 11 processor 1110 and / or other processors).
[0069] At block 1002, process 1000 may include receiving a three-dimensional image of an object. At block 1004, process 1000 may include receiving, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one portion of the object; at block 1006, process 1000 may include performing three-dimensional part segmentation of at least one portion of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0070] In some examples, the processes described herein (e.g., process 900, process 1000, and / or other processes described herein) may be performed by a computing device or apparatus (e.g., a network node such as a UE, a base station, a portion of a base station, etc.). For example, as described above, process 900 may be performed by a UE, and process 1000 may be performed by a base station or a portion of a base station. In another example, process 900 and / or process 1000 may be performed by a computing device having Figure 11 the computing system 1100 shown in. For example, a wireless communication device having Figure 11 the computing architecture shown may include components of a UE and may implement Figure 9 and / or Figure 10 the operations of.
[0071] In some cases, the computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive data, any combination thereof, and / or other components. The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to 3G, 4G, 5G, and / or other cellular standards, data according to the WiFi (802.11x) standard, data according to the Bluetooth TM standard, data according to the Internet Protocol (IP) standard, and / or other types of data.
[0072] The components of a computing device may be implemented in circuitry. For example, the components may include, and / or may be implemented using, electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0073] Process 900 and Process 1000 are illustrated as logic flowcharts, the operations of which represent a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. In general, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0074] Additionally, Process 900, Process 1000, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that execute jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program that includes a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0075] Figure 11 is a diagram illustrating an example of a system for implementing certain aspects of the disclosed technology. Specifically, Figure 11 illustrates an example of a computing system 1100, which may be any computing device, such as a component that constitutes an internal computing system, a remote computing system, a camera, or any combination thereof, where the components of the system communicate with each other using connection 1105. Connection 1105 may be a physical connection using a bus or a direct connection into a processor 1110, such as in a chipset architecture. Connection 1105 may also be a virtual connection, a networked connection, or a logical connection.
[0076] In some aspects, computing system 1100 is a distributed system, where the functions described in this disclosure can be distributed within one data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent many such components that each perform some or all of the functions the component is described for. In some aspects, a component can be a physical or virtual device.
[0077] Example system 1100 includes at least one processing unit (CPU or processor) 1110 and a connection 1105 that communicatively couples various system components including system memory 1115 (such as read-only memory (ROM) 1120 and random access memory (RAM) 1125) to the processor 1110. Computing system 1100 can include a cache 1115 that is directly connected to, in close proximity to, or integrated as part of the processor 1110 with high-speed memory.
[0078] Processor 1110 can include any general-purpose processor and hardware services or software services such as services 1132, 1134, and 1136 stored in storage device 1130, which are configured to control processor 1110 and a dedicated processor in which software instructions are incorporated into the actual processor design. Processor 1110 can be substantially a self - contained computing system that includes multiple cores or processors, buses, memory controllers, caches, etc. A multi - core processor can be symmetric or asymmetric.
[0079] To enable user interaction, computing system 1100 includes an input device 1145 that can represent any number of input mechanisms, such as a microphone for voice, a touch - sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1100 can also include an output device 1135 that can be one or more of a plurality of output mechanisms. In some cases, a multimode system can enable a user to provide multiple types of input / output to communicate with computing system 1100.
[0080] Computing system 1100 can include a communication interface 1140 that generally can govern and manage user input and system output. The communication interface can execute or facilitate receiving and / or sending wired or wireless communications using a wired and / or wireless transceiver, including using an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Apple TM Lightning TM port / plug, an Ethernet port / plug, a fiber - optic port / plug, a dedicated wired port / plug, 3G, 4G, 5G, and / or other cellular data network wireless signaling, Bluetooth TM wireless signaling, BluetoothTM Low Energy (BLE) wireless signal transmission, iBeacon TM Wireless signal transmission, Radio Frequency Identification (RFID) wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short Range Communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, Wireless Local Area Network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, Ad-hoc network signal transmission, Radio wave signal transmission, Microwave signal transmission, Infrared signal transmission, Visible light signal transmission, Ultraviolet light signal transmission, Wireless signal transmission along the electromagnetic spectrum, or those communications of some combination thereof. The communication interface 1140 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1100 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the Global Positioning System (GPS) of the United States, the Global Navigation Satellite System (GLONASS) of Russia, the Beidou Navigation Satellite System (BDS) of China, and the Galileo GNSS of Europe. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0081] The storage device 1130 can be a non-volatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as cassette tapes, flash memory cards, solid-state memory devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic strips / magnetic stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, Compact Disc Read Only Memory (CD-ROM) optical discs, Rewritable Compact Disc (CD) optical discs, Digital Video Disc (DVD) optical discs, Blu-ray Disc (BDD) optical discs, holographic optical discs, another optical medium, Secure Digital (SD) cards, Micro Secure Digital (microSD) cards, Memory Cards, smart card chips, EMV chips, subscriber identity module (SIM) cards, mini / micro / nano / pico SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., level 1 (L1) cache, level 2 (L2) cache, level 3 (L3) cache, level 4 (L4) cache, level 5 (L5) cache, other (L#) cache), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or combinations thereof.
[0082] The storage device 1130 may include software services, servers, services, etc., and when the code defining such software is executed by the processor 1110, the code causes the system to perform functions. In some aspects, the hardware services that perform specific functions may include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1110, connection 1105, output device 1135, etc.) to perform functions. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media, in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, and the code and / or machine-executable instructions may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment may be coupled to another code segment or hardware circuit. Information, arguments, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0083] Specific details are provided in the above description to provide a thorough understanding of the various aspects and examples presented herein, but those skilled in the art will recognize that this application is not limited thereto. Thus, although the illustrative aspects of this application have been described in detail herein, it is to be understood that the various inventive concepts may be embodied and employed in other various ways, and the appended claims are not to be construed as including such variations unless limited by the prior art. The various features and aspects of the above applications may be used singly or in combination. Additionally, the aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be appreciated that in alternative aspects, the methods may be performed in an order different from that described.
[0084] For clarity of explanation, in some instances, the present technology may be presented as including separate functional blocks that include devices, device components, steps, or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the aspects.
[0085] Moreover, those skilled in the art will understand that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure.
[0086] Each aspect described above may be described as a process or method that is depicted as a flowchart, a process diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations in the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operations of the process are completed, but the process may have additional steps not included in the figures. The process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When the process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0087] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessed through a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0088] In some aspects, computer-readable storage devices, media, and memories may include wires or wireless signals that contain bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0089] Those skilled in the art should understand that information and signals may be represented using any of a variety of different technologies and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may, in some cases, be represented, in part, depending on the specific application, in part, on the desired design, in part, on the corresponding technology, etc., by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0090] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take on any form factor. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may execute the necessary tasks. Examples of form factors include: laptop devices, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functionality may also be implemented on a circuit board in different chips or different processes executed on a single device.
[0091] Instructions, the medium for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0092] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices with multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium including program code that includes instructions for performing one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form a part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0093] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; however, in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.
[0094] Those of ordinary skill in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein can be replaced, respectively, with the less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this specification.
[0095] In cases where a component is described as "configured to" perform certain operations, such configuration can be implemented, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0096] The phrase "coupled to" or "communicatively coupled to" refers to any component being physically connected, directly or indirectly, to another component, and / or any component being in communication, directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).
[0097] Claim language that recites "at least one" of a set and / or "one or more" of a set, or other language indicating that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0098] Exemplary aspects of the present disclosure include:
[0099] Aspect 1. An apparatus for performing partial segmentation, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: obtain a three-dimensional capture of an object; generate one or more two-dimensional images of the object from the three-dimensional capture of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box associated with at least one part of the object; and perform three-dimensional partial segmentation of the at least one part of the object based on the one or more two-dimensional images of the object and the at least one two-dimensional bounding box.
[0100] Aspect 2. The apparatus according to aspect 1, wherein the at least one two-dimensional bounding box identifies the at least one part of the object based on training of a machine learning model.
[0101] Aspect 3. The apparatus according to aspect 2, wherein the machine learning model is a vision-language pre-training (VLP) model.
[0102] Aspect 4. The apparatus according to any one of aspects 2 or 3, wherein the machine learning model is pre-trained on two-dimensional images and text data.
[0103] Aspect 5. The apparatus according to aspect 4, wherein the text data identifies the at least one part of the object.
[0104] Aspect 6. The apparatus according to any one of aspects 1 to 5, wherein at least one processor is further configured to: receive text identifying an interesting part of the object configured; and process the one or more two-dimensional images of the object to generate the at least one two-dimensional bounding box associated with at least one part of the object based on the received text.
[0105] Aspect 7. The apparatus according to any one of aspects 1 to 6, wherein the three-dimensional capture is a three-dimensional point cloud, and wherein at least one processor is further configured to: generate a modified three-dimensional point cloud based on the three-dimensional part segmentation of at least one part of the object.
[0106] Aspect 8. The apparatus according to any one of aspects 1 to 7, wherein in order to perform the three-dimensional part segmentation of the object, the at least one processor is configured to: segment the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically label each of the plurality of superpoints based on the at least one two-dimensional bounding box; and merge superpoints of at least one subgroup of the plurality of superpoints associated with the same part.
[0107] Aspect 9. The apparatus according to aspect 8, wherein the plurality of superpoints are related to one or more sub-parts of the object based on corresponding three-dimensional normals associated with each point of the plurality of superpoints.
[0108] Aspect 10. The apparatus according to any one of aspects 8 or 9, wherein in order to semantically label each of the plurality of superpoints based on the at least one two-dimensional bounding box, the at least one processor is configured to generate a corresponding score for a corresponding part category among a plurality of part categories and a corresponding superpoint associated with the corresponding part category.
[0109] Aspect 11. The apparatus according to aspect 10, wherein the corresponding score indicates the amount of the corresponding superpoint included in one or more bounding boxes associated with the corresponding part category in at least one of the one or more two-dimensional images, and the at least one two-dimensional image includes a part associated with the corresponding part category.
[0110] Aspect 12. The apparatus according to any one of aspects 8 to 11, wherein in order to merge superpoints of at least one subgroup associated with the same part, the at least one processor is configured to merge two superpoints based on at least one of the following: the two superpoints having the same semantic label, a first superpoint of the two superpoints being adjacent to a second superpoint of the two superpoints, or whether the two superpoints are included in a corresponding bounding box of a two-dimensional image of the one or more two-dimensional images including the two superpoints.
[0111] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein, in order to perform the three-dimensional partial segmentation of the at least one part of the object, the at least one processor is configured to perform multi-view feature aggregation.
[0112] Aspect 14. The apparatus according to Aspect 13, wherein, in order to perform the multi-view feature aggregation, the at least one processor is configured to average the corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
[0113] Aspect 15. The apparatus according to Aspect 14, wherein, in order to perform the multi-view feature aggregation, the at least one processor is further configured to average the corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.
[0114] Aspect 16. An apparatus for performing partial segmentation, the apparatus comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and configured to: receive a three-dimensional image of an object; receive, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and perform three-dimensional partial segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0115] Aspect 17. A method for performing partial segmentation, the method comprising: obtaining a three-dimensional capture of an object; generating one or more two-dimensional images of the object from the three-dimensional capture of the object; processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box associated with at least one part of the object; and performing three-dimensional partial segmentation of the at least one part of the object based on the one or more two-dimensional images of the object and the at least one two-dimensional bounding box.
[0116] Aspect 18. The method according to Aspect 17, wherein the at least one two-dimensional bounding box identifies the at least one part of the object based on training of a machine learning model.
[0117] Aspect 19. The method according to Aspect 18, wherein the machine learning model is a Visual Language Pretraining (VLP) model.
[0118] Aspect 20. The method according to any one of Aspects 18 or 19, wherein the machine learning model is pre-trained on two-dimensional images and text data.
[0119] Aspect 21. The method according to aspect 20, wherein the text data identifies the at least one part of the object.
[0120] Aspect 22. The method according to any one of aspects 17 to 21, the method further comprising: receiving text identifying an interesting part of the object configured with the object; and processing the one or more two-dimensional images of the object to generate the at least one two-dimensional bounding box associated with the at least one part of the object based on the received text.
[0121] Aspect 23. The method according to any one of aspects 17 to 22, wherein the three-dimensional capture is a three-dimensional point cloud, and the method further comprises: generating a modified three-dimensional point cloud based on the three-dimensional part segmentation of the at least one part of the object.
[0122] Aspect 24. The method according to any one of aspects 17 to 23, wherein performing the three-dimensional part segmentation of the object comprises: segmenting the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box; and merging the superpoints of at least one subgroup of the plurality of superpoints associated with the same part.
[0123] Aspect 25. The method according to aspect 24, wherein the plurality of superpoints are related to one or more sub-parts of the object based on the respective three-dimensional normals associated with each point of the plurality of superpoints.
[0124] Aspect 26. The method according to any one of aspects 24 or 25, wherein semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box comprises generating a respective score for a respective part category of a plurality of part categories and a respective superpoint associated with the respective part category.
[0125] Aspect 27. The method according to aspect 26, wherein the respective score indicates the amount by which the respective superpoint is included in one or more bounding boxes associated with the respective part category in at least one of the one or more two-dimensional images, the at least one two-dimensional image including a part associated with the respective part category.
[0126] Aspect 28. The method according to any one of aspects 24 to 26, wherein merging the at least one subgroup of superpoints associated with the same part includes merging two superpoints based on at least one of the following: the two superpoints having the same semantic label, a first superpoint of the two superpoints being adjacent to a second superpoint of the two superpoints, or whether the two superpoints are included in a corresponding bounding box of a two-dimensional image in the one or more two-dimensional images including the two superpoints.
[0127] Aspect 29. The method according to any one of aspects 17 to 28, wherein performing the three-dimensional part segmentation of the at least one part of the object includes performing multi-view feature aggregation.
[0128] Aspect 30. The method according to aspect 29, wherein performing the multi-view feature aggregation includes averaging corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
[0129] Aspect 31. The method according to aspect 30, wherein performing the multi-view feature aggregation includes averaging corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.
[0130] Aspect 32. A method for performing part segmentation, the method comprising: receiving a three-dimensional image of an object; receiving, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and performing three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0131] Aspect 33. The method according to aspect 32, the method further comprising performing the operations according to any one of aspects 17 to 31.
[0132] Aspect 34. An apparatus for performing part segmentation, the apparatus comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and configured to: receive a three-dimensional image of an object; receive, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and perform three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0133] Aspect 35. A method for performing part segmentation, the method comprising: receiving a three-dimensional image of an object; receiving, from the three-dimensional image of the object, one or more two-dimensional bounding boxes associated with one or more two-dimensional images of the object generated by a model, the one or more two-dimensional bounding boxes being associated with at least one part of the object; and performing three-dimensional part segmentation of the at least one part of the object based on the three-dimensional image of the object and the one or more two-dimensional bounding boxes.
[0134] Aspect 36. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of Aspects 17 to 33 and / or 35.
[0135] Aspect 37. An apparatus for generating virtual content in a distributed system, the apparatus comprising one or more components for performing the operations according to any one of Aspects 17 to 33 and / or 35.
[0136] Aspect 38. An apparatus for performing part segmentation, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate one or more two-dimensional images of an object from a three-dimensional capture of the object; receive data identifying a part of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; perform part segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; perform semantic tagging of each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically tagged superpoints; merge at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically tagged superpoints to generate a three-dimensional point cloud; and perform three-dimensional part segmentation of the part of the object based on the three-dimensional point cloud.
[0137] Aspect 39. The apparatus according to Aspect 38, wherein the data comprises text identifying the part of the object.
[0138] Aspect 40. The apparatus according to any one of Aspects 38 to 39, wherein the plurality of superpoints comprises sub-parts of the part of the object based on three-dimensional normals at each point of the plurality of superpoints.
[0139] Aspect 41. The apparatus according to any one of aspects 38 to 40, wherein the three-dimensional capture includes an initial three-dimensional point cloud, and wherein at least one processor is further configured to: generate the three-dimensional point cloud based on the initial three-dimensional point cloud and based on performing the partial segmentation.
[0140] Aspect 42. The apparatus according to any one of aspects 38 to 41, wherein the plurality of superpoints are associated with one or more sub-parts of the object based on corresponding three-dimensional normals associated with each of the plurality of superpoints.
[0141] Aspect 43. The apparatus according to any one of aspects 38 to 42, wherein in order to semantically label each of the plurality of superpoints based on the at least one two-dimensional bounding box, the at least one processor is configured to generate a corresponding score for a corresponding part category among a plurality of part categories and a corresponding superpoint associated with the corresponding part category.
[0142] Aspect 44. The apparatus according to any one of aspects 38 to 43, wherein the corresponding score indicates the amount by which the corresponding superpoint is included in one or more bounding boxes associated with the corresponding part category in at least one of the one or more two-dimensional images, the at least one two-dimensional image including a part associated with the corresponding part category.
[0143] Aspect 45. The apparatus according to any one of aspects 38 to 44, wherein in order to merge the superpoints of at least one subgroup associated with the part of the object, the at least one processor is configured to merge two superpoints based on at least one of the following: the two superpoints having the same semantic label, a first superpoint among the two superpoints being adjacent to a second superpoint among the two superpoints, or whether the two superpoints are included in a corresponding bounding box of a two-dimensional image among the one or more two-dimensional images including the two superpoints.
[0144] Aspect 46. The apparatus according to any one of aspects 38 to 45, wherein the at least one processor is configured to obtain the three-dimensional capture of the object.
[0145] Aspect 47. The apparatus according to any one of aspects 38 to 46, wherein in order to perform the three-dimensional partial segmentation of the part of the object, the at least one processor is configured to perform multi-view feature aggregation.
[0146] Aspect 48. The apparatus according to any one of aspects 38 to 47, wherein in order to perform the multi-view feature aggregation, the at least one processor is configured to average corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
[0147] Aspect 49. The apparatus according to aspect 48, wherein, to perform the multi-view feature aggregation, the at least one processor is further configured to average corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.
[0148] Aspect 50. A method for performing partial segmentation, the method comprising: generating one or more two-dimensional images of an object from a three-dimensional capture of the object; receiving data identifying a part of the object; processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; performing partial segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merging at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and performing three-dimensional partial segmentation of the part of the object based on the three-dimensional point cloud.
[0149] Aspect 51. The method according to aspect 50, wherein the data includes text identifying the part of the object.
[0150] Aspect 52. The method according to any one of aspects 50 to 51, wherein the plurality of superpoints includes sub-parts of the part of the object based on the three-dimensional normal at each point of the plurality of superpoints.
[0151] Aspect 53. The method according to any one of aspects 50 to 52, wherein the three-dimensional capture includes an initial three-dimensional point cloud, and wherein the method further includes: generating the three-dimensional point cloud based on the initial three-dimensional point cloud and based on performing the partial segmentation.
[0152] Aspect 54. The method according to any one of aspects 50 to 53, wherein the plurality of superpoints are related to one or more sub-parts of the object based on the respective three-dimensional normals associated with each point in the plurality of superpoints.
[0153] Aspect 55. The method according to any one of aspects 50 to 54, wherein semantically labeling each of the plurality of superpoints based on the at least one two-dimensional bounding box further includes generating a respective score for a respective part category among a plurality of part categories and a respective superpoint associated with the respective part category.
[0154] Aspect 56. The method according to aspect 55, wherein the respective score indicates the amount that the respective superpoint is included in one or more bounding boxes associated with the respective part category in at least one of the one or more two-dimensional images, and the at least one two-dimensional image includes a part associated with the respective part category.
[0155] Aspect 57. The method according to any one of aspects 50 to 56, wherein merging the at least one subgroup of superpoints associated with the part of the object further includes merging two superpoints based on at least one of the following: the two superpoints having the same semantic label, a first superpoint of the two superpoints being adjacent to a second superpoint of the two superpoints, or whether the two superpoints are included in a corresponding bounding box of a two-dimensional image among the one or more two-dimensional images including the two superpoints.
[0156] Aspect 58. The method according to any one of aspects 50 to 57, the method further comprising: obtaining a three-dimensional capture of the object.
[0157] Aspect 59. The method according to aspect 58, wherein performing the three-dimensional part segmentation of the part of the object further includes performing multi-view feature aggregation.
[0158] Aspect 60. The method according to aspect 59, wherein performing the multi-view feature aggregation further includes averaging corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
[0159] Aspect 61. The method according to aspect 60, wherein performing the multi-view feature aggregation further includes averaging corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.
[0160] Aspect 62. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of aspects 50 to 61.
[0161] Aspect 63. An apparatus for generating virtual content in a distributed system, the apparatus including one or more components for performing the operations according to any one of aspects 50 to 61.
Claims
1. An apparatus for performing part segmentation, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: generate one or more two-dimensional images of an object from a three-dimensional capture of the object; receive data identifying a part of the object; process the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box that identifies the part of the object based on a vision-language pre-trained model and the data; perform part segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically label each superpoint of the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merge at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints to generate a three-dimensional point cloud based on the plurality of semantically labeled superpoints; and perform three-dimensional part segmentation of the part of the object based on the three-dimensional point cloud.
2. The apparatus according to claim 1, wherein the data comprises text identifying the part of the object.
3. The apparatus according to claim 1, wherein the plurality of superpoints comprises sub-parts of the part of the object based on the three-dimensional normal at each point of the plurality of superpoints.
4. The apparatus according to claim 1, wherein the three-dimensional capture comprises an initial three-dimensional point cloud, and wherein the at least one processor is further configured to: generate the three-dimensional point cloud based on performing the part segmentation based on the initial three-dimensional point cloud.
5. The apparatus according to claim 1, wherein the plurality of superpoints are related to one or more sub-parts of the object based on the respective three-dimensional normals associated with each point of the plurality of superpoints.
6. The apparatus according to claim 1, wherein in order to semantically label each superpoint of the plurality of superpoints based on the at least one two-dimensional bounding box, the at least one processor is configured to generate a respective score for a respective part category of a plurality of part categories and a respective superpoint associated with the respective part category.
7. The apparatus according to claim 6, wherein the respective score indicates the amount by which the respective superpoint is included in one or more bounding boxes associated with the respective part category in at least one of the one or more two-dimensional images, the at least one two-dimensional image comprising a part associated with the respective part category.
8. The apparatus according to claim 1, wherein in order to merge at least one subgroup of superpoints associated with the part of the object, the at least one processor is configured to merge two superpoints based on at least one of: the two superpoints having the same semantic label, a first superpoint of the two superpoints being adjacent to a second superpoint of the two superpoints, or whether the two superpoints are included in a respective bounding box of a two-dimensional image of the one or more two-dimensional images comprising the two superpoints.
9. The apparatus according to claim 1, wherein the at least one processor is configured to obtain the three-dimensional capture of the object.
10. The apparatus according to claim 1, wherein, in order to perform the three-dimensional partial segmentation of the part of the object, the at least one processor is configured to perform multi-view feature aggregation.
11. The apparatus according to claim 10, wherein, in order to perform the multi-view feature aggregation, the at least one processor is configured to average the corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
12. The apparatus according to claim 11, wherein, in order to perform the multi-view feature aggregation, the at least one processor is further configured to average the corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.
13. A method for performing partial segmentation, the method comprising: generating one or more two-dimensional images of an object from a three-dimensional capture of the object; receiving data identifying a part of the object; processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box, the at least one two-dimensional bounding box identifying the part of the object based on a vision-language pre-trained model and the data; performing partial segmentation of the three-dimensional capture of the object to generate a plurality of superpoints associated with the object; semantically labeling each superpoint in the plurality of superpoints based on the at least one two-dimensional bounding box to generate a plurality of semantically labeled superpoints; merging at least one subgroup of superpoints associated with the part of the object from the plurality of superpoints based on the plurality of semantically labeled superpoints to generate a three-dimensional point cloud; and performing three-dimensional partial segmentation of the part of the object based on the three-dimensional point cloud.
14. The method according to claim 13, wherein the data comprises text identifying the part of the object.
15. The method according to claim 13, wherein the plurality of superpoints comprises sub-parts of the part of the object based on the three-dimensional normal at each point of the plurality of superpoints.
16. The method according to claim 13, wherein the three-dimensional capture comprises an initial three-dimensional point cloud, and wherein the method further comprises: generating the three-dimensional point cloud based on performing the partial segmentation based on the initial three-dimensional point cloud.
17. The method according to claim 13, wherein the plurality of superpoints are related to one or more sub-parts of the object based on the respective three-dimensional normals associated with each point in the plurality of superpoints.
18. The method according to claim 13, wherein semantically labeling each superpoint in the plurality of superpoints based on the at least one two-dimensional bounding box further comprises generating a respective score for a respective part category among a plurality of part categories and a respective superpoint associated with the respective part category.
19. The method according to claim 18, wherein the corresponding score indicates the amount of the corresponding superpoint being included in one or more bounding boxes associated with the corresponding part category in at least one of the one or more two-dimensional images, the at least one two-dimensional image including parts associated with the corresponding part category.
20. The method according to claim 13, wherein merging the at least one subgroup of superpoints associated with the parts of the object further includes merging two superpoints based on at least one of the following: the two superpoints having the same semantic label, a first superpoint of the two superpoints being adjacent to a second superpoint of the two superpoints, or whether the two superpoints are included in the corresponding bounding box of a two-dimensional image in the one or more two-dimensional images including the two superpoints.
21. The method according to claim 13, the method further comprising: obtaining the three-dimensional capture of the object.
22. The method according to claim 21, wherein performing the three-dimensional part segmentation of the parts of the object further includes performing multi-view feature aggregation.
23. The method according to claim 22, wherein performing the multi-view feature aggregation further includes averaging the corresponding two-dimensional pixel features on the one or more two-dimensional images of the object for each three-dimensional point in the three-dimensional capture of the object.
24. The method according to claim 23, wherein performing the multi-view feature aggregation further includes averaging the corresponding three-dimensional point features from the three-dimensional capture of the object for each two-dimensional pixel in the one or more two-dimensional images of the object.