Systems and methods with adaptive resolution for semantic occupancy
By using an adaptive resolution generator to generate feature maps and object boundary data through a machine learning system, and combining low-resolution modules to generate hybrid occupancy maps, the high computational complexity of 3D semantic occupancy prediction is solved, and efficient and accurate 3D semantic occupancy map generation is achieved.
Patent Information
- Application Number
- CN202510564632.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2025-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies face challenges in processing 3D semantic occupancy prediction due to high computational complexity and memory requirements, making it difficult to effectively generate high-precision and high-efficiency 3D semantic occupancy maps.
An adaptive resolution generator is employed to generate feature maps and object boundary data through a machine learning system. By combining lower-resolution and higher-resolution modules, a hybrid occupancy map is generated, selectively providing higher-resolution details in important areas and optimizing the utilization of computing resources.
It achieves high-resolution details in precision-critical regions while reducing computational requirements in irrelevant regions, improving the efficiency and accuracy of 3D semantic occupancy prediction, and is suitable for real-time computer vision applications.
Smart Images

Figure CN120876237A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to digital image processing and computer vision, and more particularly to three-dimensional semantic occupancy maps. Background Technology
[0002] Generally, 3D semantic occupancy prediction is an advanced technique aimed at understanding and representing 3D environments in a semantically meaningful way. More specifically, 3D semantic occupancy prediction is a task involving identifying whether space is occupied and understanding what objects or materials are in the occupied space. However, 3D semantic occupancy is challenging due to the high computational complexity and memory requirements associated with processing such large-scale 3D data. Summary of the Invention
[0003] The following is an overview of certain embodiments described in detail below. The aspects described are provided only to give the reader a brief overview of these particular embodiments, and the description of these aspects is not intended to limit the scope of this disclosure. In fact, this disclosure may cover various aspects that may not be explicitly stated below.
[0004] According to at least one aspect, a computer-implemented method includes receiving a set of digital images. The set of digital images at least displays an environment and a set of objects. The method includes generating a feature map using the set of digital images via a first machine learning system. The method includes generating object boundary data of the object set using the feature map via a second machine learning system. The method includes generating three-dimensional (3D) feature volume data using the feature map. The method includes generating a coarse occupancy map using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range. The coarse occupancy map includes the environment and the set of objects. The method includes generating surface data of the object set using the object boundary data and the 3D feature volume data. The surface data has a second resolution over a second range. The second range differs from the first range. The method includes generating a hybrid occupancy map by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the object set at a second resolution.
[0005] According to at least one aspect, a system includes one or more processors and one or more computer memories. The one or more computer memories communicate data with the one or more processors. The one or more computer memories include computer-readable data stored thereon. The computer-readable data includes instructions that, when executed by the one or more processors, cause the one or more processors to perform a method. The method includes receiving a set of digital images. The set of digital images at least shows an environment and a set of objects. The method includes generating a feature map using the set of digital images via a first machine learning system. The method includes generating object boundary data of the set of objects using the feature map via a second machine learning system. The method includes generating three-dimensional (3D) feature volume data using the feature map. The method includes generating a coarse occupancy map using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range. The coarse occupancy map includes the environment and the set of objects. The method includes generating surface data of the set of objects using the object boundary data and the 3D feature volume data. The surface data has a second resolution over a second range. The second range is different from the first range. The method includes generating a hybrid occupancy map by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the set of objects at a second resolution.
[0006] According to at least one aspect, one or more non-transient computer-readable media have computer-readable data stored thereon. The computer-readable data includes instructions that, when executed by one or more processors, cause the one or more processors to perform a method. The method includes receiving a set of digital images. The set of digital images at least shows an environment and a set of objects. The method includes generating a feature map using the set of digital images via a first machine learning system. The method includes generating object boundary data of the object set using the feature map via a second machine learning system. The method includes generating three-dimensional (3D) feature volume data using the feature map. The method includes generating a coarse occupancy map using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range. The coarse occupancy map includes the environment and the set of objects. The method includes generating surface data of the object set using the object boundary data and the 3D feature volume data. The surface data has a second resolution over a second range. The second range differs from the first range. The method includes generating a hybrid occupancy map by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the object set at a second resolution.
[0007] These and other features, aspects, and advantages of the invention will be discussed in the following detailed description with reference to the accompanying drawings, wherein the same characters denote similar or identical parts. Furthermore, the drawings are not necessarily drawn to scale, as some features may be exaggerated or minimized to show detail of particular components. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating an example of a process for generating a hybrid occupancy map according to an exemplary embodiment of this disclosure.
[0009] Figure 2A It relates to the exemplary embodiments of this disclosure. Figure 1 A flowchart of a non-restricted data example of the lower resolution portion of the process.
[0010] Figure 2B It relates to the exemplary embodiments of this disclosure. Figure 1 A flowchart of a non-restricted data example of the higher resolution portion of the process.
[0011] Figure 2C This is a non-limiting example of a hybrid occupancy map based on exemplary embodiments of the present disclosure.
[0012] Figure 3 This is a block diagram of an example system configured to generate a hybrid occupancy map for a computer vision application, according to an example embodiment of the present disclosure. Detailed Implementation
[0013] The embodiments described herein, as illustrated and described by way of example, and their many advantages will be understood from the foregoing description, and it will be apparent that various changes can be made to the form, construction, and arrangement of the components without departing from the disclosed subject matter or sacrificing one or more of its advantages. In fact, the descriptive form of these embodiments is merely illustrative. These embodiments are susceptible to various modifications and alternatives, and the following claims are intended to encompass and include these changes, and are not limited to the specific forms disclosed, but rather cover all modifications, equivalents, and alternatives consistent with the spirit and scope of this disclosure.
[0014] Figure 1 , Figure 2A , Figure 2B and Figure 2C Examples of various aspects of the process of an adaptive resolution generator 100 according to an example embodiment are shown. More specifically, in Figure 1 In the example shown, the adaptive resolution generator 100 includes an image processing module 110, a lower resolution module 120, and a higher resolution module 130. These modules are provided to facilitate discussion of various aspects of the adaptive resolution generator 100. In this respect, there may be... Figure 1More or fewer modules are shown, as long as they are configured to perform the same functions as described herein. Furthermore, Figure 2A Non-limiting data examples are shown to illustrate various aspects related to the generation of at least one rough occupancy map 230 via image processing module 110 and lower resolution module 120. Meanwhile, Figure 2B Non-limiting data examples are shown to illustrate various aspects related to the generation of at least fine surface data 250 via image processing module 110 and higher resolution module 130. Furthermore, Figure 2C A non-limiting example of a hybrid occupancy map 260 is shown, which is generated by an adaptive resolution generator 100 by combining the output of a lower resolution module 120 (e.g., a coarse occupancy map 230) and the output of a higher resolution module 130 (e.g., high-resolution surface data 250).
[0015] As an overview, the adaptive resolution generator 100 is configured to generate one or more hybrid occupancy maps 260 for computer vision. The adaptive resolution generator 100 is configured to adjust the level of detail in at least one suggested region (e.g., at least one selected object) of interest space or environment. For example, the adaptive resolution generator 100 is advantageous in enhancing detail and improving accuracy in specific regions (e.g., specific objects in an environment or scene) selected as the most relevant. The adaptive resolution generator 100 ensures that one or more computing resources (e.g., GPU processing) are efficiently used for these specific regions. Furthermore, the adaptive resolution generator 100 reduces the computational requirements (e.g., GPU processing) for rapid processing in other unrelated regions. Therefore, the adaptive resolution generator 100 ensures efficient and effective use of computing resources (e.g., processors, memory, etc.) while generating 3D adaptive resolution semantic occupancy maps (i.e., hybrid occupancy maps 260) for real-time applications. The adaptive resolution generator 100 is advantageous in various fields requiring rapid and accurate interpretation of complex environments (e.g., navigation, at least partially autonomous driving, augmented reality, etc.).
[0016] In addition to the hybrid occupancy map 260, the adaptive resolution generator 100 is configured to provide and / or output other relevant data for downstream use. The adaptive resolution generator 100 is configured to provide multimodal predictions. In response to receiving a set of digital images, the adaptive resolution generator 100 is configured to generate at least (i) object boundary data 240 associated with detected objects (e.g., bounding boxes, etc.), (ii) object point clouds of the detected objects, and (iii) an egocentric occupancy map including the detected objects (e.g., hybrid occupancy map 260 and coarse occupancy map 230). The adaptive resolution generator 100 is configured to selectively provide greater detail and higher resolution in selected regions where accuracy is most important. The generated object point clouds are not limited by grid resolution and can be easily extended to multi-resolution grid maps and object surface construction / reconstruction. Furthermore, these multimodal outputs have proven beneficial for object detection and occupancy prediction tasks. The Adaptive Resolution Generator 100 offers advancements in 3D scene understanding while providing a balanced blueprint for enhanced accuracy and diverse outputs in a variety of computer vision-related applications, such as autonomous driving, autonomous driving, navigation, robotics, and augmented reality.
[0017] like Figure 1 As shown, the adaptive resolution generator 100 includes acquiring or receiving a set of digital images 200. The set of digital images 200 includes one or more digital images. The digital images can be two-dimensional (2D) digital images. For example, in... Figure 2A and Figure 2B In this context, the adaptive resolution generator 100 includes acquiring or receiving a set of surrounding view images as a digital image set 200. The surrounding view images provide various views of the surroundings or basic surrounding environment of the vehicle 10. More specifically, for example, in... Figure 2A and Figure 2B In this embodiment, the digital image set 200 includes (i) at least one digital image 202 showing a portion of the environment in front of or around the vehicle 10, (ii) at least one digital image 204 showing a portion of the environment in rear of or around the vehicle 10, (iii) at least one digital image 206 showing a portion of the environment in left-side or around the vehicle, and (iv) at least one digital image 208 showing a portion of the environment in right-side or around the vehicle, such that this digital image set provides a view of the surroundings of the vehicle 10. As another example (not shown), the digital image set 200 includes a sequence of images with time information, the sequence comprising two or more digital images.
[0018] The digital image set 200 is provided as input to the image processing module 110 via an image encoder. The image encoder is configured to receive and encode the digital image set 200. In an example embodiment, the image encoder includes a machine learning (ML) system 112. Figure 1 In this embodiment, the machine learning system 112 includes at least a convolutional neural network (CNN), or at least one machine learning model, configured to generate one or more feature maps 210 using one or more digital images. More specifically, the machine learning system 112 is configured to extract multiple fundamental features from the digital image set 200 using a series of convolutional layers, activation functions, and pooling layers. In this respect, the machine learning system 112 is configured to transform the raw pixel data from each digital image into a more compact and expressive feature representation, thereby capturing fundamental visual cues and patterns useful for subsequent stages of the adaptive resolution generator 100. That is, the machine learning system 112 is configured to generate one or more feature maps 210 using the digital image set.
[0019] In addition, such as Figure 1 As shown, the image processing module 110 includes a 3D generator 114 for generating 3D feature volume data 220 (e.g., voxel data) using the feature map 210. The 3D generator 114 may include hardware, software, or a combination thereof. For example, in Figure 1 In this context, the 3D generator includes software techniques (e.g., computer-readable data including instructions) executed by one or more processors of the processing system 302 to generate corresponding 3D feature volume data 220 using the feature map 210. The 3D generator 114 is configured to perform a crucial process that transforms the 2D feature map 210 into 3D feature volume data 220 intricately linked to intrinsic and extrinsic camera parameters. Parameters play a critical role in the 2D-to-3D feature projection performed by the 3D generator 114. Intrinsic parameters include and / or relate to the camera's internal characteristics (e.g., focal length, optical center, etc.). Intrinsic parameters affect how features extracted from the 2D image are transformed. Meanwhile, extrinsic parameters include and / or relate to describing the camera's position and orientation in space. Extrinsic parameters are essential for accurately locating features in a 3D coordinate system.
[0020] The synergy of these intrinsic and extrinsic parameters ensures that the projected 3D features accurately correspond to the spatial arrangement and size of the real world. The 3D generator 114 may include a bird's-eye view (BEV) encoder that receives the feature map 210 and generates BEV features as feature volume data (which is 2D BEV). The 3D generator 114 is configured to generate results or outcomes that include finely detailed 3D feature volume data 220 in space, containing rich information harvested from the digital image set 200 (e.g., one or more 2D images). This 3D feature volume data 220 is also a spatial representation full of semantic nuances, thus providing a solid foundation for 3D occupancy prediction via the machine learning system 122 (“ML system 122”). The 3D generator 114 transforms every nuance of one or more 2D images (or the digital image set 200, such as a set of surrounding view images) into 3D space with high fidelity, thereby ensuring that the downstream occupancy map (e.g., a hybrid occupancy map 260) is spatially and semantically accurate, serving as a reliable basis for various real-time applications.
[0021] The lower resolution module 120 includes generating a coarse occupancy map 230 using 3D feature volume data 220. The lower resolution module 120 includes a machine learning system 122. The machine learning system 122 includes at least one machine learning model configured to receive 3D feature volume data 220 from a 3D generator 114 and generate a low-resolution occupancy map (i.e., the coarse occupancy map 230) using the 3D feature volume information 220. The machine learning system 122 is configured to perform 3D semantic occupancy decoding. In this respect, the lower resolution module 120 is configured to use the 3D feature volume data 220 to predict occupancy and semantic labels in 3D space. More specifically, as an example, the machine learning system 122 includes at least a CNN. The CNN includes neural network layers. The neural network layers analyze 3D features to determine which spaces are occupied and what types of objects occupy these occupied spaces. This information (e.g., occupied / unoccupied spaces and object categorical types within the occupied spaces) is used to generate the coarse occupancy map 230. The coarse occupancy map 230 is a comprehensive 3D semantic occupancy map in which each point in the space is assigned a probability of being occupied, along with a semantic label classifying the occupied object or material. The coarse occupancy map 230 has a first resolution within a first range. For example, each element of the coarse occupancy map 230 is displayed at the same lower resolution within the first range.
[0022] Reference Figure 2AAs a non-limiting example, the coarse occupancy map 230 provides a 3D display of an environment with occupied voxels, each voxel having a semantic label indicating multiple objects, such as road 230A, sky 230B, building / wall 230C, traffic cone 230D, sidewalk 230E, road guardrail 230F, another vehicle 230G, another vehicle 230H, etc. The coarse occupancy map 230 may also provide unoccupied voxels with semantic labels indicating vacancies. A vacancy might represent a space considered "empty" relative to the containing object (e.g., a detected object of a predefined object category). The coarse occupancy map 230 includes voxels of the same resolution and within the same lower resolution range. The coarse occupancy map 230 is generated quickly, where each displayed object exhibits the same level of detail. Then, as... Figure 1 As shown, the rough occupancy map 230 is sent to the combiner 140.
[0023] In addition, such as Figure 1 As shown, the adaptive resolution generator 100 includes a higher resolution module 130. The higher resolution module 130 includes a machine learning system 132. The machine learning system 132 includes at least one machine learning model configured to generate object detection data, such as object boundary data 240, using feature map 210. Furthermore, the adaptive resolution generator 100 is configured to output the object boundary data 240 as one of the multimodal outputs.
[0024] In an example embodiment, the machine learning system 132 includes at least a Region Proposal Network (RPN). The machine learning system 132 may include an Object Proposal Network (OPN). The RPN or OPN is cleverly integrated to identify specific regions (e.g., areas and / or objects) deemed crucial in a scene, thereby efficiently focusing computational effort and resources on these specific regions. The machine learning system 132 (e.g., RPN, OPN, etc.) uses predefined object categories. In this respect, object categories may be specified by at least one user. For example, as a non-limiting example, object categories may include vehicles, pedestrians, road signs, animals, road obstacles, and / or one or more other objects deemed relevant or important during autonomous driving decision-making. Object boundary data 240 includes object detection corresponding to the object categories. Selecting and focusing on these specific regions (e.g., specific objects) is more advantageous than a uniform approach to the entire scene, at least because this approach concentrates computational expenditure on important areas within the scene, whereas uniform processing typically leads to unnecessary computational expenditure and inefficiency. Therefore, the adaptive resolution generator 100 is configured to selectively focus computational resources on selected objects for efficient and accurate 3D semantic prediction.
[0025] More specifically, in the example embodiment, the machine learning system 132 includes at least a DETR-style (Detection Transformer) RPN, which differs significantly from traditional anchor-based methods such as those employed in R-CNN and its variants. Instead of using anchor boxes and sliding windows, DETR applies the Transformer architecture, originally designed for natural language processing tasks, to the object detection domain. Using a DETR-style RPN, the machine learning system 132 is configured to treat object detection as a prediction problem based on direct sets, thereby eliminating the need for non-maximum suppression (NMS) and other complex intermediate steps.
[0026] Furthermore, after identifying these key or specific regions via object boundary data 240 and corresponding feature volume data 220, the adaptive resolution generator 100 shifts its attention to higher-resolution surface constructions with greater precision within the boundaries of these key or specific regions of the object set associated with object boundary data 240. In this regard, the higher resolution module 130 also includes a machine learning system 134. To specifically focus the high-resolution occupancy estimate within the boundaries of object boundary data 240 (e.g., the proposed bounding boxes of detected objects), the machine learning system 134 is configured to construct high-resolution surface data 250 for each region of interest. For example, the predicted bounding boxes... This can be defined by Equation 1, where P(object|B) is the probability of the object existing given bounding box B, and P(object|image) is the likelihood of bounding box B given input image I. Furthermore, the 900 bounding boxes can be regressed and their object classification scores calculated following the settings from DETR.
[0027]
[0028] Machine learning system 134 includes at least one machine learning model configured to generate surface data 250 using object detection data or object boundary data 240 and corresponding feature volume data 220. Surface data 250 includes high-resolution (or fine-grained) surface data. Surface data 250 is generated at a higher resolution and with more detail than the coarse occupancy map 230. More specifically, the resolution range of surface data 250 is greater than the resolution range of the coarse occupancy map 230. For example, machine learning system 134 may include at least one decoding network configured to perform the functions described herein, such as FoldingNet, Multi-Resolution Deep Implicit Function (MDIF), or a similar network. Machine learning system 134 may include a machine learning model that reconstructs objects in a scene using a hierarchical shape reconstruction approach and provides advancements in shape completion.
[0029] In an example embodiment, the machine learning system 134 includes at least a decoder for FoldingNet that uses a neural network to create surface data 250 of a 3D surface using object boundary data 240 and feature volume data 220. In this case, the surface data 250 is created from point cloud or voxel data. FoldingNet utilizes an origami-inspired folding process with several folding operations, for example, by folding layers, the point cloud gradually transforms into a higher-dimensional space, thereby creating a continuous and accurate surface representation.
[0030] The folding operation is essentially a series of transformation matrices. These transformations are learned during the training phase. The FoldingNet decoder is configured to transform each point on the 2D mesh into a point in 3D space in combination with the encoded feature vectors. The feature vectors tell how the 2D mesh should be folded to best approximate the original 3D structure. Regarding FoldingNet, the decoder is configured to generate and / or output fine surface data 250 to construct a 3D surface corresponding to each object associated with object boundary data 240. Each point in the cloud of surface data 250 corresponds to a folding point in the 2D mesh, and they collectively approximate the shape and features of the original 3D point cloud.
[0031] By precisely reconstructing the object surface within their respective 3D bounding boxes, the higher-resolution module 130 is configured to capture fine-grained geometric information, which is particularly well-suited for object-centric semantic scene filling. (Folded point cloud X) f It is calculated and represented by Equation 2, where X represents the original input point cloud or voxel data, and W and b represent the learnable weights and biases associated with the folding operation.
[0032] X f =Fold(X,W,b) [2.
[0033] As described above, the machine learning system 132 may include FoldingNet to perform simple, computationally efficient, and general decoding operations in point cloud reconstruction. For example, the FoldingNet decoder may be configured to receive cropped BEV features for each object from the estimated bounding box and then be configured to reconstruct the object in point cloud format. The FoldingNet decoder may include a multilayer perceptron (MLP) based decoder that uses the cropped BEV features. The adaptive resolution generator 100 is configured to output the point cloud data as one of the multimodal outputs. The higher resolution module 130 may be configured to further process the object point cloud into a finer occupancy format or a smooth mesh surface using Poisson surface reconstruction.
[0034] Figure 2BA visualization of a non-limiting data example associated with the higher resolution module 130 is shown. As previously described, the higher resolution module 130 includes a machine learning system 132 configured to identify a set of objects and generate object boundary data 240 for the set of objects using a feature map 210. As an example, Figure 2B A visualization of a scene including a first car and a second car on a road is shown. In this example, object boundary data 240 includes object boundary data 242 for the first car and object boundary data 244 for the second car. Furthermore, Figure 2B A visualization of feature volume data 220 is shown, which includes feature volume data 222 of the first vehicle and feature volume data 224 of the second vehicle. Furthermore, Figure 2B The machine learning system 134 is shown to generate at least higher resolution surface data 250 for a set of objects associated with object boundary data 242 and object boundary data 244 based on 3D feature volume data 252 and 3D feature volume data 254, the surface data 250 including higher resolution surface data 252 of a first vehicle and higher resolution surface data 254 of a second vehicle.
[0035] By leveraging advanced algorithms and computational techniques, the machine learning system 134 meticulously constructs high-resolution surface data 250 for the set of objects associated with object boundary data 240, thereby ensuring a higher level of detail and accuracy for this particular set of objects compared to the low-resolution portions of the mixed occupancy map 260 (e.g., background portions, unoccupied portions, etc.). Other portions and / or remaining portions of the set of objects not associated with the mixed occupancy map 260 include the low resolution of the coarse occupancy map 230. Next, the adaptive resolution generator 100 includes a combiner 140 configured to generate the mixed occupancy map 260 by combining the coarse occupancy map 230 generated from the lower resolution module 120 and the surface data 250 generated from the higher resolution module 130. The mixed occupancy map 260 combines the surface data 250 and the coarse occupancy map 230 to generate a mixed-resolution occupancy map that uses fine resolution for selected objects (e.g., foreground elements) and coarse resolution for unselected objects (e.g., example background elements).
[0036] Figure 2CA non-limiting example of a hybrid occupancy map 260 generated by an adaptive resolution generator 100 is shown. The hybrid occupancy map 260 is generated by combining a coarse occupancy map 230 with surface data 250 associated with a proposed region (e.g., a set of objects). In this respect, similar to the coarse occupancy map 230, the hybrid occupancy map 260 is a comprehensive 3D semantic occupancy map, where each point in space is assigned a probability of being occupied, along with semantic labels classifying the occupied objects or materials. The hybrid occupancy map 260 can be egocentric. For example, in… Figure 2C In the mixed occupancy map 260, self-vehicles 10 (e.g., mobile robots) can also be displayed as references.
[0037] Furthermore, compared to the coarse occupancy map 230, the hybrid occupancy map 260 provides an adaptive resolution (i.e., one or more resolution ranges) based on defined object categories. More specifically, the hybrid occupancy map 260 includes a set of objects associated with object boundary data 240, where surface data 250 is generated at a higher resolution, while the remainder of the hybrid occupancy map 260 comprises the lower resolution of the coarse occupancy map 230. The hybrid occupancy map 260 provides a 3D display of an environment with occupancy voxels, each voxel having semantic labels indicating multiple objects, such as road 230A, sky 230B, building / wall 230C, traffic cone 250D, sidewalk 230E, road guardrail 250F, another vehicle 250G, another vehicle 250H, etc. Figure 2C In the example, elements labeled 230 represent a coarse resolution, while elements labeled 250 represent a fine resolution. As a non-limiting example, in... Figure 2C In this example, the coarse resolution can correspond to voxel sizes larger than 0.2m, while the fine resolution can correspond to voxel sizes less than or equal to 0.2m. Furthermore, in this non-limiting example, the grid size of the hybrid occupancy map 260 includes 0.2m, 0.4m, 0.8m, etc. The coarse occupancy map 230 can also provide unoccupied voxels with semantic labels indicating vacancies. The coarse occupancy map 230 includes at least multiple background elements with lower resolution voxels (e.g., road 230A, sky 230B, building / wall 230C, etc.) and a set of objects with higher resolution voxels (e.g., traffic cone 250D, sidewalk 250E, road guardrail 250F, another vehicle 250G, another vehicle 250H, etc.). Utilizing adaptive resolution, the hybrid occupancy map 260 is configured to provide more detail and higher resolution for specific object sets, while rapidly generating lower resolutions for other background elements of 3D semantic occupancy in real-time computer vision applications, thereby concentrating computational resources on more relevant areas.
[0038] As described above, the adaptive resolution generator 100 provides a two-stage approach involving at least one lower-resolution module 120 and at least one higher-resolution module 130. In this respect, the higher-resolution module 130 selectively provides higher-resolution voxel data compared to the lower-resolution module 120. The adaptive resolution generator 100 is advantageous in optimizing resource utilization (e.g., GPU resource utilization, memory utilization, etc.) through selective and adaptive resolution. The adaptive resolution generator 100 also improves the accuracy of occupancy prediction, and is therefore a cornerstone for various computer vision applications that require real-time and highly accurate environmental understanding.
[0039] The Adaptive Resolution Generator 100 addresses the challenges of object-centric semantic scene completion by integrating three loss components: focus loss, DeTR loss, and chamfer loss. The Adaptive Resolution Generator 100 applies focus loss to the predicted semantic labels. Specifically, it is used for contextual understanding. For the foreground, the adaptive resolution generator 100 applies the DeTR loss (LDeTR) to detect bounding boxes. This loss selects N valid boxes from a set of bounding box candidates and estimates the position of each valid box in parallel. Finally, the chamfer loss (Lchamfer) is used as a key component for foreground object surface reconstruction, quantifying the difference between the reconstructed surface and the real ground surface.
[0040] The focus loss is a classification loss specifically designed to address problems such as class imbalance and labels that are harder to detect. In this setting, the adaptive resolution generator 100 iteratively predicts the probability that each voxel belongs to a particular object class. If the probability of each class is below a critical threshold ∈ , the adaptive resolution generator 100 assigns the voxel to the free space. The focus loss of the occupied map M is calculated via Equation 3, where p(·) represents the predicted probability of the correct class, and α and β represent hyperparameters used to balance well-classified and hard-classified examples.
[0041]
[0042] The DeTR loss comprises a focus loss and an L1 loss. It classifies the number of valid boxes from 900 bounding box candidate losses and estimates the location of each valid box separately. The focus loss is similar to Equation 4, which is configured to predict the class of all objects. If the probability of all classes is below a certain threshold ∈, the box is predicted as invalid. Furthermore, the L1 loss is a regression loss that minimizes the N predicted bounding boxes B calculated via Equation 4. pred and N corresponding ground truth bounding boxes B GT The differences between them.
[0043]
[0044] Chamfer loss is a geometric distance-based loss function used to measure the dissimilarity between two sets of points. In the context of the adaptive resolution generator 100, chamfer loss quantifies the difference between reconstructed surface points and ground truth ground points. As defined by Equation 5, where X represents the reconstructed point and Y represents the ground truth point.
[0045]
[0046] Semantic loss, particularly cross-entropy loss, is a classification loss used to measure the difference between predicted and true labels. In this case, the adaptive resolution generator 100 evaluates the difference between the predicted and actual semantic labels for the completed scene. This is represented as... The cross-entropy loss is determined and calculated by Equation 6, where C represents the number of classes, and y gt Let y represent the ground truth class probability, and y pred This refers to the predicted class probabilities. These two loss functions, namely the chamfer loss and the cross-entropy loss, together contribute to the optimization process within our semantic scene completion framework.
[0047]
[0048] Figure 3 This is a block diagram of an example system 300 according to an example embodiment, which includes an adaptive resolution generator 100 configured to generate a hybrid occupancy map 260. System 300 is configured to perform... Figure 1 The adaptive resolution generator 100 processes the data. System 300 includes at least processing system 302. Processing system 302 includes at least one processing device. For example, processing system 302 may include an electronic processor, central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), microprocessor, field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), any processing technology, or any number and combination thereof. Processing system 302 is operable to provide the functionality described herein.
[0049] System 300 includes at least one sensor system 304. Sensor system 304 includes one or more sensors. For example, sensor system 304 includes at least an image sensor, such as a camera that generates digital images. Sensor system 304 may include at least one other type of sensor (e.g., radar, LiDAR, infrared, etc.) to obtain additional sensor data, thereby enabling sensor system 304 to generate digital images based on this additional sensor information. Sensor system 304 is operable to communicate with one or more other components of system 300 (e.g., processing system 302 and storage system 306). For example, sensor system 304 may provide sensor data (e.g., digital images), which are then processed by processing system 302. Sensor system 304 is local, remote, or a combination thereof (e.g., partially local and partially remote) relative to one or more components of system 300. Upon receiving sensor data (e.g., one or more digital images), processing system 302 is configured to process the sensor data (e.g., digital images) in conjunction with adaptive resolution generator 100, other relevant data 308, computer vision application 310, or any number and combination thereof.
[0050] System 300 includes a memory system 306 operatively connected to processing system 302. In this respect, processing system 302 communicates data with memory system 306. Memory system 306 includes at least one non-transitory computer-readable storage medium configured to store and provide access to various types of data to enable at least processing system 302 to perform the operations and functions disclosed herein. Memory system 306 includes a single memory device or multiple memory devices. Memory system 306 may include electrical, electronic, magnetic, optical, semiconductor, electromagnetic, or any suitable storage technology. For example, memory system 306 may include random access memory (RAM), read-only memory (ROM), flash memory, disk drives, memory cards, optical storage devices, magnetic storage devices, storage modules, any suitable type of storage device, or any number and combination thereof.
[0051] The memory system 306 includes at least an adaptive resolution generator 100 configured to generate one or more hybrid occupancy maps 260 using a set of digital images. The adaptive resolution generator 100 includes computer-readable data that, when executed by the processing system 302, is configured to perform at least the functions disclosed in this disclosure. The computer-readable data may include instructions, code, routines, various related data, any software technology, or any number and combination thereof. For example, in an example embodiment, the adaptive resolution generator 100 includes, as referenced... Figure 1The discussion covers several software technologies and machine learning models. More specifically, the adaptive resolution generator 100 includes an image encoder that incorporates a machine learning system 112. The adaptive resolution generator 100 includes a 3D generator 114. The adaptive resolution generator 100 includes a semantic occupancy decoder that incorporates a machine learning system 122. The adaptive resolution generator 100 includes a region generator that incorporates a machine learning system 132. Furthermore, the adaptive resolution generator 100 includes a surface constructor that incorporates a machine learning system 134. Additionally, the adaptive resolution generator 100 includes a combiner 140 for generating a hybrid occupancy map 260. The adaptive resolution generator 100 is configured such that occupancy detection, object detection, and surface construction / reconstruction are tasks jointly trained given a shared spatiotemporal framework. This joint training, reflected in all three tasks, enhances the 3D understanding of the environment.
[0052] Furthermore, memory system 306 and other related data 308 provide various data (e.g., operating system, etc.) that enable system 300 and / or processing system 302 to perform the functions described herein. Additionally, memory system 306 may include computer vision application 310, which includes computer-readable data configured, when executed by processing system 302, to apply one or more hybrid occupancy maps 260 to the computer vision application. Computer-readable data may include instructions, code, routines, various related data, any software technology, or any number and combination thereof. As a non-limiting example, computer vision application 310 may relate to a navigation program that provides navigation using one or more hybrid occupancy maps 260.
[0053] System 300 may include one or more I / O devices 312 (e.g., display devices, microphones, speakers, etc.). For example, system 300 may include a display device configured to display one or more hybrid occupancy maps and / or related data. As a non-limiting example, system 300 may include a touchscreen on a mobile communication device that displays hybrid occupancy maps or related data. This feature facilitates user interaction with one or more hybrid occupancy maps.
[0054] In addition, system 300 includes other functional modules 314, such as any suitable hardware, software, or combinations thereof that assist or contribute to the functionality of system 300 and / or adaptive resolution generator 100. For example, other functional modules 314 include communication technologies (e.g., wired communication technologies, wireless communication technologies, or combinations thereof) that enable components of system 300 to communicate with each other and / or with one or more other computing devices (not shown), such as mobile communication devices, smartphones, laptops, tablets, servers, cloud computing systems, etc.
[0055] In addition, other functional modules 314 may include other components, such as actuators. In this regard, for example, when the adaptive resolution generator 100 is used for at least one vehicle (e.g., a motor vehicle, a robotic vacuum cleaner, etc.), other functional modules 314 may also include one or more actuators that are based on at least one or more hybrid occupancy maps and relate to driving, steering, stopping, and / or controlling the movement of the vehicle.
[0056] As described in this disclosure, embodiments offer numerous advantages and benefits. For example, an embodiment provides an adaptive resolution method for 3D semantic occupancy mapping. The embodiment addresses the trade-off between accuracy and computational efficiency because high-resolution prediction is performed only in specific areas where accuracy is critical. In this respect, the embodiment is selective in its focus on detail where accurate information is most important to the user and / or downstream applications. High resolution provides more information (e.g., semantic information, occupancy information, spatial information, distance representation, etc.) at the cost of higher computational load and more computing resources. In contrast, low resolution is associated with lower computational load and fewer computing resources, but provides less information. The embodiment is configured to effectively balance accuracy and resources by selectively providing higher resolution in meaningful areas and lower resolution in meaningless areas (e.g., blank areas occupied by objects not belonging to predefined object categories). For example, the embodiment includes generating higher resolution for foreground objects of predefined object categories and lower resolution for background scenes, where higher resolution is finer and more detailed than lower resolution. Balancing computational resources is a critical aspect in various applications (e.g., dynamic driving environments) because inefficient use of computing resources can hinder real-time decision-making.
[0057] Furthermore, for occupancy-based 3D maps, the embodiments provide two or more resolutions with multiple sized grids of various sizes, thus not being limited to a single general resolution (or fixed-granularity resolution) involving a fixed-size grid with uniform dimensions. Unlike other methods of 3D semantic occupancy prediction with fixed voxel sizes, these embodiments provide a new 3D semantic occupancy prediction paradigm that supports multiple voxel sizes in the same hybrid occupancy map 260. That is, a single 3D hybrid occupancy map 260 may include several voxels of different sizes. Higher resolutions (e.g., smaller voxel sizes) provide greater detail and accuracy, while lower resolutions (e.g., larger voxel sizes) reduce computational load (e.g., processing and memory requirements). The embodiments strategically optimize the computational efficiency of semantic occupancy prediction while ensuring high-resolution detail for a variety of real-time applications (e.g., autonomous driving systems, etc.). In this regard, for example, the embodiments enable vehicles to navigate safely and efficiently in real time, while responding to dynamic changes in their environment with enhanced awareness and accuracy. The embodiments generate multimodal outputs (e.g., hybrid occupancy map 260), which can be used to develop and optimize safe, reliable, and advanced autonomous driving systems.
[0058] Furthermore, the embodiments support different region suggestion patterns for high-precision occupancy prediction. For example, region suggestion patterns may include distance-based region suggestions, object-based region suggestions, motion-based region suggestions, or another type of region suggestion, or any number and combination thereof.
[0059] Furthermore, these embodiments integrate object detection and semantic segmentation within a federated learning framework, which contributes to a more comprehensive understanding of the 3D environment. These embodiments enhance 3D understanding by integrating object detection, semantic segmentation, and surface reconstruction. They possess a strategic focus by providing detailed predictions of predetermined objects around autonomous vehicles, thus paving the way for more advanced real-time navigation, multiple output formats, and enhanced safety in autonomous driving systems.
[0060] As described above, the embodiments merge object detection and surface reconstruction into a joint training pipeline, effectively improving the tasks associated with object detection and occupancy prediction by learning detailed features within each object's bounding box. The embodiments demonstrate how surface reconstruction aids in learning representations more computationally by applying high-resolution meshes to important objects such as humans, cars, animals, etc. Furthermore, joint training improves object detection, at least because the fine surface construction further refines features within densely packed regions of object features. The embodiments include joint training and an architecture that improves near-range occupancy prediction performance, at least because surface reconstruction within the bounding box enhances each feature learned within the bounding box. Moreover, applying more constraints between coarse occupancy prediction and object surface reconstruction improves consistency between the occupancy prediction task and the instance segmentation task.
[0061] Furthermore, the foregoing description is intended to be illustrative rather than limiting, and is provided in the context of a particular application and its requirements. Those skilled in the art will understand from the foregoing description that the invention can be implemented in various forms, and various embodiments can be implemented individually or in combination. Therefore, although embodiments of the invention have been described in conjunction with specific examples of the invention, the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the described embodiments, and the true scope of the embodiments and / or methods of the invention is not limited to the shown and described embodiments, as various modifications will become apparent to those skilled in the art upon studying the drawings, specification, and appended claims. Additionally, or alternatively, components and functions may be separated or combined in ways different from the various described embodiments, and may be described using different terms. These and other variations, modifications, additions, and improvements fall within the scope of disclosure as defined in the following claims.
Claims
1. A computer-implemented method, comprising: Receive a set of digital images, the set of digital images displaying at least an environment and a set of objects; A feature map is generated using the digital image set via a first machine learning system; The object boundary data of the object set is generated using the feature map via a second machine learning system; The feature map is used to generate three-dimensional (3D) feature volume data; A coarse occupancy map is generated using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range and includes the environment and the set of objects. The object boundary data and the 3D feature volume data are used to generate surface data of the object set, the surface data having a second resolution of a second range, the second range being different from the first range; and A hybrid occupancy map is generated by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the set of objects at a second resolution.
2. The computer-implemented method of claim 1, wherein the second machine learning system includes a Region Proposal Network (RPN) that generates the object boundary data using the feature map.
3. The computer-implemented method of claim 1, wherein the coarse occupancy map is generated via a third machine learning system that decodes the 3D feature volume data.
4. The computer-implemented method of claim 1, wherein the surface data is generated via another machine learning system using the object boundary data and the 3D feature volume data.
5. The computer-implemented method of claim 4, wherein the other machine learning system comprises a series of transformation matrices for generating the surface data of the set of objects.
6. The computer-implemented method of claim 1, wherein the second range is greater than the first range, such that the second resolution of the object set is greater than the first resolution of the environment.
7. The computer-implemented method according to claim 1, further comprising: The hybrid occupancy map is used to control the actuator. The actuator mentioned above is a component of a vehicle.
8. A system comprising: One or more processors; One or more computer memories that communicate with the one or more processors, the one or more computer memories having computer-readable data stored thereon, the computer-readable data including instructions that, when executed by the one or more processors, cause the one or more processors to perform a method, the method comprising... Receive a set of digital images, the set of digital images displaying at least an environment and a set of objects; A feature map is generated using the digital image set via a first machine learning system; The object boundary data of the object set is generated using the feature map via a second machine learning system; The feature map is used to generate three-dimensional (3D) feature volume data; A coarse occupancy map is generated using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range and includes the environment and the set of objects. The object boundary data and the 3D feature volume data are used to generate surface data of the object set, the surface data having a second resolution of a second range, the second range being different from the first range; and A hybrid occupancy map is generated by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the set of objects at a second resolution.
9. The system of claim 8, wherein the second machine learning system includes a Region Proposal Network (RPN) that generates the object boundary data using the feature map.
10. The system of claim 8, wherein the coarse occupancy map is generated via a third machine learning system that decodes the 3D feature volume data.
11. The system of claim 8, wherein the surface data is generated via another machine learning system using the object boundary data and the 3D feature volume data.
12. The system of claim 11, wherein the other machine learning system comprises a series of transformation matrices that generate the surface data of the set of objects.
13. The system of claim 8, wherein the second range is greater than the first range, such that the second resolution of the object set is greater than the first resolution of the environment.
14. The system of claim 8, wherein the method further comprises: The hybrid occupancy map is used to control the actuator. The actuator mentioned above is a component of a vehicle.
15. One or more non-transitory computer-readable media having computer-readable data stored thereon, the computer-readable data including instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: Receive a set of digital images, the set of digital images displaying at least an environment and a set of objects; A feature map is generated using the digital image set via a first machine learning system; The object boundary data of the object set is generated using the feature map via a second machine learning system; The feature map is used to generate three-dimensional (3D) feature volume data; A coarse occupancy map is generated using the 3D feature volume data. The coarse occupancy map has a first resolution over a first range and includes the environment and the set of objects. The object boundary data and the 3D feature volume data are used to generate surface data of the object set, the surface data having a second resolution of a second range, the second range being different from the first range; and A hybrid occupancy map is generated by combining the coarse occupancy map and the surface data. The hybrid occupancy map displays the environment at a first resolution and the set of objects at a second resolution.
16. One or more non-transient computer-readable media according to claim 15, wherein the second machine learning system includes a Region Proposal Network (RPN) that generates the object boundary data using the feature map.
17. One or more non-transient computer-readable media according to claim 15, wherein the coarse occupancy map is generated via a third machine learning system that decodes the 3D feature volume data.
18. One or more non-transient computer-readable media of claim 15, wherein the surface data is generated via another machine learning system using the object boundary data and the 3D feature volume data.
19. One or more non-transient computer-readable media of claim 18, wherein the other machine learning system includes a series of transformation matrices that generate the surface data of the set of objects.
20. One or more non-transient computer-readable media of claim 15, wherein the second range is greater than the first range, such that the second resolution of the object set is greater than the first resolution of the environment.