Methods and devices for detecting and localizing objects in three-dimensional point clouds

By applying mask and template matching on overhead views of 3D point clouds to select candidate clusters and using neural networks trained on known objects, the method addresses the challenge of object detection and localization in LiDAR-generated point clouds, improving accuracy and efficiency.

DE102025104166A1Pending Publication Date: 2025-08-14INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025104166
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-09
Filing Date
2025-02-05
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

The recognition and localization of objects in 3D point clouds generated by LiDAR sensors, particularly for objects far from the sensor, is challenging due to irregular, non-continuous patterns and decreasing density, which complicates object detection and localization.

Method used

The use of mask and template matching on overhead views of 3D point clouds to select candidate clusters, combined with neural networks trained on known objects and bounding boxes, reduces dimensionality and complexity, improving object classification and localization accuracy.

Benefits of technology

This approach enhances the accuracy and speed of object detection and localization by focusing on candidate clusters, reducing computational resources and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Example systems, devices, articles of manufacture, and methods for detecting and locating objects in three-dimensional (3D) point clouds are disclosed. The example devices disclosed herein apply at least one template or mask to a sample point of an overhead view of a 3D point cloud to identify a candidate cluster of points in the 3D point cloud, where the candidate cluster is to satisfy an occupancy objective. The disclosed example device also inputs the candidate cluster to a neural network trained to output a feature vector for the candidate cluster. The disclosed example device further processes the feature vector to output parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF REVELATION

[0001] This disclosure relates generally to computer vision, and more particularly to methods and apparatus for detecting and locating objects in three-dimensional (3D) point clouds. BACKGROUND

[0002] In recent years, the use of LiDAR (Light Detection and Ranging) sensors to implement computer vision for systems such as autonomous vehicles, robots, and so on has increased. One or more LiDAR sensors can be integrated into such systems to reflect laser beams from objects in an environment, resulting in a 3D point cloud representing a scene including the object(s). Detecting objects from the 3D point cloud can include object classification to identify objects and bounding box regression to locate objects. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram of an example object detector for detecting and localizing objects in 3D point clouds. Fig. 2 is a block diagram of an exemplary cluster selection circuit used in the exemplary object detector of Fig. 1 is included. Fig. 3 to 5 show exemplary masks and exemplary templates used by the exemplary cluster selection circuit of Fig. 1 and / or 2 can be used. Fig. 6 is a block diagram of an exemplary bounding box generation circuit and an exemplary regression calculation circuit used in the exemplary object detector of Fig. 1 is included. Fig. 7 is a block diagram of an exemplary mask generation circuit and an exemplary template generation circuit for generating the exemplary masks and exemplary templates of Fig. 3-5. Fig. Fig. 8 shows an exemplary operation of the mask generation circuit of Fig. 7 to generate an exemplary mask used by the cluster selection circuit of Fig. 1 and / or 2 is used. Fig. 9 and Fig. 10 are flowcharts illustrating example machine-readable instructions and / or example operations that may be executed, instantiated, and / or performed by example programmable circuits to implement the object detector 100 of Fig. 1 to be implemented. Fig. 11 is a block diagram of an example processing platform having programmable circuitry structured to execute the machine-readable example instructions, instantiate, and / or perform the example operations of Fig. 9 and / or 10 to separate the object detector 100 from Fig. 1 to be implemented. Fig. 12 is a block diagram of an exemplary implementation of the programmable circuit of Fig. 11. Fig. 13 is a block diagram of another exemplary implementation of the programmable circuit of Fig. 11. Fig. 14 is a block diagram of an example software / firmware / instruction distribution platform (e.g., one or more servers) for distributing software, instructions, and / or firmware (e.g., corresponding to the example machine-readable instructions of Fig. 9 and / or 10) to client devices that interact with end users and / or consumers (e.g., for licensing, sale and / or use), retailers (e.g., for sale, resale, licensing and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products distributed, for example, to retailers and / or other end users such as direct customers). Fig. Figure 15 shows exemplary patterns generated by the exemplary cluster selection circuit of Fig. 1 and / or 2 can be used to select sample points.

[0003] In general, the same reference numbers are used throughout the drawings and the accompanying written description to indicate identical or similar parts. The illustrations are not necessarily to scale. DETAILED DESCRIPTION

[0004] LiDAR sensors are used in autonomous vehicles, robots, and other systems and applications to obtain 3D point clouds of a scene. LiDAR sensors reflect laser beams off target objects in an environment, resulting in a 3D point cloud representing a scene containing the objects. The 3D point cloud is then processed to detect the object(s) for navigation, collision avoidance, sensing and moving the object(s), etc. However, the resulting 3D point clouds can contain irregular, discontinuous 3D patterns whose density decreases with increasing distance from the LiDAR sensor. Detecting and locating objects in such 3D point clouds can therefore be difficult, especially for objects located far from the sensor.

[0005] Example object detectors and / or other object detection methods, devices, and articles of manufacture (e.g., computer-readable storage media) are disclosed herein. In some examples, mask and / or template matching (e.g., on an overhead view) of a 3D point cloud is performed to select candidate clusters from the 3D point cloud as suggestions for focused object detection and localization. In the examples disclosed herein, the candidate clusters are input to one or more neural networks and / or other machine learning architectures that are trained to output object classification parameters to identify the object(s) represented by the candidate clusters.In the examples disclosed herein, the one or more neural networks and / or other machine learning architectures also output bounding box parameters to locate the detected object(s) represented by the candidate cluster(s). In some examples, an overhead view of the 3D point cloud is used for analysis. An overhead view is also referred to herein as a bird's-eye view, a top-down view, etc., and corresponds to a top-down projection of the 3D point cloud onto a two-dimensional (2D) image, a grid, etc.

[0006] Some example object detectors and / or other object detection methods, devices, and articles of manufacture (e.g., computer-readable storage media) disclosed herein also use mask and / or template matching (e.g., with overhead views, also referred to as bird's-eye views, top views, etc.) of 3D point clouds to train the neural network(s) and / or other machine learning architecture(s). In some examples, the training data used to train the neural network(s) and / or other machine learning architecture(s) includes a 3D point cloud with known, uncorrupted (e.g., labeled) objects and bounding boxes. In some examples, an overhead view of a training cluster is obtained from the 3D point cloud. Using the overhead view, a bounding box for the proposal is found that fits the points of the training cluster.The training cluster, the proposed bounding box, and a bounding box associated with the ground-truth object classification of the training cluster are then used to train the neural network(s) and / or other machine learning architecture(s) to perform object classification and bounding box regression. This training process can be repeated with other training clusters selected from the training 3D point cloud and / or with training clusters selected from other training 3D point clouds until one or more criteria for termination of training are met, e.g., a number of training iterations, an error threshold, etc.

[0007] In the preceding examples, object detection and location processing focus on candidate clusters smaller than the entire 3D point cloud. As a result, the dimensionality and complexity of the neural network(s) and / or other machine learning architectures implementing the preceding examples can be reduced compared to neural networks and / or other machine learning architectures processing the entire 3D point cloud. Such a reduction in dimensionality and complexity can reduce power consumption and computational resource requirements, improve the speed of object detection and localization, and so on. Focusing object detection and localization on candidate clusters instead of the entire 3D point cloud can also improve the accuracy of object classification and frame regression.

[0008] Fig. 1 is a block diagram of an exemplary object detector 100 for detecting and localizing objects in 3D point clouds. The object detector 100 of Fig. 1 may be instantiated (e.g., create an instance, bring it into being for an arbitrary period of time, materialize it, implement it, etc.) by programmable circuitry such as a central processing unit (CPU) that executes initial instructions. Additionally or alternatively, the object detector 100 may be Fig. 1 may be instantiated (e.g., create an instance thereof, bring it into existence for any period of time, materialize, implement, etc.) by (i) an application-specific integrated circuit (ASIC) and / or (ii) a field-programmable gate array (FPGA) structured and / or configured in response to the execution of second instructions to perform operations corresponding to the first instructions. It is understood that some or all of the circuitry of Fig. 1 can be used at the same or different times. Some or all of the circuits of Fig. 1 may, for example, be instantiated in one or more threads that execute concurrently on the hardware and / or in series on the hardware. Furthermore, in some examples, some or all of the circuitry may be Fig. 1 be implemented by microprocessor circuits executing instructions and / or FPGA circuits performing operations to implement one or more virtual machines and / or containers.

[0009] The exemplary object detector 100 of Fig. 1 includes an example cluster selection circuit 105 for selecting one or more example candidate clusters 110 from an example input 3D point cloud 115. The object detector 100 of the illustrated example also includes an example object detection circuit 120 for processing the candidate cluster(s) 110 to detect and locate one or more objects in the 3D point cloud 115. In the illustrated example, the 3D point cloud 115 is generated by one or more LiDAR sensors scanning an environment. The 3D point cloud 115 thus contains points in 3D space that are representative of reflections from one or more objects in the environment. A point can be represented by location coordinates in 3D space, e.g., a pair of horizontal coordinates (e.g., x and y) and a vertical coordinate (e.g., z). However, the 3D point cloud 115 is not limited to a LiDAR point cloud.On the contrary, the 3D point cloud 115 may be any 3D point cloud generated in any way. For example, the 3D point cloud 115 may be created using visible light sensors (e.g., a camera) and a depth sensor (e.g., an infrared sensor). In some examples, the points of the 3D point cloud 115 also have one or more texture values ​​representing the intensity, color, etc., of the various points.

[0010] The cluster selection circuit 105 of the illustrated example implements mask and / or template matching on an overhead view of the 3D point cloud 115 to select the candidate cluster(s) 110. In some examples, a given candidate cluster 110 output by the cluster selection circuit 105 is represented by a set of location coordinates corresponding to the points of the 3D point cloud 115 included in the candidate cluster 110. In some examples, the candidate cluster 110 output by the cluster selection circuit 105 also includes the intensity and / or normal values ​​of the points included in the cluster 110. Further details regarding the implementation and operation of the cluster selection circuit 105 are explained below.

[0011] The object detection circuit 120 of the illustrated example includes an exemplary neural network circuit 125, an exemplary object classification head circuit 130, and an exemplary regression head circuit 135. In the illustrated example, the neural network circuit 125 implements one or more neural networks and / or other machine learning architectures trained to output an exemplary feature vector 140 based on an input candidate cluster 110. In some examples, the neural network circuit 125 implements a PointNet++ neural network or a similar neural network and / or an extension thereof. For example, the neural network circuit 125 implements a neural network pipeline using hierarchical feature learning with multiple abstraction levels. In some examples, each set abstraction level includes a sampling layer, a grouping layer, and a mini-PointNet layer.The sampling and grouping layers generate local point sets and centroids, and the PointNet layer generates a feature vector formed by concatenating the local points. In some examples, there may be multi-scale grouping layers (denoted as PointNetMsg), where groups of points with different radii are formed around a centroid. In some examples, the neural network pipeline also includes multiple backbone layers. For example, two of the backbone layers may be PointNet set abstraction layers with multi-scale grouping (e.g., PointNetMsg layers), followed by two feature propagation layers (e.g., PointNetFP layers) and three set abstraction layers (e.g., PointNetSetAbstraction layers). The feature vector 140 output from the neural network circuit 125 of the illustrated example contains a number of feature points, e.g.,1024 feature points or another number, and is input to the object classification head circuit 130 and the regression head circuit 135. The object classification head circuit 130 processes the feature vector 140 to output exemplary object classification parameters 145 corresponding to the input candidate cluster 110. The regression head circuit 135 processes the feature vector 140 to output the bounding box parameters 150 corresponding to the input candidate cluster 110.

[0012] In the example shown, the object classification head circuit 130 implements an object classification head trained to classify detected objects based on a set of possible object classifications. Examples of such object classifications include automobile / car, truck, pedestrian, bicycle, etc. In some examples, the object classification head implemented by the object classification head circuit 130 includes a three (3) fully connected layer neural network with batch normalization and dropout, followed by a softmax operation for classification. In some examples, the object classification parameters 145 output by the object classification head circuit 130 include respective probabilities that the input candidate cluster 110 corresponds to one of the possible object classifications.In some examples, the object classification parameters 145 output by the object classification head circuit 130 include an identified object classification for the input candidate cluster 110. For example, the identified object classification for the input candidate cluster 110 may correspond to the possible object classification with the highest probability. In some examples, the object classification parameters 145 output by the object classification head circuit 130 include an indication that no object matching the set of possible object classifications was detected, e.g., if the highest object classification probability does not meet a threshold (e.g., 50%, 75%, or another value).

[0013] In the illustrated example, the regression head circuit 135 implements one or more regression heads trained to determine the bounding box parameters 150 for a detected object. In the illustrated example, the regression head circuit 135 implements a number of regression heads equal to the number of possible object classifications, with each regression head corresponding to a respective possible object classification. As explained in more detail below, a given regression head is trained to output its own set of bounding box parameters 150 based on a ground-truth bounding box associated with the given possible object classification corresponding to that regression head.In other words, a given regression head is trained to output a set of bounding box parameters 150 based on the assumption that the detected object has the possible object classification corresponding to that regression head. In some examples, there may be only one regression head, which is trained using training clusters corresponding to all possible classifications and can thus be used for inference of any of the possible classifications. In some examples, a given regression head implemented by regression head circuit 135 comprises a neural network with three (3) fully connected layers.In some examples, the set of bounding box parameters 150 output by a particular regression head implemented by regression head circuitry 135 includes regression values ​​representing differences between the ground-truth bounding box associated with that particular regression head and a predicted or suggested bounding box output by that regression head based on feature vector 140. In some examples, regression head circuitry 135 may implement the bin-based bounding box regression head and loss function used by PointRCNN. Further details regarding the implementation and operation of regression head circuitry 135 are explained below.

[0014] The exemplary object detector 100 of Fig. 1 also includes an example training circuit for training the object detection circuit 120, which includes training the neural network(s) and / or other machine learning architectures implemented by the neural network circuit 125, the object classification head circuit 130, and the regression head circuit 135 based on training data. The training circuit of the illustrated example includes an example bounding box generation circuit 155, an example regression calculation circuit 160, an example loss function evaluation circuit 165, and an example neural network update circuit 170. In the illustrated example, the bounding box generation circuit 155 generates a proposal bounding box for a training cluster selected by the cluster selection circuit 105 from a 3D point cloud included in the training data.The respective object classification for the object represented by the training cluster is known from the training data. In the example shown, the regression calculation circuit 160 compares the proposed bounding box generated by the bounding box generation circuit 155 with a ground-truth bounding box known from the training data to represent this object classification. The regression calculation circuit 160 also outputs exemplary training regression values ​​175 representing the differences between the known ground-truth bounding box for object classification and the proposed bounding box generated by the bounding box generation circuit 155.

[0015] During training, the training cluster selected from the 3D point cloud by the cluster selection circuit 105 is also input to the object detection circuit 120. As described above, the object classification head 130 outputs predicted object classification parameters 145, such as object classification probabilities, for the training cluster. As described above, the regression head circuit 135 implements one or more regression heads that output one or more sets of bounding box parameters 150 containing regression values ​​representative of the differences between the ground truth bounding box(es) and the predicted bounding box(es) determined by the regression head(s). In the illustrated example of Fig. 1, the predicted object classification parameters 145 and the predicted bounding box parameters 150 output by the object classification head circuit 130 and the regression head circuit 135 during training are shown together as an example of predicted object classification and bounding box regression values ​​180.

[0016] In the illustrated example, the loss function evaluation circuit 165 evaluates one or more loss functions based on the training regression values ​​175 output by the regression calculation circuit 160 and the predicted object classification and bounding box regression values ​​180 output by the object detection circuit 120. The loss function(s) may be any type and / or any number of loss functions capable of quantifying the error between the training regression values ​​175 and the predicted object classification and bounding box regression values ​​180. Example loss functions implemented by the loss function evaluation circuit 165 are described in more detail below.

[0017] In the illustrated example, neural network update circuit 170 uses error values ​​output by loss function evaluation circuit 165 based on the evaluation of the loss function(s) to update the neural network(s) and / or other machine learning algorithm(s) implemented by object detection circuit 120. For example, neural network update circuit 170 may implement one or more gradient descent and / or other algorithms, such as Adam gradient descent, RMSProp gradient descent, AdaGrad gradient descent, etc.that use the error value(s) output by the loss function evaluation circuit 165 to update the layer weights and / or other parameters of the neural network layers implemented by the neural network circuit 125 of the neural network, the object classification head circuit 130, and / or the regression head circuit 135 of the object detection circuit 120.

[0018] Fig. 2 is a block diagram of an exemplary implementation of the cluster selection circuit 105 used in the object detector 100 of Fig. 1. The cluster selection circuit 105 of Fig. 2 may be instantiated (e.g., create an instance, bring it into existence for an arbitrary period of time, materialize, implement, etc.) by a programmable circuit such as a central processing unit (CPU) that executes initial instructions. Additionally or alternatively, the cluster selection circuit 105 may be Fig. 2 may be instantiated (e.g., create an instance thereof, bring it into existence for any period of time, materialize, implement, etc.) by (i) an application-specific integrated circuit (ASIC) and / or (ii) a field-programmable gate array (FPGA) structured and / or configured in response to the execution of second instructions to perform operations corresponding to the first instructions. It is understood that some or all of the circuitry of Fig. 2 can be used at the same or different times. Some or all of the circuits of Fig. 2 may, for example, be instantiated in one or more threads that execute concurrently on the hardware and / or in series on the hardware. Furthermore, in some examples, some or all of the circuitry may be Fig. 2 be implemented by microprocessor circuits executing instructions and / or FPGA circuits performing operations to implement one or more virtual machines and / or containers.

[0019] The exemplary cluster selection circuit (105) in Fig. 2 includes an exemplary filter circuit (205), an exemplary overhead view projection circuit (210), an exemplary view fill circuit (215), an exemplary sample point selection circuit (220), an exemplary mask application circuit (225), and an exemplary cluster identification circuit (230). In the Fig. In the example illustrated in Figure 2, filter circuitry 205 performs one or more filtering operations on a 3D input point cloud, such as 3D input point cloud 115, to generate a filtered 3D point cloud. For example, filter circuitry 205 may perform one or more filtering operations to reduce or eliminate points from the 3D point cloud that correspond to reflections from the ground and / or the sides of the scene. In an example implementation, the filter circuit 205 (i) creates a histogram with 30 bins (or another number of bins) representing the heights (z-values) of the 3D points in the 3D point cloud, (ii) finds bin B with the maximum number of points in the histogram, (iii) rejects points whose heights are in bin B or lower bins (e.g., with a height less than that represented by bin B+1), and retains all 3D points in grid box(es) with at least one point remaining after step (iii).

[0020] The overhead view projection circuit 210 of the illustrated example projects the filtered 3D point cloud (or the input 3D point cloud directly if the filter 205 has been disabled or omitted) based on an overhead, top, or bird's-eye view projection to generate an exemplary two-dimensional (2D) overhead view of the input 3D point cloud. The resulting overhead view may be, for example, a 2D image or a grid map created by projecting the points of the filtered 3D point cloud (or the input 3D point cloud directly if the filter 205 has been disabled or omitted) downward to the lowest (e.g., ground-level) horizontal plane of 3D space corresponding to the input 3D point cloud.In some examples, the overhead view projection circuit 210 reduces the dimensionality of the overhead view by using a grid map and grouping the projected points into cells or patches of the grid. For example, the overhead view projection circuit 210 may determine that a cell (also called a patch) of the grid map forming the overhead view is occupied and fill that cell if the cell contains at least one projected point from the 3D point cloud. Fig. 2 shows an example of an overhead view 235 generated by the overhead view projection circuit 210 for the input 3D point cloud 215.

[0021] The view fill circuit 215 of the shown example fills in gaps in the overhead view generated by the overhead view projection circuit 210. In some examples, the view fill circuit 215 may apply a dilation kernel to the overhead view at various displacements. For a given displacement, the view fill circuit 215 may fill a pixel of the overhead view corresponding to the center of the dilation kernel if at least one pixel covered by the dilation kernel is not empty. For example, if the overhead view is formed by a grid map, the dilation kernel may be an N-by-N dilation kernel, such as a 9x9 dilation kernel, and the view fill circuit 215 may fill a center cell of the grid map if at least one cell covering the N-by-N dilation kernel is not empty. Fig. 2 shows an example of a padded overhead view 240 generated by the view padding circuit 215 for the input 3D point cloud 215.

[0022] The sample point selection circuit 220 of the illustrated example samples the padded overhead view output by the view padding circuit 215 (or the overhead view output by the view projection circuit 210 if the view padding circuit 215 is disabled or omitted) to generate a sampled overhead view. For example, the sample point selection circuit 220 samples the padded overhead view output by the view padding circuit 215 (or the overhead view output by the view projection circuit 210 if the view padding circuit 215 is disabled or omitted) evenly across the area of ​​the overhead view, unevenly across the area of ​​the overhead view (e.g., to avoid sample points near the edges of the view), etc.In some examples, when the padded overhead view (or the overhead view output by view projection circuit 210 when view padding circuit 215 is disabled or omitted) is represented by a grid map, sample point selection circuit 220 evenly samples the grid map by selecting sample points using a regular pattern or based on a non-uniform pattern that avoids points near the edge of the overhead view. Fig. 15 shows exemplary patterns 1505 and 1510 that can be used by the sampling point selection circuit 220. A first example of a uniform pattern 1505 shown in Fig. 15, comprises a 5x5 pixel block 1515 centered at position (2,2), with the center sample and 4 corner samples selected. In the uniform pattern 1505, the 5x5 block of selected pixels 1515 repeats at the central locations (2,6), (2,10), (6,2), (6,6), etc. A second example of a uniform pattern 1510 shown in Fig. 15 includes an exemplary 3x3 block of pixels 1520 centered at position (1,1), with the center sample and four corner samples selected. In the uniform pattern 1510, the 3x3 block of selected pixels 1520 repeats centered at (1,3), (3,1), (3,3), etc. Fig. 2 shows an exemplary sampled overhead view 245 generated by the sample point selection circuit 220 for the input 3D point cloud 215.

[0023] The mask application circuit 225 of the illustrated example applies masks (and / or templates) to the padded overhead view output by the view padding circuit 215 (or to the overhead view output by the view projection circuit 210 if the view padding circuit 215 is disabled or omitted) at sample points of the sampled overhead view output by the sample point selection circuit 220 to identify possible point groups for use in object detection. In some examples, corresponding masks (and / or templates) are generated from training data and used to identify candidate clusters that are likely to represent corresponding possible object classifications.For example, a first mask (and / or a first template) may be generated from training data and used to identify candidate clusters representative of a first possible object classification (e.g., car or motor vehicle), a second mask (and / or a first template) may be generated from the training data and used to identify candidate clusters representative of a second possible object classification (e.g., pedestrian), etc. In some examples, multiple masks (and / or templates) may be generated per class (e.g., multiple masks / templates for cars, multiple masks / templates for pedestrians, etc.). In some examples, one or more representative masks (and / or templates) may be obtained by combining individual training masks determined for multiple classes (e.g., such that one representative mask may correspond to objects with different classifications). Fig. Figure 2 shows an exemplary mask 250 for a first possible object classification, which is used by the mask application circuit 225 to identify candidate clusters. Further exemplary masks and templates, as well as examples for creating such masks and templates, are described in more detail below.

[0024] As represented in the example shown by an exemplary inset 255, the mask application circuit 225 applies the mask 250 to the padded overhead view output by the view padding circuit 215 (or to the overhead view output by the view projection circuit 210 if the view padding circuit 215 is disabled or omitted) at sample points of the sampled overhead view output by the sample point selection circuit 220. For example, the inset 255 illustrates how the mask application circuit 225 applies the mask at three different sample points by centering the mask on each of the different sample points.For a given sample point, the mask application circuit 225 includes the points of the padded overhead view (or the overhead view output by the view projection circuit 210 if the view padding circuit 215 is disabled or omitted) covered by the footprint of the mask 250 into a potential cluster, which is evaluated by the cluster identification circuit 230. The mask application circuit 225 repeats this process for different sample points and for different masks (and / or templates) corresponding to different possible object classifications.

[0025] The cluster identification circuit 230 of the shown example evaluates the potential clusters output from the mask application circuit 225 to identify candidate clusters to be used for object detection (e.g., for input to the object detection circuit 120, the bounding box generation circuit 155, etc.). In some examples, when the mask application circuit 225 applies masks to identify the potential clusters, the cluster identification circuit 230 determines that a potential cluster is a candidate cluster if the number of points of the potential cluster covered by the mask used to select the cluster meets an occupancy target. In some examples of mask-based selection, the occupancy target corresponds to a threshold occupancy fraction. For a potential cluster selected based on a particular mask (e.g.,Once a potential cluster has been identified (e.g., mask 250), cluster identification circuit 230 calculates the occupancy fraction for the potential cluster as the fraction of grid cells of the mask (e.g., mask 250) occupied by points of the potential cluster. If the occupancy fraction satisfies (e.g., is greater than, greater than, or equal to, etc.) the occupancy fraction threshold (e.g., 0.7 or another value), cluster identification circuit 230 determines that the potential cluster is a candidate cluster.

[0026] In some examples, when the mask application circuit 225 applies templates to identify the potential clusters, the cluster identification circuit 230 determines that a potential cluster is a candidate cluster if the number of points of the potential cluster covered by the template used to select the cluster meets an occupancy target. In some examples, the occupancy target for template-based selection is based on a normalized correlation coefficient. For example, for a potential cluster identified based on a particular template, the cluster identification circuit 230 calculates the normalized correlation coefficient between the potential cluster and the template. If the normalized correlation coefficient meets a threshold (e.g., 0.7 or another value) (e.g., is greater than, greater than, or equal to, etc.), the candidate cluster is selected.), the cluster identification circuit 230 determines that the potential cluster is a candidate cluster.

[0027] In the example shown, cluster identification circuitry 230 represents a given candidate cluster as the set of points from the original 3D input point cloud that match the points of the filled overhead view (or the overhead view output by view projection circuitry 210 if view fill circuitry 215 is disabled or omitted) included in the candidate cluster. After identifying a set of candidate clusters from the potential clusters output by mask application circuitry 225, cluster identification circuitry 230 implements a non-maximum suppression technique to select candidate clusters for downstream object detection processing (e.g., for input to object detection circuitry 120, etc.).For example, cluster identification circuitry 230 sorts the set of candidate clusters in descending order of score, and to break ties, in descending order of occupancy fraction or normalized correlation coefficient. In some examples, cluster identification circuitry 230 discards candidate clusters that share at least a threshold number or fraction of scores (e.g., 10% or other value) with another candidate cluster higher in the sorted set. Cluster identification circuitry 230 then outputs the sorted set of candidate clusters (e.g., after discarding candidate clusters that meet the commonality threshold) in descending order for downstream processing until the sorted set of candidate clusters is empty.

[0028] Fig. 3-5 illustrate exemplary masks and exemplary templates used by the cluster selection circuit 105 of Fig. 1 and / or 2 can be used. Fig. Figure 3 shows an example of an overhead view 305 determined from training data. The overhead view 305 contains example ground-truth objects 310-330. In the example shown, the ground-truth objects (310-330) correspond to a classification of car / automotive objects.

[0029] Fig. 4-5 show examples of masks and templates created from training data, such as the training data 305 of Fig. 3, for use by the cluster selection circuit 105 of Fig. 1 and / or 2 can be generated. Fig. For example, Figure 4 shows an exemplary template 405 and a corresponding exemplary mask 410 that are based on the information contained in the exemplary training data 305 of Fig. 3. To generate the template 405 and the mask 410, a horizontal plane covering the 3D point cloud of the training data 305 is divided into a regular grid, and points from the 3D point cloud of the training data 305 are projected downward onto this grid. The number of points in each grid cell of the grid forms a grid map, which represents an overhead view of the 3D point cloud of the training data 305.

[0030] In the example shown, a template, such as template 405, corresponds to a pattern containing the expected number of points in each grid cell for the object classification (e.g., car) represented by this template 405. In the example shown, a mask, such as mask 410, corresponds to a binary mask that is set to true (e.g., logical 1) over the entire range of template 405. Fig. 4 also shows an example template 415 and a corresponding example mask 420 corresponding to another possible object classification (e.g., pedestrian in the example shown).

[0031] Fig. Figure 5 shows another exemplary technique for generating masks from training data, such as the training data 305 of Fig. 3, for use by the cluster selection circuit 105 of Fig. 1 and / or 2. In the Fig. In the example illustrated in Figure 5, an example mask 510 is generated from the template 405 by setting values ​​of the mask 510 to true (e.g., logical 1) for grid cells of the template 405 that contain at least one point corresponding to the ground truth object. This differs from the mask 410, which has a value of true (e.g., logical 1) across the base area of ​​the template 405. Fig. 4 also shows an example template 515 and a corresponding example mask 520 created using this example technique for another possible object classification (e.g., pedestrian in the example shown).

[0032] Fig. 6 is a block diagram of an exemplary implementation of the bounding box generation circuit 155 and the regression calculation circuit 160 used in the object detector 100 of Fig. 1. The bounding box generation circuit 155 and the regression calculation circuit 160 of Fig. 6 may be instantiated (e.g., create an instance, bring it into existence for an arbitrary period of time, materialize, implement, etc.) by programmable circuitry such as a central processing unit (CPU) that executes first instructions. Additionally or alternatively, the bounding box generation circuit 155 and the regression calculation circuit 160 of Fig. 6 may be instantiated (e.g., create an instance thereof, bring it into being for any period of time, materialize, implement, etc.) by (i) an application-specific integrated circuit (ASIC) and / or (ii) a field-programmable gate array (FPGA) structured and / or configured in response to the execution of second instructions to perform operations corresponding to the first instructions. It is understood that some or all of the circuitry of Fig. 6 can be used at the same or different times. Some or all of the circuits of Fig. 6 may, for example, be instantiated in one or more threads that execute concurrently on the hardware and / or in series on the hardware. Furthermore, in some examples, some or all of the circuitry may be Fig. 6 be implemented by microprocessor circuits executing instructions and / or FPGA circuits performing operations to implement one or more virtual machines and / or containers.

[0033] The exemplary bounding box generation circuit 155 in Fig. 6 includes an exemplary overhead view projection circuit (605) and an exemplary bounding box adjustment circuit (610). The overhead view projection circuit 605 of the illustrated example accepts a 3D point cloud corresponding to a candidate cluster selected by the cluster selection circuit 105. The overhead view projection circuit 605 projects the 3D point cloud of the candidate cluster based on an overhead, top, or bird's-eye view to generate a 2D overhead view of the input 3D point cloud. The resulting overhead view may be, for example, a 2D image or a grid map created by projecting the points of the 3D point cloud of the candidate cluster downward to the lowest (e.g., ground-level) horizontal level of 3D space corresponding to the 3D point cloud.In some examples, the overhead view projection circuit 605 reduces the dimensionality of the overhead view by using a grid map and grouping the projected points into cells or patches of the grid. For example, the overhead view projection circuit 605 may determine that a cell (also called a patch) of the grid map forming the overhead view is populated and populate that cell if the cell contains at least one projected point from the candidate cluster's 3D point cloud. Fig. 6 shows an example overhead view 615 generated by the overhead view projection circuit 605 for an example 3D point cloud 620 corresponding to an input candidate cluster provided by the cluster selection circuit 105.

[0034] The bounding box adjustment circuit 610 of the illustrated example adjusts a proposal bounding box to the points of the 3D point cloud of the input candidate cluster based on the overhead view output by the overhead view projection circuit 605. The points of the overhead view correspond to the projected points of the candidate cluster. Thus, the bounding box adjustment circuit 610 can adjust the proposal bounding box to the points of the 3D point cloud of the input candidate cluster by adjusting the proposal bounding box in the horizontal dimension based on the overhead view output by the overhead view projection circuit 605 and adjusting the proposal bounding box in the vertical dimension based on the height of the 3D point cloud 620 of the input candidate cluster.In some examples, for a particular candidate cluster selected by cluster selection circuit 105, bounding box adjustment circuit 610 determines a corresponding rectangular 3D bounding box, referred to as a cuboid, to represent the position, size, and orientation of the candidate cluster. For example, bounding box adjustment circuit 610 may implement an algorithm to adjust a rectangular 2D bounding box to the overhead view corresponding to the 3D point cloud of the input candidate cluster. Bounding box adjustment circuit 610 also determines the minimum and maximum vertical coordinates of the points in the 3D point cloud of the candidate cluster. Bounding box adjustment circuit 610 then creates the cuboid (e.g.,a rectangular 3D bounding box) as a combination of the rectangular 2D bounding box in the horizontal plane and the minimum and maximum vertical coordinates in the vertical plane. As in . Fig. 6, a cuboid has a shorter and a longer horizontal dimension, denoted as S and L, respectively, and a vertical height denoted as Z. The bounding box adjustment circuit 610 also determines an angle between the shorter horizontal dimension (e.g., the shorter side of the rectangular 2D bounding box) and the forward direction, denoted as PRz, which represents an orientation of the proposal bounding box. Fig. 6 shows an example of a proposal bounding box 625 generated by the bounding box adaptation circuit 610 for the 3D point cloud 620 of the example input candidate cluster. During inference processing, the 3D point cloud 620 of the input candidate cluster and, in some examples, the proposal bounding box 625 generated from the 3D point cloud 620 of the input candidate cluster are provided to the object detection circuitry for object detection (e.g., for deriving / estimating classification and regression parameters, etc.).

[0035] During training, the overhead view projection circuit 605 and the bounding box matching circuit 610 operate similarly to determine a 2D-oriented rectangular bounding box for each ground truth cluster represented in the input training data. Furthermore, during training, the bounding box matching circuit 610 compares each input candidate cluster to the ground truth clusters present in the source point cloud to check for matches. In some examples, the bounding box matching circuit 610 identifies a match when the point intersection-over-union (IoU) of the 2D bounding boxes is greater than a threshold, such as 0.5 or another value.In some examples, the bounding box adaptation circuit 610 determines the point IoU as a fraction equal to the number of common points in the two clusters divided by the number of points in the union of the two clusters. In some examples, the bounding box adaptation circuit 610 may use a conventional IoU as the adaptation metric, where the conventional IoU is determined as a fraction equal to the intersection area of ​​bounding boxes divided by the union area of ​​bounding boxes. The matched candidate cluster and ground-truth cluster pairs are used by the regression calculation circuit 160 to generate regression parameters and train the object detection circuit 120 of the object detector 100.

[0036] For example, during training, the regression calculation circuit 160 of the illustrated example compares the proposal bounding box generated by the bounding box adaptation circuit 610 with the ground truth bounding box for the matching ground truth object cluster to generate training regression parameters to train the object detection circuit 120 of the object detector 100. In some examples, the regression calculation circuit 160 represents the proposal bounding box by determining (i) the center coordinate of the proposal bounding box, (ii) the length, width, and height of the proposal bounding box (where the length and width of the longest and longest bounding boxes are the lengths and widths, respectively), and (iii) the length, width, and height of the proposal bounding box.shortest horizontal dimension of the proposal bounding box and the height corresponds to the vertical dimension of the proposal bounding box), and (iii) the angle, denoted PRz, between the width (shortest horizontal dimension) of the proposal bounding box and the forward direction. Similarly, in some examples, the regression calculation circuit 160 represents the ground truth bounding box by using (i) the center coordinate of the ground truth bounding box, (ii) the length, width, and height of the ground truth bounding box (where the length and width correspond to the longest and shortest horizontal dimensions of the ground truth bounding box, respectively, and the height corresponds to the vertical dimension of the ground truth bounding box), and (iii) the angle, denoted GRz, between the width (shortest horizontal dimension) of the ground truth bounding box and the forward direction. . Fig. 6 shows an example ground truth bounding box 630 corresponding to the example ground truth object (or example classification of the ground truth object) represented by the example training 3D point cloud 620.

[0037] In the example shown, the regression calculation circuit 160 uses the previous parameters of the proposal bounding box and the ground-truth bounding box to calculate the training regression parameters. The parameters of the proposal bounding box and the ground-truth bounding box can be represented, for example, as follows: (Gx, Gy, Gz) represents the x, y, and z coordinates of the center of the ground truth bounding box, where x represents the forward direction in the horizontal plane, y represents the perpendicular direction in the horizontal plane, and z represents the vertical plane; (Px, Py, Pz) are the corresponding x, y, and z coordinates of the center of the proposal bounding box; (GD S , GD L , GD Z ) represents the dimensions of the ground truth bounding box, where S represents the shortest horizontal dimension, L represents the longest horizontal dimension, and Z represents the vertical dimension; (PD S , PD L , PD Z ) represents the corresponding dimensions of the proposal bounding box; GRz is the angle between the shortest horizontal dimension of the ground truth bounding box (e.g., the width) and the forward direction (e.g., x); and PRz is the angle between the shortest horizontal dimension of the proposal bounding box (e.g., width) and the forward direction (e.g., x).

[0038] Based on the preceding notation, the regression calculation circuit 160 of the illustrated example generates seven (7) exemplary training regression parameters given by equations 1-7: tx=Gx−PxPDL ty=Gy−PyPDL tz=Gz−PzPDz tw=lnGDwPDS tl=lnGDlPDL th=lnGDzPDz a=GRz−PRz

[0039] In Equations 1-3, tx, ty, and tz represent the regression differences of the bounding box center in the x, y, and z directions, respectively. In Equations 4-6, tw, tl, and th represent the regression differences of the bounding box sizes in the dimensions width, length, and height, respectively. In Equation 7, a represents a regression difference of the bounding box orientation.

[0040] As in Fig. As shown in Figure 1, during training, the neural network circuit 125 processes the 3D point cloud corresponding to the ground-truth object to generate the feature vector 140. The object classification head circuit 130 processes the feature vector 140 to output the object classification parameters 145, which contain corresponding classification probabilities for the various possible object classifications. The regression head circuit 135 implements regression heads, each corresponding to the various possible object classifications. The regression heads implemented by the regression head circuit 135 process the feature vector 140 to determine corresponding sets of bounding box parameters 150, which contain corresponding sets of inference regression parameters. (tx^, ty^, tz^, tw^, tl^, th^, a⌢), which correspond to the training regression parameters (tx, ty, tz, tw, tl, th, a) of equations 1-7. However, each regression head determines its respective set of inference regression parameters (tx^, ty^, tz^, tw^, tl^, th^, a⌢), using the feature vector and the respective possible object classification according to the regression head.

[0041] Again with reference to Fig. 1 In some examples, the loss function evaluation circuit 165 implements one or more loss functions based on the classification probabilities contained in the object classification parameters 145 output by the object classification head circuit 130, the ground truth object classification corresponding to the input training cluster (e.g., the 3D point cloud), the training regression parameters (tx, ty, tz, tw, tl, th, , a), and the inference regression parameters (tx^, ty^, tz^, tw^, tl^, th^, a⌢), contained in the respective sets of bounding box parameters 150 output by the regression heads implemented by the regression head circuit 135. For example, the loss function evaluation circuit 165 may implement an exemplary classifier loss function, an exemplary regression loss function, an exemplary angular loss function, and an exemplary total loss function based on Equations 8-11, which are: Classifier loss: C = negative logarithmic probability (ground truth class, prediction probabilities) Regression loss:R=smoothing−L1−loss((tx−tx^)+(ty−ty^)+(tz−tz^)+(tw−tw^)+(tl−tl^)+(th−th^)) Angle loss: A = smoothing loss (sin(a−a^)) Total loss = C + K * (R + A)

[0042] In Equation 11, K is a constant, e.g., the value 5 or another value. In some examples, when calculating the regression loss function R from Equation 9 and the angular loss function A from Equation 10, the loss function evaluation circuit 165 restricts the evaluation to the inference regression parameters (tx^, ty^, tz^, tw^, tl^, th^, a^), output by the regression head corresponding to the ground truth object classification.

[0043] Referring again to Fig. 1, in some examples, the neural network update circuit 170 uses the total loss error output by the loss function evaluation circuit 165 based on Equation 11 to update the layer weights and / or other parameters of the neural network layers implemented by the neural network circuit 125, the object classification head circuit 130, and / or the regression head circuit 135 of the object detection circuit 120. In some examples, when updating the layer weights and / or other parameters of the neural network layers implemented by the regression head circuit 135, the neural network update circuit 170 limits the update to those layer weights and / or other parameters that correspond to the particular regression head associated with ground-truth object classification.

[0044] Fig. 7 is a block diagram of an exemplary mask generation circuit 705 and an exemplary template generation circuit 710 for generating exemplary masks and / or exemplary templates, such as the exemplary masks and templates of Fig. 3-5, for use by the object detector 100 of Fig. 1. The mask generation circuit (705) and / or the template generation circuit (710) of Fig. 7 may be instantiated (e.g., create an instance thereof, bring it into existence for an arbitrary period of time, materialize, implement, etc.) by a programmable circuit such as a central processing unit (CPU) that executes first instructions. Additionally or alternatively, the mask generation circuit 705 and / or the template generation circuit 710 of Fig. 7 may be instantiated (e.g., create an instance thereof, bring it into existence for any period of time, materialize, implement, etc.) by (i) an application-specific integrated circuit (ASIC) and / or (ii) a field-programmable gate array (FPGA) structured and / or configured in response to the execution of second instructions to perform operations corresponding to the first instructions. It is understood that some or all of the circuitry of Fig. 7 can be used at the same or different times. Some or all of the circuits of Fig. 7 may, for example, be instantiated in one or more threads that execute concurrently on the hardware and / or in series on the hardware. Furthermore, in some examples, some or all of the circuitry may be Fig. 7 be implemented by microprocessor circuits executing instructions and / or FPGA circuits performing operations to implement one or more virtual machines and / or containers.

[0045] In some examples, the mask generation circuit 705 and the template generation circuit 710 of Fig. 7 as part of the training circuit of the object detector 100 of Fig. 1. In some examples, the training circuit of the object detector 100 may include both the mask generation circuit 705 and the template generation circuit 710. In some examples, the training circuit of the object detector 100 may include the mask generation circuit 705 or the template generation circuit 710, but not both.

[0046] In the example of Fig. 7, the mask generation circuit 705 and the template generation circuit 710 access training 3D point clouds from ground truth training data in an exemplary ground truth database 715. The mask generation circuit 705 and the template generation circuit 710 project the training 3D point clouds onto overhead views, such as the overhead grid maps described above. The mask generation circuit (705) and the template generation circuit (710) identify patterns representative of various possible object classes from the overhead grid maps. For example, the patterns identified by the template generation circuit 710 are templates represented by an array [x ij ] of size R*C, where each x ijis a number of points in the grid box. The templates are stored by the template generation circuit 710 in an exemplary template database 720. As another example, the patterns identified by the mask generation circuit 705 are binary masks derived from the templates, such that the mask has the value true (e.g., logical -1) if the number of points is non-zero, and false (logical -0) otherwise. The masks are stored by the mask generation circuit 705 in an exemplary mask database 725.

[0047] The following describes an example of a mask generation algorithm implemented by mask generation circuit 705. Mask generation circuit 705 uses the mask generation algorithm to generate representative masks for each possible object class. As described below, the mask generation algorithm generates one or more representative masks for a given possible object class as a union of a group of individual masks obtained from 3D point clouds corresponding to various instances of ground-truth objects belonging to that given possible object class. Thus, masks for each possible object class in the training dataset (e.g., car, pedestrian, etc.) are generated by merging individual masks corresponding to that possible object class.A similar template generation algorithm may be implemented by the template generation circuit 710 to generate templates instead of masks, with the difference that a normalized correlation coefficient (NCC) is used to measure the similarity of templates instead of an occupancy fraction to measure the similarity of masks.

[0048] Fig. Figure 8 shows a representative example mask 805 generated by the mask generation circuit 705 for a particular possible object class using the mask generation algorithm described below. In the example shown, the representative mask 805 is generated as the union of a first example mask 810 and a second example mask 815 determined for various instances of objects belonging to the given possible object class.

[0049] Back to Fig. 7: The exemplary mask generation algorithm implemented by the mask generation circuit 705 is described using the following terminology:

[0050] Set of masks S: This term refers to a set of masks {Mi} obtained from training data, where the set has a number of elements K(S).

[0051] Union of masks: This term refers to a mask that is obtained from a set of masks {Mi} as follows: (i) Finding the maximum of each dimension (rows, columns) over the set of masks; (ii) creating a new binary mask M with these maximum dimensions, setting all pixels to false; (iii) centering each source mask Mi in this mask M; and (iv) at each overlapping pixel in the two masks, setting the corresponding pixel in M ​​to true if the overlapping pixel in Mi is true.

[0052] Representative mask R(S) from a set of masks S: This term refers to a union of masks of the masks in the set S.

[0053] Occupancy rate of the mask Mi in relation to the representative mask R(S): This term refers to an occupancy rate determined as follows: (i) Center mask Mi in R(S); (ii) finding a clipping mask that is the intersection of the two masks, or in other words, the clipping mask is a mask of the same size as R(S) in which a pixel is true only if the overlapping pixels are true in both masks; (iii) finding the number of true pixels in the clipping mask and in R(S); and (iv) the occupancy fraction f(i) is the ratio (number of true pixels in the clipping mask) / (number of true pixels in R(S)).

[0054] Minimum occupancy fraction, f(S), of the set S: This term refers to the minimum of the occupancy fractions of all masks in the set S, which can be mathematically represented as min {f(i)}.

[0055] Occupancy fraction thresholds T(K): This term refers to a set of thresholds that depend on the size K (number of masks) of a set of masks S. A set of masks is valid if f(S) > T(K(S)).

[0056] Taking the above terminology into account, Table 1 provides an example pseudocode for a mask generation algorithm implemented by the mask generation circuit 705.

[0057] The example pseudocode in Table 1 uses an example helper function, Test_Union(), listed in Table 2.

[0058] The following explains the example mask generation algorithm, which corresponds to the pseudocode in Table 1. The algorithm greedily joins masks that are similar (and therefore closely represented by their union). In operations (i)-(iii) in Table 1, each mask is used to create a set with 1 member. Pairs of sets are tested to check whether the union of masks represents each of the elements in the set accurately enough (as evaluated by the helper function Test_Union in Table 2). Pairs that pass the test are stored as tuples (S i , S j , f) included in the Valid_Pairs_List, where f is the minimum occupancy fraction determined with the auxiliary function Test_Union from Table 2

[0059] At each iteration of operation (iv) in Table 1, the best tuple (S p , S q , f) are removed from the Valid_Pairs_List. The union set S t = Sp US q is used to S p and S q in the set list S, and is also used in any tuples in the Valid_Pairs_List where S p and S q involved. The new tuples are checked against thresholds appropriate to the size of the merged set and retained if they meet the threshold.

[0060] The example pseudocode in Table 1 and Table 2 results in sets of matching masks, each with a representative mask used for selecting candidate clusters, as described above.

[0061] In the example pseudocode of Table 1 and Table 2, the minimum occupancy fraction of each candidate set S is tested against a threshold T(K) that depends on the size of the set K(S). In some examples, T(K) decreases sublinearly with K to allow for some variation in the union mask with increasing K. The threshold can be set, for example, according to Equation 12, which is: T(K)=C1+CeaK

[0062] In Equation 12, C1, C, and a are adjusted to satisfy one or more target conditions. For example, C1 and a can be adjusted so that T(2) = 0.7 and T(200) = 0.5. Furthermore, C1 can be set to the minimum value for large K (e.g., C1 = 0.3).

[0063] In some examples, the pseudocode of Table 1 and Table 2 may form the basis for an example template generation algorithm implemented by template generation circuit 710. However, during template generation, the similarity of the templates is compared using a normalized correlation coefficient instead of the occupancy fraction. For example, the occupancy fraction calculation in the pseudocode of Table 1 and Table 2 may be replaced with the following example normalized correlation coefficient calculation.

[0064] First, each template pattern consists of an arrangement [x ij ] of size R*C, where each x ij specifies a number of points in the grid box.

[0065] Let T be the number of frames that are true in the binary mask.

[0066] Then the mean value µ is given by μ=∑xijT given.

[0067] Let yij=xij−μ and zij=yij∑y2kl.

[0068] Each normalized pattern [e.g. ij ] has a mean of 0 and a unit norm.

[0069] Finally, to obtain the normalized correlation coefficient between two normalized templates: (i) overlapping the templates so that their centers coincide and expanding as needed, padding with 0; and (ii) Calculate the normalized correlation coefficient, c ij , as c ij = < z i , e.g. j >, where <,> represents the dot product.

[0070] In addition, instead of a union of masks, the representative template of a set of templates is determined by taking the mean of the unnormalized templates and then normalizing them.

[0071] In some examples, initial K-means clustering may also be performed to group masks of similar dimensions, and the algorithms described above can be applied to each group.

[0072] In some examples, the object detector 100 includes means for selecting clusters. The means for selecting clusters may be implemented, for example, by the cluster selection circuit 105. In some examples, the cluster selection circuit 105 may be implemented by a programmable circuit, such as the exemplary programmable circuit 1112 in Fig. 11. For example, the cluster selection circuit 105 may be implemented by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in blocks 910 of Fig. 9 and Fig. 1010 of Fig. 10. In some examples, the cluster selection circuit 105 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the cluster selection circuit 105 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the cluster selection circuit 105 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0073] In some examples, the object detector 100 comprises means for processing in a neural network. The means for processing the neural network may, for example, be implemented by the neural network circuit 125. In some examples, the neural network circuit 125 may be implemented by a programmable circuit, such as the programmable circuit 1112 in Fig. 11. For example, the neural network circuit 125 may be instantiated by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in blocks 920 of Fig. 9 and Fig. 1030 of Fig. 10. In some examples, the neural network circuit 125 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the neural network circuit 125 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the neural network circuit 125 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0074] In some examples, the object detector 100 includes means for implementing an object classification head. The means for implementing an object classification head may be implemented, for example, by the object classification head circuit 130. In some examples, the object classification head circuit 130 may be implemented by a programmable circuit, such as the programmable circuit 1112 in Fig. 11. For example, the object classification head circuit 130 may be instantiated by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in blocks 925 of Fig. 9 and Fig. 1035 of Fig. 10. In some examples, the object classification head circuit 130 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the object classification header circuit 130 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the object classification header circuit 130 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0075] In some examples, the object detector 100 includes means for implementing one or more regression heads. For example, the means for implementing one or more regression heads may be implemented by the regression head circuit 135. In some examples, the regression head circuit 135 may be implemented by a programmable circuit such as the exemplary programmable circuit 1112 of Fig. 11. For example, the regression head circuit 135 may be instantiated by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in blocks 930 of Fig. 9 and Fig. 1035 of Fig. 10. In some examples, the regression head circuit 135 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the regression header circuit 135 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the regression header circuit 135 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0076] In some examples, the object detector 100 includes means for generating bounding boxes. For example, the means for implementing bounding boxes may be implemented by the bounding box generation circuit 155. In some examples, the bounding box generation circuit 155 may be instantiated by a programmable circuit, such as the programmable circuit 1112 in Fig. 11. The bounding box generation circuit 155 may be implemented, for example, by the exemplary microprocessor 1200 of Fig. 12 which executes machine-executable instructions as described at least in block 1020 of Fig. 10. In some examples, the bounding box generation circuit 155 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13 that is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the bounding box generation circuit 155 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the bounding box generation circuit 155 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0077] In some examples, the object detector 100 includes means for calculating regression values. For example, the means for calculating regression values ​​may be implemented by the regression calculation circuit 160. In some examples, the regression calculation circuit 160 may be implemented by a programmable circuit, such as the exemplary programmable circuit 1112 of Fig. 11. For example, the regression calculation circuit 160 may be implemented by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in block 1025 of Fig. 10. In some examples, the regression calculation circuit 160 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the regression calculation circuit 160 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the regression calculation circuit 160 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0078] In some examples, the object detector 100 includes means for evaluating loss functions. The means for evaluating loss functions may be implemented, for example, by the loss function evaluation circuit 165. In some examples, the loss function evaluation circuit 165 may be implemented by a programmable circuit, such as the programmable circuit 1112 in Fig. 11. For example, the loss function evaluation circuit 165 may be implemented by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in block 1040 of Fig. 10. In some examples, the loss function evaluation circuit 165 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the loss function evaluation circuit 165 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the loss function evaluation circuit 165 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0079] In some examples, the object detector 100 includes means for updating neural networks. The means for updating the neural networks may be implemented, for example, by the neural network update circuit 170. In some examples, the neural network update circuit 170 may be implemented by a programmable circuit, such as the programmable circuit 1112 in Fig. 11. For example, the neural network update circuit 170 may be instantiated by the exemplary microprocessor 1200 of Fig. 12, which executes machine-executable instructions as described at least in block 1040 of Fig. 10. In some examples, the neural network update circuit 170 may be instantiated by a hardware logic circuit implemented by an ASIC, an XPU, or the FPGA circuit 1300 of Fig. 13, which is configured and / or structured to perform operations corresponding to the machine-readable instructions. Additionally or alternatively, the neural network update circuit 170 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the neural network update circuit 170 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier, a logic circuit, etc.) configured and / or structured to execute some or all of the machine-readable instructions and / or to perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are also suitable.

[0080] While an exemplary manner of implementing the object detector 100 in Fig. 1-8, one or more of the Fig. 1-8 may be combined, divided, rearranged, omitted, eliminated and / or implemented in any other manner. Further, the exemplary cluster selection circuit 105, the exemplary object detection circuit 120, the exemplary neural network circuit 125, the exemplary object classification head circuit 130, the exemplary regression head circuit 135, the exemplary bounding box generation circuit 155, the exemplary regression calculation circuit 160, the exemplary loss function evaluation circuit 165, the exemplary neural network update circuit 170, the exemplary filter circuit 205, the exemplary overhead view projection circuit 210, the exemplary view padding circuit 215, the exemplary sample point selection circuit 220, the exemplary mask application circuit 225, the exemplary cluster identification circuit 230,The exemplary overhead view projection circuit 605, the exemplary bounding box adjustment circuit 610, the exemplary mask generation circuit 705, the exemplary template generation circuit 710, and / or more generally, the exemplary object detector 100 may be implemented by hardware alone or by hardware in combination with software and / or firmware. For example, any of the following circuits may be used: the exemplary cluster selection circuit 105, the exemplary object detection circuit 120, the exemplary neural network circuit 125, the exemplary object classification circuit 130, the exemplary regression circuit 135, the exemplary bounding box generation circuit 155, the exemplary regression calculation circuit 160, the exemplary loss function evaluation circuit 165, the exemplary neural network update circuit 170, the exemplary filter circuit 205,the example overhead view projection circuit 210, the example view padding circuit 215, the example sample point selection circuit 220, the example mask application circuit 225, the example cluster identification circuit 230, the example overhead view projection circuit 605, the example bounding box adjustment circuit 610, the example mask generation circuit 705, the example template generation circuit 710, and / or more generally, the example object detector 100 could be implemented by a programmable circuit in combination with machine-readable instructions (i.e., firmware or software), processor circuits, analog circuits, digital circuits, logic circuits, programmable processors, programmable microcontrollers, graphics processing units (GPUs), digital signal processors (DSPs), ASICs,programmable logic devices (PLDs) and / or field-programmable logic devices (FPLDs) such as FPGAs. Furthermore, the exemplary object detector 100 may include one or more elements, processes, and / or devices in addition to or instead of the elements, processes, and / or devices shown in , Fig. 1-8 and / or may include more than one or all of the elements, processes and devices shown.

[0081] Flowchart(s) illustrating example machine-readable instructions that may be executed by programmable circuits to implement and / or instantiate the object detector 100 and / or illustrating example operations that may be performed by programmable circuits to implement and / or instantiate the object detector 100 are shown in Fig. 9 and / or 10. The machine-readable instructions may be one or more executable programs or portions of one or more executable programs executed by a programmable circuit such as the programmable circuit 1112 described below in connection with Fig. 11, and / or to perform one or more functions or portions of functions provided by the exemplary processor platform 1100 discussed below in connection with Fig. 12 and / or 13 discussed in programmable example circuitry (e.g., an FPGA). In some examples, the machine-readable instructions cause a real-world operation, task, etc., to be performed and / or performed in an automated manner. "Automated," in this context, means that no human involvement is required.

[0082] The program may be embodied in instructions (e.g., software and / or firmware) stored on one or more non-transitory, computer-readable and / or machine-readable storage media, such as a cache memory, a magnetic storage device or disk (e.g., a floppy disk, a hard disk (HDD), etc.), an optical storage device or disk (e.g., a Blu-ray disk, a compact disk (CD), a digital versatile disk (DVD), etc.), a redundant array of independent disks (RAID), a register, a ROM, a solid-state drive (SSD), an SSD memory, a non-volatile memory (e.g., an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.), a volatile memory (e.g., a random access memory (RAM) of any type, etc.), and / or another storage device or disk.The instructions of the non-transitory computer-readable and / or machine-readable medium may be programmed and / or executed by programmable circuitry residing in one or more hardware devices, but the entire program and / or portions thereof could alternatively be executed and / or instantiated by one or more hardware devices that are not part of the programmable circuitry and / or embodied in special-purpose hardware. The machine-readable instructions may be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., a server and a client hardware device). The client hardware device may, for example, be implemented by an endpoint client hardware device (e.g., a hardware device connected to a human and / or machine user) or an intermediate client hardware device gateway (e.g.,a radio access network (RAN) that can facilitate communication between a server and an endpoint client hardware device. Similarly, the non-transitory computer-readable storage medium can comprise one or more media. Although the example program is described with reference to the device(s) described in . Fig. 9 and / or 10, many other methods may alternatively be used to implement the example object detector 100. For example, the order of execution of the blocks of the flowchart(s) may be changed, and / or some of the described blocks may be changed, eliminated, or combined. Additionally or alternatively, some or all of the blocks of the flowchart may be implemented by one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, an FPGA, an ASIC, a comparator, an operational amplifier, logic circuitry, etc.) structured to perform the corresponding operation without executing software or firmware. The programmable circuits may be distributed across different network locations and / or locally on one or more hardware devices (e.g.,a single-core processor (e.g., a single-core CPU), a multi-core processor (e.g., a multi-core CPU, an XPU, etc.). The programmable circuits may, for example, be a CPU and / or an FPGA residing in the same package (e.g., in the same IC package or in two or more separate packages), one or more processors in a single device, multiple processors distributed across multiple servers in a server rack, multiple processors distributed across one or more server racks, etc., and / or any combination thereof.

[0083] The machine-readable instructions described herein may be stored in one or more compressed formats, an encrypted format, a fragmented format, a compiled format, an executable format, a package format, etc. Machine-readable instructions as described herein may be embodied as data (e.g., computer-readable data, machine-readable data, one or more bits (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), a bit stream (e.g., a computer-readable bit stream, a machine-readable bit stream, etc.), or a data structure (e.g., as portion(s) of instructions, code, representations of code, etc.) that can be used to create, manufacture, and / or produce machine-executable instructions. For example, the machine-readable instructions may be fragmented and stored on one or more storage devices, disks, and / or computing devices (e.g.,servers) located at the same or different locations within a network or collection of networks (e.g., in the cloud, on edge devices, etc.). The machine-readable instructions may require one or more of the following actions: installation, modification, adaptation, updating, combination, addition, configuration, decryption, decompression, unpacking, distribution, reassignment, compilation, etc., in order to make them directly readable, interpretable, and / or executable by a computing device and / or another machine.For example, the machine-readable instructions may be stored in multiple pieces that are individually compressed, encrypted, and / or stored on separate computers, which pieces, when decrypted, decompressed, and / or combined, form a set of computer-executable and / or machine-executable instructions that implement one or more functions and / or operations that together may constitute a program such as the one described herein.

[0084] In another example, the machine-readable instructions may be stored in a state where they can be read by programmable circuitry, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., to execute the machine-readable instructions on a particular computer or other device. In another example, the machine-readable instructions may require training (e.g., storing settings, entering data, recording network addresses, etc.) before the machine-readable instructions and / or the corresponding program(s) can be executed, in whole or in part.Therefore, machine-readable, computer-readable and / or machine-readable media as used herein may contain instructions and / or programs, regardless of the particular format or state of the machine-readable instructions and / or programs.

[0085] The machine-readable instructions described herein may be represented by any past, present, or future command language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented in any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0086] As mentioned above, the example operations of Fig. 9 and / or 10 may be implemented using executable instructions (e.g., computer-readable and / or machine-readable instructions) stored on one or more non-transitory computer-readable and / or machine-readable media. As used herein, the terms non-transitory computer-readable medium, non-transitory computer-readable storage medium, non-transitory machine-readable medium, and / or non-transitory machine-readable storage medium are expressly defined to include any type of computer-readable storage device and / or storage disk and exclude the transmission of signals and transmission media.Examples of such a non-transitory computer-readable medium, a non-transitory computer-readable storage medium, a non-transitory machine-readable medium and / or a non-transitory machine-readable storage medium include optical storage devices, magnetic storage devices, a hard disk, flash memory, read-only memory (ROM), a CD, a DVD, a cache, RAM of any kind, a register and / or any other storage device or storage disk in which information is stored for any duration (e.g., for extended periods of time, permanently, for short periods of time, for temporary caching and / or for caching information).As used herein, the terms "non-volatile computer-readable storage device" and "non-volatile machine-readable storage device" are defined to include any physical (mechanical, magnetic, and / or electrical) hardware for storing information for a specific period of time, but not the transmission of signals and transmission media. Examples of non-volatile computer-readable storage devices and / or non-volatile machine-readable storage devices include random access memory of any kind, read-only memory of any kind, solid-state memory, flash memory, optical disks, magnetic disks, hard disk drives, and / or RAID (Redundant Array of Independent Disks) systems. As used herein, the term "device" refers to physical structures such as mechanical and / or electrical equipment, hardware, and / or circuitry that are embodied by computer-readable instructions, machine-readable instructions, etc.designed and / or manufactured to execute computer-readable instructions, machine-readable instructions, etc.

[0087] Fig. 9 is a flowchart illustrating example machine-readable instructions and / or example operations 900 that may be executed, instantiated, and / or performed by programmable circuitry to implement the inference processing associated with the object detector 100. The example machine-readable instructions and / or example operations 900 of Fig. 9 begin at block 905, where the interface circuitry of the cluster selection circuit 105 of the object detector 100 accesses an input 3D point cloud, such as the input 3D point cloud 115. In block 910, the cluster selection circuit 105 selects one or more candidate clusters from the input 3D point cloud based on masks and / or templates, as described above. In block 915, the cluster selection circuit 105 selects some of the candidate clusters for downstream processing, as described above.

[0088] For example, in block 920, cluster selection circuit 105 inputs a given candidate cluster to neural network circuit 125 of object detector 100, which outputs the feature vector 140 corresponding to the candidate cluster, as described above. In block 925, an object classification head implemented by object classification head circuit 130 of object detector 100 processes feature vector 140 to determine object classification parameters for the candidate cluster, as described above. In block 930, one or more regression heads implemented by regression head circuit 135 of object detector 100 process feature vector 140 to determine bounding box parameters for the candidate cluster, as described above.

[0089] In block 935, the cluster selection circuit 105 continues to input candidate clusters to the neural network circuit 125 until all candidate clusters have been processed. In block 940, the object detector 100 outputs the object classification parameter(s) and the bounding box parameter(s) determined for the respective candidate clusters. The machine-readable example instructions and / or the example operations 900 of Fig. 9 then end.

[0090] Fig. 10 is a flowchart illustrating example machine-readable instructions and / or example operations 1000 that may be executed, instantiated, and / or performed by programmable circuitry to train the object detector 100. The example machine-readable instructions and / or example operations 1000 of Fig. 10 begin in block 1005, where the interface circuit of the cluster selection circuit 105 of the object detector 100 accesses training data. In block 1010, the cluster selection circuit 105 selects one or more training clusters from the training, as described above. In block 1015, the cluster selection circuit 105 cycles through the training clusters to enable training of the object detector 100.

[0091] For example, in block 1020, the cluster selection circuit 105 passes a given training cluster to the bounding box generation circuit 155 of the object detector 100. In block 1020, the bounding box generation circuit 155 determines a proposal bounding box for the training cluster, as described above. In block 1025, the regression calculation circuit 160 of the object detector 100, as described above, determines the regression parameters for the training bounding box based on the proposal bounding box determined in block 1020 and a ground-truth bounding box associated with the training cluster. For example, in block 1025, the regression calculation circuit 160 may determine the regression parameters for training the bounding box based on Equations 1-7 above.

[0092] In block 1030, the cluster selection circuit 105 also inputs the training cluster to the neural network circuit 125 of the object detector 100, which outputs a feature vector corresponding to the training cluster, as described above. In block 1035, the feature vector is processed by the object classification head, implemented by the object classification head circuit 130 of the object detector, to determine parameters for the inference object classification, as described above. In block 1035, the feature vector is also processed by one or more regression heads, implemented by the regression head circuit 135 of the object detector, to determine the parameters of the inference bounding box regression, as described above.

[0093] In block 1040, the loss function evaluation circuit 165 and the neural network update circuit 170 of the object detector 100 operate as described above to train the neural network circuit 125, the object classification head circuit 130, and the regression head circuit 135 based on the inference object classification parameters and the inference bounding box regression parameters determined in block 1035, as well as the training bounding box regression parameters determined in block 1025. For example, in block 1040, the loss function evaluation circuit 165 may evaluate the loss functions of Equation 8-11 above, and the neural network update circuit 170 may update the weights and / or parameters of the neural network circuit 125, the object classification head circuit 130, and the regression head circuit 135 based on the loss function outputs, as described above.

[0094] In block 1045, the cluster selection circuit 105 continues to select training clusters to be used for training the object detector 100 until all training clusters have been processed, or until one or more other termination criteria have been met. The machine-readable example instructions and / or the example operations 1000 of Fig. 10 then end.

[0095] Fig. 11 is a block diagram of an example programmable circuit platform 1100 structured to implement the example machine-readable instructions and / or the example operations of the Fig. 9 and / or 10 to implement the object detector 100. The programmable circuit platform 1100 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an internet device, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset (e.g., an augmented reality (AR) headset, a virtual reality (VR) headset, etc.), or another wearable device or any other type of computing and / or electronic device.

[0096] The programmable circuit platform 1100 of the illustrated example includes programmable circuits 1112. The programmable circuit 1112 of the illustrated example is hardware. For example, the programmable circuit 1112 may be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers of any family or manufacturer. The programmable circuit 1112 may be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, the programmable circuit 1112 implements the object detector 100 and / or, more specifically, one or more of the following circuits: the exemplary cluster selection circuit 105, the exemplary object detection circuit 120, the exemplary neural network circuit 125, the exemplary object classification head circuit 130,the example regression head circuit 135, the example bounding box generation circuit 155, the example regression calculation circuit 160, the example loss function evaluation circuit 165, the example neural network update circuit 170, the example filter circuit 205, the example overhead view projection circuit 210, the example view padding circuit 215, the example sample point selection circuit 220, the example mask application circuit 225, the example cluster identification circuit 230, the example overhead view projection circuit 605, the example bounding box adjustment circuit 610, the example mask generation circuit 705, and / or the example template generation circuit 710.

[0097] The programmable circuit 1112 of the illustrated example includes a local memory 1113 (e.g., a cache, registers, etc.). The programmable circuit 1112 of the illustrated example is connected via a bus 1118 to the main memory 1114, 1116, which includes a volatile memory 1114 and a non-volatile memory 1116. The volatile memory 1114 can be implemented by a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), a RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of RAM device. The non-volatile memory 1116 can be implemented by a flash memory and / or any other desired type of memory device. Access to the main memory 1114, 1116 of the illustrated example is controlled by a memory controller 1117.In some examples, the memory controller 1117 may be implemented by one or more integrated circuits, logic circuits, microcontrollers of any desired family or manufacturer, or any other type of circuitry to manage the flow of data to and from the main memory 1114, 1116.

[0098] The programmable circuit platform 1100 of the illustrated example also includes interface circuitry 1120. The interface circuitry 1120 may be implemented in hardware in accordance with any interface standard, such as an Ethernet interface, a Universal Serial Bus (USB) interface, a Bluetooth® interface, a Near Field Communication (NFC) interface, a Peripheral Component Interconnect (PCI) interface, and / or a Peripheral Component Interconnect Express (PCIe) interface. In some examples, the interface circuitry 1120 accesses the input 3D point cloud, such as the 3D point cloud 115, and / or the training data to be processed by the object detector 100.

[0099] In the illustrated example, one or more input devices 1122 are connected to the interface circuit 1120. The input device(s) 1122 enable a user (e.g., a human user, a machine user, etc.) to input data and / or instructions into the programmable circuit 1112. The input device(s) 1122 may be implemented, for example, by an audio sensor, a microphone, a camera (still image or video), a keyboard, a button, a mouse, a touchscreen, a trackpad, a trackball, an isopoint device, and / or a speech recognition system.

[0100] One or more output devices 1124 are also connected to the interface circuit 1120 of the illustrated example. The output device(s) 1124 can be implemented, for example, by display devices (e.g., a light-emitting diode (LED), an organic light-emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching display (IPS), a touchscreen, etc.), a tactile output device, a printer, and / or a speaker. The interface circuit 1120 of the illustrated example therefore typically includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit such as a graphics processor.

[0101] The interface circuit 1120 of the illustrated example also includes a communication device, such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface, to facilitate data exchange with external machines (e.g., computing devices of any kind) over a network 1126. Communication may occur, for example, via an Ethernet connection, a digital subscriber line (DSL), a telephone line, a coaxial cable system, a satellite system, a wireless outside line-of-sight system, a wireless line-of-sight system, a cellular phone system, an optical connection, etc.

[0102] The programmable circuit platform 1100 of the illustrated example also includes one or more mass storage disks or devices 1128 for storing firmware, software, and / or data. Examples of such mass storage disks or devices 1128 include magnetic storage devices (e.g., floppy disks, drives, HDDs, etc.), optical storage devices (e.g., Blu-ray disks, CDs, DVDs, etc.), RAID systems, and / or solid-state storage disks or devices such as flash memory devices and / or SSDs.

[0103] The machine-readable instructions 1132, which are replaced by the machine-readable instructions of the Fig. 9 and / or 10 may be stored in mass storage 1128, volatile memory 1114, non-volatile memory 1116, and / or on at least one non-volatile, computer-readable storage medium such as a CD or DVD, which may be removable.

[0104] Fig. 12 is a block diagram of an exemplary implementation of the programmable circuit 1112 of Fig. 11. In this example, the programmable circuit 1112 is Fig. 11 is implemented by a microprocessor 1200. The microprocessor 1200 may be, for example, a general-purpose microprocessor (e.g., a general-purpose microprocessor circuit). The microprocessor 1200 executes some or all of the machine-readable instructions of the flowcharts of Fig. 9 and / or 10 to switch the circuits of Fig. 2 as logical circuits to perform operations corresponding to these machine-readable instructions. In some of these examples, the circuit is Fig. 1 is instantiated by the hardware circuitry of the microprocessor 1200 in combination with the machine-readable instructions. The microprocessor 1200 may, for example, be implemented by multi-core hardware circuitry such as a CPU, a DSP, a GPU, an XPU, etc. Although it may include any number of example cores 1202 (e.g., 1 core), the microprocessor 1200 of this example is a multi-core semiconductor device having N cores. The cores 1202 of the microprocessor 1200 may operate independently of one another or cooperate to execute machine-readable instructions. For example, machine code corresponding to a firmware program, an embedded software program, or a software program may be executed by one of the cores 1202 or by multiple cores 1202 at the same or different times.In some examples, the machine code corresponding to the firmware program, the embedded software program, or the software program is divided into threads and executed in parallel by two or more cores 1202. The software program may correspond to some or all of the machine-readable instructions and / or operations described in the flowcharts of . Fig. 9 and / or 10 are shown.

[0105] The cores 1202 may communicate via a first example bus 1204. In some examples, the first bus 1204 may be implemented by a communication bus to enable communication in association with one or more cores 1202. For example, the first bus 1204 may be implemented by at least one of the following buses: Inter-Integrated Circuit (I2C) bus, Serial Peripheral Interface (SPI) bus, PCI bus, or PCIe bus. Additionally or alternatively, the first bus 1204 may be implemented by any other type of computer or electrical bus. The cores 1202 may receive data, instructions, and / or signals from one or more external devices via an example interface circuit 1206. The cores 1202 may output data, instructions, and / or signals to one or more external devices via the interface circuit 1206. Although the cores 1202 of this example may include an example local memory 1220 (e.g.,Level 1 (L1) cache, which may be divided into an L1 data cache and an L1 instruction cache), the microprocessor 1200 also includes an exemplary shared memory 1210 that may be shared by the cores (e.g., Level 2 (L2 cache)) for high-speed access to data and / or instructions. Data and / or instructions may be transferred (e.g., shared) by writing to and / or reading from the shared memory 1210. The local memory 1220 of each of the cores 1202 and the shared memory 1210 may be part of a hierarchy of memory devices that may include multiple levels of cache memories and main memory (e.g., main memory 1114, 1116 of FIG. Fig. 11). Typically, higher memory levels in the hierarchy have lower access times and smaller storage capacities than lower memory levels. Changes at the different levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherence policy.

[0106] Each core 1202 may be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuitry. Each core 1202 includes a control unit 1214, an arithmetic and logic (AL) circuit (sometimes referred to as an ALU) 1216, a plurality of registers 1218, the local memory 1220, and a second example bus 1222. Other structures may be present. For example, each core 1202 may include circuitry for a vector unit, a single instruction multiple data (SIMD) unit, a load / store unit (LSU) circuit, a branch / jump unit, a floating point unit (FPU) circuit, etc. The control unit circuitry 1214 includes semiconductor-based circuitry configured to control (e.g., coordinate) data movement within the corresponding core 1202.The AL circuit 1216 includes semiconductor-based circuitry structured to perform one or more mathematical and / or logical operations on the data in the corresponding core 1202. In some examples, the AL circuit 1216 performs integer operations. In other examples, the AL circuit 1216 also performs floating-point operations. In further examples, the AL circuit 1216 may include a first AL circuit that performs integer operations and a second AL circuit that performs floating-point operations. In some examples, the AL circuit 1216 may be referred to as an arithmetic logic unit (ALU).

[0107] Registers 1218 are semiconductor-based structures for storing data and / or instructions, such as results of one or more operations performed by the AL circuitry 1216 of the corresponding core 1202. Registers 1218 may include, for example, vector registers, SIMD registers, general-purpose registers, flag registers, segment registers, machine-specific registers, instruction pointer registers, control registers, debug registers, memory management registers, machine check registers, etc. Registers 1218 may be arranged in a bank, as shown in Fig. 12. Alternatively, the registers 1218 may be organized in a different arrangement, format, or structure, e.g., by distributing them across the core 1202 to reduce access time. The second bus 1222 may be implemented by at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.

[0108] Each core 1202 and / or, more generally, the microprocessor 1200 may include additional and / or different structures than those shown and described above. For example, one or more clock circuits, one or more power supplies, one or more power gates, one or more cache home agents (CHAs), one or more converged / common mesh stops (CMSs), one or more shifters (e.g., barrel shifters), and / or other circuits may be present. The microprocessor 1200 is a semiconductor device fabricated to include many transistors interconnected to implement the structures described above in one or more integrated circuits (ICs) contained in one or more packages.

[0109] Microprocessor 1200 may include and / or cooperate with one or more accelerators (e.g., accelerator circuits, hardware accelerators, etc.). In some examples, accelerators are implemented by logic circuitry to perform certain tasks faster and / or more efficiently than is possible with a general-purpose processor. Examples of accelerators include ASICs and FPGAs, as described herein. A GPU, DSP, and / or other programmable device may also be an accelerator. Accelerators may be located on microprocessor 1200, in the same die package as microprocessor 1200, and / or in one or more packages separate from microprocessor 1200.

[0110] Fig. 13 is a block diagram of another exemplary implementation of the programmable circuit 1112 of Fig. 11. In this example, the programmable circuit 1112 is implemented by an FPGA circuit 1300. For example, the FPGA circuit 1300 may be implemented by an FPGA. The FPGA circuit 1300 may, for example, be used to perform operations that would otherwise be performed by the exemplary microprocessor 1200 of Fig. 12, which executes corresponding machine-readable instructions. However, once configured, the FPGA circuit 1300 instantiates the operations and / or functions corresponding to the machine-readable instructions in hardware and can therefore often execute the operations / functions faster than would be possible with a general-purpose microprocessor executing the corresponding software.

[0111] More precisely, in contrast to the above-described microprocessor 1200 from Fig. 12 (which is a general-purpose device that can be programmed to execute any or all of the machine-readable instructions represented by the flowchart(s) of Fig. 9 and / or 10, but whose connections and logic circuits are fixed after manufacture), the FPGA circuit 1300 of the example of Fig. 13 includes connections and logic circuits that may be formed, structured, programmed and / or connected in various ways after manufacture, for example, to instantiate some or all of the operations / functions corresponding to the machine-readable instructions represented by the flowchart(s) of Fig. 9 and / or 10. In particular, the FPGA circuit 1300 can be viewed as an arrangement of logic gates, interconnects, and switches. The switches can be programmed to change the way the logic gates are connected to each other by the interconnects, effectively forming one or more dedicated logic circuits (unless the FPGA circuit 1300 is reprogrammed). The configured logic circuits allow the logic gates to cooperate in different ways to perform various operations on the data received from the input circuits. These operations may correspond to some or all of the instructions (e.g., software and / or firmware) described in the flowchart(s) of Fig. 9 and / or 10. Thus, the FPGA circuit 1300 may be configured and / or structured to perform some or all of the operations / functions corresponding to the machine-readable instructions of the flowchart(s) of Fig. 9 and / or 10, are effectively instantiated as dedicated logic circuits to perform the operations / functions corresponding to these software instructions in a dedicated manner analogous to an ASIC. Therefore, the FPGA circuit 1300 can perform the operations / functions corresponding to some or all of the machine-readable instructions of the Fig. 9 and / or 10 faster than the general-purpose microprocessor can execute them.

[0112] In the example of Fig. 13, the FPGA circuit 1300 is configured and / or structured to be programmed (and / or reprogrammed one or more times) based on a binary file. In some examples, the binary file may be compiled and / or generated based on instructions in a Hardware Description Language (HDL) such as Lucid, Very High Speed ​​Integrated Circuits (VHSIC) Hardware Description Language (VHDL), or Verilog. For example, a user (e.g., a human user, a machine user, etc.) may write code or a program corresponding to one or more operations / functions in an HDL; the code / program may be translated into a low-level language if necessary; and the code / program (e.g., the code / program in the low-level language) may be converted (e.g., by a compiler, a software application, etc.) into the binary file. In some examples, the FPGA circuit 1300 may be Fig. 13 access and / or load the binary file to the FPGA circuit 1300 of Fig. 13, to be configured and / or structured to perform the one or more operations / functions. For example, the binary file may be implemented by a bit stream (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), data (e.g., computer-readable data, machine-readable data, etc.), and / or machine-readable instructions to which the FPGA circuit 1300 of Fig. 13 to configure and / or structure the FPGA circuit 1300 of Fig. 13 or one or more parts thereof.

[0113] In some examples, the binary file is compiled, generated, transformed, and / or otherwise output by a unified software platform used to program FPGAs. For example, the unified software platform may translate first instructions (e.g., code or a program) corresponding to one or more operations / functions in a high-level language (e.g., C, C++, Python, etc.) into second instructions corresponding to the one or more operations / functions in an HDL. In some of these examples, the binary file is compiled, generated, and / or otherwise output by the unified software platform based on the second instructions. In some examples, the FPGA circuit 1300 may Fig. 13 access and / or load the binary file to the FPGA circuit 1300 of Fig. 13, to be configured and / or structured to perform the one or more operations / functions. For example, the binary file may be implemented by a bit stream (e.g., one or more computer-readable bits, one or more machine-readable bits, etc.), data (e.g., computer-readable data, machine-readable data, etc.), and / or machine-readable instructions to which the FPGA circuit 1300 of Fig. 13 to configure and / or structure the FPGA circuit 1300 of Fig. 13 or one or more parts thereof.

[0114] The FPGA circuit 1300 from Fig. 13 includes an example input / output circuit 1302 for receiving and / or outputting data from / to an example configuration circuit 1304 and / or external hardware 1306. For example, the configuration circuit 1304 may be implemented by an interface circuit that may receive a binary file that may be implemented by a bit stream, data, and / or machine-readable instructions to configure the FPGA circuit 1300 or portions thereof. In some examples, the configuration circuit 1304 may receive the binary file from a user, a machine (e.g., a hardware circuit (e.g., a programmable or dedicated circuit) that may implement an artificial intelligence / machine learning (AI / ML) model to generate the binary file), etc., and / or any combination thereof. In some examples, the external hardware 1306 may be implemented by external hardware circuits.The external hardware 1306 can be implemented, for example, by the microprocessor 1200 of . Fig. 12 will be implemented.

[0115] The FPGA circuit 1300 also includes an array of example logic gate circuits 1308, a plurality of configurable example interconnections 1310, and example memory circuits 1312. The logic gate circuit 1308 and the configurable interconnections 1310 are configurable to instantiate one or more operations / functions corresponding to at least some of the machine-readable instructions of Fig. 9 and / or 10 and / or other desired operations. The Fig. The logic gate circuit 1308 shown in Figure 13 is fabricated in blocks or groups. Each block contains semiconductor-based electrical structures that can be formed into logic circuits. In some examples, the electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that represent basic building blocks for logic circuits. Electrically controllable switches (e.g., transistors) are present in each of the logic gate circuits 1308 to enable the configuration of the electrical structures and / or the logic gates to form circuits for performing desired operations / functions. The logic gate circuit 1308 may include other electrical structures such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.

[0116] The configurable connections 1310 of the example shown are conductive paths, traces, vias, or the like that may include electrically controllable switches (e.g., transistors) whose state may be changed by programming (e.g., using an HDL instruction language) to enable or disable one or more connections between one or more of the logic gate circuits 1308 to program desired logic circuits.

[0117] The memory circuit 1312 of the example shown is configured to store the results of one or more of the operations performed by the corresponding logic gates. The memory circuit 1312 can be implemented using registers or the like. In the example shown, the memory circuits 1312 are distributed among the logic gate circuits 1308 to facilitate access and increase execution speed.

[0118] The example FPGA circuit 1300 from Fig. 13 also includes an example dedicated operations circuit 1314. In this example, dedicated circuit 1314 includes special circuitry 1316 that can be called to implement frequently used functions so that these functions do not need to be programmed in-place. Examples of such special-purpose circuitry 1316 include memory (e.g., DRAM) controller circuitry, PCIe controller circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of special-purpose circuitry may be present. In some examples, FPGA circuitry 1300 may also include a programmable general-purpose circuitry 1318, such as an example CPU 1320 and / or an example DSP 1322. Additionally or alternatively, other programmable circuitry 1318 may be present, such as a GPU, an XPU, etc., that can be programmed to perform other operations.

[0119] Although Fig. 12 and Fig. 13 two exemplary implementations of the programmable circuit 1112 of Fig. 11, many other approaches are conceivable. For example, the FPGA circuit may include an integrated CPU, such as one or more of the exemplary CPUs 1320 in Fig. 12. Therefore, the programmable circuit 1112 of Fig. 11 additionally by combining at least the exemplary microprocessor 1200 of Fig. 12 and the exemplary FPGA circuit 1300 of Fig. 13. In some such hybrid examples, one or more cores 1202 of Fig. 12 execute a first part of the machine-readable instructions represented by the flowchart(s) of Fig. 9 and / or 10 to perform a first operation(s) / function(s), the FPGA circuit 1300 of Fig. 13 may be configured and / or structured to perform second operations / functions corresponding to a second portion of the machine-readable instructions represented by the flowcharts of Fig. 9 and / or 10, and / or an ASIC may be designed and / or structured to perform third operations / functions corresponding to a third portion of the machine-readable instructions represented by the flowcharts of Fig. 9 and / or 10 are shown.

[0120] It goes without saying that some or all of the circuits of Fig. 1 can therefore be instantiated at the same or different times. For example, the same and / or different parts of the microprocessor 1200 can Fig. 12 may be programmed to execute portions of the machine-readable instructions at the same and / or different times. In some examples, the same and / or different portions of the FPGA circuit 1300 may be Fig. 13 be designed and / or structured to perform operations / functions corresponding to sections of machine-readable instructions at the same and / or different times.

[0121] In some examples, some or all of the circuits may be Fig. 1, e.g., in one or more threads that execute concurrently and / or in series. For example, the microprocessor 1200 may be Fig. 12 machine-readable instructions in one or more threads that run concurrently and / or sequentially. In some examples, the FPGA circuit 1300 may be Fig. 13 may be designed and / or structured to perform operations / functions simultaneously and / or sequentially. Furthermore, in some examples, some or all of the circuits of Fig. 1 be implemented in one or more virtual machines and / or containers running on the microprocessor 1200 of Fig. 12 are executed.

[0122] In some examples, the programmable circuit 1112 may be Fig. 11 may be housed in one or more housings. For example, the microprocessor 1200 of Fig. 12 and / or the FPGA circuit 1300 of Fig. 13 may be housed in one or more packages. In some examples, an XPU may be implemented by the programmable circuit 1112 of Fig. 11, which can be located in one or more packages. For example, the XPU can integrate a CPU (e.g., the 1200 microprocessor from Fig. 12, the CPU 1320 from Fig. 13 etc.) in one housing, a DSP (e.g. the DSP 1322 from Fig. 13) in another package, a GPU in another package and an FPGA (e.g. the FPGA circuit 1300 of Fig. 13) contained in another housing.

[0123] A block diagram illustrating an example software distribution platform 1405 for distributing software such as the example machine-readable instructions 1132 of Fig. 11 to other hardware devices (e.g. hardware devices owned by third parties and / or operated by the owner and / or operator of the software distribution platform) is permitted in Fig. 14. The exemplary software distribution platform 1405 may be implemented by any computer server, data facility, cloud service, etc., capable of storing and transmitting software to other computing devices. The third parties may be customers of the company that owns and / or operates the software distribution platform 1405. For example, the company that owns and / or operates the software distribution platform 1405 may be a developer, a vendor, and / or a licensor of software, such as the machine-readable instructions 1132 of Fig. 11. Third parties may be consumers, users, retailers, OEMs, etc., who acquire and / or license the software for use and / or resale and / or sublicensing. In the illustrated example, the software distribution platform 1405 includes one or more servers and one or more storage devices. The storage devices store the machine-readable instructions 1132, which correspond to the machine-readable instruction examples described above. Fig. 9 and / or 10. The one or more servers of the example software distribution platform 1405 are in communication with an example network 1410, which may correspond to one or more of the Internet and / or example networks described above. In some examples, the server(s) respond to requests to transfer the software to a requesting party as part of a commercial transaction. Payment for the delivery, sale, and / or licensing of the software may be processed through the one or more servers of the software distribution platform and / or through a third party payment entity. The servers enable purchasers and / or licensors to download the machine-readable instructions 1132 from the software distribution platform 1405. For example, the software corresponding to the example machine-readable instructions of Fig. 9 and / or 10, may be downloaded to the exemplary programmable circuit platform 1100, which is to execute the machine-readable instructions 1132 to implement the object detector 100. In some examples, one or more servers of the software distribution platform 1405 regularly provide, transmit, and / or enforce updates to the software (e.g., the machine-readable instructions 1132 of Fig. 11) to ensure that improvements, patches, updates, etc., are distributed and applied to the software on end-user devices. Although referred to as software above, the distributed "software" could alternatively be firmware.

[0124] The terms "including" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, when a claim uses any form of "comprising" or "containing" (e.g., comprises, includes, contains, with, including, with, etc.) as a preamble or in any claim language, it is understood that additional elements, terms, etc. may be present without exceeding the scope of the relevant claim or language. When the term "at least," for example, is used in the preamble of a claim as a transitional term, it is as open-ended as the terms "comprising" and "including." The term "and / or," when used, for example,in a form such as A, B and / or C refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A with B, (5) A with C, (6) B with C or (7) A with B and with C. As used herein in the context of describing structures, components, items, objects and / or things, the phrase "at least one of A and B" is intended to refer to implementations that include (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects and / or things, the phrase "at least one of A or B" is intended to refer to implementations that include (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.As used herein in the context of describing the performance or execution of processes, instructions, acts, activities, etc., the phrase "at least one of A and B" is intended to refer to implementations that include (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, the phrase "at least one of A or B" as used herein in the context of describing the performance or execution of processes, instructions, actions, activities, etc., is intended to refer to implementations that include (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.

[0125] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not preclude a plurality. The term "a" or "an" object, as used herein, refers to one or more of such objects. The terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. Moreover, a plurality of means, elements, or actions, even if listed individually, may be performed, e.g., by the same device or object. Even if individual features are included in different examples or claims, they may possibly be combined, and inclusion in different examples or claims does not imply that a combination of features is not possible and / or advantageous.

[0126] Unless otherwise specified, the term "above" describes the relationship between two parts relative to the Earth. A first part is above a second part if the second part has at least one part between the Earth and the first part. Likewise, a first part is "below" a second part if the first part is closer to the Earth than the second part. As previously stated, a first part can be above or below a second part, with one or more of the following characteristics: other parts between them, with no other parts between them, with the first and second parts touching, or without the first and second parts being in direct contact with each other.

[0127] Notwithstanding the foregoing, when referring to a semiconductor device (e.g., a transistor), a semiconductor chip containing a semiconductor device, and / or a package of an integrated circuit (IC) containing a semiconductor chip during manufacture or fabrication, the term "above" refers not to the ground, but to an underlying substrate on which the components in question are manufactured, assembled, mounted, supported, or otherwise provided. Unless otherwise stated or clear from the context, a first component in a semiconductor chip (e.g., a transistor or other semiconductor device) is located "above" a second component in the semiconductor chip if, during manufacture / fabrication, the first component is farther from a substrate (e.g., a semiconductor wafer) than the second component on which the two components are manufactured or otherwise provided.Unless otherwise specified or apparent from the context, a first component within an IC package (e.g., a semiconductor die) is said to be "above" a second component within the IC package during manufacture if the first component is farther from a printed circuit board (PCB) on which the IC package is to be mounted or attached. It is understood that semiconductor devices are often used in a different orientation than when they were manufactured. Therefore, when referring to a semiconductor device (e.g., a transistor), a semiconductor die containing a semiconductor device, and / or a package for an integrated circuit (IC) containing a semiconductor die, the definition of "above" in the preceding paragraph will likely apply, depending on the context of use (i.e., the term "above" describes the relationship between two parts relative to ground).

[0128] In this patent, to state that a part (e.g., a layer, film, area, region, or plate) is in any way on top of another part (e.g., positioned on, located on, disposed on, or formed on, etc.) means that the part in question is either in contact with the other part or that the part in question is overlying the other part with one or more intermediate parts therebetween.

[0129] As used herein, the terms "connection" (e.g., "attached," "coupled," "connected," and "joined") may include intermediate links between the elements referred to by the term "connection" and / or relative movement between those elements, unless otherwise noted. References to connections do not necessarily imply that two elements are directly connected and / or in a fixed relationship. When one part is said to be in "contact" with another part, it means that there is no intermediate part between the two parts.

[0130] Unless expressly stated otherwise, descriptors such as "first", "second", "third", etc. are used herein without implying or otherwise indicating any meaning of priority, physical order, arrangement in a list, and / or order, but serve merely as labels and / or arbitrary names to distinguish elements for easier understanding of the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in the detailed description, while the same element may be referred to in a claim with a different descriptor such as "second" or "third". In such cases, such descriptors should be used merely to uniquely identify the elements in the context of discussion (e.g., within a claim), where the elements might otherwise have the same name, for example.

[0131] As used herein, "approximately" and "about" change their themes / values ​​to acknowledge the potential existence of variations that occur in real-world applications. For example, the terms "approximately" and "about" may alter dimensions that are not exact due to manufacturing tolerances and / or other imperfections, as those familiar with the subject matter will understand. For example, "approximately" and "about" may indicate that dimensions are within a tolerance range of + / - 10% unless otherwise noted herein.

[0132] As used herein, "essentially real-time" refers to near-instantaneous occurrence, although real-world delays may occur in computing time, transmission, etc. Unless otherwise noted, "essentially real-time" refers to real time + 1 second.

[0133] As used herein, the term "in communication," including variations thereof, includes direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or constant communication, but additionally includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.

[0134] As used herein, the term "programmable circuits" includes (i) one or more special-purpose electrical circuits (e.g., an application-specific integrated circuit (ASIC)) structured to perform specific operations, and one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based electrical circuits that can be programmed with instructions to perform specific functions and / or operations, and one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of programmable circuits include programmable microprocessors such as central processing units (CPUs) that can execute first instructions to perform one or more operations and / or functions, field-programmable gate arrays (FPGAs),which can be programmed with second instructions to cause the configuration and / or structuring of the FPGAs to instantiate one or more operations and / or functions according to the first instructions, graphics processing units (GPUs) that can execute first instructions to perform one or more operations and / or functions, digital signal processors (DSPs) that can execute first instructions to perform one or more operations and / or functions, XPUs, network processing units (NPUs), one or more microcontrollers that can execute first instructions to perform one or more operations and / or functions, and / or integrated circuits such as application-specific integrated circuits (ASICs). An XPU can, for example, be implemented by a heterogeneous computing system with multiple types of programmable circuits (e.g., one or more FPGAs, one or more CPUs, one or more GPUs,one or more NPUs, one or more DSPs, etc.), and / or any combination thereof), and an orchestration technology (e.g., application programming interface(s) (API(s)) that can assign the computational task(s) to the respective type(s) of the multiple types of programmable circuits that is / are suitable and available to perform the computational task(s).

[0135] Integrated circuits are defined here as one or more semiconductor packages containing one or more circuit elements such as transistors, capacitors, inductors, resistors, current paths, diodes, etc. An integrated circuit can be implemented, for example, as an ASIC, FPGA, chip, microchip, programmable circuit, semiconductor substrate connecting multiple circuit elements, system on chip (SoC), etc.

[0136] From the foregoing, it can be seen that exemplary systems, devices, articles of manufacture, and methods have been disclosed that detect and locate objects in 3D point clouds. The presented systems, devices, articles of manufacture, and methods improve the efficiency of a computing device by focusing object detection and localization on candidate clusters rather than the entire 3D point cloud. As a result, the dimensionality and complexity of the neural network(s) and / or other machine learning architectures implementing object detection and localization can be reduced compared to neural networks and / or other machine learning architectures that process the entire 3D point cloud.Such a reduction in dimensionality and complexity can reduce power consumption, reduce the need for computational resources, improve the speed at which objects are detected and located, etc. Faster object detection can be extremely important in many applications, such as autonomous vehicle navigation and / or other applications where collision avoidance and / or object interception is desired. By focusing object detection and localization on candidate clusters rather than the 3D point cloud as a whole, object classification and frame regression accuracy can also be improved. The presented systems, devices, articles of manufacture, and methods accordingly aim at one or more improvements in the operation of a machine such as a computer or other electronic and / or mechanical device.

[0137] Further examples and combinations thereof are the following. Example 1 includes an object detection apparatus comprising interface circuitry to obtain a three-dimensional point cloud of a scene, machine-readable instructions, and at least one processor circuit to be programmed by the machine-readable instructions to apply at least one of a template or a mask to a sample point of an overhead view of the three-dimensional point cloud to identify a candidate cluster of points in the three-dimensional point cloud, input the candidate cluster to a neural network, the neural network outputting a feature vector for the candidate cluster, and process the feature vector to output parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster.

[0138] Example 2 includes the apparatus of Example 1, wherein one or more of the at least one processor circuit(s) apply the mask centered on the sample point of the overhead view of the three-dimensional point cloud, the mask being associated with a particular object class, the mask corresponding to a grid map having grid positions having a first value or a second value, some of the grid positions corresponding to the object classification being to have the first value and some of the grid positions not corresponding to the object classification being to have the second value, the occupancy target corresponding to a target number of points of the candidate cluster to be covered by the mask.

[0139] Example 3 includes the apparatus of Example 1 or Example 2, wherein one or more of the at least one processor circuit(s) applies the template centered on the sample point of the overhead view of the three-dimensional point cloud, the template being associated with a particular object class, the template corresponding to a grid map having grid positions with respective values ​​representative of the number of points expected for object classification at the grid positions.

[0140] Example 4 includes the apparatus of any of Examples 1 to 3, wherein one or more of the at least one processor circuit(s) implements an object classification header and a plurality of regression heads to output the parameters, the parameters comprising respective prediction values ​​for a plurality of possible object classifications and respective sets of regression values ​​associated with respective bounding boxes corresponding to the plurality of possible object classifications, the object classification header to output the respective prediction values ​​for the plurality of possible object classifications, the plurality of regression heads to output the respective sets of regression values ​​associated with the respective bounding boxes corresponding to the plurality of possible object classifications.

[0141] Example 5 includes the apparatus of any of Examples 1 to 4, wherein a first of the sets of regression values ​​output by a first of the regression heads corresponding to a first of the possible object classifications includes values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications.

[0142] Example 6 includes the apparatus of any of Examples 1 to 5, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, the parameters are first parameters, the interface circuit is to obtain a second three-dimensional point cloud, and one or more of the at least one processor circuit(s) is to select a second candidate cluster from the second three-dimensional point cloud, generate a proposal bounding box to cover a volume of the second candidate cluster, one or more processor circuits select a second candidate cluster from the second three-dimensional point cloud, generate a proposal bounding box to cover a volume of the second candidate cluster, determine a first set of regression values,representative of differences between the proposal bounding box and a ground truth bounding box, input the second candidate cluster to the neural network, wherein the neural network outputs a second feature vector for the second candidate cluster, process the second feature vector with a plurality of heads to output parameters including a second set of regression values, determine an output of a loss function based on the first set of regression values ​​and the second set of regression values, and update the neural network and the plurality of heads based on the output of the loss function.

[0143] Example 7 includes the apparatus of any of Examples 1 to 6, wherein the first set of regression values ​​includes a plurality of regression values ​​representing differences between a center point of the ground truth bounding box and a center point of the proposal bounding box, a plurality of regression values ​​representing differences between the dimensions of the ground truth bounding box and the dimensions of the proposal bounding box, and a regression value representing a difference between an orientation of the ground truth bounding box and an orientation of the proposal bounding box.

[0144] Example 8 includes the apparatus of any one of Examples 1 to 7, wherein the object classification is one of a plurality of possible object classifications, and one or more of the at least one processor circuit(s) determines the mask based on a union of valid masks determined for a first of the possible object classifications from ground truth data, or determines the mask based on an average of valid masks determined for the first of the possible object classifications from ground truth data, wherein the valid masks are selected based on a calculation of the normalized correlation coefficient.

[0145] Example 9 includes at least one non-transitory computer-readable storage medium containing instructions to cause at least one processor circuit to generate at least one overhead view of a three-dimensional point cloud, identify a candidate cluster of points in the three-dimensional point cloud based on at least one template or a mask applied to a sample point of the overhead view, wherein the candidate cluster is to satisfy an occupancy objective, and output parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster, the parameters based on the candidate cluster.

[0146] Example 10 includes the at least one non-transitory computer-readable storage medium of Example 9, wherein the instructions are to cause one or more of the at least one processor circuitry to implement a neural network trained to output a feature vector based on the candidate cluster, wherein the parameters are based on the feature vector.

[0147] Example 11 includes the at least one non-transitory computer-readable storage medium of Example 9 or Example 10, wherein the parameters include object classification probabilities for respective ones of a plurality of possible object classifications, and the instructions are to cause one or more of the at least one processor circuitry to implement an object classification header to process the feature vector to determine the object classification probabilities.

[0148] Example 12 includes the at least one non-transitory computer-readable storage medium of any of Examples 9 to 11, wherein the parameters comprise sets of bounding box regression values ​​corresponding to respective ones of the plurality of possible object classifications, and the instructions are to cause one or more of the at least one processor circuit to implement a plurality of regression heads corresponding to the plurality of possible object classifications, wherein respective ones of the regression heads process the feature vector to determine respective ones of the sets of bounding box regression values.

[0149] Example 13 includes the at least one non-transitory computer-readable storage medium of any one of Examples 9 to 12, wherein a first of the regression heads corresponds to a first of the possible object classifications, and the instructions are to cause one or more of the at least one processor circuitry to implement the first of the regression heads to process the feature vector to determine a first of the sets of bounding box regression values, the first of the sets of bounding box regression values ​​including values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications.

[0150] Example 14 includes the at least one non-transitory computer-readable storage medium of any one of Examples 9 to 13, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, and the instructions are to cause one or more of the at least one processor circuitry to select a second candidate cluster from a second three-dimensional point cloud, generate a proposal bounding box to cover a volume of the second candidate cluster, determine a set of regression values, represent the differences between the proposal bounding box and a ground truth bounding box, and train one or more machine learning algorithms based on the set of regression values, wherein the one or more machine learning algorithms are to determine the parameters based on the candidate cluster.

[0151] Example 15 includes a method comprising: identifying, by at least one processor circuit programmed by at least one instruction, a candidate cluster of points in the three-dimensional point cloud based on at least one template or mask applied to a sample point of an overhead view of the three-dimensional point cloud, wherein the candidate cluster is to satisfy an occupancy objective, processing the candidate cluster with a neural network to output a feature vector for the candidate cluster, and outputting parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster, the parameters based on the feature vector.

[0152] Example 16 includes the method of Example 15, wherein the feature vector is processed with an object classification head to determine a plurality of object classification probabilities, each corresponding to a plurality of possible object classifications.

[0153] Example 17 includes the method of Example 15 or Example 16, further comprising processing the feature vector with a plurality of regression heads to determine a plurality of sets of bounding box regression values ​​each corresponding to a plurality of possible object classifications, wherein the respective regression heads correspond to the respective possible object classifications.

[0154] Example 18 includes the method of any one of Examples 15 to 17, wherein a first of the regression heads corresponds to a first of the possible object classifications, and processing the feature vector with the plurality of regression heads comprises processing the feature vector with the first of the regression heads to determine a first of the sets of bounding box regression values, the first of the sets of bounding box regression values ​​including values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications.

[0155] Example 19 includes the method of any one of Examples 15 to 18, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, and further comprises selecting a second candidate cluster from a second three-dimensional point cloud, generating a proposal bounding box to cover a volume of the second candidate cluster, determining a set of regression values ​​representative of differences between the proposal bounding box and a ground truth bounding box, and training the neural network based on the set of regression values.

[0156] Example 20 includes the method of any of Examples 15 to 19, wherein the set of regression values ​​comprises a plurality of regression values ​​representing differences between a center point of the ground truth bounding box and a center point of the proposal bounding box, a plurality of regression values ​​representing differences between dimensions of the ground truth bounding box and dimensions of the proposal bounding box, and a regression value representing a difference between an orientation of the ground truth bounding box and an orientation of the proposal bounding box.

[0157] The following claims are hereby incorporated by reference into this Detailed Description. Although certain exemplary systems, devices, articles of manufacture, and methods have been disclosed herein, the scope of this patent is not so limited. Rather, this patent covers all systems, devices, articles of manufacture, and methods falling within the scope of the claims of this patent.

Claims

[1] An object detection device comprising: an interface circuit to obtain a three-dimensional point cloud of a scene; machine-readable instructions; and at least one processor circuit programmed by the machine-readable instructions to cause: applying at least one template or a mask to a sample point of an overhead view of the three-dimensional point cloud to identify a candidate cluster of points in the three-dimensional point cloud, the candidate cluster to satisfy an occupancy objective; Inputting the candidate cluster into a neural network, wherein the neural network outputs a feature vector for the candidate cluster; and Processing the feature vector to output parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster. [2] The apparatus of claim 1, wherein one or more of the at least one processor circuit(s) apply the mask centered on the sampling point of the overhead view of the three-dimensional point cloud, the mask being associated with a particular object class, the mask corresponding to a grid map having grid positions with a first value or a second value, some of the grid positions corresponding to the object classification having the first value and some of the grid positions not corresponding to the object classification having the second value, the occupancy target corresponding to a target number of points of the candidate cluster to be covered by the mask. [3] The apparatus of claim 1, wherein one or more of the at least one processor circuit(s) apply the template centered on the sample point of the overhead view of the three-dimensional point cloud, the template being associated with a particular object class, the template corresponding to a grid map having grid positions with respective values ​​representative of the number of points expected for object classification at the grid positions. [4] Apparatus according to any one of claims 1 to 3, wherein one or more of the at least one processor circuit(s) implement an object classification header and a plurality of regression heads to output the parameters, the parameters comprising respective prediction values ​​for a plurality of possible object classifications and respective sets of regression values ​​associated with respective bounding boxes corresponding to the plurality of possible object classifications, the object classification header to output the respective prediction values ​​for the plurality of possible object classifications, the plurality of regression heads to output the respective sets of regression values ​​associated with the respective bounding boxes corresponding to the plurality of possible object classifications. [5] The apparatus of claim 4, wherein a first of the sets of regression values ​​output by a first of the regression heads corresponding to a first of the possible object classifications includes values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications. [6] Apparatus according to any one of claims 1 to 4, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, the parameters are first parameters, the interface circuit is for obtaining a second three-dimensional point cloud, and one or more of the at least one processor circuit is for: Selecting a second candidate cluster from the second three-dimensional point cloud; Generating a proposal bounding box covering the volume of the second candidate cluster; Determining a first set of regression values ​​representing the differences between the proposal bounding box and a ground truth bounding box; inputting the second candidate cluster to the neural network, wherein the neural network outputs a second feature vector for the second candidate cluster; processing the second feature vector with a plurality of heads to output parameters including a second set of regression values; Determining an output of a loss function based on the first set of regression values ​​and the second set of regression values; and Updating the neural network and the plurality of heads based on the output of the loss function. [7] The apparatus of claim 6, wherein the first set of regression values ​​includes: a plurality of regression values ​​representing the differences between a center point of the ground truth bounding box and a center point of the proposal bounding box; a plurality of regression values ​​representative of the differences between the dimensions of the ground-truth bounding box and the dimensions of the proposal bounding box; and a regression value that represents a difference between an orientation of the ground truth bounding box and an orientation of the proposal bounding box. [8] Device according to one of claims 1 to 3, wherein the object classification is one of a plurality of possible object classifications and one or more of the at least one processor circuit belongs to at least one of the following: Determining the mask based on a union of valid masks determined for a first of the possible object classifications from ground truth data, wherein the valid masks are selected based on an occupancy fraction calculation; or Determining the template based on an average of valid templates determined for the first of the possible object classifications from ground truth data, wherein the valid templates are selected based on a calculation of the normalized correlation coefficient. [9] At least one non-transitory, computer-readable storage medium containing instructions for causing at least one processor circuit to do at least the following: Creating an overhead view of a three-dimensional point cloud; Identifying a candidate cluster of points in the three-dimensional point cloud based on a template and / or a mask applied to a sample point of the overhead view, wherein the candidate cluster is to satisfy an occupancy objective; and Outputting parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster, the parameters being based on the candidate cluster. [10] At least one non-transitory computer-readable storage medium according to claim 9, wherein the instructions are to cause one or more of the at least one processor circuits to implement a neural network trained to output a feature vector based on the candidate cluster, the parameters being based on the feature vector. [11] At least one non-transitory computer-readable storage medium according to claim 10, wherein the parameters include object classification probabilities for respective ones of a plurality of possible object classifications, and the instructions are to cause one or more of the at least one processor circuit to implement an object classification header to process the feature vector to determine the object classification probabilities. [12] At least one non-transitory computer-readable storage medium according to claim 11, wherein the parameters comprise sets of bounding box regression values ​​each corresponding to one of the plurality of possible object classifications, and the instructions are to cause one or more of the at least one processor circuit to implement a plurality of regression heads each corresponding to the plurality of possible object classifications, the respective regression heads processing the feature vector to determine the respective sets of bounding box regression values. [13] At least one non-transitory computer-readable storage medium according to claim 12, wherein a first of the regression heads corresponds to a first of the possible object classifications, and the instructions are to cause one or more of the at least one processor circuitry to implement the first of the regression heads to process the feature vector to determine a first of the sets of bounding box regression values, the first of the sets of bounding box regression values ​​including values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications. [14] At least one non-transitory computer-readable storage medium according to any one of claims 9 to 12, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, and the instructions are to cause one or more of the at least one processor circuit to: Selecting a second candidate cluster from a second three-dimensional point cloud; Generating a proposal bounding box covering the volume of the second candidate cluster; Determining a set of regression values ​​representing the differences between the proposal bounding box and a ground-truth bounding box; and Training one or more machine learning algorithms based on the set of regression values, wherein the one or more machine learning algorithms determine the parameters based on the candidate cluster. [15] Procedure comprising: Identifying, by at least one processor circuit programmed by at least one instruction, a candidate cluster of points in a three-dimensional point cloud based on at least one template or mask applied to a sample point of an overhead view of the three-dimensional point cloud, the candidate cluster satisfying an occupancy objective; Processing the candidate cluster with a neural network to output a feature vector for the candidate cluster; and Outputting parameters associated with an object classification and a bounding box for an object corresponding to the candidate cluster, the parameters being based on the feature vector. [16] The method of claim 15, further comprising processing the feature vector with an object classification head to determine a plurality of object classification probabilities, each corresponding to a plurality of possible object classifications. [17] The method of claim 15, further including processing the feature vector with a plurality of regression heads to determine a plurality of sets of bounding box regression values ​​each corresponding to a plurality of possible object classifications, the respective regression heads corresponding to the respective possible object classifications. [18] The method of claim 17, wherein a first of the regression heads corresponds to a first of the possible object classifications, and processing the feature vector with the plurality of regression heads comprises processing the feature vector with the first of the regression heads to determine a first of the sets of bounding box regression values, the first of the sets of bounding box regression values ​​including values ​​representing differences between (i) a first of the bounding boxes predicted by the first of the regression heads based on the feature vector and (ii) a ground truth bounding box corresponding to the first of the possible object classifications. [19] The method of any one of claims 15 to 17, wherein the three-dimensional point cloud is a first three-dimensional point cloud, the candidate cluster is a first candidate cluster, and further comprising: Selecting a second candidate cluster from a second three-dimensional point cloud; Generating a proposal bounding box to cover a volume of the second candidate cluster; Determining a set of regression values ​​representative of differences between the proposal bounding box and a ground-truth bounding box; and Training the neural network based on the set of regression values. [20] The method of claim 19, wherein the set of regression values ​​includes: a plurality of regression values ​​representing the differences between a center point of the ground truth bounding box and a center point of the proposal bounding box; a plurality of regression values ​​representative of the differences between the dimensions of the ground-truth bounding box and the dimensions of the proposal bounding box; and a regression value that represents a difference between an orientation of the ground truth bounding box and an orientation of the proposal bounding box.