Object tracking

WO2026180804A1PCT designated stage Publication Date: 2026-09-03OXA AUTONOMY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2026/050276
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-14
Filing Date
2026-02-25
Publication Date
2026-09-03

Smart Images

  • Figure GB2026050276_03092026_PF_FP_ABST
    Figure GB2026050276_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method of computer- implemented method of tracking an object from a scene in which an autonomous vehicle operates. The computer-implemented method comprises: receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle; estimating a ground surface using the depth map and the LiDAR point cloud; filtering the LiDAR point cloud to discard points of the LiDAR point cloud below the ground surface and retain points of the LiDAR point cloud at or above the ground surface; and generating an occupancy grid for tracking any objects in the scene using the filtered LiDAR point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] OBJECT TRACKING

[0002] FIELD

[0003]

[0001] The subject-matter of the present disclosure relates to computer-implemented methods of tracking objects in a scene in which an autonomous vehicle operates, methods of moving an autonomous vehicle, transitory, or non-transitory, computer-readable media, and autonomous vehicles.

[0004] BACKGROUND

[0005]

[0002] Typically, autonomous vehicles include a plurality of sensors including a LiDAR sensor, a RADAR sensor, and a camera. Sensor data from the plurality of sensors capture a scene in which the autonomous vehicle operates. Due to the myriad different permutations of scenes, tracking objects in the scene is a difficult challenge.

[0006]

[0003] It is an aim of the present invention to address such problems and improve on the prior art.

[0007] SUMMARY

[0008]

[0004] According to an aspect of the present disclosure, there is provided a computer-implemented method of tracking an object from a scene in which an autonomous vehicle operates, the computer-implemented method comprising: receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle; estimating a ground surface using the depth map and the LiDAR point cloud; filtering the LiDAR point cloud to discard points of the LiDAR point cloud below the ground surface and retain points of the LiDAR point cloud at or above the ground surface; and generating an occupancy grid for tracking any objects in the scene using the filtered LiDAR point cloud.

[0009]

[0005] In this way, the method reduces an influence of false detections by a LiDAR sensor. More specifically, ground water may prove difficult for a LiDAR sensor because an emitted LiDAR beam may bounce of the ground, or surface, water. The return beam may not reach the LiDAR sensor or may be delayed due to it bouncing off another object after the surface water. In this way, the LiDAR point cloud may include points below the water surface, giving the impression that the surface water is covering an extremely deep hole, e.g. a cliff edge. By filtering the LiDAR point cloud in the claimed manner, such false detections may be discarded.

[0006] In an embodiment, the method also comprises tracking the object from the scene using the generated occupancy grid. In another embodiment, the method comprises generating a three-dimensional representation of a scene using the generated occupancy grid. The term three-dimensional representation may be a visual representation or a non-visual representation, e.g. a format of documenting objects which can be used by downstream models for tasks such as trajectory generation.

[0010]

[0007] In an embodiment, the sensors include a LiDAR sensor, a RADAR sensor, and an image sensor, wherein receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle, comprises: receiving, from the LiDAR sensor, the RADAR sensor, and the image sensor a LiDAR point cloud, a RADAR point cloud, and images, respectively; generating a depth map using and the LiDAR point cloud, the RADAR point cloud and the images.

[0011]

[0008] In an embodiment, generating a depth map using and the LiDAR point cloud, the RADAR point cloud and the images comprises: inputting the LiDAR point cloud, the RADAR point cloud, and the images to a machine learning model trained to generate a depth map.

[0012]

[0009] In an embodiment, the computer-implemented method further comprises: optimising depth values of points of the depth map using points from each of the LiDAR point cloud and the RADAR point cloud.

[0013]

[0010] In an embodiment, verifying depth values of points of the depth map includes, for each point of the depth map: comparing a distance to points of the LiDAR and RADAR point clouds; and ignoring any points of the LiDAR and RADAR point clouds that are over a threshold distance to the depth map point; and optimising the depth map point using at least one point from the LiDAR and RADAR point clouds that is less than or equal to the threshold distance to the depth map point.

[0014]

[0011] In an embodiment, verifying the depth map point using at least one point from the LiDAR and RADAR point clouds comprises: adjusting the depth value of the depth map point to match a depth value of the at least one point from the LiDAR and RADAR point clouds.

[0015]

[0012] In an embodiment, estimating a ground surface using the depth map and the LiDAR point cloud comprises: inputting the depth map and the LiDAR point cloud to a machine learning model trained to estimate a ground surface.

[0013] In an embodiment, the occupancy grid comprises a plurality of cells, each cell including at least one label, wherein the at least one label is selected from a list of labels including a label of occupied or non-occupied, a label of occluded or non-occluded, a label of velocity, and a label of heading.

[0016]

[0014] In an embodiment, the occupancy grid is a first occupancy grid, wherein the computer-implemented method further comprising: generating a second occupancy grid based on the RADAR point cloud; generating a third occupancy grid based on the depth map; and fusing the first, second, and third occupancy grids.

[0017]

[0015] In an embodiment, the computer-implemented method further comprises: generating a three-dimensional object perimeter for the object based on the LiDAR point cloud, the RADAR point cloud, and the images; and semantically annotating the fused occupancy grid using the three-dimensional object perimeter.

[0018]

[0016] In an embodiment, the computer-implemented method further comprises: generating a three-dimensional representation of the scene using the semantically annotated occupancy grid and the three-dimensional object perimeter.

[0019]

[0017] In an embodiment, generating the three-dimensional representation of the scene comprises inputting the semantically annotated occupancy grid and the three-dimensional object perimeter to a machine learning model trained to generate the three-dimensional representation of the scene.

[0020]

[0018] According to an aspect of the present disclosure, there is provided a computer-implemented method of moving an autonomous vehicle, the computer-implemented method comprising: tracking an object in a scene in which the autonomous vehicle operates using the computer-implemented of any preceding aspect or embodiment; generating a trajectory for the autonomous vehicle based on the occupancy grid; and moving the autonomous vehicle along the trajectory.

[0021]

[0019] According to an aspect of the present disclosure, there is provided a transitory, or non-transitory, computer-readable medium, having instructions stored thereon that when executed by at least one processor, causes the at least one processor to perform the computer-implemented method of any preceding aspect or embodiment.

[0022]

[0020] According to an aspect of the present disclosure, there is provided an autonomous vehicle, comprising: a LiDAR sensor; a RADAR sensor; a camera; at least one processor; and storage having instructions stored thereon that when executed bythe at least one processor, causes the at least one processor to perform the computer-implemented method of any preceding aspect or embodiment.

[0023] BRIEF DESCRIPTION OF DRAWINGS

[0024]

[0021] The subject-matter of the present disclosure is best described with reference to the accompanying figures, in which:

[0025]

[0022] Figure 1 shows a side schematic diagram of an autonomous vehicle according to at least one embodiment;

[0026]

[0023] Figure 2 shows a block diagram of a perception model used by the autonomous vehicle of Figure 1 to track objects, according to at least one embodiment;

[0027]

[0024] Figure 3 shows a fundamental functions part of the perception model from Figure 2, according to at least one embodiment;

[0028]

[0025] Figure 4 shows a block diagram of an occupancy classification functions part of the perception model from Figure 2, according to at least one embodiment;

[0029]

[0026] Figure 5 shows a schematic diagram representing a method of filtering a point cloud using a ground surface estimate, by the occupancy classification functions part, according to at least one embodiment;

[0030]

[0027] Figure 6 shows a schematic of an occupancy grid generated by the occupancy classification functions part, according to at least one embodiment;

[0031]

[0028] Figure 7 shows an object detection model of the perception model from Figure 2, according to at least one embodiment;

[0032]

[0029] Figure 8 shows a semantic fusion functions part of the perception model from Figure 2, according to at least one embodiment;

[0033]

[0030] Figure 9 shows a schematic diagram in birds’-eye-view, of an object being tracked by the perception model, according to at least one embodiment;

[0034]

[0031] Figures 10A to 10C show a projected object perimeter, a bounding box, and a joined object perimeter, respectively, as generated by the semantic fusion functions part, according to at least one embodiment;

[0035]

[0032] Figures 11A to 11C show an occupancy grid, an object perimeter, and an adjusted object perimeter, respectively, as generated by the semantic fusion functions part, according to at least one embodiment;

[0033] Figure 12 is a three-dimensional perspective view of a scene 40 generated by the perception model of Figure 2, according to at least one embodiment;

[0036]

[0034] Figures 13 to 18 show flow charts summarising computer implemented methods of tracking objects in a scene in which an autonomous vehicle operates, according to at least one embodiment; and

[0037]

[0035] Figure 19 shows a flow chart summarising a computer-implemented method of moving the autonomous vehicle.

[0038] DESCRIPTION OF EMBODIMENTS

[0039]

[0036] Although the example embodiments have been described with reference to the components, modules and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements. Various combinations of optional features have been described herein, and it will be appreciated that described features may be combined in any suitable combination. In particular, the features of any one example embodiment may be combined with features of any other embodiment, as appropriate, except where such combinations are mutually exclusive. Throughout this specification, the term “comprising” or “comprises” means including the component(s) specified but not to the exclusion of the presence of others.

[0040]

[0037] The embodiments described herein may be embodied as sets of instructions stored as electronic data in one or more storage media. Specifically, the instructions may be provided on a transitory or non-transitory computer-readable media. When executed by the processor, the processor is configured to perform the various methods described in the following embodiments. In this way, the methods may be computer-implemented methods. In particular, the processor and a storage including the instructions may be incorporated into a vehicle. The vehicle may be an autonomous vehicle (AV).

[0041]

[0038] Whilst the following embodiments provide specific illustrative examples, those illustrative examples should not be taken as limiting, and the scope of protection is defined by the claims. Features from specific embodiments may be used in combination with features from other embodiments without extending the subject-matter beyond the content of the present disclosure.

[0042]

[0039] With reference to Figure 1, an AV 10 may include a plurality of sensors 12. The sensors 12 may be mounted on a roof of the AV 10, or integrated into the bumpers, grill, bodywork, etc. The sensors 12 may be communicatively connected to a computer 14.The computer 14 may be onboard the AV 10. The computer 14 may include a processor 16 and storage 18. The memory may include the non-transitory computer-readable media described above. Alternatively, the non-transitory computer-readable media may be located remotely and may be communicatively linked to the computer 14 via the cloud 20. The computer 14 may be communicatively linked to one or more actuators 22 for control thereof to move the AV 10. The actuators may include, for example, a motor, a braking system, a power steering system, etc.

[0043]

[0040] The computer 14 includes an autonomy stack for controlling the AV 10 stored on the storage 18. The autonomy stack may control the AV 10 in response to the sensor data. To achieve this, the autonomy stack may include one or more models. The one or more models may include an end-to-end machine learning model that is trained to provide control commands to actuators of the AV 10 in response to the sensor data. The one or more models may include models responsible for each function, namely perception, planning, and control. This may be in addition to the end-to-end model or as an alternative to the end-to-end model. Perception functions may include object detection and classification based on sensor data. Planning functions may include object tracking and trajectory generation. Control functions including setting control instructions for one or more actuators 22 of the AV 10 to move the AV 10 according to the trajectory.

[0044]

[0041] The sensors 12 may include various sensor types. Examples of sensor types include LiDAR sensors, RADAR sensors, and cameras. Each sensor type may be referred to as a sensor modality. Each sensor type may record data associated with the sensor modality. For example, the LiDAR sensor may record LiDAR modality data.

[0045]

[0042] The data may capture various scenes that the AV 10 encounters. For example, a scene may be a visible scene around the AV 10 and may include roads, buildings, weather, objects (e.g. other vehicles, pedestrians, animals, etc.), etc.

[0046]

[0043] With reference to Figure 2, the sensors 12 include a LiDAR sensor 12_1, a RADAR sensor 12_2, and a camera 12_3. The sensors 12 also include odometry 12_4. A platform configuration 12_5 is also provided.

[0047]

[0044] A perception model 30 is configured to generate a local scene 40 from sensor data from the respective sensors 12. The perception model 30 includes a fundamental functions part 32, an occupancy classification functions part 34, an object detection model 36, and a semantic fusion functions part 38, which collectively generate the local scene 40.

[0045] With reference to Figure 12, the local scene 40 may be a three-dimensional representation of the scene in which the AV 10 operates. The three-dimensional representation of the scene may include a plurality of object 160 from the scene, features of the scene, such as roads, sidewalks, buildings, road signs, lights, etc. and the AV 10.

[0048]

[0046] With further reference to Figure 2, the raw sensor data is input to the fundamental functions part 32. The fundamental functions part 32 outputs various data to the occupancy classification functions part 34 and the object detection model 36. The occupancy classification functions part 34 and the object detection model 36 output various data to the semantic fusion functions part 38, which generates the local scene 40. The specific data passed between the various models is described in more detail below.

[0049]

[0047] The fundamental functions part 32 includes a LiDAR combiner 42, a RADAR combiner 44, a camera combiner 46, a depth estimation model 48, and a ground surface estimation model 52.

[0050]

[0048] The LiDAR combiner 42 receives LiDAR sensor signals from the LiDAR sensor 12_1. The LiDAR combiner 42 generates a LiDAR point cloud 43 from the LiDAR sensor signals. Those LiDAR point clouds 43 are constructed for each time interval over a plurality of time intervals. The LiDAR combiner 42 combines the LiDAR point clouds to create a stream of LiDAR point clouds 43 over the plurality of time intervals.

[0051]

[0049] The RADAR combiner 44 receives RADAR sensor signals from the RADAR sensor 12_2. The RADAR combiner 44 generates a RADAR point cloud 45 form the RADAR sensor signals. Those RADAR point clouds 45 are constructed for each time interval over a plurality of time intervals. The RADAR combiner 44 combines the RADAR point clouds 45 to create a stream of RADAR point clouds 45 over the plurality of time intervals.

[0052]

[0050] The camera combiner 46, or image combiner, receives images from the camera 12_3, or cameras, and outputs the images 47 in a temporal, or spatial, sequence. In other words, the images are captured at each time interval over the plurality of time intervals. The images 47 may be two-dimensional images, or three-dimensional images. The images may be in three colour channel, e.g. red, green, and blue (RGB).

[0053]

[0051] The depth estimation model 48 receives the LiDAR point cloud 43, the RADAR point cloud 45, and the images 47 from the LiDAR combiner 42, the RADAR combiner 44, and the camera combiner 46, respectively. The depth estimation model 48 includesa learned model, e.g. a neural network. The learned model receives the images as an input, and outputs a depth map. In other words, the method may comprise generating a depth map using the LiDAR point cloud, the RADAR point cloud, and the images. The learned model is trained using supervised learning. For example, training data includes images which are paired with depth maps. The machine learning model is trained using back propagation and an optimisation algorithm, e.g. gradient descent. For example, the machine learning model may be trained by generating a depth map for each image, then optimising the weights of the machine learning model to reduce a loss between the generated depth maps and the paired depth maps of the training data.

[0054]

[0052] The depth estimation model 48 also includes a rules-based model. The term rules-based may mean a computer program which a computer programmer has coded, and includes steps that are executed by a processor. The rules-based model is used to verify a position of each point of the depth map 49 to see if it matches points of the LiDAR and RADAR point clouds. The position being verified may be a depth value of the depth map points. However, the horizontal and vertical values may also be verified in certain embodiments. The respective point clouds are sparser than the depth map. Therefore, the points of the depth map are compared positionally to the points of the point clouds. If there are no point cloud points within a threshold distance to the depth map point, the depth map point is ignored. In other words, those points are retained without being modified, because those points are not considered to correspond to one another. However, if a distance between a depth map point and a point cloud point is less than or equal to the threshold distance, the depth map point is corrected using the point cloud points. These respective points are considered to correspond to one another. The correction may be to move the depth map point to the position of the LiDAR and RADAR point clouds. In certain circumstances, the depth map points may be moved to a position between the point cloud point and the depth map point. The movement may be smoothed taking account of positions of adjacent points, e.g. fitting to a polynomial curve.

[0055]

[0053] The machine learning model of the depth estimation model is also used to obtain depth features 50. The depth features may also be called a feature map. The depth features may be features of objects represented in the depth map. For example, a car may have several edges, those edges may be captured as features in the feature map. The feature map may be extracted from a hidden layer of the neural network.

[0056]

[0054] Therefore, the output of the depth estimation model 48 is the depth map 49 and the depth features 50.

[0055] The ground surface estimation model 52 includes, or is, a machine learning model. The machine learning model may include a neural network. The machine learning model receives the verified depth map and the LiDAR point cloud as inputs. Therefore, the neural network may include two encoders; one for encoding the verified depth map and the other for encoding the LiDAR point cloud. Two encoders may be required because the input dimensions are different since the depth map is denser and has more points than the LiDAR point cloud. The encodings from the respective encoders may be combined, e.g. concatenated. The combined encoding may be input to a decoder to estimate a ground surface 53. The two encoders and the decoder may each be neural networks. The machine learning model may be trained end to end in a similar way to the supervised method described above, using LiDAR point clouds, corresponding depth maps, and corresponding, or matched, ground surfaces. The ground surface may be a mesh. The mesh may be a three-dimensional mesh. The mesh may be a coarse mesh, and may map contours of the ground surface. The supervised training may be end-to-end and may use similar, or the same, algorithms to those described above, such as back propagation and gradient descent. The output from the ground surface estimation model 52 may therefore be an estimate of a ground surface 53.

[0057]

[0056] The LiDAR point cloud 43, the RADAR point cloud 45, and the images 47 may be input to the object detection model 36. The depth map 49 and depth features 50 may be input to the object detection model 36. The depth map 49 and the ground surface 53 may be input to the occupancy classification functions part 34.

[0058]

[0057] With reference to Figure 4, the occupancy classification functions part 34 overall is responsible for generating an occupancy grid. The output is an occupancy grid that is the result of fusing three occupancy grids, those three occupancy grids generated from the respective input modalities. By fusing those occupancy grids, false positives may be suppressed using a voting procedure. False positives are particularly troublesome for LiDAR detections including smoke / vapour exhausted from a vehicle.

[0059]

[0058] The occupancy classification functions part 34 includes three occupancy grid estimation models. Those include a LiDAR occupancy grid estimation model 54 (or first occupancy grid estimation model), a RADAR occupancy grid estimation model 56 (or second occupancy grid estimation model), and an image occupancy grid estimation model 58 (or third occupancy grid estimation model). The occupancy grid estimation model 34 also includes an occupancy grid fusion model 60.

[0059] The first occupancy grid estimation model 54 is model free. In other words, it is a rules-based model, and includes no learned models. The first occupancy grid estimation model 54 receives the LiDAR point cloud 43 and the ground surface 53 as inputs.

[0060]

[0060] With reference to Figure 5, the points of the LiDAR point cloud 43 are compared positionally to the ground surface 53. Any points of the LiDAR point cloud 43 below the ground surface 53 are discarded. Discarded points are shown ringed 62 in Figure 5. Any points at or above the ground surface 53 are retained. Points associated with the smoke / vapour 61 are also shown ringed 63.

[0061]

[0061] LiDAR is particularly prone to false detections when surface water is present. This is because the LiDAR beam reflects off the surface of the water. Therefore, reflections from such points may take longer to return to the LiDAR sensor than would be the case if they were to reflect off the ground beneath the surface water. By filtering the LiDAR point cloud in this way, the occupancy grid is less likely to include erroneous points and so less likely to suffer false detections. In this way, LiDAR points that correspond to reflections from surface water are discarded.

[0062]

[0062] After filtering, an occupancy grid is generated.

[0063]

[0063] With reference to Figure 6, the occupancy grid may be a 2D grid in birds’ eye view and includes an array of cells. The occupancy grid is a low-resolution occupancy grid. In other words, it does not contain any semantic information about the objects occupying any cells. The occupancy grid may include a plurality of labels, which may include a binary label of occupied 64 or not occupied 66, and also may include a label relating to the heading direction 68 of the object occupying the cell, and a label of a velocity of the occupying object. Another label type may be occluded 70.

[0064]

[0064] The binary label may be determined based on a proximity of a point of the filtered LiDAR point cloud to the cell. For example, if the point is exactly in a centre of a cell, the probability of that cell being labelled occupied may be 1 and the probability of the cell being not occupied may be 0. Similarly, if a point is at an edge of the cell, the probability of being occupied may be between 0.5 and 1, and if the point is within an adjacent cell, the probability may be between 0 and 0.5. The probability may be fit to relationship of probability versus proximity. That relationship may be linear, logarithmic, a polynomial, etc., or may be a probability distribution, e.g. a gaussian distribution having a mean as a centre of a cell and decaying further from the centre of the cell. The binary label of occupancy status may be based on the intensity of a LiDAR return signal, additionally, or alternatively, to the proximity. In other words, a high intensity provides a higherprobability of being occupied, and a lower intensity provides a lower probability of being occupied.

[0065]

[0065] With further reference to Figure 4, the second occupancy grid estimation model 56 may be a learned model. The learned model may be a machine learning model and may include a neural network. The machine learning model 56 takes the RADAR point cloud(s) and the ground surface as inputs. The second occupancy grid estimation model 56 outputs an occupancy grid labelled in the same way as described above for the first occupancy grid.

[0066]

[0066] The machine learning model may include two encoders, one for each input. In other words, one encoder is for receiving and encoding the RADAR point cloud, the other encoder is for receiving and encoding the ground surface 53. The encodings may be combined. For example, the encodings may be concatenated. The machine learning model may also include a decoder to decode the combined encoding to produce the occupancy grid. The occupancy grid may be the same format as described above.

[0067]

[0067] The third occupancy grid estimation model 58 is a learned model. The learned model may be a machine learning model and may include a neural network. The machine learning model 58 takes the depth map 49 and the ground surface 53 as inputs. In this way, generating a third occupancy grid may be based on the depth map 49 and / or the ground surface 53. The third occupancy grid estimation model 58 outputs an occupancy grid labelled in the same way as described above for the first and second occupancy grid estimation models 56, 58.

[0068]

[0068] The machine learning model may include two encoders, one for each input. In other words, one encoder is for receiving and encoding the depth map 49, the other encoder is for receiving and encoding the ground surface 53. The encodings may be combined. For example, the encodings may be concatenated. The machine learning model may also include a decoder to decode the combined encoding to produce the occupancy grid. The occupancy grid may be the same format as described above.

[0069]

[0069] Each of the two encoders and the decoder may be neural network models and may be trained end-to-end. In other words, the training data may include depth maps and ground surfaces, respectively matched with occupancy grids. The training may be supervised and may use back propagation and an optimisation algorithm such as gradient descent. For example, the machine learning model may be run to estimate an occupancy grid using the respective depth map and ground surface. A loss may be calculated between the estimated occupancy grids and the ground truth, namely thematched occupancy grids. The weights of the neural network models may be optimised to minimise the loss.

[0070]

[0070] The labels of the occupancy grid are determined as a probability. For example, an output node of a decoder predicts that the label occupied has 0.86 probability and the label of not-occupied has 0.14 probability, the label of occupied will be selected because it is higher than the label of not occupied. This comparison approach may also be used for the first and second occupancy grids where the probability of a cell being occupied or not occupied is based on proximity to a point cloud point. The label with a highest probability can be selected for the label of occupied or not occupied.

[0071]

[0071] The occupancy grids may be called first, second, and third occupancy grids 72, 74, 76, from the first, second, and third occupancy grid estimation models 54, 56, 58, respectively.

[0072]

[0072] The three occupancy grids 72, 74, 76, are input to the occupancy grid fusion model 60.

[0073]

[0073] The fusion involves voting on a cell-by-cell approach. The probability associated with a label may be taken as a confidence of the label. The voting may be weighted. The weights may be calibration factors associated with the respective sensors. For example, LiDAR may have a lower calibration factor than RADAR and the camera because LiDAR is known to be prone to errors such as false positive detections.

[0074]

[0074] As an example, water vapour or smoke from an exhaust of a vehicle may be detected by LiDAR as a positive detection of an object. This false detection may be present in the first occupancy grid 72. However, the second, and third, occupancy grids 74, 76, may not have a label of occupied on those corresponding cells but may instead have a label for those cells that is not occupied. Since the calibration factor for LiDAR may be lower than that for RADAR and cameras, the occupied label will be weights lower and so the fused occupancy grid 78 may have not occupied labels in those affected cells.

[0075]

[0075] The occupancy grid may be a dynamic occupancy grid when including a velocity for an occupied cell. For example, the RADAR occupancy grid and the fused occupancy grid may always be a dynamic occupancy grid.

[0076]

[0076] In a more detailed example, a cell of the first (LiDAR) occupancy grid 72 may have a occupied label whereas the corresponding cell of the second (RADAR), and third (camera) occupancy grids 74, 76, may have a not occupied label. The associated confidences producing those labels may be as follows:

[0077] Occupied Confidences

[0077] LiDAR 0.4

[0078] RADAR 0 (no data)

[0079] Camera 0

[0080]

[0078] Not occupied Confidences

[0081] LiDAR 0.2

[0082] RADAR 0 (no data)

[0083] Camera 0.86

[0084]

[0079] The weighted voting to obtain a fused label confidence (FusedLabelConfidence) may use a formula such as:

[0085] ConfidenceLiDAR*CalibrationFactorLiDAR + ConfidenceRADAR*CalibrationFactorRADAR + Confidencecamera*CalibrationFactorcamera = FusedLabelConfidence

[0086] This formula may be used twice; once for the occupied label and once for the nonoccupied label.

[0087]

[0080] The calibration factor for the occupied and not occupied labels may be different. This is to take into account the fact that the probability may be higher for occupied than not occupied labels for phantom objects such as the water vapour or smoke example provided above. In particular, the LiDAR calibration factor may be higher for the not-occupied label than the occupied label. The RADAR and camera calibration factors may be the same for each label.

[0088]

[0081] As an illustrative example, occupied calibration factors for LiDAR, RADAR, and camera, may be 0.8, 1 , and 1 , respectively. The not-occupied calibration factors forthose modalities may respectively be 1.2, 1, and 1.

[0089]

[0082] Using the equation above, this would equate to 0.4*0.8 + 0 + 0 = 0.32 for an occupied label.

[0090]

[0083] Using the equation above, this would equate to 0.2*1.2 + 0 + 0.86*1 = 1.1 for a not-occupied label. Therefore, the not-occupied label is used for that cell of the fused model.

[0091]

[0084] In this way, a weight for each occupancy grid may be a calibration factor associated with the respective sensor. The calibration factor for the first occupancy grid(LiDAR occupancy grid) may be higher for the non-occupied label than the occupied label. The calibration factor for the second and third occupancy grids may be the same for both the non-occupied and occupied labels.

[0092]

[0085] With reference to Figure 7, the object detection model 36 takes the depth map 49, the images 47, the depth features 50, the RADAR point cloud 45, and the LiDAR point cloud 43, as inputs. The object detection model 36 uses those inputs to generate a projected object perimeter 126 (which may be three-dimensional), an object embedding 82, a bounding box 84 (which may be three-dimensional), and a birds’-eye-view (BEV) semantic segmentation map 86, which may be input to the semantic fusion functions part 38.

[0093]

[0086] The object detection model 36 performs a plurality of functions when generating the respective outputs using those inputs.

[0094]

[0087] The plurality of functions includes a camera features extraction 88, a RADAR features extraction 90, a LiDAR features extraction 92, camera BEV transformation 94, RADAR BEV transformation 96, LiDAR BEV transformation 98, feature fusion 100, bounding box generation 102, BEV segmentation generation 104, two-dimensional image segmentation 106, and three-dimensional projection 108.

[0095]

[0088] The feature extractions 88, 90, 92, the BEV transformations 94, 96, 98, and respective encodings 116, 118, 120, for each modality (LiDAR, RADAR, and camera) may be achieved using a respective encoder. For example, there may be three encoders of the object detection model 36. Those include a camera encoder, a RADAR encoder, and a LiDAR encoder. The image, or camera, encoder receives the images 47 and the depth features as inputs. More specifically, the encoder receives the images 47 as an input vector. A hidden layer of the encoder represents the feature map 110. In other words, the method may comprise generating a feature map 110 using the received images 47. Therefore, the depth features 50 can be combined with the feature map 110 for input to the next hidden layer. The image encoder thus encodes those inputs to generate an image, or camera, encoding 116. In other words, the method may comprise generating, using an image encoder, an image encoding from the images 47.

[0096]

[0089] The LiDAR encoder receives the LiDAR point cloud 43 as an input, and encodes the LiDAR point cloud into a LiDAR encoding 120. Similarly, the RADAR encoder receives the RADAR point cloud 45 as an input, and encodes the RADAR point cloud into a RADAR encoding 118. The LiDAR and RADAR feature maps 114, 112, may thus represent values at a hidden layer of the respective encoder.

[0090] The object detection model 36 also performs a fusion function 100 for fusing the three encodings. To perform the fusion function 100, the object detection model 36 uses a fusion encoder. The fusion encoder fuses the camera encoding 116, the LiDAR encoding 120, and the RADAR encoding 118, and outputs a fused encoding, also called an embedding 122.

[0097]

[0091] The embedding can be extracted and input to the semantic fusion part 38, and also input to two decoders of the object detection model 36. The two encoders perform respective functions of bounding box generation 102 and semantic segmentation in bird’s eye view (BEV) 104. The decoders receive the embedding 122 as an input and output respective bounding box 84 and semantic segmentation map in BEV 86.

[0098]

[0092] The bounding box model 102 may be trained to generate a bounding box for any objects in the scene. The objects are those objects that the features relate to. The bounding box 84 may be a three-dimensional bounding box which surrounds the object and includes a semantic label for the class of the object. The classes of objects may include classes such as car, bus, truck, pedestrian, dog, cat, bird, pram, bicycle, etc. The training includes estimating a bounding box (e.g. its size, orientation, position, and semantic label) from an embedding, as part of an end-to-end training process described below.

[0099]

[0093] The object detection model 36 also performs the functions of 2D semantic segmentation and 3D projection. The 2D semantic segmentation is performed by a learned semantic segmentation model 106, which receives the image feature map 110 and outputs a 2D segmentation map 124. In this way, the method comprises generating a two-dimensional semantic segmentation map 124 based on the images 47. In other words, the method may comprise generating a two-dimensional semantic segmentation map using the feature map. More broadly, the method may comprise generating a two-dimensional semantic segmentation map using the received images 47. The segmentation map 124 takes its usual meaning in the field, and may be in two-dimensions. Those two-dimensions may by the vertical and horizontal dimensions of the image feature map, which correspond to the vertical and horizontal dimensions from the corresponding images 47.

[0100]

[0094] The 3D projection may be performed by a 3D projection model 108 which generates a 3D projected object perimeter 126 from the 2D segmentation map, by using the depth map 49. In other words, the method may comprise generating a three-dimensional projected object perimeter based on the two-dimensional semanticsegmentation map and the depth map. This may be a learned approach or a rules-based approach whereby depth values are taken from the depth map 49 to estimate a 3D projection of an object perimeter from the 2D segmentation map 124. For example, the two-dimensional coordinates of the two-dimensional segmentation maps 124 are complemented with a depth coordinate from the depth map 49. In other words, a point on the depth map may be matched with a feature of the two-dimensional segmentation maps 124, and the depth coordinate value may be added to the horizontal and vertical coordinate value to provide a three-dimensional coordinate value for that point. Since the depth values of the projected object perimeter are estimates, they are prone to errors. These errors may be corrected downstream by the semantic fusion model 38 when generating a fused object perimeter by taking account of the bounding box dimensions.

[0101]

[0095] In the learned approach, the 3D projection model 108 may be a machine learning model, e.g. including a neural network, and may generate the projected object perimeter 126 using the 2D semantic segmentation map 124 and the depth map 49 as inputs. In other words, the 3D projection model 108 may include two encoders, one for the 2D semantic segmentation map 124 and the other for the depth map 49, and those encoders may generate respective encodings. The encodings may be combined, e.g. by concatenation, and input to a decoder to generate the projected object perimeter 126.

[0102]

[0096] The encoders and decoders described in relation to the object detection model 36 may be neural networks.

[0103]

[0097] The object detection model 36 may be trained end-to-end. In other words, the various encoders, decoders, and models may be trained end-to-end using training data. For example, the training data may include images, LiDAR point cloud, RADAR point cloud, depth feature map 50 and a depth map 49 as inputs, paired with outputs project object perimeter, embedding, bounding box, and BEV semantic segmentation map. In other words, the individual encoders, and decoders, may not be trained individually, each using its own designated training data. The training may be performed in a supervised manner such as estimating the outputs, computing a loss compared to the ground truth of those outputs in the training data, and optimising the respective encoders, models, and decoders, to minimise the loss. This may be achieved using backpropagation and an optimisation algorithm such as gradient descent.

[0104]

[0098]

[0105]

[0099] With reference to Figure 8, the semantic fusion functions part 38 takes inputs from the object detection model 36 and the occupancy classification functions part 34.The inputs from the object detection model 36 are the three-dimensional bounding box 84, the three-dimensional projected object perimeter 126, and the embeddings 122. The input from the occupancy classification functions part 34 is the fused occupancy grid 78.

[0106]

[0100] The output from the semantic fusion functions part 38 is a three-dimensional representation of the scene, as would be commonly understood by the term in this field.

[0107]

[0101] The semantic fusion functions part 38 includes a perimeter generation model 130, a semantic annotation model 132, a spatial alignment model 134, a state estimate prediction model 136, and a temporal alignment model 138.

[0108]

[0102] The perimeter generation model 130 may be a learned model, and may be a machine learning model, e.g. including a neural network. The machine learning model receives three-dimensional bounding boxes 84 and three-dimensional projected object perimeters 126 as inputs, and generates an object perimeter as an output. The object perimeter may be a three-dimensional object perimeter. The training may be supervised. The training data may includes three-dimensional bounding boxes 84, projected object perimeters 126, matched with three-dimensional object perimeters. The supervised training may take the same approach described above with other models, where weights are optimised to reduce a loss between an object perimeter estimate and those of the ground truth in the training data.

[0109]

[0103] The machine learning model may include two encoders, one to encode the bounding boxes 84 to an encoding and other to encode the segmentation maps 126 to an encoding. Once combined, the encodings are decoded by a decoder to generate the three-dimensional joined object perimeter 140. The term joined may be replaced with enhanced.

[0110]

[0104] The three-dimensional joined object perimeter 140 may be similar to a bounding box in that it is a polygon surrounding the object. However, it may not a box shape, i.e. it may not be a rectangular cuboid or a cube. It may have more edges and faces than a bounding box. Therefore, the object perimeter could, theoretically, include bounding boxes (in cases where no adjustments are made to the bounding box size and shape), and polygons of other shapes.

[0111]

[0105] The perimeter generation model 130 in other embodiments may be a rules-based model. In such embodiments, the rules-based model may adjust the bounding box 84 to cover points occupied by the object in the projected object perimeter 126. The adjustment may be to adjust a size and shape of the bounding box.

[0106] With reference to Figure 9, an illustrative example may include a car 142 having a car door 144 which is open. The car door 144 may be an edge case that is not in the training data for training the bounding box model 102. Therefore, the car door will protrude outwardly from the bounding box 84, if the car were superimposed on the bounding box.

[0112]

[0107] With reference to Figure 10A, the three-dimensional projected object perimeter 126 is shown, albeit from a top, or BEV, for ease of illustration. The projected object perimeter 126 shows the car 142 and the car door 144. The projected object perimeter 126 is deficient because its starting point is a 2D semantic segmentation. For example, points behind the car door 144 may be erroneous because the 2D semantic segmentation map will not have been able to see past the car door. Therefore, it needs to be corrected using the bounding box.

[0113]

[0108] With reference to Figure 10B, the bounding box 84 is also shown in two-dimensions from a BEV for ease of illustration. However, only the car overlaps with the bounding box 84. The car door does not overlap the bounding box 84.

[0114]

[0109] With reference to Figure 10C, the three-dimensional joined object perimeter 140 is also shown from a BEV, and so is shown in two-dimensions, for illustrative purposes. The three-dimensional joined object perimeter 140 is constructed by replacing an edge at a lateral side of the car to cover the car door 144. In this way, the three-dimensional joined object perimeter is the same as a bounding box but is not limited to a rectangular cuboid or a cube shape. Therefore, it has a higher resolution than the bounding box and a lower resolution than the three-dimensional projected object perimeter.

[0115]

[0110] In this way, any edge cases such as car doors which are not in the training data can be accommodated when generating the three-dimensional representation of the scene 40.

[0116]

[0111] With further reference to Figure 8, the semantic annotation model 132 may be a rules-based model. The semantic annotation model 132 may receive the three-dimensional joined object perimeter 140 and the occupancy grid (the fused occupancy grid) 78 as inputs. The semantic annotation model 132 may be configured to annotate the occupancy grid 78 using the three-dimensional object perimeter. To do this, the label of the object in the three-dimensional object perimeter 140 is added to cells of the occupancy that positionally correspond. In other words, if a cell positionally falls within, or overlaps, the object perimeter 140 of the object, that cell can be labelled with the semantic label associated with that object.

[0112] The output from the semantic annotation model 132 is an annotated occupancy grid 146.

[0117]

[0113] The spatial alignment model 134 receives the annotated occupancy grid 146 and the three-dimensional object perimeter 140 as inputs. The spatial alignment model 134 may be a rules-based model. Operation of the spatial alignment model 134 may be best described with reference to a specific example.

[0118]

[0114] The output from the spatial alignment model 134 are filtered occupancy grids 147 and adjusted object perimeters 150.

[0119]

[0115] With reference to Figure 9, the car 142, which may be construed broadly enough to include a lorry or truck, e.g. a flat-bed truck, may be carrying cargo 148. The cargo 148 may be dimensioned such that it overlaps the object perimeter 140. For example, the cargo may be elongated and may be too long for the object perimeter 140, e.g. the cargo 148 may be, or include, logs.

[0120]

[0116] With reference to Figure 11 A, the annotated occupancy grid 146 may include cells that are labelled as occupied which correspond to the car 142, and have been semantically labelled as the car. There may be other cells that are labelled as occupied and correspond to the cargo 148, but which are not semantically labelled as the car.

[0121]

[0117] With reference to Figure 11 B, those latter cells do not overlap the object perimeter 140. It will be appreciated that the object perimeter could be the same as that shown in Figure 9C in the event that the car door was open on a car carrying cargo 148, as in this example.

[0122]

[0118] With reference to Figure 11 C, the joined object perimeter 140 may be adjusted to cover the latter cells, i.e. those that relate to the cargo 148, to form an adjusted three-dimensional object perimeter 150. It will be appreciated that most of the time the object perimeters are not adjusted. The adjustments are to take account of edge cases such as the cargo example provided above.

[0123]

[0119] The annotated occupancy grid 146 is adjusted to remove occupied labels from cells that overlap positionally with the object in the adjusted joined object perimeter 140. In this way, the object is only tracked once.

[0124]

[0120] In order to identify the cells associated with cargo 148, the spatial alignment model 134 creates a perimeter surrounding the object, e.g. the car 142. This perimeter may be a threshold distance, and may be a number of cells. The number of cells may be predetermined. The spatial alignment model then compares the heading and velocitylabels of any occupied cells within the threshold distance to those of the occupied cells associated with the object, e.g. the car 142. If the heading and velocity are the same, or within a threshold difference, those cells are identified as being related to the object. They may be labelled in the same way, e.g. semantically labelled as the car. Any other cells within the threshold distance are ignored. Then, any cells that are semantically labelled as the object, and do not fall within the object perimeter 140, may be identified and the object perimeter 150 may be adjusted to cover those identified cells. The adjustment may be to adjust the size and shape of the object perimeter.

[0125]

[0121] It will again be appreciated that not all object perimeters are adjusted, as they may not relate to an edge case. Therefore, it is possible that the adjusted object perimeter may be, in effect, a bounding box. Therefore, the term adjusted object perimeter, may be merely an object perimeter in some case, and so could be called such in those cases, and so may also be called a finalised, or verified, three-dimensional object perimeter.

[0126]

[0122] With further reference to Figure 8, the state estimate prediction model 136 receives the adjusted object perimeter 150, the annotated occupancy grid 146, and the embedding 122 as inputs. Those inputs are received at each time period over a plurality of time periods.

[0127]

[0123] The state estimate prediction model 136 may be a learned model, e.g. a machine learning model, and may include neural networks. Those neural networks may be split into three encoders, and two decoders. There is one encoder for each input type. The encoders encode those inputs to an encoding. Those encodings are combined, e.g. by concatenation. The combined encodings may be decoded by a first of the two decoders to predict an occupancy grid 152, where the occupancy grid may be semantically annotated. The second decoder may decode the combined encodings to predict a three-dimensional object perimeter 154.

[0128]

[0124] The embeddings provide additional information that may not be captured by the three-dimensional projected object perimeters, the bounding boxes, or the occupancy grids. For example, hand gestures from a driver of a vehicle can be included in the embeddings which may be missed in the other inputs to the semantic fusion functions part. In this way, the predicted occupancy grid 152 and the predicted object perimeter 154 may be less conservative.

[0129]

[0125] The temporal alignment model 138 may be a machine learning model, which may include neural networks. The temporal alignment model 138 may receive the predictedoccupancy grid 152 from the state estimate prediction model 136, the predicted three-dimensional object perimeter 154 from the state estimate prediction model 136, filtered occupancy grid 147 from the spatial alignment model 134, and the adjusted three-dimensional object perimeter 150 from the spatial alignment model 134.

[0130] The predicted occupancy grid 152 and the predicted 3D Object perimeter 154 are computed using past estimates, evaluated using kinematic models to predict their current state. These predictions can then be compared (same time step) in Temporal Alignment model 138 with the filtered occupancy grid 147 and the adjusted 3D Object perimeters 150.

[0131]

[0126] The machine learning model may include four encoders, each being a neural network. The encoders respectively receive one of the four inputs. Those encoders encode those inputs to generate encodings. The encodings are combined. The combined encodings are the decoded by a decoder to generate the scene 40. Training the model may be done in the same way as those models described above. The training data may include examples of those four inputs, matched with scenes. The training may be end-to-end. The weights of the neural networks are optimised to minimise a loss between scene predictions and the ground truth scenes from the training data.

[0132]

[0127] In this way, the scene will be less conservative than if the state estimate prediction model 136 were not included.

[0133]

[0128] The state estimate prediction model 136 may also help address another issue. For example, some objects in a scene may be temporarily occluded. For example, a bus moving between the AV 10 and a pedestrian on a sidewalk. To address this issue, the state estimate prediction model 136 may predict the occupancy grid 152 and the three-dimensional object perimeter 154 for the current time step, using an annotated occupancy grid 146 and an adjusted three-dimensional object perimeter 150 from a previous time step (using the specific kinematic model for the object semantic class). To address this issue, the embedding may be used but is optional as it is not necessarily required.

[0134]

[0129] The scene 40 may be used by the planning and control models to, respectively, generate a trajectory for the AV 10 and to control the AV 10 to move the AV 10 along the trajectory. The trajectory may be generated based on, or using, the scene 40. For example, the trajectory may be generated to avoid collisions with any objects in the scene 40. If no trajectories are available that avoid collisions, then a minimal risk manoeuvre (MRM) may be performed by the AV 10.

[0130] The perception model 30, or perception architecture 30, may be used to solve various problems. Those problems may be understood best with reference to different use cases. Those use cases are not mutually exclusive and can operate simultaneously, independently, or in combination. Therefore, features described below in relation to one use case may be performed equally in another use case.

[0135]

[0131] With reference to Figure 13, the perception model 30 may be used to help overcome issues with edge cases, where objects may have dimensions that are larger than the generate bounding boxes. For example, a vehicle with an open door, e.g. a car door opened, as described above, may not be included in the training data for generating the bounding boxes, so the car door may overlap a bounding box associated with the car. The perception model may be used to solve this issue as a computer-implemented method of tracking an object from a scene in which an AV 10 operates. The computer-implemented method may be summarised as including: receiving 202 a LiDAR point cloud, a RADAR point cloud, and images, from a LiDAR sensor, a RADAR sensor, and a camera, respectively, of the autonomous vehicle; generating 204 a three-dimensional bounding box around the object identified based on the LiDAR point cloud, the RADAR point cloud, and the images; generating 206 a three-dimensional projected object perimeter based on the LiDAR point cloud, the RADAR point cloud, and the images; generating 208 a three-dimensional object perimeter using the three-dimensional bounding box and the three-dimensional projected object perimeter; and generating 210 a three-dimensional representation of the scene including the object, using the three-dimensional object perimeters.

[0136]

[0132] Generating the three-dimensional bounding box around the object comprises: extracting features from the LiDAR point cloud, the RADAR point cloud, and the images; fusing the extracted features; and generating the three-dimensional bounding box using the fused extracted features. This may be done using the object detection model 36. More particularly, the end-to-end model may generate the bounding box based on the sensor data of those modalities.

[0137]

[0133] Extracting features from the LiDAR point cloud, the RADAR point cloud, and the images, comprises: extracting features from the LiDAR point cloud, the RADAR point cloud, and the images; and transforming the respective extractive features into a birds’ eye view, and wherein fusing the extracted features comprises fusing the extractedfeatures in the birds’ eye view. This may all be done using the end-to-end model. For example, the fusing may be done using the fusion function, or fusion model, 100.

[0138]

[0134] Extracting features from the LiDAR point cloud, the RADAR point cloud, and the images comprises inputting the LiDAR point cloud, the RADAR point cloud, and the image to respective encoders trained to extract features. The features are formed as the encodings from the encoders. In this way, the method may be said to generate encodings from the respective input modalities, where the encodings represent extracted features, optionally in BEV.

[0139]

[0135] Fusing the extracted features comprises inputting the extracted features in birds’ eye view to a machine learning model trained to fuse extracted features. The machine learning model may be the fusion encoder. The fusion encoder may fuse the encodings associated with each of the input modalities, to form a combined encoding. The combined encoding may represent the fused extracted features. The combined encoding may be considered an embedding.

[0140]

[0136] Generating the three-dimensional bounding box around the object comprises inputting the embedding to a machine learning model trained to decode the embedding into a three-dimensional bounding box semantically labelled with the object. The machine learning model may be the bounding box model 102.

[0141]

[0137] Generating a three-dimensional object perimeter using the three-dimensional bounding box and the three-dimensional projected object perimeter comprises: adjusting a shape and size of the three-dimensional bounding box to encompass the object in the three-dimensional projected object perimeter map. This process is described above with reference to Figures 10A to 10C.

[0142]

[0138] The method may also comprise generating an occupancy grid based on the LiDAR point cloud, the RADAR point cloud and the images. This may be achieved using the occupancy grid estimation model 34, as described above.

[0143]

[0139] The method may also include semantically annotating the occupancy grid using the three-dimensional object perimeter. This is described in more detail above in relation to Figures 11A to 11C.

[0144]

[0140] Generating a three-dimensional representation of the scene including the object, also uses the semantically annotated occupancy grid 146. In this way, the temporal alignment model 138 may use the semantically annotated occupancy grid 146 directly, or indirectly when the gird is filtered. In other words, the semantically annotatedoccupancy grid 146 or the filtered occupancy grid 147 may be input to the temporal alignment model 138 to generate the scene 40.

[0145]

[0141] In this way, generating the three-dimensional representation of the scene 40 including the object comprises inputting the semantically annotated occupancy grid 146 (or filtered occupancy grid 147) and the three-dimensional object perimeter 140, 150, to a machine learning model trained to generate the three-dimensional representation of the scene. The machine learning model may be the temporal alignment model 138.

[0146]

[0142]

[0147]

[0143] With reference to Figure 14, the perception model 30 may be used to help address issues with particular sensor modalities that are vulnerable to false detections in certain scenarios. For example, LiDAR sensor may exhibit erroneous detections when encountering surface water. One reason for this is that the LiDAR pulse may bounce off the surface of the water and reflect from another element, or not at all. This could be interpreted as the surface water being a very deep hole, due to the delay time to receive the return signal being extended.

[0148]

[0144] Therefore, the perception model may be used as part of a computer-implemented method of tracking an object from a scene in which the AV 10 operations, where the method may be summarised as including: receiving 302 a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle; estimating 304 a ground surface using the depth map and the LiDAR point cloud; filtering 306 the LiDAR point cloud to discard points of the LiDAR point cloud below the ground surface and retain points of the LiDAR point cloud at or above the ground surface; and generating 308 an occupancy grid for tracking any objects in the scene using the filtered LiDAR point cloud.

[0149]

[0145] Receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle, comprises: receiving, from the LiDAR sensor, the RADAR sensor, and the image sensor a LiDAR point cloud, a RADAR point cloud, and images, respectively; generating a depth map 49 using and the LiDAR point cloud, the RADAR point cloud and the images, generating a depth map using and the LiDAR point cloud, the RADAR point cloud and the images comprises: inputting the LiDAR point cloud, the RADAR point cloud, and the images to a machine learning model trained to generate a depth map 49. The machine learning model may be included in the depth estimation model 48. The depth map 49 may be generated using the depth estimation model 48, as described above.

[0146] The method may further comprise optimising depth values of points of the depth map using points from each of the LiDAR point cloud and the RADAR point cloud, verifying depth values of points of the depth map includes, for each point of the depth map: comparing a distance to points of the LiDAR and RADAR point clouds; and ignoring any points of the LiDAR and RADAR point clouds that are over a threshold distance to the depth map point; and optimising the depth map point using at least one point from the LiDAR and RADAR point clouds that is less than or equal to the threshold distance to the depth map point, verifying the depth map point using at least one point from the LiDAR and RADAR point clouds comprises: adjusting the depth value of the depth map point to match a depth value of the at least one point from the LiDAR and RADAR point clouds.

[0150]

[0147] Estimating a ground surface using the depth map and the LiDAR point cloud comprises: inputting the depth map and the LiDAR point cloud to a machine learning model trained to estimate a ground surface. The machine learning model may be included in the ground surface estimation model 52, as described above.

[0151]

[0148] With brief reference to Figure 6, the occupancy grid comprises a plurality of cells, each cell including a label. The label may include a binary label of occupied or nonoccupied, and may include another binary label of occluded or non-occluded, and may include a label of velocity and heading.

[0152]

[0149] The occupancy grid may be a first occupancy grid, and the method may further comprise generating a second occupancy grid based on the RADAR point cloud; generating a third occupancy grid based on the LiDAR point cloud; and fusing the first, second, and third occupancy grids. This may be achieved using the first, second, and third, occupancy grid estimation models 54, 56, 58, and the occupancy grid fusion model 60, as described above.

[0153]

[0150] The method may further comprise generating a three-dimensional object perimeter for the object based on the LiDAR point cloud, the RADAR point cloud, and the images; and semantically annotating the fused occupancy grid using the three-dimensional object perimeter. The object perimeter may be the projected object perimeter 126 described above. The semantic annotation may be achieved using the semantic annotation model 132 described above.

[0154]

[0151] Generating a three-dimensional representation of the scene 40 may be achieved using the semantically annotated occupancy grid and the three-dimensional joined object perimeter. In more detail, this may be achieved using the spatial alignment model 134,the temporal alignment model 138, and optionally the state estimate prediction model 136, as described above.

[0155]

[0152] Generating the three-dimensional representation of the scene 40 comprises inputting the semantically annotated occupancy grid and the three-dimensional object perimeter to a machine learning model trained to generate the three-dimensional representation of the scene. The three-dimensional object perimeter may be a joined object perimeter and the semantically annotated occupancy grid may be the filtered occupancy grid. The machine learning model may refer to the temporal alignment model 138.

[0156]

[0153] With reference to Figure 15, another problem to solve is to eliminate false positive detections of the LiDAR sensor. For example, smoke or vapour (e.g. water vapour such as steam) emitted from a vehicle can be misinterpreted as a solid object, e.g. exhaust gases may be misinterpreted as a trailer towed by the vehicle. Therefore, there is provided a computer-implemented method of tracking an object from a scene in which the AV operates, which may be summarised as including: receiving 402 a LiDAR point cloud from a LiDAR sensor, a RADAR point cloud from a RADAR sensor, and images from a camera, of the autonomous vehicle; generating 404 a first occupancy grid based on the LiDAR point cloud; generating 406 a second occupancy grid based on the RADAR point cloud; generating a third occupancy grid based on the images; and fusing 408 the first, second, and third occupancy grids.

[0157]

[0154] Fusing the first, second, and third occupancy grids comprises: generating a fused occupancy grid by applying voting to the first, second, and third occupancy grids. This may be achieved using the occupancy grid fusion model 60. Voting between the first, second, and third occupancy grids comprises applying weighted voting to the first, second, and third occupancy grids.

[0158]

[0155] Applying weight voting comprises applying voting where the first, second, and third occupancy grids are weighted differently depending on a confidence associated with the LiDAR sensor, RADAR sensor, and camera, respectively. A weight of the first occupancy grid is lower than a weight of each of the second and third occupancy grids. The weighted voting is described above with reference to the occupancy grid fusion model 60.

[0159]

[0156] Each of the first to third, and the fused, occupancy grids, may include at least one label, or a plurality of labels. The label may include a label indicating if the cell is occupied or non-occupied. The label may include a label indicating if the cell is occluded or non-occluded. The label may include a label indicating the velocity and a label indicating a heading.

[0160]

[0157] The method may further comprise: generating a three-dimensional object perimeter for the object based on the LiDAR point cloud, the RADAR point cloud, and the images; and semantically annotating the fused occupancy grid using the three-dimensional object perimeter. This may be achieved using the semantic annotation model 132. The object perimeter used may be a joined object perimeter 140.

[0161]

[0158] The method may further comprise: generating a three-dimensional representation of the scene 40 using the semantically annotated occupancy grid and the three-dimensional object perimeter. In more detail, this may be achieved using the spatial alignment model 134, the temporal alignment model 138, and optionally the state estimate prediction model 136, as described above.

[0162]

[0159] More specifically, generating the three-dimensional representation of the scene comprises inputting the semantically annotated occupancy grid (or the filtered occupancy grid) and the three-dimensional object perimeter (e.g. the projected, enhanced, or adjusted) to a machine learning model trained to generate the three-dimensional representation of the scene. The machine learning model may be the temporal alignment model 138.

[0163]

[0160] With reference to Figure 16, the perception model 30 may be used to track objects in a scene, even for edge cases where objects are protruding from a vehicle. This is similar to the scenario envisaged in relation to Figure 13, but the protruding objects are handled differently. The protruding objects may be for objects that move, and may include things like cargo protruding from a vehicle, e.g. logs protruding from a flatbed truck.

[0164]

[0161] For example, there may be a computer-implemented method of tracking an object from a scene in which an AV 10 operates, which may be summarised as including: receiving 502 a LiDAR point cloud, a RADAR point cloud, and images, from a LiDAR sensor, a RADAR sensor, and a camera, respectively, of the autonomous vehicle; generating 504 an occupancy grid using the LiDAR point cloud the RADAR point cloud, and the images; generating 506 a three-dimensional object perimeter of the object based on the LiDAR point cloud, the RADAR point cloud, and the images; modifying 508 the three-dimensional object perimeter to cover corresponding positions of cells of the occupancy grid that are associated with the object; and generating 510 a three-dimensional representation of the scene using the occupancy grid and the modified three-dimensional object perimeter.

[0165]

[0162] Modifying the three-dimensional object perimeter to cover corresponding positions of additional cells of the occupancy grid that are associated with the object comprises: semantically annotating cells of the occupancy grid with a label of the object when a position of the cell corresponds to the three-dimensional object perimeter. This may be achieved using the semantic annotation model 132 described above.

[0166]

[0163] Modifying the three-dimensional object perimeter to cover corresponding positions of additional cells of the occupancy grid that are associated with the object comprises: identifying additional cells in the occupancy grid that are not annotated as the object, and have a same movement as cells labelled as the object; and adjusting the three-dimensional object perimeter to cover the additional cells. A same movement of the cells labelled as the object includes a same speed of the object and a same heading as the object, identifying additional cells in the occupancy grid that are not annotated as the object, and have a same movement as cells labelled as the object comprises: determining a distance of the cells labelled as the object to each identified additional cell; discarding additional cells where the distance is greater than a cell distance threshold; and retaining additional cells where the distance is less than or equal to a cell distance threshold.

[0167]

[0164] Generating the three-dimensional representation of the scene using the occupancy grid and the modified three-dimensional object perimeter comprises inputting the annotated occupancy grid and the adjusted three-dimensional object perimeter to a machine learning model trained to generate a three-dimensional representation of the scene 40. The machine learning model may be the temporal alignment model 138. The annotated occupancy grid may be the filter occupancy grid.

[0168]

[0165] This is all described above with reference to Figures 11A through 110.

[0169]

[0166] Generating the three-dimensional object perimeter comprises: generating a three-dimensional projected object perimeter based on the LiDAR point cloud, the RADAR point cloud, and the images; generating a three-dimensional bounding box based on the LiDAR point cloud, the RADAR point cloud, and the images; and generating the three-dimensional object perimeter using the three-dimensional projected object perimeter and the three-dimensional bounding box. The generated three-dimensional object perimeter may thus be the joined, or enhanced, object perimeter described above. This may be achieved using the perimeter generation model 130.

[0167] Generating the three-dimensional object perimeter comprises adjusting a size and shape of the three-dimensional bounding box to cover the object in the three-dimensional projected object perimeter. The object perimeter may be a polygon surrounding the object and semantically labelled as the object.

[0170]

[0168] The generated occupancy grid includes a plurality of cells, each cell including at least one label. The at least one label includes a label of occupied or non-occupied, a label of occluded or non-occluded, a label of velocity, and / or a label of heading.

[0171]

[0169] With reference to Figure 17, another use case helps address an issue with missing cues. A cue may include a cue such as a hand gesture or signal from a driver of another vehicle in the scene. The hand signal may be to yield or give way to the AV 10. Such cues may be missed when using only object perimeters and occupancy grids, for example.

[0172]

[0170] To address this issue, there may be provided a computer-implemented method of tracking an object in a scene in which the AV 10 operates, the computer-implemented method comprising: receiving 602 a LiDAR point cloud, a RADAR point cloud, and images, from a LiDAR sensor, a RADAR sensor, and a camera, respectively; generating 604 an occupancy grid and a three-dimensional object perimeter using the LiDAR point cloud, the RADAR point cloud, and the images; extracting 606 features from the LiDAR point cloud, the RADAR point cloud, and the images; generating 608 an embedding using the extracted features; and generating 610 a three-dimensional representation of the scene including the object using the generated occupancy grid, the generated object perimeter, and the embedding.

[0173]

[0171] Generating the three-dimensional object perimeter comprises: generating a three-dimensional projected object perimeter based on the LiDAR point cloud, the RADAR point cloud, and the images; generating a three-dimensional bounding box based on the LiDAR point cloud, the RADAR point cloud, and the images; and generating the three-dimensional object perimeter using the three-dimensional projected object perimeter and the three-dimensional bounding box. The three-dimensional object perimeter may thus be the joined, or enhanced, three-dimensional object perimeter, and may be generated using the perimeter generation model 130.

[0174]

[0172] The three-dimensional object perimeter may be a polygon sounding the object. Generating the three-dimensional object perimeter comprises changing a shape and size of the three-dimensional bounding box to cover the object in the three-dimensional projected object perimeter.

[0173] The occupancy grid may be the fused occupancy grid, and may include a plurality of cells each having at least one label. The at least one label may include a label of occupied or non-occupied, a label of occluded or non-occluded, a label of velocity, and a label of heading.

[0175]

[0174] Extracting features comprises extracting features from each of the LiDAR point cloud, the RADAR point cloud, and the images, and fusing the respective extracted features. This may be done using the object detection model 36, as described above.

[0176]

[0175] The method may further comprise transforming the extracted features from each of the LiDAR point cloud, the RADAR point cloud, and the images, into a birds’ eye view. Fusing the respective extracted features comprises fusing the birds’ eye view representations of the respective extracted features. This may all be done using the encoders described above, and the fused features will be the embedding.

[0177]

[0176] The method also comprises semantically annotating the occupancy grid using a semantic label associated with the three-dimensional object perimeter. This may be achieved using the semantic annotation model 132.

[0178]

[0177] Generating the embedding comprises: inputting the extracted features to a machine learning model trained to generate an embedding. The machine learning model may be the fusion encoder described above. In other words, the extracted features may be in the form of encodings, which may be fused to form the fused encoding, or embedding.

[0179]

[0178] Generating the three-dimensional representation of the scene may be performed using the state estimate prediction model 136 and the temporal alignment model 138, as described above.

[0180]

[0179] With reference to Figure 18, the perception model 30 may be used to address issues associated with temporary occlusions. For example, a pedestrian becomes temporarily occluded by a bus. When the pedestrian reappears once the bus moves on, there will be no prediction of the reappearance if using only the occupancy grid and object perimeter.

[0181]

[0180] To address this issue, there may be provided a computer-implemented method of tracking an object in a scene in which the AV 10 operates, where the computer-implemented method may be summarised as including, for each time step over a plurality of time steps: receiving 702 a LiDAR point cloud, a RADAR point cloud, and images, from a LiDAR sensor, a RADAR sensor, and a camera, respectively; generating 704 anoccupancy grid and a three-dimensional object perimeter using the LiDAR point cloud, the RADAR point cloud, and the images; extracting 706 features from the LiDAR point cloud, the RADAR point cloud, and the images; predicting 708 an occupancy grid and a three-dimensional object perimeter surrounding the object for a current time step using the occupancy grid and the three-dimensional object perimeter of a previous time step; and generating 710 a three-dimensional representation of the scene including the object using the predicted occupancy grid and the predicted three-dimensional object perimeter of the current time step and a predicted occupancy grid and a three-dimensional point cloud of the current time step.

[0182]

[0181] Generating the three-dimensional object perimeter comprises: generating a three-dimensional projected object perimeter based on the LiDAR point cloud, the RADAR point cloud, and the images; generating a three-dimensional bounding box based on the LiDAR point cloud, the RADAR point cloud, and the images; and generating the three-dimensional object perimeter (joined, or enhanced) using the three-dimensional projected object perimeter and the three-dimensional bounding box. This is described in relation to Figures 10A to 10C.

[0183]

[0182] The three-dimensional object perimeter is a polygon sounding the object. Generating the three-dimensional object perimeter comprises changing a shape and size of the three-dimensional bounding box to cover the object in the three-dimensional projected object perimeter.

[0184]

[0183] The occupancy grid includes a plurality of cells each having at least one label. The at least one label may include a label denoting occupied, or non-occupied, a label denoting occluded, or non-occluded, a label denoting velocity, and / or a label denoting heading.

[0185]

[0184] Extracting features comprises extracting features from each of the LiDAR point cloud, the RADAR point cloud, and the images, and fusing the respective extracted features. The method may further comprise transforming the extracted features from each of the LiDAR point cloud, the RADAR point cloud, and the images, into a birds’ eye view. Fusing the respective extracted features comprises fusing the birds’ eye view representations of the respective extracted features. The fused extracted features may correspond to the embedding described above.

[0186]

[0185] The method may further comprise semantically annotating the occupancy grid using a semantic label associated with the three-dimensional object perimeter.

[0186] Predicting an occupancy grid and a three-dimensional object perimeter surrounding the object for a current time step comprises inputting the three-dimensional object perimeter and the semantically annotated occupancy grid to a machine learning model trained to predict the occupancy grid and the three-dimensional object perimeter surrounding the object. The inputs may be from a previous time step. This may be achieved using the state estimate prediction model 136.

[0187]

[0187] Generating a three-dimensional representation of the scene 40 including the object comprises: inputting the semantically annotated occupancy grid and the three-dimensional object perimeter, and the predicted occupancy grid and the predicted three-dimensional object perimeter to a machine learning model trained to generate the three-dimensional representation of the scene. This may be achieved using the temporal alignment model 138.

[0188]

[0188] With reference to Figure 19, with any of these use cases, there may also be provided a computer-implemented method of moving the AV 10 which may be summarised as including: tracking 802 an object in the scene using the computer-implemented method of any of the foregoing scenarios; generating 804 a trajectory for the autonomous vehicle using the tracked object; and moving 806 the autonomous vehicle along the trajectory. Generating the trajectory may take account of the tracked objects in the three-dimensional scene 40, to reduce the risk of collisions with those objects. Moving the trajectory may involve generating control signals to control the actuators of the AV 10 to move the AV 10 along the trajectory.

[0189]

[0189] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments.

[0190]

[0190] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfil the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measured cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.

Claims

CLAIMS1. A computer-implemented method of tracking an object from a scene in which an autonomous vehicle operates, the computer-implemented method comprising:receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle;estimating a ground surface using the depth map and the LiDAR point cloud;filtering the LiDAR point cloud to discard points of the LiDAR point cloud below the ground surface and retain points of the LiDAR point cloud at or above the ground surface; andgenerating an occupancy grid for tracking any objects in the scene using the filtered LiDAR point cloud.

2. The computer-implemented method of Claim 1 , wherein the sensors include a LiDAR sensor, a RADAR sensor, and an image sensor, wherein receiving a LiDAR point cloud and a depth map, each generated using sensor data from sensors of the autonomous vehicle, comprises:receiving, from the LiDAR sensor, the RADAR sensor, and the image sensor a LiDAR point cloud, a RADAR point cloud, and images, respectively; generating the depth map using and the LiDAR point cloud, the RADAR point cloud and the images.

3. The computer-implemented method of Claim 2, wherein generating a depth map using and the LiDAR point cloud, the RADAR point cloud and the images comprises:inputting the LiDAR point cloud, the RADAR point cloud, and the images to a machine learning model trained to generate a depth map.

4. The computer-implemented method of Claim 2 or Claim 3, further comprising:optimising depth values of points of the depth map using points from each of the LiDAR point cloud and the RADAR point cloud.

5. The computer-implemented method of Claim 4, wherein verifying depth values of points of the depth map includes, for each point of the depth map: comparing a distance to points of the LiDAR and RADAR point clouds; andignoring any points of the LiDAR and RADAR point clouds that are over a threshold distance to the depth map point; andoptimising the depth map point using at least one point from the LiDAR and RADAR point clouds that is less than or equal to the threshold distance to the depth map point.

6. The computer-implemented method of Claim 5, wherein verifying the depth map point using at least one point from the LiDAR and RADAR point clouds comprises:adjusting the depth value of the depth map point to match a depth value of the at least one point from the LiDAR and RADAR point clouds.

7. The computer-implemented method of any preceding claim, wherein estimating a ground surface using the depth map and the LiDAR point cloud comprises:inputting the depth map and the LiDAR point cloud to a machine learning model trained to estimate a ground surface.

8. The computer-implemented method of any preceding claim, wherein the occupancy grid comprises a plurality of cells, each cell including at least one label, wherein the at least one label is selected from a list of labels including a label of occupied or non-occupied, a label of occluded or non-occluded, a label of velocity, and a label of heading.

9. The computer-implemented method of any preceding claim, wherein the occupancy grid is a first occupancy grid, wherein the computer-implemented method further comprising:generating a second occupancy grid based on the RADAR point cloud;generating a third occupancy grid based on the depth map; and fusing the first, second, and third occupancy grids.

10. The computer-implemented method of Claim 9, further comprising: generating a three-dimensional object perimeter for the object based on the LiDAR point cloud, the RADAR point cloud, and the images; and semantically annotating the fused occupancy grid using the three-dimensional object perimeter.

11. The computer-implemented method of Claim 10, further comprising: generating a three-dimensional representation of the scene using the semantically annotated occupancy grid and the three-dimensional object perimeter.

12. The computer-implemented method of Claim 11, wherein generating the three-dimensional representation of the scene comprises inputting the semantically annotated occupancy grid and the three-dimensional object perimeter to a machine learning model trained to generate the three-dimensional representation of the scene.

13. A computer-implemented method of moving an autonomous vehicle, the computer-implemented method comprising:tracking an object in a scene in which the autonomous vehicle operates using the computer-implemented of any preceding claim;generating a trajectory for the autonomous vehicle based on the occupancy grid; andmoving the autonomous vehicle along the trajectory.

14. A transitory, or non-transitory, computer readable medium, having instructions stored thereon that when executed by at least one processor, causes the at least one processor to perform the computer-implemented method of any preceding claim.

15. An autonomous vehicle including:a LiDAR sensor;a RADAR sensor;a camera;at least one processor; andstorage having instructions stored thereon that when executed by the at least one processor, causes the at least one processor to perform the computer-implemented method of any of Claims 1 to 13.