Machine learning-based placeholder grid generation

Through the combination of machine learning models and loss functions, the sensor data is aggregated to generate a placeholder grid, which solves the problem of inaccuracy of placeholder grids caused by sensor limitations and improves the safety and user experience of the autonomous driving system.

CN120303583APending Publication Date: 2025-07-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084779.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-19
Filing Date
2023-10-31
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When generating a placeholder grid, the prior art is limited by various limitations of the sensor, resulting in insufficient accuracy and reliability of the placeholder grid, which may lead to vehicles misjudging road drivingability, increasing collision risk and poor user experience.

Method used

Using machine learning model combined with loss function, aggregated frames are generated by aggregating sensor data, and only the loss of known placeholder states is calculated. The machine learning model is trained to predict the probability of placeholder states and improve the accuracy of the placeholder raster.

Benefits of technology

It improves the accuracy and reliability of the placeholder grid, reduces misjudgment, reduces collision risks, and improves the safety and user experience of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303583A_ABST
    Figure CN120303583A_ABST
Patent Text Reader

Abstract

In some aspects, a device may receive sensor data associated with a vehicle and a set of frames. The device may aggregate sensor data associated with the set of frames using the first gesture to generate an aggregated frame, where the aggregated frame is associated with the set of cells. The device may obtain an indication of a respective placeholder flag from each cell of the set of cells, where the respective placeholder flag includes a first placeholder flag or a second placeholder flag, and where the set of cells from the set of cells is associated with the first placeholder flag. The device may train a machine learning model using data associated with the aggregated frame to generate a placeholder grid based on a loss function that calculates only losses from respective cells of the set of cells. Numerous other aspects are described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This patent application claims priority to U.S. Patent Application No. 18 / 067,798, titled "MACHINE LEARNING BASED OCCUPANCY GRID GENERATION", filed on December 19, 2022, and assigned to its assignee. The disclosure of the prior application is considered part of this patent application and is incorporated herein by reference. Technical Field

[0003] Aspects of the present disclosure generally relate to occupancy grid generation, and for example, to machine - learning - based occupancy grid generation. Background Art

[0004] Occupancy grid mapping can be used for road - scene understanding in autonomous driving. Occupancy grid mapping can summarize information about drivable areas and road obstacles in the environment traversed by an autonomous vehicle. Summary of the Invention

[0005] Some aspects described herein relate to a device. The device can include one or more memories and one or more processors coupled to the one or more memories. The one or more processors can be configured to receive sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The one or more processors can be configured to aggregate the one or more sensor detections into an aggregated frame using a first pose, where the aggregated frame is associated with a set of cells. The one or more processors can be configured to obtain an indication of a corresponding occupancy label for each cell from the set of cells, where the corresponding occupancy label includes a first occupancy label indicating a known occupancy state or a second occupancy label indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy label. The one or more processors can be configured to use data associated with the aggregated frame to train a machine - learning model to generate an occupancy grid, where training the machine - learning model is associated with a loss function that calculates the loss for corresponding cells from the subset of cells, and where the machine - learning model is trained to predict the probability of the occupancy state of corresponding cells from the set of cells. The one or more processors can be configured to provide the machine - learning model to another device.

[0006] Some aspects described herein relate to a device. The device may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to receive sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The one or more processors may be configured to aggregate the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with the one or more sensor detections. The one or more processors may be configured to generate an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells. The one or more processors may be configured to perform an action based on the occupancy grid.

[0007] Some aspects described herein relate to a method. The method may include: receiving, by a device, sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The method may include aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame is associated with a set of cells. The method may include obtaining, by the device, an indication of a corresponding occupancy label for each cell from the set of cells, where the corresponding occupancy label includes a first occupancy label indicating a known occupancy state or a second occupancy label indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy label. The method may include training, using data associated with the aggregated frame, a machine learning model to generate an occupancy grid, where training the machine learning model is associated with a loss function that calculates the loss for corresponding cells from the subset of cells, and where the machine learning model is trained to predict probabilities of occupancy states for corresponding cells from the set of cells. The method may include providing, by the device, the machine learning model to another device.

[0008] Some aspects described herein relate to a method. The method may include: receiving, by a device, sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The method may include aggregating, using a first pose, the sensor data associated with the set of frames to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections. The method may include generating, using a machine learning model, an occupancy grid, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells. The method may include performing, by the device, an action based on the occupancy grid.

[0009] Some aspects described herein relate to a non-transitory computer-readable medium storing a set of instructions. The set of instructions, when executed by one or more processors of a device, may cause the device to receive sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The set of instructions, when executed by one or more processors of the device, may cause the device to aggregate, using a first pose, the sensor data associated with the set of frames to generate an aggregated frame, where the aggregated frame is associated with a set of cells. The set of instructions, when executed by one or more processors of the device, may cause the device to obtain an indication of a corresponding occupancy label for each cell from the set of cells, where the corresponding occupancy label includes a first occupancy label indicating a known occupancy state or a second occupancy label indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy label. The set of instructions, when executed by one or more processors of the device, may cause the device to train, using data associated with the aggregated frame, a machine learning model to generate an occupancy grid, where training the machine learning model is associated with a loss function that calculates the loss for corresponding cells from the subset of cells, and where the machine learning model is trained to predict probabilities of occupancy states for corresponding cells from the set of cells. The set of instructions, when executed by one or more processors of the device, may cause the device to provide the machine learning model to another device.

[0010] Some aspects described herein relate to a non-transitory computer-readable medium storing a set of instructions. The set of instructions, when executed by one or more processors of a device, can cause the device to receive sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The set of instructions, when executed by one or more processors of the device, can cause the device to aggregate the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections. The set of instructions, when executed by one or more processors of the device, can cause the device to generate an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells. The set of instructions, when executed by one or more processors of the device, can cause the device to perform an action based on the occupancy grid.

[0011] Some aspects described herein relate to a device. The device can include means for receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The device can include means for aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame is associated with a set of cells. The device can include means for obtaining an indication of a corresponding occupancy label for each cell from the set of cells, where the corresponding occupancy label includes a first occupancy label indicating a known occupancy state or a second occupancy label indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy label. The device can include means for training a machine learning model to generate an occupancy grid using data associated with the aggregated frame, where training the machine learning model is associated with a loss function that calculates the loss for corresponding cells from the subset of cells, and where the machine learning model is trained to predict probabilities of occupancy states for corresponding cells from the set of cells. The device can include means for providing the machine learning model to another device.

[0012] Some aspects described herein relate to an apparatus. The apparatus can include components for receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections. The apparatus can include components for aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections. The apparatus can include components for generating an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells. The apparatus can include components for performing an action based on the occupancy grid.

[0013] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer-readable media, user equipment, user gear, wireless communication devices, and / or processing systems substantially as described and illustrated in connection with the accompanying figures and the specification.

[0014] The features and technical advantages of examples in accordance with the present disclosure have been outlined rather broadly above in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. When considered in conjunction with the accompanying figures, the characteristics (both its organization and method of operation) of the concepts disclosed herein, as well as the associated advantages, will be better understood. Each of the figures is provided for the purpose of illustration and description, and not as a definition of the limits of the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To enable a more particular understanding of the above-described features of the present disclosure, a more specific description may be obtained by reference to the aspects, some of which are illustrated in the accompanying figures. It should be noted, however, that the figures illustrate only certain typical aspects of the present disclosure and should not be considered limiting of its scope, as the description may admit to other equally effective aspects. Like reference numerals in the different figures may identify the same or similar elements.

[0016] Figure 1 is a diagram of an example environment in which systems and / or methods described herein can be implemented in accordance with the present disclosure.

[0017] Figure 2 is a diagram illustrating example components of a device in accordance with the present disclosure.

[0018] Figures 3A to 3C is a diagram showing an example associated with machine learning-based occupancy grid generation according to the present disclosure.

[0019] Figure 4 is a flowchart of an example process associated with machine learning-based occupancy grid generation according to the present disclosure.

[0020] Figure 5 is a flowchart of an example process associated with machine learning-based occupancy grid generation according to the present disclosure. Detailed Description

[0021] Aspects of the present disclosure are described more fully hereinafter with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout this disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art. Those skilled in the art should understand that the scope of the present disclosure is intended to cover any aspect of the present disclosure disclosed herein, whether implemented independently of any other aspect of the present disclosure or in combination with any other aspect of the present disclosure. For example, any number of the aspects set forth herein may be used to implement a device or practice a method. Additionally, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functions, or a combination of structures and functions in addition to or different from the aspects of the present disclosure set forth herein. It should be understood that any aspect of the present disclosure disclosed herein may be embodied by one or more elements of the claims.

[0022] A vehicle may include a system configured to control the operation of the vehicle (e.g., an electronic control unit (ECU) and / or an autonomous driving system). The system may use data obtained by one or more sensors of the vehicle to perform occupancy mapping to determine the occupancy state of the environment around the vehicle (e.g., unoccupied space, occupied space, and / or drivable space). For example, the system may use data obtained by a global navigation satellite system (GNSS) / inertial measurement unit (IMU), a camera, a light detection and ranging (LIDAR) scanner, and / or a radar scanner, among other examples, to determine the occupancy state of the environment around the vehicle. The system may detect the drivable space that the vehicle can occupy based on the occupancy state of the environment around the vehicle. The system may be configured to identify the occupancy state of the environment around the vehicle in real time and determine the drivable space that the vehicle can occupy based on the occupancy state of the environment. To perform occupancy and free space detection when using sensors configured to obtain point data of an object (e.g., a radar sensor, a LIDAR sensor, and / or a camera), the system may divide a region of interest (e.g., the area around the vehicle) into a plurality of evenly spaced square grids (e.g., occupancy grids). Based on radar returns, the occupancy state of each grid is determined to generate an occupancy grid. The occupancy grid may include a static occupancy grid (e.g., associated with relatively static or non-moving objects in the environment around the vehicle) and / or a dynamic occupancy grid (e.g., associated with dynamic or moving objects in the environment around the vehicle). However, the system may not account for various limitations of one or more sensors that may have a negative impact on the system's ability to detect drivable space.

[0023] For example, the GNSS / IMU may provide data indicating the position of the vehicle in the environment. The system may couple the data obtained by the GNSS / IMU with a high-resolution map to determine the exact position of the vehicle on the map and may use the map to estimate the occupancy state of the environment around the vehicle and / or estimate the drivable space within the environment. However, the map may not include information associated with the most recent changes to the environment. For example, the map may not include information associated with construction being performed on the road, other vehicles traveling along the road, and / or objects, people, and / or animals located on or adjacent to the road, among other examples.

[0024] The camera can obtain images of the vehicle's surrounding environment. The system can perform object detection to identify objects within the images and can determine the occupancy status of the vehicle's surrounding environment at least in part based on the detected objects within the images. However, the camera can be a two-dimensional sensor that by itself cannot measure the distance of an object from the vehicle. Instead, the system and / or the camera can use one or more algorithms to estimate the distance of the objects depicted in the images from the vehicle. Since the distance is estimated rather than measured, the estimation of the speed of the objects can be prone to errors and noise. Additionally, the camera can be sensitive to the environment in which the camera is operating, and environmental conditions such as rain, fog, and / or snow, among other examples, can affect the quality of the images captured by the camera.

[0025] When the LIDAR scanner rotates, the LIDAR scanner can use light in the form of pulsed lasers to obtain point data. The point data can correspond to the reflection of the light from the objects and can be used to perform three-dimensional (3D) object detection and determine the speed of the objects. However, radiation safety requirements may limit the amount of energy emitted by the LIDAR scanner. The limitation on the amount of energy that the LIDAR scanner can emit can cause the LIDAR scanner to use a scanning scheme that focuses all of the energy that will be emitted by the LIDAR scanner in a limited number of directions (e.g., rotation of the laser head and / or rotation of a current mirror). The use of the scanning scheme can make the speed measurement prone to errors caused by the trailing of the LIDAR signals across the individual segments of the objects (e.g., due to scanning). Additionally, the LIDAR scanner can be sensitive to the environment in which the LIDAR scanner is operating, and environmental conditions such as rain, fog, and / or snow, among other examples, can affect the quality of the point data obtained by the LIDAR scanner.

[0026] The radar scanner can emit one or more pulses of electromagnetic waves. The one or more pulses can be reflected by an object in the path of the one or more pulses. The reflection can be received by the radar scanner. The radar scanner can determine one or more characteristics associated with the reflected pulses (e.g., amplitude and / or frequency) and can determine point data indicative of the position of the object based on the one or more characteristics. However, the radar scanner can be associated with poor angular resolution, sparse detection, and / or unreliable behavior, among other examples, which can reduce the effectiveness of the object detection data collected by the radar scanner for occupancy grid generation.

[0027] In some cases, the system can generate an occupancy grid based on acquired sensor data (e.g., camera data, LIDAR data, and / or radar data). The system can analyze the sensor data and use a deterministic method (e.g., using a deterministic algorithm such as a Bayesian-derived algorithm (e.g., Bayesian algorithm)) to generate the occupancy grid. However, due to various limitations of one or more sensors used to acquire the sensor data (e.g., as described above), using a deterministic method to generate the occupancy grid may result in an inaccurate occupancy grid. For example, the system can determine the occupancy status of a cell in the occupancy grid based on cell-specific information (e.g., sensor data). For example, information from one cell may not affect the occupancy status of another cell. Thus, noisy or inaccurate sensor data (or sparse sensor data) may lead to an incorrect determination of the occupancy status of a cell in the occupancy grid.

[0028] As an example, the system can detect a bridge (e.g., overpass) above the road on which the vehicle is traveling. Because the system uses a deterministic method and cell-specific information, the system may determine that one or more cells in the occupancy grid in which the bridge is detected are occupied (e.g., non-drivable due to sensor detection). However, the road under the bridge may be drivable. However, because the system determines that the occupancy status of one or more cells in the occupancy grid in which the bridge is detected is occupied (e.g., non-drivable due to sensor detection), the system may incorrectly determine that the road under the bridge is non-drivable. This may cause the system to perform an action to avoid a passable or drivable road, resulting in an increased risk of collision (e.g., caused by the vehicle performing unnecessary and / or unexpected maneuvers to avoid a passable or drivable road) and / or a poor user experience.

[0029] Some aspects described herein implement machine learning-based occupancy grid generation. For example, a system (e.g., of a vehicle) can use acquired sensor data and a machine learning model to generate an occupancy grid (e.g., where the sensor data is provided as an input to the machine learning model). In some aspects, the machine learning model can output the probability of the corresponding occupancy status class for each cell of the occupancy grid. The system can determine the class associated with the occupancy status of each cell in the occupancy grid (e.g., based on the output of the machine learning model).

[0030] In some aspects, the system can use aggregated frames (e.g., aggregated data) as input to a machine learning model. For example, as described elsewhere herein, sensor detections can be sparse over time (e.g., for radar sensor detections). Thus, to increase the data available for analysis by the machine learning model, the system can aggregate one or more sensor detections into an aggregated frame. For example, the environment around a vehicle can be divided into 2D grids (e.g., frames). The system can aggregate sensor data (e.g., obtained over time) into the frames to generate aggregated frames. The aggregated frames can be used to train the machine learning model and / or can be used as input to the machine learning model.

[0031] In some aspects, one or more loss functions can be used to train the machine learning model. The loss function can compute the loss for only the cells of the aggregated frame associated with known occupancy states. For example, the training data for the machine learning model can include one or more aggregated frames. Each aggregated frame can be associated with one or more cells labeled with a known occupancy state (e.g., indicating that the one or more cells are associated with an occupancy state of a given class) and one or more cells labeled with an unknown occupancy state (e.g., indicating that the occupancy state of the one or more cells is unknown). The system or device can train the machine learning model by using a loss function that computes the loss for the one or more cells labeled with a known occupancy state. The system or device can ignore (or suppress the computation of the loss for) the one or more cells labeled with an unknown occupancy state. In some aspects, the system or device can apply a penalty weight to one or more cells associated with an incorrect occupancy state determination and / or an incorrect class determination made by the machine learning model. In some aspects, the system or device can update one or more weights of the machine learning model based on the computed loss.

[0032] Thus, the system or device can use the machine learning model to generate an occupancy grid (e.g., a static occupancy grid), where the machine learning model enables the determination of the occupancy state of each cell using information associated with each cell and information associated with other cells of the occupancy grid. This can improve the determination of the occupancy state and / or the class of the occupancy state of each cell of the occupancy grid. For example, the machine learning model can be configured or trained to output the probability of the occupancy state and / or class of each cell of the occupancy grid (e.g., regardless of whether the cell is associated with sensor data obtained as input to the machine learning model). This improves the accuracy of the occupancy grid because the determination of the occupancy state of each cell takes into account information associated with other cells of the occupancy grid.

[0033] Additionally, the accuracy of occupancy status determination can be improved by aggregating sensor data into an aggregation frame. For example, radar data collected by a vehicle may be sparse in time. Thus, aggregating sensor data (e.g., radar data) into an aggregation frame that includes sensor data collected over multiple time intervals results in additional available data being analyzed by a machine learning model. This can improve the training and / or performance of the machine learning model. Further, using a loss function that only computes the loss for cells marked with a known occupancy status improves the accuracy of training of the machine learning model. For example, a marker indicating a known occupancy status (e.g., a ground truth marker indicating a known occupancy status) may have improved reliability compared to a marker indicating an unknown occupancy status (e.g., a missing ground truth marker). In other words, the presence of a positive ground truth marker (e.g., such as a marker indicating a known occupancy status) can be more reliable than the absence of a positive ground truth marker (e.g., such as a marker indicating an unknown occupancy status). Thus, a marker of a known occupancy status can be more reliable than a marker of an unknown occupancy status. Therefore, only computing the loss for cells associated with known occupancy status markers can improve the reliability and / or accuracy of the computed loss and / or the training of the machine learning model.

[0034] Figure 1 is a diagram of an example environment 100 in which the systems and / or methods described herein can be implemented according to the present disclosure. As Figure 1 shown, the environment 100 can include a vehicle 110, and the vehicle 110 includes an ECU 112, a wireless communication device 120, a server device 130, and a network 140. The devices of the environment 100 can be interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0035] The vehicle 110 can include any vehicle capable of sending and / or receiving data associated with machine learning-based occupancy grid generation based on camera data, radar data, and / or LIDAR data, and other examples, as described herein. For example, the vehicle 110 can be a consumer vehicle, an industrial vehicle, and / or a commercial vehicle, and other examples. The vehicle 110 can be capable of traveling on public roads and / or providing transportation, and / or can be capable of being used in operations associated with a worksite (e.g., a construction worksite), and other examples. The vehicle 110 can include a sensor system, where the sensor system includes one or more sensors for generating and / or providing vehicle data associated with the vehicle 110 and / or a radar scanner and / or a LIDAR scanner for obtaining point data for road scene understanding in autonomous driving.

[0036] Vehicle 110 can be controlled by ECU 112, which can include one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with occupancy clustering based on point data (e.g., data obtained by a radar scanner, LIDAR scanner, and / or camera) and / or road scene understanding described herein. For example, ECU 112 can be associated with an autonomous driving system and / or can include a communication and / or computing device and / or be a component of a communication and / or computing device, such as an in-vehicle computer, console, operator station, or similar type of device. ECU 112 can be configured to communicate with the autonomous driving system of vehicle 110, ECUs of other vehicles, and / or other devices. For example, advancements in communication technologies have enabled vehicle-to-everything (V2X) communication, which can include vehicle-to-vehicle (V2V) communication and / or vehicle-to-pedestrian (V2P) communication and other examples. In some aspects, ECU 112 can receive vehicle data associated with vehicle 110 (e.g., location information, sensor data, radar data, and / or LIDAR data) and perform machine learning-based occupancy grid generation to determine the occupancy status of the environment around vehicle 110, and determine the drivable space that the vehicle can occupy based on the occupancy status of the environment according to the vehicle data, as described herein.

[0037] Wireless communication device 120 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with machine learning-based occupancy grid generation, as described elsewhere herein. For example, wireless communication device 120 can include a base station and / or access point and other examples. Additionally or alternatively, wireless communication device 120 can include a communication and / or computing device, such as a mobile phone (e.g., a smartphone, and / or a wireless phone), a laptop computer, a tablet computer, a handheld computer, a desktop computer, a gaming device, a wearable communication device (e.g., a smartwatch, and / or a pair of smart glasses), and / or similar types of devices.

[0038] Server device 130 can include one or more devices capable of receiving, generating, storing, processing, providing, and / or routing information associated with machine learning-based occupancy grid generation, as described elsewhere herein. Server device 130 can include a communication device and / or a computing device. For example, server device 130 can include a server, such as an application server, a client server, a web server, a database server, a host server, a proxy server, a virtual server (e.g., executed on computing hardware), or a server in a cloud computing system. In some aspects, server device 130 can include computing hardware used in a cloud computing environment. In some aspects, server device 130 can include one or more devices capable of training a machine learning model associated with occupancy grid generation, as described in more detail elsewhere herein.

[0039] Network 140 includes one or more wired and / or wireless networks. For example, network 140 may include a peer-to-peer (P2P) network, a cellular network (e.g., a Long-Term Evolution (LTE) network, a Code Division Multiple Access (CDMA) network, an Open Radio Access Network (O-RAN), a New Radio (NR) network, a 3G network, a 4G network, a 5G network, or another type of next-generation network), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad-hoc network, an intranet, the Internet, a fiber-based network, and / or a cloud computing network, among other examples, and / or a combination of these or other types of networks. In some aspects, network 140 may include and / or be a P2P communication link directly between one or more devices of environment 100.

[0040] Figure 1 The number and arrangement of the devices and networks shown are provided as examples. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks compared to those shown Figure 1 in. Additionally, Figure 1 two or more of the devices shown may be implemented within a single device, or Figure 1 a single device shown may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices of environment 100 (e.g., one or more devices) may perform one or more functions described as being performed by another set of devices of environment 100.

[0041] Figure 2 is a diagram showing example components of device 200 according to the present disclosure. Device 200 may correspond to vehicle 110, ECU 112, wireless communication device 120, and / or server device 130. In some aspects, vehicle 110, ECU 112, wireless communication device 120, and / or server device 130 may include one or more device 200 and / or one or more components of device 200. As Figure 2 shown, device 200 may include a bus 205, a processor 210, a memory 215, a storage component 220, an input component 225, an output component 230, a communication interface 235, one or more sensors 240, a radar scanner 245, and / or a LIDAR scanner 250.

[0042] The bus 205 includes components that permit communication among the components of the device 200. The processor 210 is implemented in hardware, firmware, or a combination of hardware and software. The processor 210 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some aspects, the processor 210 includes one or more processors that can be programmed to perform functions. The memory 215 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 210.

[0043] The storage component 220 stores information and / or software related to the operation and use of the device 200. For example, the storage component 220 can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-transitory computer-readable medium, as well as a corresponding drive.

[0044] The input component 225 includes components that permit the device 200 to receive information such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally or alternatively, the input component 225 can include components for determining the location or position of the device 200 (e.g., a global positioning system (GPS) component or a GNSS component) and / or sensors for sensing information (e.g., an accelerometer, a gyroscope, an actuator, or another type of position or environmental sensor). The output component 230 includes components that provide output information from the device 200 (e.g., a display, a speaker, a haptic feedback component, and / or an audio or visual indicator).

[0045] The communication interface 235 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 200 to communicate with other devices such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 235 can permit the device 200 to receive information from and / or provide information to another device. For example, the communication interface 235 can include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency interface, a universal serial bus (USB) interface, a wireless local area interface (e.g., a Wi-Fi interface), and / or a cellular network interface.

[0046] One or more sensors 240 may include one or more devices capable of sensing characteristics associated with the device 200. The sensors 240 may include one or more integrated circuits (e.g., on a packaged silicon die) and / or one or more passive components of an integrated (flex) circuit to enable communication with one or more components of the device 200. The sensors 240 may include an optical sensor having a field of view in which the sensors 240 may determine one or more characteristics of the environment of the device 200. In some aspects, the sensors 240 may include a camera. For example, the sensors 240 may include a low-resolution camera (e.g., Video Graphics Array (VGA)) capable of capturing images of less than one million pixels, images of less than 1216×912 pixels, and other examples. The sensors 240 may be low-power devices (e.g., devices consuming less than ten milliwatts (mW) of power) that have always-on capabilities while the device 200 is powered on. Additionally or alternatively, the sensors 240 may include magnetometers (e.g., Hall effect sensors, anisotropic magnetoresistive (AMR) sensors, and / or giant magnetoresistive (GMR) sensors), position sensors (e.g., GPS receivers and / or local positioning system (LPS) devices (e.g., using triangulation and / or multilateration)), gyroscopes (e.g., microelectromechanical systems (MEEMS) gyroscopes or similar types of devices), accelerometers, speed sensors, motion sensors, infrared sensors, temperature sensors, and / or pressure sensors, and other examples.

[0047] The radar scanner 245 may include one or more devices that use radio waves to determine the range, angle, and / or velocity of an object based on radar data obtained by the radar scanner 245. The radar scanner 245 may provide the radar data to the ECU 112 such that the ECU 112 can perform machine learning-based occupancy grid generation based on the radar data, as described herein.

[0048] The LIDAR scanner 250 may include one or more devices that use light in the form of pulsed lasers to measure the distance of an object from the LIDAR scanner based on LIDAR data obtained by the LIDAR scanner 250. The LIDAR scanner 250 may provide the LIDAR data to the ECU 112 such that the ECU 112 can perform machine learning-based occupancy grid generation based on the LIDAR data, as described herein.

[0049] Device 200 may perform one or more of the processes described herein. Device 200 may perform these processes based on software instructions stored by a non-transitory computer-readable medium (e.g., memory 215 and / or storage component 220) and executed by processor 210. A computer-readable medium is defined herein as a non-transitory memory device. A memory device includes memory space within a single physical storage device or memory space distributed across multiple physical storage devices.

[0050] The software instructions may be read into memory 215 and / or storage component 220 from another computer-readable medium or from another device via communication interface 235. When executed, the software instructions stored in memory 215 and / or storage component 220 may cause processor 210 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more of the processes described herein. Accordingly, aspects described herein are not limited to any particular combination of hardware circuitry and software.

[0051] In some aspects, device 200 includes components for performing one or more of the processes described herein and / or for performing one or more operations of the processes described herein. For example, device 200 may include: components for receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections; components for aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame is associated with a set of cells; components for obtaining an indication of corresponding occupancy tokens for each cell from the set of cells, where the corresponding occupancy tokens include a first occupancy token indicating a known occupancy state or a second occupancy token indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy token; components for training a machine learning model using data associated with the aggregated frame to generate an occupancy grid, where training the machine learning model is associated with a loss function that calculates the loss for corresponding cells from the subset of cells, and where the machine learning model is trained to predict the probability of the occupancy state of corresponding cells from the set of cells; and / or components for providing the machine learning model to another device and other examples. Additionally or alternatively, device 200 may include: components for receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections; components for aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections; components for generating an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss only using cells associated with known occupancy states, and where the output of the machine learning model includes the probability of one or more occupancy states for corresponding cells from the set of cells; and / or components for performing an action based on the occupancy grid and other examples. In some aspects, such components may include one or more components of device 200 described in conjunction with Figure 2 as described, such as bus 205, processor 210, memory 215, storage component 220, input component 225, output component 230, communication interface 235, one or more sensors 240, radar scanner 245, and / or LIDAR scanner 250.

[0052] Figure 2 The number and arrangement of the components shown in Figure 2fewer components, different components, or components in a different arrangement than the components shown. Additionally or alternatively, a set of components (e.g., one or more components) of device 200 may perform one or more functions described as being performed by another set of components of device 200.

[0053] Figures 3A to 3C is a diagram illustrating example 300 associated with machine learning-based occupancy grid generation in accordance with the present disclosure. As Figures 3A to 3C shown, example 300 includes server device 130, vehicle 110, and ECU 112. Server device 130 may include a machine learning (ML) model trainer 325. Vehicle 110 and / or ECU 112 may include an occupancy grid generator 355. ML model trainer 325 and occupancy grid generator 355 are described in more detail elsewhere herein.

[0054] Figure 3A and Figure 3B depicts an example associated with training a machine learning model (e.g., ML model 330) to predict an occupancy state of a corresponding cell associated with an occupancy grid. For example, ML model 330 may be trained to output a probability indicating the likelihood that a given cell is associated with a given class or occupancy state. For example, classes of occupancy states may include a drivable occupancy state or a non-drivable occupancy state and other examples. Additionally or alternatively, classes of occupancy states may include a road class and a "non-road" class (e.g., associated with cells including objects that are not roads (e.g., not drivable areas)). The "non-road" class may also be referred to as a "prop" class. As another example, classes of occupancy states may include a background class (e.g., associated with background objects), a parking class (e.g., associated with areas for parking rather than driving), a railroad or railway track class, a sidewalk class, a bike lane class, a lane marking class, a pedestrian class, and / or a crosswalk or sidewalk class and other examples.

[0055] In some aspects, as shown in example 300, ML model 330 may be trained by server device 130. For example, ML model 330 may be "offline" trained by server device 130 before being provided to vehicle 110 and / or ECU 112. In some other aspects, ML model 330 may be trained by a device associated with vehicle 110 (such as ECU 112 or a controller / processor of vehicle 110) in a manner similar to that described herein.

[0056] As Figure 3AAs shown, the server device 110 can obtain sensor data associated with one or more vehicles (e.g., such as vehicle 110). The sensor data can include radar data, LIDAR data, and / or camera data, among other examples. The sensor data can be data collected or obtained by one or more vehicles. In some aspects, the sensor data can be historical sensor data used by the server device 130 to generate a training data set for the ML model 330.

[0057] The sensor data can be associated with a set of frames. For example, a frame can be associated with a grid. The grid can define a set of cells associated with the frame. For example, the grid information can include information associated with a static fixed coordinate system and a vehicle-fixed coordinate system. The static fixed coordinate system can remain unchanged during a period in which the vehicle 110 travels along a route (e.g., during a period that starts when the vehicle 110 moves from an initial position of the vehicle 110 and ends when the vehicle 110 reaches a destination, the ignition of the vehicle 110 moves to the off position, and / or the vehicle 110 transitions to a parked state). In some aspects, the static fixed coordinate system includes an origin corresponding to the initial position of the vehicle 110 on a map, and each axis of the fixed coordinate system can extend in a respective direction perpendicular to the direction in which each other axis extends. For example, the static fixed coordinate system can be in an East-North-Up (ENU) format, and the first axis can be aligned in the east-west direction (e.g., the value of the coordinate of the first axis increases as the vehicle 110 travels east and decreases as the vehicle 110 travels west), the second axis can be aligned in the north-south direction (e.g., the value of the coordinate of the second axis increases as the vehicle 110 travels north and decreases as the vehicle 110 travels south), and / or the third axis can be aligned in the up-down direction (e.g., the value of the coordinate of the third axis increases as the vehicle 110 travels up (e.g., along a ramp in a parking garage) and decreases as the vehicle 110 travels down). The ENU coordinate system is provided as an example, and multiple other coordinate systems can be similarly applicable as described herein.

[0058] In some aspects, a static fixed coordinate system can be divided into a grid of multiple cells corresponding to respective regions on a map. In some aspects, each of the multiple cells can have the same size as other cells among the multiple cells. In some other aspects, each cell can not have the same size. The size of the multiple cells can be at least partially based on the rate at which a sensor or scanner (e.g., radar scanner 245 and / or LIDAR scanner 250) obtains a point data frame (e.g., the size of the multiple cells can be inversely proportional to the rate at which the sensor or scanner obtains a frame of point data, or the size of the multiple cells can be proportional to the rate at which the sensor or scanner obtains a frame of point data), and / or the type of region associated with the environment around vehicle 110 (e.g., rural or urban), among other examples. The rate at which the sensor or scanner obtains a frame of point data can define the duration of a frame, e.g., the time interval between two subsequent frames.

[0059] In some aspects, a vehicle fixed coordinate system can have an origin located at the current position of vehicle 110 (e.g., the position of the origin changes as the position of vehicle 110 changes). Each axis of the vehicle fixed coordinate system can be aligned with a corresponding axis of the static fixed vehicle system. For example, the vehicle fixed coordinate system can be in ENU format, and the first axis can be aligned with the first axis of the static fixed coordinate system in the east-west direction, the second axis can be aligned with the second axis of the static fixed coordinate system in the north-south direction, and / or the third axis can be aligned with the third axis of the static fixed coordinate system in the up-down direction.

[0060] The vehicle fixed coordinate system can be divided into a grid of multiple cells corresponding to respective regions on a map. In some aspects, the size of the multiple cells of the vehicle fixed coordinate system is the same as the size of the multiple cells of the static fixed coordinate system. In some aspects, the boundaries of the multiple cells of the vehicle fixed coordinate system can be aligned with the boundaries of the multiple cells of the static fixed system. In some aspects, when vehicle 110 travels along a route, the vehicle fixed coordinate system is shifted by an integer number of cells to eliminate any offset between the cell boundaries of the cells of the static fixed coordinate system and the cells of the vehicle fixed coordinate system.

[0061] As shown by reference numeral 305, the server device 130 can obtain a training dataset (e.g., for ML model 330) by aggregating multiple frames of vehicle sensor data using a common pose (e.g., a common coordinate reference system). For example, as Figure 3AAs shown, the server device 130 may aggregate N frames of sensor data to generate an aggregated frame. Each point in the frame may indicate a sensor detection (e.g., a radar detection or a LIDAR detection). For example, the vehicle 110 and / or the ECU 112 may receive point data from sensors (e.g., a radar scanner 245 and / or a LIDAR scanner 250). The sensors may emit energy pulses (e.g., radio waves and / or light waves) in a first direction and may obtain a first frame of point data based at least in part on the reflection of the energy pulses from an object. The sensors may emit energy pulses in a second direction and may obtain a second frame of point data. The sensors may continue in a similar manner to obtain a series of frames of point data corresponding to one or more objects in the environment surrounding the vehicle 110.

[0062] The sensors may provide one or more of the frames of point data to the ECU 112. In some aspects, the sensors may provide the frames of point data to the ECU 112 at least in part based on obtaining the frames of point data. In some aspects, the sensors provide a set of frames of point data to the ECU 112. In some aspects, the set of frames of point data includes each frame of point data obtained by the sensors as the sensors rotate 360 degrees (e.g., a complete point cloud of point data).

[0063] In some aspects, the frames of point data include one or more instances of point data. "Point data" and "sensor data" may be used interchangeably herein. Each instance of point data included in a frame of point data (referred to herein as a "point" or a "point of point data") may include one or more characteristics of an object associated with the point of point data. For example, a point of point data may include a set of coordinates corresponding to the location of the object (e.g., an x coordinate, a y coordinate, and / or a z coordinate in a Cartesian coordinate system), a set of velocities associated with the object (e.g., a velocity in a direction corresponding to a first axis of the coordinate system (e.g., V x ), a velocity in a direction corresponding to a second axis of the coordinate system (e.g., V y ), and / or a velocity in a direction corresponding to a third axis of the coordinate system (e.g., V z), an indication of the probability of existence associated with the object, and / or an indication of the size of the object represented by a set of points (e.g., the radar cross-section associated with the object), and other examples. Additionally, multiple auxiliary parameters can be provided in the point data, e.g., additional information regarding the dynamic properties of the points. For example, each cell of a frame can be associated with point data or sensor data. The sensor data of the cell can include one or more radar cross-section values (e.g., the minimum radar cross-section value, the maximum radar cross-section value, and / or the average radar cross-section value), one or more object speed values associated with one or more objects associated with the cell (e.g., the minimum speed value associated with one or more objects, the maximum speed value associated with one or more objects, and / or the average speed value associated with one or more objects), the coordinate position of at least one sensor detection (e.g., associated with one or more objects), the number of sensor detections (e.g., the number of points within the cell and / or the number of radar or LIDAR detections within the cell), and / or the ego speed value (e.g., the speed of vehicle 110), and other examples.

[0064] The server device 130 can obtain a set of frames (e.g., one or more frames) of sensor data and / or point data (e.g., which are generated or obtained by the vehicle 110 and / or the ECU 112 in a similar manner as described above). The server device 130 can aggregate the set of frames to generate an aggregated frame. The server device 130 can use a first pose or a first coordinate reference system to aggregate the set of frames. For example, the server device 130 can use a common coordinate reference system of the set of frames (such as the ENU coordinate system) to aggregate the set of frames. For example, the server device 130 can use a static fixed coordinate system to aggregate the set of frames. This can ensure that the sensor detections and / or points are accurately placed within the aggregated frame with respect to a common coordinate reference system (e.g., the static fixed coordinate system). Aggregating the set of frames can increase the amount of point data and / or sensor data available for training the ML model 330. This can improve the reliability and / or accuracy of the training data associated with the ML model 330. For example, providing more data to the ML model 330 can enable the ML model 330 to make improved inferences and / or predictions associated with the occupancy status of the corresponding cells included in the aggregated frame.

[0065] In some aspects, as shown by reference numeral 310, the server device 130 may convert an aggregated frame (or one or more aggregated frames) from a first pose to a second pose (e.g., a second coordinate reference system). For example, after aggregating sensor data and / or point data from a set of frames (e.g., using the first pose or the first coordinate reference system), the server device 130 may place the sensor data and / or point data onto a grid (e.g., of the aggregated frame) using the second pose or the second coordinate reference system. For example, the second pose may be a coordinate reference system of the current position of the vehicle 110. For example, the second pose or the second coordinate reference system may be a vehicle-fixed coordinate system. In some aspects, the second pose or the second coordinate reference system may be a unified localization and mapping (ULM) pose. For example, the vehicle 110 and / or the ECU 112 may track the movement of the vehicle 110 such that the vehicle 110 and / or the ECU 112 can track the position of the vehicle 110 (e.g., the pose of the vehicle 110) during the aggregation of the set of frames. This enables the server device 130, the vehicle 110, and / or the ECU 112 to aggregate the set of frames using a common reference system. The aggregated frame may be input into the ML model 330 in a different reference system (e.g., the second pose or the second coordinate reference system). In some aspects, the server device 130 may use the aggregated frame in the second pose as an input to train the ML model 330. For example, one or more aggregated frames (e.g., converted to the second pose or the second coordinate reference system) may be provided as an input to the ML model 330. Converting the aggregated frame to the second pose or the second coordinate reference system provides additional flexibility to the server device 130 (or any other device training the ML model 330) to aggregate sensor data using the first reference system and train the ML model 330 using the second reference system (e.g., the server device 130, the vehicle 110, and / or the ECU 112 are not limited to a specific predefined reference system for the data provided to the ML model 330).

[0066] As shown by reference numeral 315, the server device 130 may obtain labels for a subset of cells of one or more aggregated frames. For example, the server device 130 may obtain an indication of known labels for a subset of cells from a set of cells associated with the grid of the aggregated frame. In some aspects, the server device 130 may obtain an indication of corresponding placeholder labels for each cell from the set of cells. The corresponding placeholder labels may include a first placeholder label indicating a known placeholder state (e.g., a "known label" as shown in Figure 3A or a second placeholder label indicating an unknown placeholder state (e.g., as shown in Figure 3Athe "unknown marker" shown in). For example, the placeholder marker(s) can be ground truth markers for training the ML model 330. In some aspects, known placeholder status markers can be associated with positive ground truth markers (e.g., values greater than zero). Unknown placeholder status markers can be associated with ground truth marker values of zero.

[0067] Known placeholder status markers can indicate that a cell is associated with a placeholder status of a known category. For example, if a cell is associated with sensor detections associated with a given category (e.g., road category, non - road category, background category, parking category, railroad or track category, sidewalk category, bike lane category, lane marking category, pedestrian category, and / or crosswalk or sidewalk category, and other examples), the cell can be marked with a known placeholder marker of the given category (e.g., the ground truth marker value "1" of the given category). If a cell is not associated with sensor detections and / or is not associated with a known category, the cell can be marked with an unknown placeholder marker (e.g., the true marker value "0"). For example, an array indicating the ground truth markers of the aggregated frame and the placeholder status of a given category can be generated for each category. A value of "1" in an entry of the array can indicate that the cell corresponding to the entry is associated with a known placeholder marker of the category associated with the array. A value of "0" in an entry of the array can indicate that the cell corresponding to the entry is associated with an unknown placeholder marker of the category associated with the array. The server device 130 can obtain the placeholder status markers of the aggregated frames of one or more categories. For example, the server device 130 can obtain the placeholder status markers of the aggregated frames of the road category and the non - road category, and other examples.

[0068] In some aspects, the server device 130 can obtain the placeholder status markers based on user input. For example, a user can mark cells to facilitate training the ML model 330 to classify the placeholder status of each cell of the aggregated frame, as described in more detail elsewhere in this document. In some aspects, the server device 130 can determine the placeholder status markers of one or more categories of the aggregated frame based on analyzing the sensor detections and / or points included in the aggregated frame.

[0069] As Figure 3BAs shown, and by reference numeral 320, the server device 130 may train an ML model 330 to classify the categories of the occupancy states of all cells of the aggregated frame. The server device 130 (and / or the ML model trainer 325) may use data associated with the aggregated frame to train the ML model 330 to generate an occupancy grid. As described elsewhere herein, the occupancy grid may be a static occupancy grid or a dynamic occupancy grid. For example, the ML model trainer 325 may be a component of the server device 130 that is configured to train the ML model 330 based at least in part on data associated with the aggregated frame. For example, the aggregated frame may be associated with sensor data or point data for each cell of the aggregated frame that includes sensor detections (e.g., radar detections and / or LIDAR detections).

[0070] For example, as shown by reference numeral 335, the model trainer 325 may provide data for one or more aggregated frames as input to the ML model 330. For example, the cells of the aggregated frame may be associated with data. In some aspects, a subset of cells from a set of cells of the aggregated frame may be associated with data (e.g., sensor data). The data may include one or more radar cross-section values (e.g., minimum radar cross-section value, maximum radar cross-section value, and / or average radar cross-section value), one or more object velocity values associated with one or more objects associated with the cell (e.g., minimum velocity value associated with one or more objects, maximum velocity value associated with one or more objects, and / or average velocity value associated with one or more objects), the coordinate position of at least one sensor detection (e.g., associated with one or more objects), the number of sensor detections (e.g., the number of points within the cell and / or the number of radar or LIDAR detections within the cell), and / or a self-velocity value (e.g., the velocity of the vehicle 110), among other examples. The ML model 330 may predict the probability of the category of the occupancy state of each cell of the aggregated frame based on the data of the aggregated frame.

[0071] As shown by reference numeral 340, the ML model 330 can output one or more inferences. For example, one or more inferences can include the probability of a class that aggregates the occupancy states of each cell. For example, the ML model 330 can output an indication of a class (e.g., road, non-road, and / or another class described herein) associated with a respective cell from a set of cells included in the aggregated frame. As shown by reference numeral 345, the ML model trainer 325 can calculate a loss using a loss function based only on cells associated with known occupancy markers. For example, as described elsewhere herein, a subset of cells included in the aggregated frame can be associated with known occupancy markers. Training a machine learning model can be associated with a loss function that calculates the loss for respective cells from the subset of cells (e.g., rather than cells associated with unknown occupancy state markers).

[0072] For example, the ML model trainer 325 can use a loss function to calculate the loss for respective cells from the subset of cells. The ML model trainer 325 can suppress calculating the loss for respective cells associated with a second occupancy marker (e.g., cells associated with an unknown occupancy state and / or ground truth "0"). In some aspects, the ML model trainer 325 can use a loss function for the occupancy state of a respective class to calculate the loss. For example, the loss function can include a first loss function associated with a first class (e.g., drivable occupancy state) and a second loss function associated with a second class (e.g., non-drivable occupancy state), among other examples. For example, the ML model trainer 325 compares the ground truth of each class with the output of the ML model 330. The ML model trainer 325 can calculate the loss for the occupancy state of each class using only cells associated with the positive ground truth of a given class.

[0073] For example, a ground truth marker indicating a known occupancy state or class can be more reliable in classifying the occupancy state of a given cell than a ground truth marker indicating the absence of a known occupancy state or class (e.g., an unknown occupancy state marker). For example, a cell marked with an unknown occupancy state marker can be associated with a class of occupancy state, but the sensor data may not detect any sensor detections in the set of frames associated with the aggregated frame of the cell. Thus, relying on the cell not being associated with an occupancy state or class can lead to inaccurate training of the ML model 330. Thus, using one or more loss functions that calculate the loss for only the cells of the aggregated frame associated with known occupancy state markers can improve the accuracy of training of the ML model 330.

[0074] In some aspects, the ML model trainer 325 can identify, from the output of a machine learning model, one or more cells from a set of cells that are associated with an incorrect occupancy state or an incorrect category within a known occupancy state category. For example, the ML model trainer 325 can identify, based on the output of the ML model 330, whether any cells are classified as a first category, and based on the ground truth label, identify whether any cells are classified as a second category. The ML model trainer 325 can apply a penalty weight for one or more cells to a loss function for the second category. For example, when calculating the loss for the second category, the ML model trainer 325 can apply a penalty to cells associated with an incorrect occupancy state or an incorrect category within a known occupancy state category (e.g., apply a penalty to one or more cells classified as the road category when the ground truth label indicates the "non-road" category, and vice versa). For example, since the loss is only calculated for cells associated with a known occupancy state label(s), training of the ML model 330 can result in cells located near cells associated with a known occupancy state being classified as the category of the nearby cells. However, if a cell located near a cell associated with a known occupancy state is actually associated with a different occupancy state category, this classification may be incorrect and may degrade the performance and / or accuracy of the ML model 330. Thus, to mitigate the risk of incorrect classification or incorrect labeling caused by using the loss functions described herein, the ML model trainer 325 can apply a penalty to cells associated with an incorrect occupancy state or an incorrect category within a known occupancy state category.

[0075] As shown by reference numeral 350, the ML model trainer 325 and / or the server device 130 can update one or more weights associated with the ML model 330 based on the losses of the respective cells from a subset of cells. For example, the ML model 330 can include a neural network, and the ML model trainer 325 and / or the server device 130 can update one or more weights of the neural network based on the losses of the respective cells from a subset of cells to improve the performance of the ML model 330. The ML model trainer 325 can continue training the ML model 330 by providing data for one or more aggregated frames as input to the ML model 330 and calculating the losses for the occupancy states of one or more categories in a similar manner as described above until a training criterion is met.

[0076] As Figure 3CAs shown, vehicle 110 and / or ECU 112 may include an occupancy grid generator 355. The occupancy grid generator 355 may be a component of vehicle 110 and / or ECU 112 configured to generate an occupancy grid based on the output of ML model 330. For example, as indicated by reference numeral 360, vehicle 110 and / or ECU 112 may obtain ML model 330 (e.g., after ML model 330 has been trained, as described in more detail elsewhere herein). For example, vehicle 110 and / or ECU 112 may obtain ML model 330 from server device 130 (e.g., vehicle 110 and / or ECU 112 may download the trained ML model 330 from server device 130). In other aspects, vehicle 110 and / or ECU 112 may not obtain ML model 330. For example, ML model 330 may be maintained by another device such as server device 130. Vehicle 110 and / or ECU 112 may provide data (e.g., data associated with an aggregated frame) to another device. The other device may input the data into ML model 330. In such an example, the other device may send the output of ML model 330, and vehicle 110 and / or ECU 112 may receive the output of ML model 330.

[0077] As indicated by reference numeral 365, vehicle 110 may obtain sensor data and / or point data collected by radar scanner 245 and / or LIDAR scanner 250, among other examples. The sensor data may indicate one or more sensor detections (e.g., one or more radar detections and / or one or more LIDAR detections). For example, ECU 112 may receive sensor data and / or point data from radar scanner 245 and / or LIDAR scanner 250. The sensor data may identify a plurality of points corresponding to one or more objects located in the physical environment of vehicle 110. For example, radar scanner 245 may emit one or more pulses of electromagnetic waves. The one or more pulses may be reflected by an object in the path of the one or more pulses. The reflection may be received by radar scanner 245. Radar scanner 245 may determine one or more characteristics associated with the reflected pulse (e.g., amplitude, frequency, etc.) and may determine point data indicative of the location of the object based on the one or more characteristics. Radar scanner 245 may provide point data indicative of the radar detection to ECU 112.

[0078] Additionally or alternatively, the LIDAR scanner 250 may emit one or more light pulses. The one or more pulses may be reflected by an object in the path of the one or more pulses. The reflection may be received by the LIDAR scanner 250. The LIDAR scanner 250 may determine one or more characteristics associated with the reflected pulses and may determine point data indicative of the location of the object based on the one or more characteristics. The LIDAR scanner 250 may provide the point data indicative of the LIDAR detection to the ECU 112.

[0079] As shown by reference numeral 370, the vehicle 110 and / or the ECU 112 may aggregate sensor data associated with a set of frames using a first pose (e.g., a common coordinate reference system) to generate an aggregated frame. For example, the vehicle 110 and / or the ECU 112 may aggregate sensor data in a manner similar to that described in more detail elsewhere herein (such as in connection with Figure 3A ). For example, the vehicle 110 and / or the ECU 112 may obtain a set of frames of sensor data and / or point data (e.g., one or more frames). The vehicle 110 and / or the ECU 112 may aggregate the set of frames to generate an aggregated frame. The vehicle 110 and / or the ECU 112 may use the first pose or the first coordinate reference system to aggregate the set of frames. For example, the vehicle 110 and / or the ECU 112 may use a common coordinate reference system of the set of frames (such as the ENU coordinate system) to aggregate the set of frames. This may ensure that sensor detections and / or points are accurately placed within the aggregated frame relative to the common coordinate reference system (e.g., a static fixed coordinate system). Aggregating the set of frames may increase the amount of point data and / or sensor data available for analysis by the ML model 330. This may improve the reliability and / or accuracy of the output of the ML model 330. For example, providing more data to the ML model 330 may enable the ML model 330 to make improved inferences and / or predictions associated with the occupancy status of the corresponding cells included in the aggregated frame.

[0080] In some aspects, vehicle 110 and / or ECU 112 may transform an aggregated frame (or one or more aggregated frames) from a first pose to a second pose (e.g., a second coordinate reference system). For example, after aggregating sensor data and / or point data from a set of frames (e.g., using the first pose or the first coordinate reference system), vehicle 110 and / or ECU 112 may place the sensor data and / or point data onto a grid (e.g., of the aggregated frame) using the second pose or the second coordinate reference system. For example, the second pose may be a coordinate reference system of the current position of vehicle 110. For example, the second pose or the second coordinate reference system may be a vehicle-fixed coordinate system. In some aspects, the second pose or the second coordinate reference system may be a ULM pose. For example, vehicle 110 and / or ECU 112 may track the movement of vehicle 110 such that vehicle 110 and / or ECU 112 can track the position of vehicle 110 (e.g., pose 110) during the aggregation of the set of frames. This enables vehicle 110 and / or ECU 112 to use a common reference system to aggregate the set of frames. The aggregated frame may be input into ML model 330 in a different reference system (e.g., the second pose or the second coordinate reference system).

[0081] As shown by reference numeral 375, occupancy grid generator 355 may provide data associated with the aggregated frame to ML model 330 as an input. ML model 330 may determine probabilities associated with respective categories of occupancy states. For example, based on the data associated with the aggregated frame, ML model 330 may be trained (e.g., as described in more detail elsewhere herein) to predict probabilities associated with respective categories of the occupancy state of each cell of the aggregated frame. As an example, for a given cell, the output of ML model 330 may indicate that the cell has a 75% probability of being associated with a road category and a 25% probability of being associated with a "non-road" category. ML model 330 may predict probabilities for each cell of the aggregated frame in a similar manner. These categories are provided as examples, and ML model 330 may be trained to predict probabilities of additional and / or different categories, as described elsewhere herein.

[0082] As shown by reference numeral 380, occupancy grid generator 355 may obtain classification probabilities for each cell of the (one or more) aggregated frames. This may enable occupancy grid generator 355 to determine the occupancy state of each cell of the occupancy grid. For example, compared to a deterministic method for generating an occupancy grid (e.g., where sensor data may be required to identify the occupancy state of a given cell), occupancy grid generator 355 may be enabled to determine the occupancy state of each cell of the occupancy grid based on the output of ML model 330. For example, occupancy grid generator 355 may determine that a given cell is associated with the occupancy state or category associated with the highest probability, as indicated by the output of ML model 330.

[0083] As shown by reference numeral 385, the occupancy grid generator 355 can generate an occupancy grid at least in part based on the classification probabilities indicated by the ML model 330. For example, the occupancy grid generator 355 can determine the occupancy status and / or category associated with each cell of the occupancy grid based on the output of the ML model 330. As shown by reference numeral 390, the vehicle 110 and / or the ECU 112 can perform actions at least in part based on the generated occupancy grid. For example, the vehicle 110 and / or the ECU 112 can control the vehicle 110 according to the occupancy grid (e.g., cause the vehicle 110 to avoid an area associated with a cell associated with a non-drivable occupancy status). The ECU 112 can perform actions associated with controlling the vehicle 110 (e.g., accelerating, decelerating, stopping, and / or changing lanes) based on the location information associated with the cell, where the cell is associated with a non-drivable occupancy status indicated by the occupancy grid. In some aspects, the location information indicates a grid position associated with one or more of the cells included in the occupancy grid. The ECU 112 can convert the grid information into a region of the physical environment, the region of the physical environment corresponding to the region of the grid that includes one or more cells associated with a non-drivable occupancy status. In some aspects, the location information indicates an area of the physical environment occupied by one or more objects. The region of the physical environment can correspond to the region of the grid that includes one or more cells associated with a non-drivable occupancy status.

[0084] In some aspects, performing the actions includes the ECU 112 indicating, via a user interface of the vehicle 110 and at least in part based on the location information, a location of one or more cells associated with a non-drivable occupancy status relative to the vehicle 110. For example, the user interface can display a map of the physical environment of the vehicle 110. The origin of the grid can correspond to the current position of the vehicle 110. The ECU 112 can cause information associated with one or more cells associated with a non-drivable occupancy status (e.g., an icon and / or another type of information corresponding to the category of one or more cells) to be displayed on the map at a position corresponding to the position of one or more cells in the grid in combination with information associated with the current position of the vehicle 110.

[0085] As above, Figures 3A to 3C is provided as an example. Other examples can be different from the examples described with respect to Figures 3A to 3C the examples.

[0086] Figure 4 is a flow chart of an example process 400 associated with occupancy grid generation based on machine learning according to the present disclosure. In some aspects, Figure 4One or more processing blocks are performed by a server device (e.g., server device 130). In some aspects, Figure 4 One or more processing blocks are performed by another device or a set of devices (such as vehicle 110, ECU 112, and / or wireless communication device 120) that are separate from or include the server device. Additionally or alternatively, Figure 4 One or more process blocks may be performed by one or more components of device 200 (such as processor 210, memory 215, storage component 220, input component 225, output component 230, communication interface 235, one or more sensors 240, radar scanner 245, and / or LIDAR scanner 250).

[0087] As Figure 4 shown, process 400 may include receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections (block 410). For example, the server device may receive sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections, as described above. In some aspects, the sensor data indicates one or more sensor detections.

[0088] As Figure 4 further shown, process 400 may include aggregating sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame is associated with a set of cells (block 420). For example, the server device may use a first pose to aggregate sensor data associated with a set of frames to generate an aggregated frame, where the aggregated frame is associated with a set of cells, as described above. In some aspects, the aggregated frame is associated with a set of cells.

[0089] As Figure 4 further shown, process 400 may include obtaining an indication of corresponding occupancy markers for each cell from the set of cells, where the corresponding occupancy markers include a first occupancy marker indicating a known occupancy state or a second occupancy marker indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy marker (block 430). For example, the server device may obtain an indication of corresponding occupancy markers for each cell from the set of cells, where the corresponding occupancy markers include a first occupancy marker indicating a known occupancy state or a second occupancy marker indicating an unknown occupancy state, and where a subset of cells from the set of cells is associated with the first occupancy marker, as described above. In some aspects, the corresponding occupancy markers include a first occupancy marker indicating a known occupancy state or a second occupancy marker indicating an unknown occupancy state. In some aspects, a subset of cells from the set of cells is associated with the first occupancy marker.

[0090] AsFigure 4 As further shown, process 400 may include using data associated with an aggregated frame to train a machine learning model to generate an occupancy grid, wherein training the machine learning model is associated with a loss function that calculates a loss for corresponding cells from a subset of cells, and wherein the machine learning model is trained to predict a probability of an occupancy state for corresponding cells from a set of cells (block 440). For example, a server device may use data associated with an aggregated frame to train a machine learning model to generate an occupancy grid, wherein training the machine learning model is associated with a loss function that calculates a loss for corresponding cells from a subset of cells, and wherein the machine learning model is trained to predict a probability of an occupancy state for corresponding cells from a set of cells, as described above. In some aspects, training the machine learning model is associated with a loss function that calculates a loss for corresponding cells from a subset of cells. In some aspects, the machine learning model is trained to predict a probability of an occupancy state for corresponding cells from a set of cells.

[0091] As Figure 4 As further shown, process 400 may include providing the machine learning model to another device (block 450). For example, a server device may provide the machine learning model to another device, as described above.

[0092] Process 400 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in combination with one or more other process descriptions elsewhere herein.

[0093] In a first aspect, training the machine learning model includes using a loss function to calculate a loss for corresponding cells from a subset of cells, suppressing calculation of a loss for corresponding cells associated with a second occupancy marker, and updating one or more weights associated with the machine learning model based on the loss for corresponding cells from the subset of cells.

[0094] In a second aspect, separate or in combination with the first aspect, the first occupancy marker is associated with indicating an occupancy state of a subset of cells, and the occupancy state includes a drivable occupancy state or a non-drivable occupancy state.

[0095] In a third aspect, separate or in combination with one or more of the first and second aspects, training the machine learning model includes identifying, from an output of the machine learning model, one or more cells from a set of cells associated with an incorrect occupancy state among a drivable occupancy state or a non-drivable occupancy state, applying a penalty weight for the one or more cells to the loss function, and updating one or more weights associated with the machine learning model based on an output of the loss function.

[0096] In a fourth aspect, either alone or in combination with one or more of the first to third aspects, the loss function includes a first loss function associated with a drivable occupancy state and a second loss function associated with a non-drivable occupancy state.

[0097] In a fifth aspect, either alone or in combination with one or more of the first to fourth aspects, process 400 includes transforming an aggregate frame from a first pose to a second pose, where the first pose is a common coordinate reference system for a set of frames and where the second pose is a coordinate reference system of the current position of the vehicle.

[0098] In a sixth aspect, either alone or in combination with one or more of the first to fifth aspects, the occupancy grid includes at least one of a static occupancy grid or a dynamic occupancy grid.

[0099] In a seventh aspect, either alone or in combination with one or more of the first to sixth aspects, the sensor data includes at least one of radar data, LIDAR data, or camera data.

[0100] Although Figure 4 example boxes of process 400 are shown, in some aspects, process 400 includes additional boxes, fewer boxes, different boxes, or boxes arranged differently compared to those depicted in Figure 4 . Additionally or alternatively, two or more boxes of process 400 may be executed in parallel.

[0101] Figure 5 is a flowchart of an example process 500 associated with machine learning-based occupancy grid generation according to the present disclosure. In some aspects, Figure 5 one or more processing boxes of Figure 5 are executed by a vehicle (e.g., vehicle 110 and / or ECU 112). In some aspects, Figure 5 one or more processing boxes of

[0102] are executed by another device or set of devices (such as ECU 112, server device 130, and / or wireless communication device 120) that is separate from or includes the vehicle. Additionally or alternatively, Figure 5As shown, process 500 may include receiving sensor data associated with a vehicle and a set of frames, where the sensor data indicates one or more sensor detections (block 510). For example, a vehicle may receive sensor data associated with the vehicle and a set of frames, where the sensor data indicates one or more sensor detections, as described above. In some aspects, the sensor data indicates one or more sensor detections.

[0103] As Figure 5 further shown, process 500 may include aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections (block 520). For example, a vehicle may aggregate the sensor data associated with the set of frames using a first pose to generate an aggregated frame, where the aggregated frame includes a set of cells, and where the aggregated frame includes data for a subset of cells associated with one or more sensor detections, as described above. In some aspects, the aggregated frame includes a set of cells. In some aspects, the aggregated frame includes data for a subset of cells associated with one or more sensor detections.

[0104] As Figure 5 further shown, process 500 may include generating an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells (block 530). For example, a vehicle may generate an occupancy grid using a machine learning model, where the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and where the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells, as described above. In some aspects, the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states. In some aspects, the output of the machine learning model includes probabilities of one or more occupancy states for corresponding cells from the set of cells.

[0105] As Figure 5 further shown, process 500 may include performing an action based on the occupancy grid (block 540). For example, a vehicle may perform an action based on the occupancy grid, as described above.

[0106] Process 500 may include additional aspects, such as any individual aspect or any combination of aspects described below and / or in combination with one or more other process descriptions described elsewhere herein.

[0107] In a first aspect, process 500 includes obtaining a machine learning model from another device.

[0108] In a second aspect, either alone or in combination with the first aspect, the device is a control unit of a vehicle.

[0109] In a third aspect, either alone or in combination with one or more of the first and second aspects, process 500 includes transforming an aggregated frame from a first pose to a second pose, where the first pose is a common coordinate reference system for a set of frames and where the second pose is a different coordinate reference system.

[0110] In a fourth aspect, either alone or in combination with one or more of the first through third aspects, generating an occupancy grid includes using the aggregated frame in the second pose as an input to the machine learning model to generate the occupancy grid.

[0111] In a fifth aspect, either alone or in combination with one or more of the first through fourth aspects, one or more occupancy states include a drivable occupancy state or a non-drivable occupancy state.

[0112] In a sixth aspect, either alone or in combination with one or more of the first through fifth aspects, data for a subset of cells includes at least one of one or more radar cross-section values, one or more object velocity values, a coordinate position of at least one sensed detection, a number of sensor detections, or a self-velocity value.

[0113] Although Figure 5 example boxes of process 500 are shown, in some aspects, process 500 includes additional boxes, fewer boxes, different boxes, or boxes in a different arrangement compared to those depicted Figure 5 herein. Additionally or alternatively, two or more of the boxes of process 500 may be executed in parallel.

[0114] An overview of some aspects of the present disclosure is provided below:

[0115] Aspect 1: A method, comprising: receiving, by a device, sensor data associated with a vehicle and a set of frames, wherein the sensor data indicates one or more sensor detections; aggregating, using a first pose, the sensor data associated with the set of frames to generate an aggregated frame, wherein the aggregated frame is associated with a set of cells; obtaining, by the device, an indication of respective occupancy tokens for each cell from the set of cells, wherein the respective occupancy tokens include a first occupancy token indicating a known occupancy state or a second occupancy token indicating an unknown occupancy state, and wherein a subset of cells from the set of cells is associated with the first occupancy token; training, using data associated with the aggregated frame, a machine learning model to generate an occupancy grid, wherein training the machine learning model is associated with a loss function that calculates a loss for a respective cell from the subset of cells, and wherein the machine learning model is trained to predict a probability of an occupancy state of a respective cell from the set of cells; and providing, by the device, the machine learning model to another device.

[0116] Aspect 2: The method according to aspect 1, wherein training the machine learning model comprises: calculating, using the loss function, a loss for a respective cell from the subset of cells; suppressing the calculation of the loss for a respective cell associated with the second occupancy token; and updating one or more weights associated with the machine learning model based on the loss for a respective cell from the subset of cells.

[0117] Aspect 3: The method according to any one of aspects 1 to 2, wherein the first occupancy token is associated with an indication of an occupancy state of the subset of cells, and wherein the occupancy state includes a drivable occupancy state or a non-drivable occupancy state.

[0118] Aspect 4: The method according to aspect 3, wherein training the machine learning model comprises: identifying, from an output of the machine learning model, one or more cells from the set of cells associated with an incorrect occupancy state among the drivable occupancy state or the non-drivable occupancy state; applying a penalty weight for the one or more cells to the loss function; and updating one or more weights associated with the machine learning model based on an output of the loss function.

[0119] Aspect 5: The method according to any one of aspects 3 to 4, wherein the loss function includes a first loss function associated with the drivable occupancy state and a second loss function associated with the non-drivable occupancy state.

[0120] Aspect 6: The method according to any one of aspects 1 to 5, further comprising: transforming the aggregated frame from the first pose to a second pose, wherein the first pose is a common coordinate reference system of the set of frames, and wherein the second pose is a coordinate reference system of a current position of the vehicle.

[0121] Aspect 7: The method according to any one of Aspects 1 to 6, wherein training the machine learning model includes using the aggregated frame in the second pose as input to train the machine learning model.

[0122] Aspect 8: The method according to any one of Aspects 1 to 7, wherein the data associated with the aggregated frame includes detection data from each cell included in the subset of cells, and wherein the detection data includes at least one of the following: one or more radar cross-section values, one or more object velocity values, coordinate positions detected by at least one sensor, the number of sensor detections, or a self-velocity value.

[0123] Aspect 9: The method according to any one of Aspects 1 to 8, wherein the occupancy grid includes at least one of a static occupancy grid or a dynamic occupancy grid.

[0124] Aspect 10: The method according to any one of Aspects 1 to 9, wherein the sensor data includes at least one of radar data, LIDAR data, or camera data.

[0125] Aspect 11: A method, comprising: receiving, by a device, sensor data associated with a vehicle and a set of frames, wherein the sensor data indicates one or more sensor detections; aggregating, using a first pose, the sensor data associated with the set of frames to generate an aggregated frame, wherein the aggregated frame includes a set of cells, and wherein the aggregated frame includes data of a subset of cells associated with the one or more sensor detections; generating, using a machine learning model, an occupancy grid, wherein the machine learning model is trained using a loss function that calculates the loss only using cells associated with known occupancy states, and wherein the output of the machine learning model includes probabilities of one or more occupancy states of corresponding cells from the set of cells; and performing, by the device, an action based on the occupancy grid.

[0126] Aspect 12: The method according to Aspect 11, comprising: obtaining the machine learning model from another device.

[0127] Aspect 13: The method according to any one of Aspects 11 to 12, wherein the device is a control unit of the vehicle.

[0128] Aspect 14: The method according to any one of Aspects 11 to 13, further comprising: converting the aggregated frame from the first pose to a second pose, wherein the first pose is a common coordinate reference system of the set of frames, and wherein the second pose is a different coordinate reference system.

[0129] Aspect 15: The method according to Aspect 14, wherein generating the occupancy grid includes: using the aggregated frame in the second pose as input to the machine learning model to generate the occupancy grid.

[0130] Aspect 16: The method according to any one of aspects 11 to 15, wherein one or more occupancy states include a drivable occupancy state or a non-drivable occupancy state.

[0131] Aspect 17: The method according to any one of aspects 11 to 16, wherein the data of the subset of cells includes at least one of the following items: one or more radar cross-section values, one or more object velocity values, the coordinate positions of at least one sensed detection, the number of sensor detections, or the ego velocity value.

[0132] Aspect 18: The method according to any one of aspects 11 to 17, wherein the occupancy grid includes at least one of a static occupancy grid or a dynamic occupancy grid.

[0133] Aspect 19: The method according to any one of aspects 11 to 18, wherein the sensor data includes at least one of radar data, LIDAR data, or camera data.

[0134] Aspect 20: A system configured to perform one or more operations recited in one or more of aspects 1 to 10.

[0135] Aspect 21: An apparatus comprising components for performing one or more operations recited in one or more of aspects 1 to 10.

[0136] Aspect 22: A non-transitory computer-readable medium storing an instruction set, the instruction set including one or more instructions that, when executed by a device, cause the device to perform one or more operations recited in one or more of aspects 1 to 10.

[0137] Aspect 23: A computer program product including instructions or code for performing one or more operations recited in one or more of aspects 1 to 10.

[0138] Aspect 24: A system configured to perform one or more operations recited in one or more of aspects 11 to 19.

[0139] Aspect 25: An apparatus comprising components for performing one or more operations recited in one or more of aspects 11 to 19.

[0140] Aspect 26: A non-transitory computer-readable medium storing an instruction set, the instruction set including one or more instructions that, when executed by a device, cause the device to perform one or more operations recited in one or more of aspects 11 to 19.

[0141] Aspect 27: A computer program product including instructions or code for performing one or more operations recited in one or more of aspects 11 to 19.

[0142] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the aspects to the precise forms disclosed. Modifications and variations can be made in light of the above disclosure, or can be obtained from the practice of the aspects.

[0143] As used herein, the term "component" is intended to be broadly construed as a combination of hardware and / or hardware and software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other terms, "software" should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, processes, and / or functions, among other examples. As used herein, a "processor" is implemented as a combination of hardware and / or hardware and software. It will be apparent that the systems and / or methods described herein can be implemented in different forms of combinations of hardware and / or hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit these aspects. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, as those skilled in the art will understand that software and hardware can be designed to implement the systems and / or methods at least in part based on the description herein.

[0144] As used herein, depending on the context, "meeting a threshold" can refer to a value that is greater than the threshold, greater than or equal to the threshold, less than the threshold, less than or equal to the threshold, equal to the threshold, not equal to the threshold, and so on.

[0145] Even if a particular combination of features is recited in the claims and / or disclosed in the specification, such combinations are not intended to limit the disclosure of the aspects. Many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. The disclosure of the aspects includes the combination of each dependent claim with every other claim in the set of claims. As used herein, the phrase "at least one" in reference to a list of items means any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover a, b, c, a + b, a + c, b + c, and a + b + c, as well as any combination with multiple identical elements (e.g., a + a, a + a + a, a + a + b, a + a + c, a + b + b, a + c + c, b + b, b + b + b, b + b + c, c + c, and c + c + c, or any other ordering of a, b, and c).

[0146] Unless expressly described as such, elements, acts, or instructions used herein should not be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be interchangeable with "one or more." Additionally, as used herein, the article "the" is intended to include one or more items referred to in conjunction with the article "the" and may be interchangeable with "one or more." Additionally, as used herein, the terms "set" and "group" are intended to include one or more items and may be interchangeable with "one or more." In the case of only Figure 1 one item, the phrase "only one" or similar language is used. Additionally, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms that do not limit the elements they modify (e.g., an element "having" A may also have B). Additionally, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on." Additionally, as used herein, the term "or" is intended to be inclusive when used in series and may be interchangeable with "and / or" unless expressly stated otherwise (e.g., if used in combination with "either" or "only one of").

Claims

1. A device, comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to: receive sensor data associated with a vehicle and a set of frames, wherein the sensor data indicates one or more sensor detections; aggregate the one or more sensor detections into an aggregated frame using a first pose, wherein the aggregated frame is associated with a set of cells; obtain an indication of a corresponding occupancy token for each cell from the set of cells, wherein the corresponding occupancy token includes a first occupancy token indicating a known occupancy state or a second occupancy token indicating an unknown occupancy state, and wherein a subset of cells from the set of cells is associated with the first occupancy token; use data associated with the aggregated frame to train a machine learning model to generate an occupancy grid, wherein training the machine learning model is associated with a loss function that calculates a loss for corresponding cells from the subset of cells, and wherein the machine learning model is trained to predict a probability of an occupancy state of corresponding cells from the set of cells; and provide the machine learning model to another device.

2. The device according to claim 1, wherein, For training the machine learning model, the one or more processors are configured to: use the loss function to calculate the loss for the corresponding cells from the subset of cells; suppress calculation of the loss for corresponding cells associated with the second occupancy token; and update one or more weights associated with the machine learning model based on the loss for the corresponding cells from the subset of cells.

3. The device according to claim 1, wherein The first occupancy token is associated with indicating the occupancy state of the subset of cells, and wherein the occupancy state includes a drivable occupancy state or a non-drivable occupancy state.

4. The apparatus according to claim 3, wherein, For training the machine learning model, the one or more processors are configured to: identify, from an output of the machine learning model, one or more cells from the set of cells associated with an incorrect occupancy state among the drivable occupancy state or the non-drivable occupancy state; apply a penalty weight to the one or more cells for the loss function; and update one or more weights associated with the machine learning model based on an output of the loss function.

5. The device according to claim 3, wherein, The loss function includes a first loss function associated with the drivable occupancy state and a second loss function associated with the non-drivable occupancy state.

6. The device according to claim 1, wherein, The one or more processors are further configured to: convert the aggregated frame from the first pose to a second pose, wherein the first pose is a common coordinate reference system of the set of frames, and wherein the second pose is a coordinate reference system of a current position of the vehicle.

7. The apparatus according to claim 6, wherein, For training the machine learning model, the one or more processors are configured to: use the aggregated frame in the second pose as an input to train the machine learning model.

8. The apparatus according to claim 1, wherein The data associated with the aggregated frame includes detection data for each cell included in the subset of cells, and wherein the detection data includes at least one of the following: One or more radar cross-section values, One or more object velocity values, Coordinate positions detected by at least one sensor, The number of sensor detections, or Self-velocity values.

9. An apparatus, comprising: One or more memories; And One or more processors, coupled to the one or more memories, configured to: Receive sensor data associated with a vehicle and a set of frames, Wherein the sensor data indicates one or more sensor detections; Use a first pose to aggregate the sensor data associated with the set of frames to generate an aggregated frame, Wherein the aggregated frame includes a set of cells, and Wherein the aggregated frame includes data of a subset of cells associated with the one or more sensor detections; Use a machine learning model to generate an occupancy grid, Wherein the machine learning model is trained using a loss function that calculates the loss using only cells associated with known occupancy states, and Wherein the output of the machine learning model includes probabilities of one or more occupancy states of corresponding cells from the set of cells; and Perform an action based on the occupancy grid.

10. The device according to claim 9, wherein, The one or more processors are further configured to: Convert the aggregated frame from the first pose to a second pose, Wherein the first pose is a common coordinate reference system of the set of frames, and Wherein the second pose is a different coordinate reference system.

11. The apparatus according to claim 10, wherein, To generate the occupancy grid, the one or more processors are configured to: Use the aggregated frame in the second pose as an input to the machine learning model to generate the occupancy grid.

12. The device according to claim 9, wherein, The one or more occupancy states include a drivable occupancy state or a non-drivable occupancy state.

13. The device according to claim 9, wherein, The data of the subset of cells includes at least one of the following: One or more radar cross-section values, One or more object velocity values, Coordinate positions detected by at least one sensor, The number of sensor detections, or Self-velocity values.

14. The apparatus according to claim 9, wherein, The occupancy grid includes at least one of a static occupancy grid or a dynamic occupancy grid.

15. The device according to claim 9, wherein, The sensor data includes at least one of radar data, LIDAR data, or camera data.

16. A method, comprising: Receiving, by a device, sensor data associated with a vehicle and a set of frames, Wherein the sensor data indicates one or more sensor detections; Using a first pose to aggregate the sensor data associated with the set of frames to generate an aggregated frame, Wherein the aggregated frame is associated with a set of cells; Obtaining, by the device, an indication of a corresponding occupancy label for each cell from the set of cells, Wherein the corresponding occupancy label includes a first occupancy label indicating a known occupancy state or a second occupancy label indicating an unknown occupancy state, and Wherein a subset of cells from the set of cells is associated with the first occupancy label; Training, using data associated with the aggregated frame, a machine learning model to generate an occupancy grid, Wherein training the machine learning model is associated with a loss function that calculates the loss of corresponding cells from the subset of cells, and Wherein, the machine learning model is trained to predict the probability of the occupancy status of corresponding cells from the set of cells; and The machine learning model is provided by the device to another device.

17. The method according to claim 16, wherein, Training the machine learning model includes: Using the loss function to calculate the loss of the corresponding cells from the subset of cells; Suppressing the calculation of the loss of the corresponding cells associated with the second occupancy marker; and Updating one or more weights associated with the machine learning model based on the loss of the corresponding cells from the subset of cells.

18. The method according to claim 16, wherein The first occupancy marker is associated with indicating the occupancy status of the subset of cells, and Wherein, the occupancy status includes a drivable occupancy status or a non-drivable occupancy status.

19. The method according to claim 18, wherein Training the machine learning model includes: Identifying, from the output of the machine learning model, one or more cells from the set of cells associated with an incorrect occupancy status among the drivable occupancy status or the non-drivable occupancy status; Applying a penalty weight to the one or more cells for the loss function; and Updating one or more weights associated with the machine learning model based on the output of the loss function.

20. The method according to claim 18, wherein The loss function includes a first loss function associated with the drivable occupancy status and a second loss function associated with the non-drivable occupancy status.

21. The method according to claim 16, further comprising: Converting the aggregated frame from the first pose to a second pose, Wherein, the first pose is a common coordinate reference system of the set of frames, and Wherein, the second pose is a coordinate reference system of the current position of the vehicle.

22. The method according to claim 16, wherein, The occupancy grid includes at least one of a static occupancy grid or a dynamic occupancy grid.

23. The method according to claim 16, wherein, The sensor data includes at least one of radar data, LIDAR data, or camera data.

24. A method, comprising: Receiving, by a device, sensor data associated with a vehicle and a set of frames, Wherein, the sensor data indicates one or more sensor detections; Aggregating the sensor data associated with the set of frames using a first pose to generate an aggregated frame, Wherein, the aggregated frame includes a set of cells, and Wherein, the aggregated frame includes data of a subset of cells associated with the one or more sensor detections; Generating an occupancy grid using a machine learning model, Wherein, the machine learning model is trained using a loss function that calculates the loss only using cells associated with known occupancy statuses, and Wherein, the output of the machine learning model includes probabilities of one or more occupancy statuses of corresponding cells from the set of cells; and Performing an action by the device based on the occupancy grid.

25. The method according to claim 24, comprising: Obtaining the machine learning model from another device.

26. The method according to claim 24, wherein, The device is a control unit of the vehicle.

27. The method according to claim 24, further comprising: Converting the aggregated frame from the first pose to a second pose, Wherein, the first pose is a common coordinate reference system of the set of frames, and Wherein, the second pose is a different coordinate reference system.

28. The method according to claim 27, wherein Generating the occupancy grid includes: Using the aggregated frame in the second pose as an input to the machine learning model to generate the occupancy grid.

29. The method according to claim 24, wherein The one or more occupancy states include a drivable occupancy state or a non-drivable occupancy state.

30. The method according to claim 24, wherein, The data of the cell subset includes at least one of the following: One or more radar cross-section values, One or more object speed values, Coordinate positions detected by at least one sensor, The number of sensor detections, or A self-speed value.