Processing map data

By processing perception data to generate and refine high-definition maps with 3D data from vehicles, the system addresses limitations of real-time perception data, enhancing map accuracy and safety in autonomous driving.

WO2026090771A1PCT designated stage Publication Date: 2026-05-07QUALCOMM INC +3
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Real-time perception data for autonomous driving systems has limitations that can compromise vehicle safety, as it may not provide comprehensive information about the environment, such as detailed road features and dynamic changes.

Method used

The system processes perception data to generate three-dimensional data, determines location information, transmits this data to a map server, receives a map based on plan-view-image data and 3D data, and adjusts vehicle operations accordingly, using machine-learning models to refine and update high-definition maps with 3D data from multiple vehicles.

Benefits of technology

This approach enhances map data accuracy and freshness, providing detailed and up-to-date information for autonomous driving systems, improving safety and decision-making capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127612_07052026_PF_FP_ABST
    Figure CN2024127612_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and techniques are described herein for processing map data. For example, a method for using map data is provided. The method may include processing perception data representative of an area to generate three-dimensional (3D) data; determining location information associated with the 3D data; transmitting the 3D data and the location information to a map server; receiving a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjusting an operating parameter of a vehicle based on the map. As another example, a method for generating map data is provided. The method may include generating a map of an area based on a plan-view image of the area; obtaining 3D data based on perception data captured in the area; and refining the map based on the 3D data to generate a refined map.
Need to check novelty before this filing date? Find Prior Art

Description

PROCESSING MAP DATATECHNICAL FIELD

[0001] The present disclosure generally relates to map data. For example, aspects of the present disclosure include systems and techniques for processing (e.g., using, generating, updating, and / or refining) map data (e.g., high-definition (HD) map data) .BACKGROUND

[0002] Perception data informs a driving systems (e.g., autonomous, semi-autonomous, or assisted driving systems, such as an advanced driver assistance system (ADAS) ) what area is drivable and what objects (e.g., road users, other vehicles, bikes, pedestrian, etc. ) are present and / or are moving in the environment around the vehicle. The driving system then makes decisions about how to move (e.g., slower, faster, stop, changing lanes, turning, a path to take, etc. ) . But real-time perception data has limitations which might influence the vehicle safety for autonomous driving.

[0003] Map data can provide additional information beyond what perception data can ‘see. ’ For example, map data (e.g., a high-definition (HD) map) may include map points –three-dimensional coordinates of surfaces of roads at a sub-meter granularity. Map data (e.g., an HD map) may also include additional features such as lane markers, road signs, traffic lights, traffic signs, poles, etc. Driving systems may use map data (e.g., HD maps) to make determinations about steering, accelerating, braking, path planning, and / or to provide information to a driver, etc.SUMMARY

[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.

[0005] Systems and techniques are described for using map data. According to at least one example, a method is provided for using map data. The method includes: processing perception  data representative of an area to generate three-dimensional (3D) data; determining location information associated with the 3D data; transmitting the 3D data and the location information to a map server; receiving a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjusting an operating parameter of a vehicle based on the map.

[0006] In another example, an apparatus for using map data is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: process perception data representative of an area to generate three-dimensional (3D) data; determine location information associated with the 3D data; transmit the 3D data and the location information to a map server; receive a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjust an operating parameter of a vehicle based on the map.

[0007] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: process perception data representative of an area to generate three-dimensional (3D) data; determine location information associated with the 3D data; transmit the 3D data and the location information to a map server; receive a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjust an operating parameter of a vehicle based on the map.

[0008] In another example, an apparatus for using map data is provided. The apparatus includes: means for processing perception data representative of an area to generate three-dimensional (3D) data; means for determining location information associated with the 3D data; means for transmitting the 3D data and the location information to a map server; means for receiving a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and means for adjusting an operating parameter of a vehicle based on the map.

[0009] In another example, a method is provided for generating map data. The method includes: generating a map of an area based on a plan-view image of the area; obtaining three-dimensional (3D) data based on perception data captured in the area; and refining the map based on the 3D data to generate a refined map of the area.

[0010] In another example, an apparatus for generating map data is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: generate a map of an area based on a plan-view image of the area; obtain three-dimensional (3D) data based on perception data captured in the area; and refine the map based on the 3D data to generate a refined map of the area.

[0011] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: generate a map of an area based on a plan-view image of the area; obtain three-dimensional (3D) data based on perception data captured in the area; and refine the map based on the 3D data to generate a refined map of the area.

[0012] In another example, an apparatus for generating map data is provided. The apparatus includes: means for generating a map of an area based on a plan-view image of the area; means for obtaining three-dimensional (3D) data based on perception data captured in the area; and means for refining the map based on the 3D data to generate a refined map of the area.

[0013] In some aspects, one or more of the apparatuses described herein is, can be part of, or can include an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a vehicle (or a computing device, system, or component of a vehicle) , a mobile device (e.g., a mobile telephone or so-called “smart phone” , a tablet computer, or other type of mobile device) , a smart or connected device (e.g., an Internet-of-Things (IoT) device) , a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television) , a robotics device or system, or other device. In some aspects, each apparatus can include an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each apparatus can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each apparatus can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each apparatus can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of the apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state) , and / or for other purposes.

[0014] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject  matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0015] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Illustrative examples of the present application are described in detail below with reference to the following figures:

[0017] FIG. 1 is a block diagram illustrating an example system for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure;

[0018] FIG. 2 is a block diagram illustrating another example system for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure;

[0019] FIG. 3 is a block diagram illustrating yet another example system for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure;

[0020] FIG. 4 includes an example image overlaid with a representation 3D data, according to various aspects of the present disclosure;

[0021] FIG. 5 includes an example plan-view image of an area, according to various aspects of the present disclosure;

[0022] FIG. 6 is a block diagram illustrating an example system for generating map data based on plan-view images, according to various aspects of the present disclosure;

[0023] FIG. 7 is a block diagram illustrating an example system for generating map data based on map data and 3D data, according to various aspects of the present disclosure;

[0024] FIG. 8 is a flow diagram illustrating an example process for using map data, in accordance with aspects of the present disclosure;

[0025] FIG. 9 is a flow diagram illustrating an example process for using map data, in accordance with aspects of the present disclosure;

[0026] FIG. 10 is a block diagram illustrating an example of a deep learning neural network that can be used to perform various tasks, according to some aspects of the disclosed technology;

[0027] FIG. 11 is a block diagram illustrating an example of a convolutional neural network (CNN) , according to various aspects of the present disclosure; and

[0028] FIG. 12 is a block diagram of an example transformer in accordance with some aspects of the disclosure;

[0029] FIG. 13 is a block diagram illustrating an example computing-device architecture of an example computing device which can implement the various techniques described herein.DETAILED DESCRIPTION

[0030] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.

[0031] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects will provide those skilled in the art with an enabling description for implementing an exemplary aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0032] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.

[0033] As mentioned above, perception data informs a driving systems (e.g., autonomous, semi-autonomous, or assisted driving systems, such as an advanced driver assistance system (ADAS) ) what area is drivable and what objects (e.g., road users, other vehicles, bikes, pedestrian, etc. ) are present and / or are moving in the environment around the vehicle. The driving system then makes decisions about how to move (e.g., slower, faster, stop, changing lanes,  turning, a path to take, etc. ) . But real-time perception data has limitations which might influence the vehicle safety for autonomous driving.

[0034] Map data can provide additional information beyond what perception data can ‘see. ’ For example, map data (e.g., a high-definition (HD) map) may include map points –three-dimensional coordinates of surfaces of roads at a sub-meter granularity. Map data (e.g., an HD map) may also include additional features such as lane markers, road signs, traffic lights, traffic signs, poles, etc. Driving systems may use map data (e.g., HD maps) to make determinations about steering, accelerating, braking, path planning, and / or to provide information to a driver, etc.

[0035] In the context of HD maps, the term “high” typically refers to the level of detail and accuracy of the map data. In some cases, an HD map may have a higher spatial resolution and / or level of detail as compared to a non-HD map. While there is no specific universally accepted quantitative threshold to define “high” in HD maps, several factors contribute to the characterization of the quality and level of detail of an HD map. Some key aspects considered in evaluating the “high” quality of an HD map include resolution, geometric accuracy, semantic information, dynamic data, and coverage. With regard to resolution, HD maps generally have a high spatial resolution, meaning they provide detailed information about the environment. The resolution can be measured in terms of meters per pixel or pixels per meter, indicating the level of detail captured in the map. With regard to geometric accuracy, an accurate representation of road geometry, lane boundaries, and other features can be important in an HD map. High-quality HD maps strive for precise alignment and positioning of objects in the real world. Geometric accuracy is often quantified using metrics such as root mean square error (RMSE) or positional accuracy. With regard to semantic information, HD maps include not only geometric data but also semantic information about the environment. This may include lane-level information, traffic signs, traffic signals, road markings, building footprints, and more. The richness and completeness of the semantic information contribute to the level of detail in the map. With regard to dynamic data, some HD maps incorporate real-time or near real-time updates to capture dynamic elements such as traffic flow, road closures, construction zones, and temporary changes. The frequency and accuracy of dynamic updates can affect the quality of the HD map. With regard to coverage, the extent of coverage provided by an HD map is another important factor. Coverage refers to the geographical area covered by the map. An HD map can cover a significant  portion of a city, region, or country. In general, an HD map may exhibit a rich level of detail, accurate representation of the environment, and extensive coverage.

[0036] A vehicle (or a computing system of the vehicle) may capture perception data (e.g., image data, light detection and ranging (LIDAR) data, and / or radio detection and ranging (RADAR) data) . The vehicle (e.g., in a process that may be referred to as online map generation) , or another computing device, such as a server (e.g., in a process that may be referred to as offline map generation) , may generate a map (e.g., an HD map) based on the perception data. For example, while in an area, the vehicle may capture images, LIDAR data and / or RADAR data including representations of objects (such as lane boundaries, road boundaries, pedestrian crossings, lane dividers, traffic lights, traffic signs, etc. ) in an area. The vehicle, or the other computing device, may determine points, polylines, and / or polygons to represent the objects in the area. The vehicle, or the other computing device, may store the points, polylines, and / or polygons in a map of the area (e.g., an HD map of the area) .

[0037] Additionally or alternatively, a vehicle (or a computing system of the vehicle) , or another computing system, may refine and / or update a map of an area (e.g., an HD map of the area) based on perception data captured in the area. For example, while in an area, the vehicle may capture perception data including representations of objects. The vehicle, or the other computing system, may determine points, polylines, and / or polygons to represent the objects in the area. The vehicle, or the other computing system, may compare the points, polylines, and / or polygons to points, polylines, and / or polygons of a previously-generated map. Where there are differences between the positions of the objects in the map and the positions of the objects as represented by the perception data, the vehicle, may refine and / or update the points, polylines, and / or polygons of the map.

[0038] Systems, apparatuses, methods (also referred to as processes) , and computer-readable media (collectively referred to herein as “systems and techniques” ) are described herein for generating, refining, updating, and / or using map data (e.g., HD map data) . For example, the systems and techniques described herein may generate maps (e.g., HD maps) by generating an initial map based on plan-view images (e.g., satellite and / or aerial images) and fusing this map with three-dimensional (3D) points based on perception data from one or more vehicles.

[0039] For example, a machine-learning model (e.g., a deep-learning neural network) may be trained to generate maps based on plan-view images and altitude information related to the plan- view images. In the present disclosure, the term “plan-view image” may refer to an orthographic projection of a 3D object from the position of a horizontal plane through the object. A plan-view image may be referred to as a “bird’s-eye-view” image. Images captured from a satellite, aerial vehicle, drone, weather balloon, etc. may be examples of plan-view images. In the present disclosure, the terms “image data, ” “plan-view image data, ” and like terms may include one or more images or plan-view images. The machine-learning model may be trained to generate maps (e.g., HD maps) , including elements such as lane boundaries, lane markings, crosswalks, etc. Once trained, the machine-learning model may be used to generate a map of an area based on one or more plan-view images of the area (e.g., plan-view image data) and altitude information for the area. The altitude information may indicate an altitude of the area, for example, relative to sea level.

[0040] Over time, such a map may become out of date. For example, the road may be changed by construction. Additionally, such a map may have gaps, for example, where trees or other obstacles occlude a bird’s-eye-view of the road. Further, such a map may lack some elements, such as street signs, based on the plan-view images being captured from above.

[0041] The systems and techniques may use 3D points based on perception data from one or more vehicles to update and / or enhance such a map of an area. For example, on a given day, a number of vehicles may travel in a given area. The number of vehicles may include sensors (e.g., cameras, LIDAR systems, and / or RADAR systems) . The vehicles may capture perception data in the area.

[0042] The vehicles may process the perception data and generate 3D data. The 3D data may represent points in the environment. The 3D data may include 3D meshes (e.g., representing surfaces, such as the road) , 3D polylines (e.g., representing lane lines and / or lane boundaries) , 3D polygons (e.g., representing traffic signs, crosswalks, or other lane markings) , 3D bounding boxes (e.g., representing traffic lights or poles) , and / or 3D points.

[0043] Additionally, the vehicles may tag the 3D data with location information (e.g., latitude and longitude coordinates) . In some cases, the vehicles may label the 3D data with semantic labels, for example, indicating what the 3D points represent (e.g., labels such as road surface, lane line, lane boundary, cross walk, traffic light, traffic sign, etc. ) .

[0044] The vehicles may provide the 3D data to a cloud computing devices. For example, the vehicles may transmit the perception data to the cloud computing device while the vehicles are travelling and / or in a batch at the end of the given day.

[0045] The systems and techniques may obtain the 3D data from the one or more vehicles that travelled in the area. In some cases, the systems and techniques may obtain a number (e.g., tens, hundreds, thousands, or more) of sets of 3D data for a given area from a number of respective vehicles that travelled in the given area. In other cases, the systems and techniques may obtain 3D data from one vehicle that travelled in the area.

[0046] The systems and techniques may align 3D data with the map (e.g., the map generated based on the plan-view image) based on the location metadata of the 3D data. Additionally or alternatively, the systems and techniques may align the perception data with the map based on elements common to the map and the 3D data. For example, the systems and techniques may determine a first alignment (e.g, a rough alignment) between the 3D data and the map based on the location metadata. Further, the systems and techniques may determine a second alignment (e.g., a more precise alignment) based on a correspondence between elements of the map (e.g., lane boundaries) and the elements as the elements are represented in the 3D data.

[0047] The systems and techniques may use a machine-learning model (e.g., a deep-learning machine-learning model) to fuse the 3D data with the map to refine, update and / or enhance the map with elements represented in the 3D data (such as virtual lanes and boundaries, traffic signs, traffic lights, common speed, and / or observed common behaviors for a given location, etc. ) .

[0048] The machine-learning model may include a perception branch that may encode 3D data into a 3D feature space. Additionally, the machine-learning model may include a plan-view branch that may encode the plan-view images, or a map generated based on plan-view images, into a 3D feature space. The machine-learning model may include a combiner that may combine (e.g., fuse) the perception-based 3D features (in the 3D feature space) with the plan-view -based features (in the 3D feature space) . Further, the machine-learning model may include a decoder that may decode the features (in the 3D feature space) to generate a final map. The final map may be updated and / or enhanced relative to the initial map generated based on the plan-view images.

[0049] Various aspects of the application will be described with respect to the figures below.

[0050] FIG. 1 is a block diagram illustrating an example system 100 for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure. system 100 include a vehicle 102 and a map system 112. In general, a perception module 104 of vehicle 102 may capture perception data 106. A harvesting module 108 may generate 3D data 110 based on perception data 106. Vehicle 102 may provide 3D data 110 to map system 112. Map system 112 may generate map data 118 based on 3D data 110 and provide map data 118 to vehicle 102. A map module 120 of map module 120 may provide map data 118 to a localization module 122 and drive-policy module 126. Localization module 122 may provide location data 124 to drive-policy module 126. Drive-policy module 126 may make determinations regarding driving of vehicle 102 based on map data 118 and location data 124.

[0051] Vehicle 102 may be, or may include, a vehicle including a computing system, such as an autonomous, semi-autonomous, or assisted driving system. In the present disclosure, references to operations performed by a vehicle, (e.g., vehicle 102) may be performed by computing system of the vehicle. Elements of vehicle 102 may be part of and / or be implemented by computing system of vehicle 102.

[0052] Perception module 104 may generate perception data 106. Perception module 104 may be, or may include, a number of cameras that may capture image data, a LIDAR system that may capture LIDAR data, and / or a RADAR system that may capture RADAR data. Perception module 104 may generate perception data 106 to be, or to include, image data, LIDAR-based point-cloud data , and / or RADAR-based point-cloud data.

[0053] Harvesting module 108 may generate 3D data 110 based on perception data 106. For example, harvesting module 108 may take perception data 106 as input and reconstruct objects represented in the perception data 106 into a 3D space. In some aspects, harvesting module 108 may employ curve fitting (e.g., to generate polylines) . In some aspects, harvesting module 108 may use object-detection techniques (e.g., a 3D object detector) and / or image-segmentation techniques.

[0054] 3D data 110 may include 3D points, 3D bounding boxes, 3D meshes, 3D polygons, and / or 3D polylines, sizes (e.g., of bounding boxes) , orientations (e.g., of bounding boxes) , etc. 3D data 110 may include labels (e.g., semantic labels) identifying the 3D points, 3D bounding boxes, 3D meshes, 3D polygons, and / or 3D polylines, for example, as points of a road surface, a traffic light, a traffic sign, a crosswalk, a stop line in the road, an arrow in the road, a pole, a lane  boundary, a lane line, a lane divider, etc. Additionally, 3D data 110 may include location metadata, for example, describing a point of origin of the 3D points in a common reference system (e.g., latitude and longitude) . In the present disclosure, the “term 3D points” may include 3D points (e.g., individual point or points of collections) , 3D bounding boxes, 3D meshes, 3D polygons, and / or 3D polylines, sizes (e.g., of bounding boxes) , orientations (e.g., of bounding boxes) .

[0055] Vehicle 102 may transmit 3D data 110 to map system 112. For example, vehicle 102 may include a communication module (not illustrated in FIG. 1) that vehicle 102 may use to transmit data to map system 112, for example, through a wireless communication network.

[0056] Map system 112 may, according to various aspects of the present disclosure, process 3D data 110 to generate, refine, update, or enhance map data 132 of map system 112 to generate map data 118. In general, server 114 may generate and / or update map data 118 and cloud module may determine how to provide map data 118 to vehicles including vehicle 102. Map system 112 may provide (e.g., transmit via the wireless communication network) map data 118 to vehicle 102.

[0057] Map module 120 may receive and store map data 118. Additionally, map module 120 may provide map data 118 to localization module 122. Localization module 122 may determine a location (e.g., in a reference coordinate system such as latitude and longitude) of vehicle 102. For example, localization module 122 may include a global positioning system (GPS) or global navigation satellite system (GNSS) system that may determine the location of vehicle 102. In some aspects, localization module 122 may provide location data 124 to harvesting module 108 and harvesting module 108 may use location data 124 as location metadata in generating 3D data 110.

[0058] Drive-policy module 126 may control vehicle 102 and / or provide information to a driver of vehicle 102 based on perception data 106, map data 118, and / or location data 124. Perception module 104 may provide perception data 106 to drive-policy module 126, map module 120 may provide map data 118 to drive-policy module 126, and localization module 122 may provide location data 124 to drive-policy module 126. Drive-policy module 126 may adjust driving parameters of vehicle 102, such as parameters relating to a path of the vehicle, steering parameters, braking parameters, and / or lane-change parameters. Additionally or alternatively,  drive-policy module 126 may cause vehicle 102 to present information (either visually, audibly, or haptically) to a driver of vehicle 102.

[0059] Maps (such as map data 118) play an important role in autonomous driving and ADAS areas (e.g., by drive-policy module 126) . Hence creating maps with high degrees of accuracy and freshness is an important task for autonomous driving. The systems and techniques may take advantage of 3D data from a number of vehicles to generate, update, refine, and / or enhance map data (e.g., map data 118) .

[0060] FIG. 2 is a block diagram illustrating an example system 200 for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure. In general, a perception module 208 of a vehicle 202 may process image data 204 and sensor data 206 to generate objects 210. A keypoint module 212 may process image data 204 to generate keypoints 214. A harvesting module 216 may process objects 210, keypoints 214, location data 218, inertial data 220, sensor data 206, and map 238 to generate map snippets 222. Vehicle 202 may transmit map snippets 222 to map system 224. Map system 224 may process map snippets 222 to generate map updates 234. Map system 224 may transmit map updates 234 to vehicle 202. Provisioning module 236 of vehicle 202 may generate map 238 based on map updates 234 and sensor data 206.

[0061] Vehicle 202 may be the same as, may be substantially similar to, and / or may perform the same, or substantially the same, operations as vehicle 102 of FIG. 1.

[0062] Image data 204 may be, or may include, one or more images captured by one or more cameras of a vehicle (e.g., vehicle 102) . Image data 204 may represent an area in which the vehicle is traveling. Sensor data 206 may be, or may include, LIDAR data, RADAR data, and / or other type of sensor data. Image data 204 and / or sensor data 206 may the same as, or may be substantially similar to, perception data 106 of FIG. 1.

[0063] Perception module 208 may process image data 204 and sensor data 206 to generate objects spatial point pattern (SPP) 210. Objects 210 may be, or may include, a number of 3D points representing objects represented by image data 204 and / or sensor data 206. Perception module 208 may be, or may include, a 3D object detector.

[0064] Keypoint module 212 may process image data 204 to generate keypoints 214. Keypoints 214 may be, or may include, 3D points based on keypoints of image data 204. For example,  keypoint module 212 may identify keypoints (e.g., visually distinctive points) of image data 204 and project the keypoints to generate keypoints 214.

[0065] Location data 218 may be, or may include, information indicative of a location of vehicle 202. Location data 218 may be, or may include, coordinate in a reference coordinate system (e.g., latitude and longitude) .

[0066] Inertial data 220 may be, or may include, data indicative of acceleration measured, for example, by an inertial measurement unit of vehicle 202. Inertial data 220 may include acceleration data relative to three independent axes (e.g., an x-axis, a y-axis, and a z-axis) .

[0067] Harvesting module 216 may process objects 210, keypoints 214, location data 218, inertial data 220, and sensor data 206 to generate map snippets 222. Harvesting module 216 may be the same as, may be substantially similar to, and / or may perform the same, or substantially the same, operations as harvesting modules 108 of FIG. 1.

[0068] Map snippets 222 may be, or may include, a number of 3D points, 3D meshes, 3D polygons, and / or 3D polylines. Map snippets 222 may the same as, or may be substantially similar to, 3D data 110 of FIG. 1.

[0069] Vehicle 202 may provide map snippets 222 to map system 224. Map system 224 may update map data 230 of map system 224. For example, server 226 may generate map updates 234 based on map snippets 222. Map system 224 may be the same as, may be substantially similar to, and / or may perform the same, or substantially the same, operations as map system 112 of FIG. 1.

[0070] Additionally, map system 224 may determine configuration data 232 and server 226 may provide configuration data 232 to harvesting module 216. Harvesting module 216 may capture, generate, and / or provide map snippets 222 to map system 224 based on configuration data 232. For example, configuration data 232 may instruct vehicle 202 to capture image and / or sensor data in a specific area and to generate map snippets 222 based on the specific area and to provide the map snippets to map system 224. As another example, configuration data 232 may instruct harvesting module 216 to process longer data for some scenarios.

[0071] Server 226 may provide map updates 234 to provisioning module 236. Provisioning module 236 may store map updates 234 and / or update a map of provisioning module 236 based on map updates 234 and sensor data 206. Provisioning module 236 may be the same as, may be  substantially similar to, and / or may perform the same, or substantially the same, operations as map module 120 of FIG. 1. Provisioning module 236 may provide map 238 to harvesting module 216.

[0072] A driving system of vehicle 202 may control or adjust driving parameters of vehicle 202 based on map 238. For example, drive-policy module 126 may adjust driving parameters of vehicle 202 based on map 238.

[0073] FIG. 3 is a block diagram illustrating an example system 300 for generating, enhancing, updating, refining, and / or using map data, according to various aspects of the present disclosure. In general, a number of vehicles (e.g., vehicle 302, vehicle 312, and vehicle 322) may each capture perception data (e.g., perception data 304, perception data 314, and perception data 324) , generate 3D data (e.g., 3D data 308, 3D data 318, and 3D data 328) based on the perception data, and provide the 3D data to map system 332 (e.g., via network 330) . Map system 332 may aggregate the received 3D data as vehicle reports 334. Additionally, map system 332 may obtain plan-view images 336. Map generator 338 may generate map 340 based on vehicle reports 334 and plan-view images 336.

[0074] Each of vehicle 302, vehicle 312, and vehicle 322 may be the same as, may be substantially similar to, and / or may perform the same, or substantially the same, operations as vehicle 102 of FIG. 1 and / or vehicle 202 of FIG. 2. Each of vehicle 302, vehicle 312, and vehicle 322 may generate and provide 3D data in substantially the way described with regard to vehicle 102 of FIG. 1 generating 3D data 110. For example, each of vehicle 302, vehicle 312, and vehicle 322 may include or implement perception module 104 and harvesting module 108. Additionally each of vehicle 302, vehicle 312, and vehicle 322 may generate and provide 3D data in substantially the way described with regard to vehicle 202 of FIG. 1 generating map snippets 222. For example, each of vehicle 302, vehicle 312, and vehicle 322 may include or implement perception module 208, keypoint module 212, and harvesting module 216.

[0075] Map system 332 may provide map 340 to each of vehicle 302, vehicle 312, and vehicle 322. Each of vehicle 302, vehicle 312, and vehicle 322 may use map 340 to adjust driving parameters (e.g., as described with regard to drive-policy module 126 of FIG. 1) .

[0076] Map system 332 may be the same as, may be substantially similar to, and / or may perform the same, or substantially the same, operations as map system 112 of FIG. 1 and / or map  system 224 of FIG. 2. For example, map system 332 may received 3D data 308, 3D data 318, and 3D data 328 and generate map 340 based on 3D data 308, 3D data 318 and 3D data 328.

[0077] Map 340 may be a 3D map of the area (e.g., an HD map) represented by plan-view images 336, perception data 304, perception data 314, and perception data 324. For example, map 340 may include 3D points, 3D bounding boxes, 3D meshes, 3D polylines, and / or 3D polygons representing the area.

[0078] Map generator 338 of map system 332 may obtain plan-view images 336. Plan-view images 336 may be, or may include, images of an area captured from above the area. Plan-view images 336 may be captured by a satellite orbiting the Earth. Additionally or alternatively, plan-view images 336 may be captured by an airplane, drone, weather balloon, etc. above the area.

[0079] Map generator 338 may generate an initial map of the area based on plan-view images 336. For example, map generator 338 may include a machine-learning model trained to generate maps based on plan-view images.

[0080] Map generator 338 may enhance, refine, or update the initial map based on vehicle reports 334. For example, plan-view images 336 may represent objects as viewed from above. Traffic lights and / or traffic signs, as examples, may be viewed better from level of the traffic lights and / or traffic signs than from above. For example, it may not be possible to read a traffic sign from above. Additionally, objects in the area may be occluded by other objects when viewed from above. For example, a tree near a road may occlude a bird’s-eye view of lane boundaries and / or sidewalks. Map generator 338 may enhance the initial map by adding 3D points from vehicle reports 334 (which includes 3D data 308, 3D data 318, and 3D data 328) .

[0081] Additionally, a resolution of perception data 304 (and a spatial resolution of 3D data 308 which is based on perception data 304) may be higher than a resolution of plan-view images 336. Accordingly, map generator 338 may increase a spatial resolution of map 340 using 3D data 308.

[0082] Additionally, plan-view images 336 may be relatively old. For example, plan-view images 336 may be captured one day. A newer plan-view image may not be available for a year or more. Vehicle reports 334 may include 3D data from vehicles that arrives more frequently, for example, daily. Map generator 338 may update map data by adjusting map 340 to reflect changes to the road (e.g., resulting from construction) .

[0083] Vehicle reports 334 may include 3D data from any number of vehicles. Vehicle reports 334 may be continually (or often) updated with new 3D data. Vehicle reports 334 may include data from much of the area (e.g., based on vehicles travelling in much of the area) . Thus, vehicle reports 334 may have good coverage of the area.

[0084] Map generator 338 may fuse 3D data based on plan-view images 336 with 3D data of vehicle reports 334 to generate map 340. Thus, map generator 338 may generate map 340 that is fresh (e.g., reflects current or near-current) .

[0085] Additionally, map generator 338 may be able to generate map 340 in a way that produces more accurate results and / or is less computationally expensive (and / or otherwise expensive) than other techniques for generating maps (e.g., HD maps) . For example, other techniques may use perception data to directly generate maps. The maps may be only as good as the perception data. In contrast, map generator 338 may generate map 340 based on 3D data from a number of vehicles and may take advantage of the perception data of the number of vehicles, which may result in a more accurate map. Additionally, it may be computationally expensive (e.g., in terms of power and / or processing time) to generate a map based on perception data directly. It may be less computationally expensive to generate a map based on plan-view images 336 and to refine the map based on vehicle reports 334. Thus, map generator 338 may generate fresh, low-cost maps (e.g., HD maps) .

[0086] FIG. 4 includes an example image 400 overlaid with a representation (e.g., a projection) of 3D data, according to various aspects of the present disclosure. Image 400 is an example of perception data. Image 400 may be captured by a vehicle traveling on a road in an area. Image 400 is overlaid with 3D data representative of features of the area. For example, polyline 402 is a polyline representing a road edge, polyline 404 is a polyline representing a lane edge, polyline 406 is a polyline representing a lane boundary, polyline 408 is a polyline representing a lane center (e.g., of the lane in which the vehicle is traveling) , polyline 410 is a polyline representing a lane boundary, polyline 412 is a polyline representing a road edge, polygon 414 is a polygon representing a lane divider.

[0087] FIG. 5 includes an example plan-view image 500 of an area, according to various aspects of the present disclosure. Plan-view image 500 may be captured from a bird's-eye-view (BEV) perspective. Plan-view image 500 may be highly accurate. A system, such as map generator 338  may align 3D data based on perception data (e.g., 3D data 308) with a map based on plan-view image 500.

[0088] FIG. 6 is a block diagram illustrating an example system 600 for generating map data 610 based on plan-view images 602, according to various aspects of the present disclosure. In general, a map encoder 604 may process plan-view images 602 to generate features 606 and a map decoder 608 may process features 606 to generate map data 610. Map generator 338 of FIG. 3 may include or implement system 600.

[0089] Plan-view images 602 may be, or may include, one or more images of an area. Plan-view images 602 may be the same as, or may be substantially similar to, plan-view images 336 of FIG. 3 and / or plan-view image 500 of FIG. 5.

[0090] Map encoder 604 may be, or may include, an image-encoder network, such as a deep residual network (ResNet) . Map encoder 604 may be trained to generate features based on image data and altitude information (not illustrated in FIG. 6) . Features 606 may be a latent-space representation of plan-view images 602 generated by map encoder 604.

[0091] Map decoder 608 may be, or may include, a decoder network, such as a transformer. Map decoder 608 may be trained to generate map data based on features encoded by an encoder. Map data 610 may be 3D map of the area (e.g., an HD map) represented by plan-view images 602.

[0092] FIG. 7 is a block diagram illustrating an example system 700 for generating map data 716 based on map data 702 and 3D data 704, according to various aspects of the present disclosure. In general, a coarse-alignment module 706 may align map data 702 with 3D data 704 and fine-alignment module 708 may further align map data 702 with 3D data 704. Map fuser 710 may combine (e.g., fuse) map data 702 with 3D data 704 to generate features 712. Map decoder 714 may decode features 712 to generate map data 716. Map generator 338 of FIG. 3 may include or implement system 700.

[0093] Map data 702 may be the same as, or may be substantially similar to, map data 610 of FIG. 6. For example, map generator 338 may implement system 600 to generate map data 610 based on plan-view images 336. Further, map generator 338 may provide map data 610 to system 700 as map data 702.

[0094] 3D data 704 may be the same as, or may be substantially similar to, at least one of vehicle reports 334 of FIG. 3. For example, 3D data 704 may be the same as, or may be substantially similar to, any of 3D data 308, 3D data 318, or 3D data 328 of FIG. 3.

[0095] Coarse-alignment module 706 may align map data 702 with 3D data 704 based on location metadata associated with 3D data 704 and location metadata associated with map data 702. For example, map data 702 may include coordinates (e.g., in a reference coordinate system, such as latitude and longitude) . Similarly, 3D data 704 may include or be associated with metadata including coordinates (e.g., in the reference coordinate system) . The coordinates may indicate a location from which perception data (on which 3D data 704 is based) was captured. Coarse-alignment module 706 may determine a position of 3D points, 3D polylines, 3D polygons, 3D meshes, and / or 3D bounding boxes (which may be collectively referred to as 3D points) of 3D data 704 relative to map data 702 based on the location information of each of map data 702 and 3D data 704.

[0096] Fine-alignment module 708 may align map data 702 with 3D data 704 based on 3D points of map data 702 and 3D points of 3D data 704 (e.g., based on road elements and / or structure information) . For example, coarse-alignment module 706 may compare 3D points of 3D data 704 with 3D points of map data 702 to determine where the 3D points of 3D data 704 fit within map data 702. Fine-alignment module 708 may begin the comparison based on the alignment determined by coarse-alignment module 706.

[0097] Map fuser 710 may combine map data 702 with 3D data 704 to generate features 712. In some aspects, map fuser 710 may insert 3D points of 3D data 704 into map data 702. Map fuser 710 may be a machine-learning model trained to combine 3D points based on perception data with 3D points based on plan-view images.

[0098] In some aspects, map fuser 710 may include a branch that encodes the 3D data 704 (e.g., which may be based on perception data) into a 3D feature space, and a branch that encodes map data 702 into a 3D feature space. Map fuser 710 may then fuse these two together to generate map features 712 and provide map features 712 to map decoder 714.

[0099] Map decoder 714 may decode features 712 to generate map data 716. Map decoder 714 may be, or may include, a transformer machine-learning model.

[0100] To provide a high-level description of the systems and techniques of the present disclosure, map system 332 may obtain plan-view images 336. System 600 (which may be  included or implemented in map generator 338) may generate map data 610 based on plan-view images 336 and altitude information. Map data 610 may include elements on the road surface and can be used for aligning perception results from vehicles. As plan-view images 336 may be out of date, map data 610 has a possibility to be not fresh.

[0101] Map system 332 may collect vehicle reports 334 based on perception data from a number of vehicles (e.g., vehicle 302, vehicle 312, vehicle 322) . Coarse-alignment module 706 of system 700 (which may be included or implemented in map generator 338) may using coordinate information to align map data 702 (which may be an example of map data 610) with 3D data 704 (which may be an example instance of 3D data of vehicle reports 334) . Additionally, fine-alignment module 708 of system 700 may align map data 702 with 3D data 704 based on the road structure information from 3D data 704 and map data 702. Thus, fine-alignment module 708 may precisely align 3D data 704 with map data 702.

[0102] Map fuser 710, which may be, or may include, a deep-learning model, may fuse map data 702 with 3D data 704 (which may be an example instance of vehicle reports 334, which may be based on fresh perception results) . Map decoder 714 may decode features 712 to generate map data 716. Thus, map data 716 may update (e.g., regularly) and / or enhance a 3D map. For example, system 700 may generate map data 716 with additional elements (e.g., behavior, virtual lanes and boundaries, traffic signs, traffic lights, common speed, etc. ) , where the additional elements are not part of map data 702.

[0103] FIG. 8 is a flow diagram illustrating an example process 800 for using map data, in accordance with aspects of the present disclosure. One or more operations of process 800 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the one or more operations of process 800. The one or more operations of process 800 may be implemented as software components that are executed and run on one or more processors.

[0104] At block 802, a computing device (or one or more components thereof) may process perception data representative of an area to generate three-dimensional (3D) data. For example,  harvesting module 108 may process perception data 106 to generate 3D data 110. As another example, harvesting module 216 may process objects 210, keypoints 214, location data 218, inertial data 220, and / or sensor data 206 to generate map snippets 222.

[0105] In some aspects, the perception data may be, or may include, at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data. For example, perception data 106 and / or sensor data 206 may be, or may include, RADAR data and / or LIDAR data.

[0106] At block 804, the computing device (or one or more components thereof) may determine location information associated with the 3D data. For example, localization module 122 may determine location data 124.

[0107] At block 806, the computing device (or one or more components thereof) may transmit the 3D data and the location information to a map server. For example, vehicle 102 may transmit 3D data 110, tagged with location data 124 to map system 112. As another example, vehicle 202 may transmit map snippets 222, which may include location data, to map system 224.

[0108] In some aspects, the computing device (or one or more components thereof) may label the 3D data; and transmit labels of the 3D data to the map server. For example, vehicle 102 may label 3D data 110 (e.g., with semantic labels) before providing 3D data 110 to map system 112.

[0109] At block 808, the computing device (or one or more components thereof) may receive a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data. For example, vehicle 102 may obtain map data 118 from map system 112. As another example, vehicle 202 may obtain map updates 234 from map system 224.

[0110] In some aspects, the map may be based on a plurality of sets of 3D data generated based on a corresponding plurality of sets of perception data captured by a corresponding plurality of vehicles. For example, map system 332 may obtain 3D data 308, 3D data 318, and 3D data 328 from vehicle 302, vehicle 312, and vehicle 322 respectively. Map system 332 may generate map 340 based on 3D data 308, 3D data 318, and 3D data 328.

[0111] In some aspects, the map server may generate map data based on the plan-view-image data of the area; and refine the map data based on the 3D data. For example, system 600 may  generate map data 610 based on plan-view images 602. Further, system 700 may refine map data 702 based on 3D data 704.

[0112] In some aspects, the map server may align the 3D data with map data based on the location information to generate aligned 3D data; and combine the aligned 3D data with the map data. For example, coarse-alignment module 706 and / or fine-alignment module 708 may align 3D data 704 with map data 702.

[0113] At block 810, the computing device (or one or more components thereof) may adjust an operating parameter of a vehicle based on the map. For example, vehicle 102 may adjust an operating parameter of vehicle 102. As another example, vehicle 202 may adjust an operating parameter of vehicle 202.

[0114] In some aspects, the computing device (or one or more components thereof) may be, or may include, a computing system of a vehicle. For example, process 800 may be performed by a computing system of a vehicle, such as vehicle 102.

[0115] In some aspects, the operating parameter is associated with at least one of a path for the vehicle to travel, a braking parameter for operating one or more brakes of the vehicle, a steering parameter for steering the vehicle, a lane-change parameter for causing the vehicle to navigate from a first lane to a second lane, or displaying information using a user interface of the vehicle.

[0116] FIG. 9 is a flow diagram illustrating an example process 900 for using map data, in accordance with aspects of the present disclosure. One or more operations of process 900 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the one or more operations of process 900. The one or more operations of process 900 may be implemented as software components that are executed and run on one or more processors.

[0117] At block 902, a computing device (or one or more components thereof) may generate a map of an area based on a plan-view image of the area. For example, system 600 may generate map data 610 based on plan-view images 602.

[0118] At block 904, the computing device (or one or more components thereof) may obtain three-dimensional (3D) data based on perception data captured in the area. For example, system 700 may obtain 3D data 704.

[0119] In some aspects, the perception data may be, or may include, at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data. For example, map data 702 may be, or may include, LIDAR data and / or RADAR data.

[0120] At block 906, the computing device (or one or more components thereof) may refine the map based on the 3D data to generate a refined map of the area. Further, system 700 may refine map data 702 based on 3D data 704.

[0121] In some aspects, to refine the map, the computing device (or one or more components thereof) may align the 3D data with the map based on location information associated with the 3D data to generate aligned 3D data; and combine the aligned 3D data with the map. For example, system 700 may align map data 702 with 3D data 704 based on location data associated with 3D data 704. Further, map fuser 710 may combine map data 702 with 3D data 704.

[0122] In some aspects, the 3D data are aligned with the map further based on a comparison of the 3D data with map points of the map. For example, coarse-alignment module 706 may align map data 702 with 3D data 704 based on a comparison of map data 702 and 3D data 704.

[0123] In some aspects, the computing device (or one or more components thereof) may obtain a plurality of sets of 3D data from a corresponding plurality of vehicles; and refine the map based on the plurality of sets of 3D data. For example, map system 332 may obtain 3D data 308, 3D data 318, and 3D data 328 from vehicle 302, vehicle 312, and vehicle 322 respectively. Map system 332 may generate map 340 based on 3D data 308, 3D data 318, and 3D data 328.

[0124] In some examples, as noted previously, the methods described herein (e.g., process 800 of FIG. 8, process 900 of FIG. 9, and / or other methods described herein) can be performed, in whole or in part, by a computing device or apparatus. In one example, one or more of the methods can be performed by system 100 of FIG. 1, vehicle 102 of FIG. 1, map system 112 of FIG. 1, system 200 of FIG. 2, vehicle 202 of FIG. 2, map system 224 of FIG. 2, system 300 of FIG. 3, vehicle 302 of FIG. 3, vehicle 312 of FIG. 3, vehicle 322 of FIG. 3, map system 332 of FIG. 3, system 600 of FIG. 6, system 700 of FIG. 7, or by another system or device. In another example, one or more of the methods (e.g., process 900, and / or other methods described herein) can be performed, in whole or in part, by the computing-device architecture 1300 shown in FIG. 13. For  instance, a computing device with the computing-device architecture 1300 shown in FIG. 13 can include, or be included in, the components of the system 100 of FIG. 1, vehicle 102 of FIG. 1, map system 112 of FIG. 1, system 200 of FIG. 2, vehicle 202 of FIG. 2, map system 224 of FIG. 2, system 300 of FIG. 3, vehicle 302 of FIG. 3, vehicle 312 of FIG. 3, vehicle 322 of FIG. 3, map system 332 of FIG. 3, system 600 of FIG. 6, system 700 of FIG. 7, and can implement the operations of process 800, process 900, and / or other process described herein. In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component (s) that are configured to carry out the steps of processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component (s) . The network interface can be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.

[0125] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs) , digital signal processors (DSPs) , central processing units (CPUs) , and / or other suitable electronic circuits) , and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.

[0126] Process 800, process 900, and / or other process described herein are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0127] Additionally, process 800, process 900, and / or other process described herein can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0128] As noted above, various aspects of the present disclosure can use machine-learning models or systems.

[0129] FIG. 10 is an illustrative example of a neural network 1000 (e.g., a deep-learning neural network) that can be used to implement machine-learning based image encoding, map encoding, feature decoding, feature segmentation, implicit-neural-representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc. ) , feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, neural network 1000 may be an example of, or can implement, perception module 208 of FIG. 2, keypoint module 212 of FIG. 2, map encoder 604 of FIG. 6, map decoder 608 of FIG. 6, map fuser 710 of FIG. 7, and / or map decoder 714 of FIG. 7.

[0130] An input layer 1002 includes input data. In one illustrative example, input layer 1002 can include data representing image data 204 of FIG. 2, sensor data 206 of FIG. 6 plan-view images 602 of FIG. 6, features 606 of FIG. 6, features 712 of FIG. 7. Neural network 1000 includes multiple hidden layers, for example, hidden layers 1006a, 1006b, through 1006n. The hidden layers 1006a, 1006b, through hidden layer 1006n include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. Neural network 1000 further includes an output layer 1004 that provides an output resulting from the processing performed by the hidden layers 1006a, 1006b, through 1006n. In one illustrative example, output layer 1004 can provide objects 210 of FIG. 2, keypoint module 212 of FIG. 2, features 606 of FIG. 6, map data 610 of FIG. 6, features 712 of FIG. 7, map data 716 of FIG. 7.

[0131] Neural network 1000 may be, or may include, a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated  with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, neural network 1000 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some cases, neural network 1000 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.

[0132] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of input layer 1002 can activate a set of nodes in the first hidden layer 1006a. For example, as shown, each of the input nodes of input layer 1002 is connected to each of the nodes of the first hidden layer 1006a. The nodes of first hidden layer 1006a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 1006b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 1006b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 1006n can activate one or more nodes of the output layer 1004, at which an output is provided. In some cases, while nodes (e.g., node 1008) in neural network 1000 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.

[0133] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of neural network 1000. Once neural network 1000 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset) , allowing neural network 1000 to be adaptive to inputs and able to learn as more and more data is processed.

[0134] Neural network 1000 may be pre-trained to process the features from the data in the input layer 1002 using the different hidden layers 1006a, 1006b, through 1006n in order to provide the output through the output layer 1004. In an example in which neural network 1000 is used to identify features in images, neural network 1000 can be trained using training data that includes both images and labels, as described above. For instance, training images can be input into the network, with each training image having a label indicating the features in the images  (for the feature-segmentation machine-learning system) or a label indicating classes of an activity in each image. In one example using object classification for illustrative purposes, a training image can include an image of a number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0 0] .

[0135] In some cases, neural network 1000 can adjust the weights of the nodes using a training process called backpropagation. As noted above, a backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until neural network 1000 is trained well enough so that the weights of the layers are accurately tuned.

[0136] For the example of identifying objects in images, the forward pass can include passing a training image through neural network 1000. The weights are initially randomized before neural network 1000 is trained. As an illustrative example, an image can include an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like) .

[0137] As noted above, for a first training iteration for neural network 1000, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1) . With the initial weights, neural network 1000 is unable to determine low-level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a cross-entropy loss. Another example of a loss function includes the mean squared error (MSE) , defined as Etotal = Σ 1 / 2 (target -output) 2. The loss can be set to be equal to the value of Etotal.

[0138] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. Neural network 1000 can  perform a backward pass by determining which inputs (weights) most contributed to the loss of the network and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as w = wi -η dL  / dW, where w denotes a weight, wi denotes the initial weight, and η denotes a learning rate. The learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.

[0139] Neural network 1000 can include any suitable deep network. One example includes a convolutional neural network (CNN) , which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling) , and fully connected layers. Neural network 1000 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs) , a Recurrent Neural Networks (RNNs) , among others.

[0140] FIG. 11 is an illustrative example of a convolutional neural network (CNN) 1100. The input layer 1102 of the CNN 1100 includes data representing an image or frame. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like) . The image can be passed through a convolutional hidden layer 1104, an optional non-linear activation layer, a pooling hidden layer 1106, and fully connected layer 1108 (which fully connected layer 1108 can be hidden) to get an output at the output layer 1110. While only one of each hidden layer is shown in FIG. 11, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 1100. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.

[0141] The first layer of the CNN 1100 can be the convolutional hidden layer 1104. The convolutional hidden layer 1104 can analyze image data of the input layer 1102. Each node of  the convolutional hidden layer 1104 is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 1104 can be considered as one or more filters (each filter corresponding to a different activation or feature map) , with each convolutional iteration of a filter being a node or neuron of the convolutional hidden layer 1104. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28×28 array, and each filter (and corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 1104. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptive field in the input image. Each node of the convolutional hidden layer 1104 will have the same weights and bias (called a shared weight and a shared bias) . For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for an image frame example (according to three color components of the input image) . An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.

[0142] The convolutional nature of the convolutional hidden layer 1104 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 1104 can begin in the top-left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 1104. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top-left corner of the input image array) . The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 1104. For example, a filter can be moved by a step amount (referred to as a stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 1104.

[0143] The mapping from the input layer to the convolutional hidden layer 1104 is referred to as an activation map (or feature map) . The activation map includes a value for each node representing the filter results at each location of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a stride of 1) of a 28 x 28 input image. The convolutional hidden layer 1104 can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 11 includes three activation maps. Using three activation maps, the convolutional hidden layer 1104 can detect three different kinds of features, with each feature being detectable across the entire image.

[0144] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 1104. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f (x) = max (0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the non-linear properties of the CNN 1100 without affecting the receptive fields of the convolutional hidden layer 1104.

[0145] The pooling hidden layer 1106 can be applied after the convolutional hidden layer 1104 (and after the non-linear hidden layer when used) . The pooling hidden layer 1106 is used to simplify the information in the output from the convolutional hidden layer 1104. For example, the pooling hidden layer 1106 can take each activation map output from the convolutional hidden layer 1104 and generates a condensed activation map (or feature map) using a pooling function. Max-pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 1106, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 1104. In the example shown in FIG. 11, three pooling filters are used for the three activation maps in the convolutional hidden layer 1104.

[0146] In some examples, max-pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a stride (e.g., equal to a dimension of the filter, such as a stride of 2) to an activation map output from the convolutional hidden layer 1104. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around.  Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes in the previous layer (with each node being a value in the activation map) . For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max-pooling filter is applied to an activation filter from the convolutional hidden layer 1104 having a dimension of 24x24 nodes, the output from the pooling hidden layer 1106 will be an array of 12x12 nodes.

[0147] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2×2 region (or other suitable region) of an activation map (instead of computing the maximum values as is done in max-pooling) and using the computed values as an output.

[0148] The pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image and discards the exact positional information. This can be done without affecting results of the feature detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max-pooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 1100.

[0149] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 1106 to every one of the output nodes in the output layer 1110. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 1104 includes 3×24×24 hidden feature nodes based on application of a 5×5 local receptive field (for the filters) to three activation maps, and the pooling hidden layer 1106 includes a layer of 3×12×12 hidden feature nodes based on application of max-pooling filter to 2×2 regions across each of the three feature maps. Extending this example, the output layer 1110 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 1106 is connected to every node of the output layer 1110.

[0150] The fully connected layer 1108 can obtain the output of the previous pooling hidden layer 1106 (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 1108  can determine the high-level features that most strongly correlate to a particular class and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 1108 and the pooling hidden layer 1106 to obtain probabilities for the different classes. For example, if the CNN 1100 is being used to predict that an object in an image is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the object, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person) .

[0151] In some examples, the output from the output layer 1110 can include an M-dimensional vector (in the prior example, M=10) . M indicates the number of classes that the CNN 1100 has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the M-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0] , the vector indicates that there is a 5%probability that the image is the third class of object (e.g., a dog) , an 80%probability that the image is the fourth class of object (e.g., a human) , and a 15%probability that the image is the sixth class of object (e.g., a kangaroo) . The probability for a class can be considered a confidence level that the object is part of that class.

[0152] FIG. 12 is a block diagram of an example transformer in accordance with some aspects of the disclosure. In a convolutional neural network (CNN) model, the number of operations required to relate signals from two arbitrary input or output positions grows in the distance between positions, which makes learning dependencies at different distant positions challenging for a CNN model. A transformer 1200 reduces the operations of learning dependencies by using an encoder 1210 and a decoder 1230 that implement an attention mechanism at different positions of a single sequence to compute a representation of that sequence. An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function of the query with the corresponding key.

[0153] In one example of a transformer, the encoder 1210 is composed of a stack of six identical layers and each layer has two sub-layers. The first sub-layer is a multi-head self-attention engine  1212, and the second sub-layer is a fully-connected feed-forward network 1214. A residual connection (not shown) connects around each of the sub-layers followed by normalization.

[0154] In this example transformer 1200, the decoder 1230 is also composed of a stack of six 6 identical layers. The decoder also includes a masked multi-head self-attention engine 1232, a multi-head attention engine 1234 over the output of the encoder 1210, and a fully-connected feed-forward network 1226. Each layer includes a residual connection (not shown) around the layer, which is followed by layer normalization. The masked multi-head self-attention engine 1232 is masked to prevent positions from attending to subsequent positions and ensures that the predictions at position i can depend only on the known outputs at positions less than i (e.g., auto-regression) .

[0155] In the transformer, the queries, keys, and values are linearly projected by a multi-head attention engine into learned linear projects, and then attention is performed in parallel on each of the learned linear projects, which are concatenated and then projected into final values.

[0156] The transformer also includes a positional encoder 1240 to encode positions because the model does not contain recurrence and convolution and relative or absolute position of the tokens is needed. In the transformer 1200, the positional encodings are added to the input embeddings at the bottom layer of the encoder 1210 and the decoder 1230. The positional encodings are summed with the embeddings because the positional encodings and embeddings have the same dimensions. A corresponding position decoder 1250 is configured to decode the positions of the embeddings for the decoder 1230.

[0157] In some aspects, the transformer 1200 uses self-attention mechanisms to selectively weigh the importance of different parts of an input sequence during processing and allows the model to attend to different parts of the input sequence while generating the output. The input sequence is first embedded into vectors and then passed through multiple layers of self-attention and feed-forward networks. The transformer 1200 can process input sequences of variable length, making it well-suited for natural language processing tasks where input lengths can vary greatly. Additionally, the self-attention mechanism allows the transformer 1200 to capture long-range dependencies between words in the input sequence, which is difficult for RNNs and CNNs. The transformer with self-attention has achieved results in several natural language processing tasks that are beyond the capabilities of other neural networks and has become a popular choice for language and text applications. For example, the various large language models, such as a  generative pretrained transformer (e.g., ChatGPT, etc. ) and other current models are types of transformer networks.

[0158] FIG. 13 illustrates an example computing-device architecture 1300 of an example computing device which can implement the various techniques described herein. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a personal computer, a laptop computer, a video server, a vehicle (or computing device of a vehicle) , or other device. For example, the computing-device architecture 1300 may include, implement, or be included in any or all of system 100 of FIG. 1, vehicle 102 of FIG. 1, map system 112 of FIG. 1, system 200 of FIG. 2, vehicle 202 of FIG. 2, map system 224 of FIG. 2, system 300 of FIG. 3, vehicle 302 of FIG. 3, vehicle 312 of FIG. 3, vehicle 322 of FIG. 3, map system 332 of FIG. 3, system 600 of FIG. 6, system 700 of FIG. 7 and / or other devices, modules, or systems described herein. Additionally or alternatively, computing-device architecture 1300 may be configured to perform process 800, process 900, and / or other process described herein.

[0159] The components of computing-device architecture 1300 are shown in electrical communication with each other using connection 1312, such as a bus. The example computing-device architecture 1300 includes a processing unit (CPU or processor) 1302 and computing device connection 1312 that couples various computing device components including computing device memory 1310, such as read only memory (ROM) 1308 and random-access memory (RAM) 1306, to processor 1302.

[0160] Computing-device architecture 1300 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1302. Computing-device architecture 1300 can copy data from memory 1310 and / or the storage device 1314 to cache 1304 for quick access by processor 1302. In this way, the cache can provide a performance boost that avoids processor 1302 delays while waiting for data. These and other modules can control or be configured to control processor 1302 to perform various actions. Other computing device memory 1310 may be available for use as well. Memory 1310 can include multiple different types of memory with different performance characteristics. Processor 1302 can include any general-purpose processor and a hardware or software service, such as service 1 1316, service 2 1318, and service 3 1320 stored in storage device 1314, configured to control processor 1302 as well as a special-purpose processor where software instructions are  incorporated into the processor design. Processor 1302 may be a self-contained system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.

[0161] To enable user interaction with the computing-device architecture 1300, input device 1322 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. Output device 1324 can also be one or more of a number of output mechanisms known to those of skill in the art, such as a display, projector, television, speaker device, etc. In some instances, multimodal computing devices can enable a user to provide multiple types of input to communicate with computing-device architecture 1300. Communication interface 1326 can generally govern and manage the user input and computing device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.

[0162] Storage device 1314 is a non-volatile memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile discs (DVDs) , cartridges, random-access memories (RAMs) 1306, read only memory (ROM) 1308, and hybrids thereof. Storage device 1314 can include services 1316, 1318, and 1320 for controlling processor 1302. Other hardware or software modules are contemplated. Storage device 1314 can be connected to the computing device connection 1312. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1302, connection 1312, output device 1324, and so forth, to carry out the function.

[0163] The term “substantially, ” in reference to a given parameter, property, or condition, may refer to a degree that one of ordinary skill in the art would understand that the given parameter, property, or condition is met with a small degree of variance, such as, for example, within acceptable manufacturing tolerances. By way of example, depending on the particular parameter, property, or condition that is substantially met, the parameter, property, or condition may be at least 90%met, at least 95%met, or even at least 99%met.

[0164] Aspects of the present disclosure are applicable to any suitable electronic device (such as security systems, smartphones, tablets, laptop computers, vehicles, drones, or other devices)  including or coupled to one or more active depth sensing systems. While described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to specific devices.

[0165] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, one processing system and so on) . As used herein, a device may be any electronic device with one or more parts that may implement at least some portions of this disclosure. While the below description and examples use the term “device” to describe various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. While the below description and examples use the term “system” to describe various aspects of this disclosure, the term “system” is not limited to a specific configuration, type, or number of objects.

[0166] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by one of ordinary skill in the art that the aspects may be practiced without these specific details. For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blocks including devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.

[0167] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a  subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.

[0168] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.

[0169] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction (s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD) , flash memory, magnetic or optical disks, USB devices provided with non-volatile memory, networked storage devices, any suitable combination thereof, among others. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0170] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.

[0171] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any  combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor (s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.

[0172] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.

[0173] In the foregoing description, aspects of the application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.

[0174] One of ordinary skill will appreciate that the less than ( “<” ) and greater than ( “>” ) symbols or terminology used herein can be replaced with less than or equal to ( “≤” ) and greater than or equal to ( “≥” ) symbols, respectively, without departing from the scope of this description.

[0175] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g.,  microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.

[0176] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.

[0177] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on) , or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.

[0178] Claim language or other language reciting “at least one processor configured to, ” “at least one processor being configured to, ” “one or more processors configured to, ” “one or more processors being configured to, ” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation (s) . For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.

[0179] Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be  performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.

[0180] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method) , the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function) .

[0181] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0182] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general-purposes computers, wireless communication  device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium including program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random-access memory (RAM) such as synchronous dynamic random-access memory (SDRAM) , read-only memory (ROM) , non-volatile random-access memory (NVRAM) , electrically erasable programmable read-only memory (EEPROM) , flash memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0183] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs) , general-purpose microprocessors, an application specific integrated circuits (ASICs) , field programmable logic arrays (FPGAs) , or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor, ” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.

[0184] Illustrative aspects of the disclosure include:

[0185] Aspect 1. An apparatus for using map data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: process perception data representative of an area to generate three-dimensional (3D) data; determine location information associated with the 3D data; transmit the 3D data and the location  information to a map server; receive a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjust an operating parameter of a vehicle based on the map.

[0186] Aspect 2. The apparatus of aspect 1, wherein the map is based on a plurality of sets of 3D data generated based on a corresponding plurality of sets of perception data captured by a corresponding plurality of vehicles.

[0187] Aspect 3. The apparatus of any one of aspects 1 or 2, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.

[0188] Aspect 4. The apparatus of any one of aspects 1 to 3, wherein the map server is configured to: generate map data based on the plan-view-image data of the area; and refine the map data based on the 3D data.

[0189] Aspect 5. The apparatus of aspect 4, wherein the map server is configured to: align the 3D data with map data based on the location information to generate aligned 3D data; and combine the aligned 3D data with the map data.

[0190] Aspect 6. The apparatus of any one of aspects 1 to 5, wherein the at least one processor is configured to: label the 3D data; and transmit labels of the 3D data to the map server.

[0191] Aspect 7. The apparatus of any one of aspects 1 to 6, wherein the apparatus comprises a computing system of a vehicle.

[0192] Aspect 8. The apparatus of any one of aspects 1 to 7, wherein the operating parameter is associated with at least one of a path for the vehicle to travel, a braking parameter for operating one or more brakes of the vehicle, a steering parameter for steering the vehicle, a lane-change parameter for causing the vehicle to navigate from a first lane to a second lane, or displaying information using a user interface of the vehicle.

[0193] Aspect 9. An apparatus for generating map data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate a map of an area based on a plan-view image of the area; obtain three-dimensional (3D) data based on perception data captured in the area; and refine the map based on the 3D data to generate a refined map of the area.

[0194] Aspect 10. The apparatus of aspect 9, wherein, to refine the map, the at least one processor is configured to: align the 3D data with the map based on location information associated with the 3D data to generate aligned 3D data; and combine the aligned 3D data with the map.

[0195] Aspect 11. The apparatus of aspect 10, wherein the 3D data are aligned with the map further based on a comparison of the 3D data with map points of the map.

[0196] Aspect 12. The apparatus of any one of aspects 9 to 11, wherein the at least one processor is configured to: obtain a plurality of sets of 3D data from a corresponding plurality of vehicles; and refine the map based on the plurality of sets of 3D data.

[0197] Aspect 13. The apparatus of any one of aspects 9 to 12, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.

[0198] Aspect 14. A method for using map data, the method comprising: processing perception data representative of an area to generate three-dimensional (3D) data; determining location information associated with the 3D data; transmitting the 3D data and the location information to a map server; receiving a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; and adjusting an operating parameter of a vehicle based on the map.

[0199] Aspect 15. The method of aspect 14, wherein the map is based on a plurality of sets of 3D data generated based on a corresponding plurality of sets of perception data captured by a corresponding plurality of vehicles.

[0200] Aspect 16. The method of any one of aspects 14 or 15, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.

[0201] Aspect 17. The method of any one of aspects 14 to 16, wherein the map server is configured to: generate map data based on the plan-view-image data of the area; and refine the map data based on the 3D data.

[0202] Aspect 18. The method of aspect 17, wherein the map server is configured to: align the 3D data with map data based on the location information to generate aligned 3D data; and combine the aligned 3D data with the map data.

[0203] Aspect 19. The method of any one of aspects 14 to 18, further comprising: labeling the 3D data; and transmitting labels of the 3D data to the map server.

[0204] Aspect 20. The method of any one of aspects 14 to 19, wherein the operating parameter is associated with at least one of a path for the vehicle to travel, a braking parameter for operating one or more brakes of the vehicle, a steering parameter for steering the vehicle, a lane-change parameter for causing the vehicle to navigate from a first lane to a second lane, or displaying information using a user interface of the vehicle.

[0205] Aspect 21. A method for generating map data, the method comprising: generating a map of an area based on a plan-view image of the area; obtaining three-dimensional (3D) data based on perception data captured in the area; and refining the map based on the 3D data to generate a refined map of the area.

[0206] Aspect 22. The method of aspect 21, wherein refining the map comprises: aligning the 3D data with the map based on location information associated with the 3D data to generate aligned 3D data; and combining the aligned 3D data with the map.

[0207] Aspect 23. The method of aspect 22, wherein the 3D data are aligned with the map further based on a comparison of the 3D data with map points of the map.

[0208] Aspect 24. The method of any one of aspects 21 to 23, further comprising: obtaining a plurality of sets of 3D data from a corresponding plurality of vehicles; and refining the map based on the plurality of sets of 3D data.

[0209] Aspect 25. The method of any one of aspects 21 to 24, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.

[0210] Aspect 26. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations according to any of aspects 14 to 25.

[0211] Aspect 27. An apparatus for providing virtual content for display, the apparatus comprising one or more means for perform operations according to any of aspects 14 to 25.

Claims

1.An apparatus for using map data, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:process perception data representative of an area to generate three-dimensional (3D) data;determine location information associated with the 3D data;transmit the 3D data and the location information to a map server;receive a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; andadjust an operating parameter of a vehicle based on the map.2.The apparatus of claim 1, wherein the map is based on a plurality of sets of 3D data generated based on a corresponding plurality of sets of perception data captured by a corresponding plurality of vehicles.3.The apparatus of claim 1, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.4.The apparatus of claim 1, wherein the map server is configured to:generate map data based on the plan-view-image data of the area; andrefine the map data based on the 3D data.5.The apparatus of claim 4, wherein the map server is configured to:align the 3D data with map data based on the location information to generate aligned 3D data; andcombine the aligned 3D data with the map data.6.The apparatus of claim 1, wherein the at least one processor is configured to:label the 3D data; andtransmit labels of the 3D data to the map server.7.The apparatus of claim 1, wherein the apparatus comprises a computing system of a vehicle.8.The apparatus of claim 1, wherein the operating parameter is associated with at least one of a path for the vehicle to travel, a braking parameter for operating one or more brakes of the vehicle, a steering parameter for steering the vehicle, a lane-change parameter for causing the vehicle to navigate from a first lane to a second lane, or displaying information using a user interface of the vehicle.9.An apparatus for generating map data, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:generate a map of an area based on a plan-view image of the area;obtain three-dimensional (3D) data based on perception data captured in the area; andrefine the map based on the 3D data to generate a refined map of the area.10.The apparatus of claim 9, wherein, to refine the map, the at least one processor is configured to:align the 3D data with the map based on location information associated with the 3D data to generate aligned 3D data; andcombine the aligned 3D data with the map.11.The apparatus of claim 10, wherein the 3D data are aligned with the map further based on a comparison of the 3D data with map points of the map.12.The apparatus of claim 9, wherein the at least one processor is configured to:obtain a plurality of sets of 3D data from a corresponding plurality of vehicles; andrefine the map based on the plurality of sets of 3D data.13.The apparatus of claim 9, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.14.A method for using map data, the method comprising:processing perception data representative of an area to generate three-dimensional (3D) data;determining location information associated with the 3D data;transmitting the 3D data and the location information to a map server;receiving a map representative of the area from the map server, wherein the map is based on plan-view-image data of the area and the 3D data; andadjusting an operating parameter of a vehicle based on the map.15.The method of claim 14, wherein the map is based on a plurality of sets of 3D data generated based on a corresponding plurality of sets of perception data captured by a corresponding plurality of vehicles.16.The method of claim 14, wherein the perception data comprises at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data.17.The method of claim 14, wherein the map server is configured to:generate map data based on the plan-view-image data of the area; andrefine the map data based on the 3D data.18.The method of claim 17, wherein the map server is configured to:align the 3D data with map data based on the location information to generate aligned 3D data; andcombine the aligned 3D data with the map data.19.The method of claim 14, further comprising:labeling the 3D data; andtransmitting labels of the 3D data to the map server.20.The method of claim 14, wherein the operating parameter is associated with at least one of a path for the vehicle to travel, a braking parameter for operating one or more brakes of the vehicle, a steering parameter for steering the vehicle, a lane-change parameter for causing the vehicle to navigate from a first lane to a second lane, or displaying information using a user interface of the vehicle.

Citation Information

Patent Citations

  • Information processing device

    US20190025071A1

  • Architecture for map change detection in autonomous vehicles

    US20220146277A1

  • Unified framework and tooling for lane boundary annotation

    US20240127603A1