Map information for pose determination
By determining map constraints and requesting downscaled maps, devices optimize resource usage and maintain accurate localization, addressing inefficiencies in providing highly-detailed maps.
Patent Information
- Application Number
- PCT/CN2024/080329
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-11
AI Technical Summary
Existing localization systems waste computational resources by providing highly-detailed maps to devices that do not require or cannot efficiently process them, leading to inefficiencies in memory, processing, and bandwidth usage.
Devices determine map constraints based on their capabilities and intended use, requesting and receiving downscaled maps that match their resource availability and needs, reducing the size of maps shared by map servers.
This approach conserves computational resources and bandwidth while maintaining accurate localization, making it suitable for autonomous and semi-autonomous driving systems.
Smart Images

Figure CN2024080329_12092025_PF_FP_ABST
Abstract
Description
MAP INFORMATION FOR POSE DETERMINATIONTECHNICAL FIELD
[0001] The present disclosure generally relates to map information. For example, aspects of the present disclosure include systems and techniques for requesting, generating, selecting, and / or providing map information for pose determination.BACKGROUND
[0002] A device (e.g., an ego vehicle) may capture images of its environment, compare the images with other images captured from other positions within the environment, and determine the location of the device within the environment based on the comparison. In the present disclosure, the process of determining a location of a device within an environment may be referred to herein as “localization. ”SUMMARY
[0003] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0004] Systems and techniques are described for obtaining map information. According to at least one example, a method is provided for obtaining map information. The method includes: determining a map constraint related to a vehicle; transmitting a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and receiving the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0005] In another example, an apparatus for obtaining map information is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: determine a map constraint related to a vehicle; cause at least one transmitter to transmit a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and receive the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0006] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: determine a map constraint related to a vehicle; cause at least one transmitter to transmit a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and receive the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0007] In another example, an apparatus for obtaining map information is provided. The apparatus includes: means for determining a map constraint related to a vehicle; means for transmitting a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and means for receiving the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0008] In another example, a method is provided for providing map information. The method includes: receiving a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; obtaining a map of the location constrained according to the map constraint; and transmitting the map of the location to the vehicle.
[0009] In another example, an apparatus for providing map information is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: receive a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; obtain a map of the location constrained according to the map constraint; and cause at least one transmitter to transmit the map of the location to the vehicle.
[0010] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; obtain a map of the location constrained according to the map constraint; and cause at least one transmitter to transmit the map of the location to the vehicle.
[0011] In another example, an apparatus for providing map information is provided. The apparatus includes: means for receiving a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; means for obtaining a map of the location constrained according to the map constraint; and means for transmitting the map of the location to the vehicle.
[0012] In some aspects, one or more of the apparatuses described herein is, can be part of, or can include an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a vehicle (or a computing device, system, or component of a vehicle) , a mobile device (e.g., a mobile telephone or so-called “smart phone” , a tablet computer, or other type of mobile device) , a smart or connected device (e.g., an Internet-of-Things (IoT) device) , a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television) , a robotics device or system, or other device. In some aspects, each apparatus can include an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each apparatus can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each apparatus can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each apparatus can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of the apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state) , and / or for other purposes.
[0013] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0014] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Illustrative examples of the present application are described in detail below with reference to the following figures:
[0016] FIG. 1 is a diagram illustrating an example system in which a vehicle may request a map, and a map server may provide a map, according to various aspects of the present disclosure;
[0017] FIG. 2 illustrates an example of a wireless communication network, according to various aspects of the present disclosure;
[0018] FIG. 3 is a diagram illustrating an example of an image including a keypoint according to various aspects of the present disclosure;
[0019] FIG. 4 is a diagram illustrating an example of relative-pose determination using keypoints from an image at a camera according to various aspects of the present disclosure;
[0020] FIG. 5 is a block diagram illustrating data of the map of FIG. 1, according to various aspects of the present disclosure;
[0021] FIG. 6 includes an image to illustrate dimensionality reduction and clustering therein according to various aspects of the present disclosure; FIG. 6 also includes a 3D plot to illustrate an example showcasing 3 clusters;
[0022] FIG. 7 illustrates an example scenario in which a vehicle sends a first request for a map at a first time and a second request for a map at a second time, according to various aspects of the present disclosure;
[0023] FIG. 8 illustrates an example scenario in which a first vehicle sends a first request for a map and second vehicle sends a second request for a map, according to various aspects of the present disclosure;
[0024] FIG. 9 is a flow diagram illustrating another example process for obtaining map information, in accordance with aspects of the present disclosure;
[0025] FIG. 10 is a flow diagram illustrating another example process for providing map information, in accordance with aspects of the present disclosure;
[0026] FIG. 11 is a block diagram illustrating an example of a deep learning neural network that can be used to perform various tasks, according to some aspects of the disclosed technology;
[0027] FIG. 12 is a block diagram illustrating an example of a convolutional neural network (CNN) , according to various aspects of the present disclosure; and
[0028] FIG. 13 is a block diagram illustrating an example computing-device architecture of an example computing device which can implement the various techniques described herein.DETAILED DESCRIPTION
[0029] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
[0030] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects will provide those skilled in the art with an enabling description for implementing an exemplary aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.
[0031] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.
[0032] As described above, a device (e.g., a vehicle) may determine its location within an environment (e.g., the device may localize itself) by comparing images captured by the device with a map of the environment. For example, the device may determine keypoints within the images captured by the device and compare descriptors of the keypoints with descriptors of map points in the map of the environment.
[0033] As an example, a device may receive a map including a point cloud made up of a number of map points. Each of the map points may include three-dimensional coordinates referencing a point in the environment.
[0034] The map may also include descriptors descriptive of each map point. The descriptors may be, or may include, a feature-space representation of images of the points of the environment. The descriptors of a given map point may include feature-space representations of the point in the environment from different poses within the environment. For example, the descriptors of a given map point may include a number of descriptors, each of the descriptors corresponding to a different pose from which the map point may be viewed.
[0035] The map of the environment may be related to a number of keyframes within the environment. A keyframe may define a pose within the environment. In the present disclosure, the term pose may refer to a position in three degrees of freedom (e.g., according to three perpendicular axes, such as an x-axis, a y-axis, and a z-axis) and an orientation in three degrees of freedom (e.g., according to three rotational axis, such as roll, pitch, and yaw) .
[0036] The map may also include keyframe indices for each map point (and / or for the descriptors) . A keyframe index of a given map point may be indicative of keyframes (poses) within the environment from which the given map point can be observed. Additionally or alternatively, the keyframe index may indicate which descriptors of a given map point correspond to which keyframes.
[0037] The device may capture an image of the environment from its location within the environment and determine keypoints (e.g., visually distinct points) of the image. The device may further determine descriptors of each of the keypoints. The device may compare the determined descriptors with the descriptors received in the map to determine matching descriptors. The device may use the matching descriptors to correlate keypoints of the captured image with map points of the map. The device may determine the location of the device based on the correlated keypoints and map points (e.g., through triangulation) .
[0038] Systems and techniques are described for map sharing for localization. The systems and techniques may be implemented at a map-sharing device (e.g., a server) to share a map with a localizing device (e.g., an ego vehicle) to enable the localizing device to perform localization. Additionally or alternatively, the systems and techniques may be implemented in a localizing device, for example, to request a map and / or to modify a received map.
[0039] A map-sharing device (e.g., a server) may have a highly-detailed map of an environment for example, including thousands of map points, with hundreds of descriptors for each of the map points, and with information regarding thousands of keyframes within an environment from which the map points may be observable.
[0040] In some cases, a localizing device, (e.g., an ego vehicle) may not efficiently use a highly-detailed map of the environment. For example, the localizing device may not have sufficient memory to store and / or processing capacity to process a highly-detailed map. Additionally or alternatively, the localizing device may not need to use all the details of a highly-detailed map of the environment. For example, in some cases, the localizing device may use relatively few (e.g., 1 out of 10 or fewer) of the descriptors of a highly-detailed map.
[0041] For example, the localizing device may need only a coarse localization solution (e.g., accurate to 1.5 meters (m) ) and may be able to determine the coarse localization solution using tens of map points rather than the thousands of map points of the highly-detailed map. As another example, many of the keyframes of the highly-detailed map may be relatively close to one another, for example, so close that descriptors of a map point may be substantially similar when viewed from two close keyframes. As such, the localizing device may be able to match map points of a received map with keypoints of images captured by the localizing device using fewer than all of the keyframes of the highly-detailed map.
[0042] In many cases, the map-sharing device may waste computational resources (including, as examples, computing time, power, transmission power, and / or transmission time) by providing a highly-detailed map to the localizing device. Additionally or alternatively, the localizing device may waste computational resources (including memory space, receiving time, processing time, and / or power) receiving, storing, and / or using the highly-detailed map.
[0043] A localizing device, according to various aspects of the present disclosure, may determine a map constraint and request a map according to the map constraint. The map constraint may be based on available computational resources of the localizing device and / or an intended use of the map. For example, the map constraint may be based on available memory and / or processing bandwidth of the localizing device. As another example, if the localizing device intends to determine a coarse localization solution (e.g., accurate to 1.5 meters) , the localizing device may request a less-detailed map than if the localizing device intended to determine fine localization solution (e.g., accurate to 10 centimeters) . A map-sharing device, according to various aspects of the present disclosure, may receive the request, and provide a map to the localizing device responsive to the request. The provided map may be constrained according to the map constraint.
[0044] Additionally or alternatively, a map-sharing device, according to various aspects of the present disclosure, may determine and / or provide a downscaled map that is smaller than the most-detailed map possessed by the map-sharing device to localizing devices. For example, the map-sharing device may determine a downscaled map (that is smaller than the most-detailed map possessed by the map-sharing device) to share with localizing devices and share the downscaled map with the localizing devices.
[0045] In some cases, the map-sharing device may determine a downscaled map that is appropriate to the map constraints and provide the downscaled map to the localizing device. In some cases, the map-sharing device may generate and store several downscaled maps of a variety of sizes and select the downscaled map to share from among the several downscaled maps. In some cases, the map-sharing device may generate a downscaled map to satisfy the map constraints responsive to a request from a localizing device. For example, in some cases, a localizing device may request a map no larger than a particular size and the map-sharing device may downscale a highly-detailed map to the particular size or smaller and provide the downscaled map to the localizing device responsive to the request.
[0046] In some cases, the map-sharing device may share (e.g., broadcast, for example, responsive to a request or not) one or more downscaled maps having one or more different respective sizes. For example, based on a location of the map-sharing device, the map-sharing device may determine one or more downscaled maps appropriate to its location and share the one or more downscaled maps. For instance, a map-sharing device located proximate to a highway (where it is expected that cars will travel quickly through one of a limited number of predefined paths) may have several downscaled maps with sizes selected to be appropriate for expected use cases for the downscaled maps. The map-sharing device may share the several downscaled maps.
[0047] To generate a downscaled map based on a highly-detailed map, the map-sharing device may reduce the number of map points of the highly-detailed map, the number of descriptors of the highly-detailed map, and / or the number of keyframes from which descriptors are stored for map points for the highly-detailed map. For example, the systems and techniques may reduce the size of a highly-detailed map by reducing the number of map points of point clouds of the map. For example, a highly-detailed map may include 1000 map points and a downscaled map may include 100 map points.
[0048] As another example, the map-sharing device may reduce a number of descriptors in sets of descriptors of the map points of the highly-detailed map. For example, a highly-detailed map may include, on average, 100 descriptors for each map point of the map and a downscaled map may include, on average, 10 descriptors for each map point of the map. In such cases, along with the downscaled map, the map-sharing device may provide a matching threshold to the localizing device. The localizing device may use the matching threshold when determining matches between descriptors of keypoints in images captured by the localizing device and the descriptors of the map.
[0049] Additionally or alternatively, the map-sharing device may reduce a size of data used to represent descriptors by encoding the descriptors. For example, the map-sharing device may generate a codebook. The codebook may include, for example, 100 common descriptors of the map and an index to each descriptor. The common descriptors may be determined based on a clustering of descriptors of the map. The codebook may include a unique index for each of the common descriptors. For example, a descriptor may be represented, by default, by 50 bytes of data. An index into a codebook may be represented by 1 byte of data. When deployed, the map-sharing device may provide the codebook to a localizing device. Then, rather than providing the descriptors in full, the map-sharing device may provide indices of the codebook corresponding to the descriptors. The localizing device may look up the descriptors in the codebook.
[0050] As yet another example, in other localization techniques, a map-sharing device may share map points and keyframe indices. The localizing device may determine how to use the keyframes. For example, a localizing device may determine a respective volume (e.g., a sphere, a cone, or a frustrum) around each keyframe and make localization determinations based on such volumes. For example, the localizing device may determine an estimated location of the localizing device relative to the spheres around the keyframes and determine a closest keyframe. The localizing device may select descriptors related to the determined closest keyframe and try to match keypoints of images captured by the localizing device with the selected descriptors.
[0051] The map-sharing device, rather than providing keyframe indices, may generate observation regions which may define locations in the environment from which map points may be observable. The observation regions may be relatively large, for example, encompassing several keyframes. The observation regions may have any three-dimensional shape. Rather than providing keyframe indices related to each map point, the systems and techniques may provide an observation-region definition descriptive of a plurality of observation regions within the environment. The localizing device may make localization determinations based on the observation regions.
[0052] For example, the localizing device may determine an estimated location of the localizing device relative to the observation regions. For example, the localizing device may estimate which observation region the localizing device is in. The localizing device may select descriptors of the map that are related to the observation region and use match the descriptors to descriptors of keypoints in images captured by the localizing device.
[0053] It may be useful for driving systems (e.g., autonomous, semi-autonomous, or assisted driving systems, such as an advanced driver assistance system (ADAS) ) of vehicles to have accurate location information. These capabilities may become even more important for higher levels of autonomy, such as autonomy levels 3 and higher. For example, autonomy level 0 requires full control from the driver as the vehicle has no autonomous driving system, and autonomy level 1 involves basic assistance features, such as cruise control, in which case the driver of the vehicle is in full control of the vehicle. Autonomy level 2 refers to semi-autonomous driving, where the vehicle can perform functions, such as drive in a straight path, stay in a particular lane, control the distance from other vehicles in front of the vehicle, or other functions own. Autonomy levels 3, 4, and 5 include much more autonomy. For example, autonomy level 3 refers to an on-board autonomous driving system that can take over all driving functions in certain situations, where the driver remains ready to take over at any time if needed. Autonomy level 4 refers to a fully autonomous experience without requiring a user’s help, even in complicated driving situations (e.g., on highways and in heavy city traffic) . With autonomy level 4, a person may still remain in the driver’s seat behind the steering wheel. Vehicles operating at autonomy level 4 can communicate and inform other vehicles about upcoming maneuvers (e.g., a vehicle is changing lanes, making a turn, stopping, etc. ) . Autonomy level 5 vehicles fully autonomous, self-driving vehicles that operate autonomously in all conditions. A human operator is not needed for the vehicle to take any action. Thus, autonomous, semi-autonomous, or assisted driving systems are an example of where the systems and techniques described may be employed. Also, the systems and techniques described herein may be employed in non-autonomous (e.g., human controlled) vehicles. For example, the systems and techniques may provide location information to a driver.
[0054] Robust localization with multiple layers that complement each other may benefit autonomous, semi-autonomous, or assisted driving systems. Semantic camera-based localization making use of lane markers and / or traffic signs is used to enable autonomous driving solutions. Alternate localization layers that are based on non-semantic point cloud generated from different sensors are also evolving and gaining traction. A 2D or 3D point cloud maybe generated using ranging sensors like Lidar, Radar or even using camera. These solutions can be highly effective when semantic features or GPS based solutions are unavailable (for instance at urban intersection or toll ways wherein lane markers are absent, or underground tunnels wherein GPS is challenged) or to provide a gracious degradation in system when underlying assumptions on weather, road conditions, etc. fail.
[0055] Another aspect in enabling accurate map-based localization is ability to update and maintain real-time maps. Having dedicated fleets to collect mapping data is one method, however another alternative method is crowdsourced mapping wherein consumer vehicles contribute to providing the backend mapping servers data to generate maps. Considering non-semantic layers (NSL) in localization has implications on such mapping techniques too not only in terms of algorithms needed for map generation but also data exchange between the vehicles and the server (s) .
[0056] This disclosure discloses, among other things, techniques based on camera-based non-semantic layer (NSL) . In such solutions 2D point cloud features are first extracted from camera images. For mapping, these features are triangulated by tracking to get 3D point cloud in world frame. For localization, the generated 2D features are matched with 3D point cloud from the map after projecting the 3D points onto the camera image to trigger measurement updates using Bayesian filtering methods like extended Kalman filter or particle filter. One challenge is to ensure repeatability of the point cloud from different viewing angles of the vehicle or different weather conditions or different camera sensors. To ensure accurate matching, usually descriptors and keyframes are attached to the 3D map points and similarly descriptors are attached to 2D features. The descriptors usually describe properties of local neighborhood of a feature point extracted from camera image and helps in providing a unique signature of each NSL point and can be high dimensional (like >40 or 50 dimensions) . There can be multiple descriptors per 3D map point given different viewing angles or weather conditions that can lead to observability of the map point. The keyframes help to identify given an ego pose what are the relevant descriptors that can be used for matching. This helps reduce runtime complexity at the cost of larger memory requirement.
[0057] However, one challenge is that the camera NSL map is a heavy data structure due to the descriptors and keyframes involved. Depending on reliance of camera NSL in a proprietary localization system it may not be worth it to deal with large quantities of data that not only may add to bandwidth requirements on the wireless link between map server and vehicle, but also compute and memory storage requirements. This disclosure includes aspects to allow for real-time and / or non-real-time reduction in map size at the loss of some accuracy in camera NSL localization to benefit in other key performance indicators like bandwidth requirement for map exchange, on-target complexity and memory consumption. The map-size reduction can have implication on localization algorithm changes necessary as well.
[0058] Various aspects of the application will be described with respect to the figures below.
[0059] FIG. 1 is a diagram illustrating an example system 100 in which a vehicle 102 may request a map, and a map server 106 may provide a map 112, according to various aspects of the present disclosure. For example, vehicle 102 may determine a map constraint 110 and transmit a request 108, based on map constraint 110, to map server 106. Map server 106 may obtain map 112 (e.g., by generating map 112 or selecting map 112 from previously generated maps) of environment 114 of vehicle 102. Map 112 may be constrained based on map constraint 110. Map server 106 may provide map 112 to vehicle 102.
[0060] Vehicle 102 is provided as an example of a localizing device. Vehicle 102 may include an advanced driver assistance system (ADAS) 104. ADAS 104 may determine map constraint 110.
[0061] In some cases, ADAS 104 may determine map constraint 110 based on an intended use of map 112. For example, ADAS 104 may determine a use case for map 112 based on a speed of vehicle 102, based on an environment of vehicle 102, and / or based on how map 112 will be used by ADAS 104. ADAS 104 may determine map constraint 110 based on the use case.
[0062] In some cases, ADAS 104 may determine map constraint 110 based on computational resources of ADAS 104. For example, ADAS 104 may determine a size constraint for map 112 based on memory, processing, and / or communication resources of ADAS 104. For example, to conserve power, computation time, transmission and / or reception time and / or power, ADAS 104 may determine map constraint 110 to limit a size of map 112. In any case, vehicle 102 may transmit request 108 (including an indication of map constraint 110) to map server 106.
[0063] In some cases, map server 106 may generate map 112 based on map constraint 110. For example, map server 106 may generate map 112 to be no bigger that a size described by map constraint 110. In some cases, map server 106 may store one or more maps of varying sizes. In such cases, map server 106 may select map 112 from among the stored maps. In any case, map server 106 may provide (e.g., transmit) map 112 to vehicle 102.
[0064] Map 112 may be, or may include, a non-semantic layer (NSL) map of environment 114. For example, map 112, may be, or may include, a camera NSL map of an environment and may be related to non-semantic features extracted from two-dimensional images captured by a camera. Additionally or alternatively, in some aspects, map 112 may be based on radio detection and ranging (RADAR) information. For example, map 112 may be, or may include, non-semantic features extracted from RADAR information. In any case, map 112 may include map points, descriptors corresponding to the map points, and indications of poses from which images on which the descriptors are based were captured.
[0065] ADAS 104 may localize vehicle 102 (e.g., determine a location of vehicle 102) based on map 112 and one or more images captured by cameras of vehicle 102.
[0066] FIG. 2 illustrates an example of a wireless communication network 200, according to various aspects of the present disclosure. Wireless communication networks (e.g, wireless communication network 200) are deployed to provide various communication services such as voice, video, packet data, messaging, broadcast, and the like. Wireless communication network 200 may support both access links and sidelinks for communication between wireless devices. An access link may refer to any communication link between a client device (e.g., a user equipment (UE) , such as UE 214 and / or UE 216, a vehicle 202 (which may be, or may include, a UE) , or other client device) , and a base station (e.g., a 3GPP gNB, a 3GPP eNB, a Wi-Fi access point (AP) , or other base station) . For example, an access link may support uplink signaling, downlink signaling, connection procedures, etc.
[0067] Uplink and / or downlink signaling may allow client devices to communicate with a server 218. Server 218 may provide various services for the client devices. For example, vehicle 202 may communicate with server 218 via base station 210 in what may be referred to as Car-to-Cloud (C2C) communications. As such server 218 may be referred to as a C2C server.
[0068] A sidelink may refer to any communication link between client devices (e.g., vehicle 202, vehicle 204, UE 214, UE 216, etc. ) . For example, a sidelink may support device-to-device (D2D) communications, vehicle-to-everything (V2X) and / or vehicle-to-vehicle (V2V) communications, message relaying, discovery signaling, beacon signaling, or any combination of these or other signals transmitted over-the-air from one UE to one or more other UEs. In some examples, sidelink communications may be transmitted using a licensed frequency spectrum or an unlicensed frequency spectrum (e.g., 5 GHz or 6 GHz) . As used herein, the term sidelink may refer to 3GPP sidelink (e.g., using a PC5 sidelink interface) , Wi-Fi direct communications (e.g., according to a Dedicated Short-Range Communication (DSRC) protocol) , or using any other direct device-to-device communication protocol.
[0069] V2X communications may include communications between vehicles (e.g., vehicle-to-vehicle (V2V) ) , communications between vehicles and infrastructure (e.g., vehicle-to-infrastructure (V2I) ) , communications between vehicles and pedestrians (e.g., vehicle-to-pedestrian (V2P) ) , and / or communications between vehicles and network severs (vehicle-to-network (V2N) ) . For V2V, V2P, and V2I communications, data packets may be sent directly (e.g., using a PC5 interface, using an 802.11 DSRC interface, etc. ) between vehicles without going through the network, eNB, or gNB. V2X-enabled vehicles, for instance, may use a short-range direct-communication mode that provides 160° non-line-of-sight (NLOS) awareness, complementing onboard line-of-sight (LOS) sensors, such as cameras, radio detection and ranging (RADAR) , Light Detection and Ranging (LIDAR) , among other sensors. The combination of wireless technology and onboard sensors enables V2X vehicles to visually observe, hear, and / or anticipate potential driving hazards (e.g., at blind intersections, in poor weather conditions, and / or in other scenarios) . V2X vehicles may also understand alerts or notifications from other V2X-enabled vehicles (based on V2V communications) , from infrastructure systems (based on V2I communications) , and from user devices (based on V2P communications) . Infrastructure systems may include roads, stop lights, road signs, bridges, toll booths, and / or other infrastructure systems that may communicate with vehicles using V2I messaging.
[0070] Depending on the desired implementation, sidelink communications may be performed according to 3GPP communication protocols sidelink (e.g., using a PC5 sidelink interface according to LTE, 5G, etc. ) , Wi-Fi direct communication protocols (e.g., DSRC protocol) , or using any other device-to-device communication protocol. In some examples, sidelink communication may be performed using one or more Unlicensed National Information Infrastructure (U-NII) bands. For instance, sidelink communications may be performed in bands corresponding to the U-NII-4 band (5.850 –5.925 GHz) , the U-NII-5 band (5.925 –6.425 GHz) , the U-NII-6 band (6.425 –6.225 GHz) , the U-NII-7 band (6.225 –6.875 GHz) , the U-NII-8 band (6.875 –7.125 GHz) , or any other frequency band that may be suitable for performing sidelink communications.
[0071] In some examples, sidelink communication may include D2D or V2X communication. V2X communication involves the wireless exchange of information directly between not only vehicles (e.g., vehicle 202 and vehicle 204) themselves, but also directly between vehicle 202 and / or vehicle 204 and infrastructure, for example, roadside units (e.g., roadside unit 206) , such as streetlights, buildings, traffic cameras, tollbooths or other stationary objects. V2X communication may also include the wireless exchange of information directly between vehicle 202 and / or vehicle 204, pedestrians (e.g., a UE of pedestrian 208) , wireless communication networks (e.g., base station 210) , UE 214, and / or UE 216. In some examples, V2X communication may be implemented in accordance with the New Radio (NR) cellular V2X standard defined by 3GPP, Release 16, or other suitable standard.
[0072] V2X communication enables vehicle 202 and / or vehicle 204 to obtain information related to the weather, nearby accidents, road conditions, activities of nearby vehicles and pedestrians, objects nearby the vehicle, and other pertinent information that may be utilized to improve the vehicle driving experience and increase vehicle safety. For example, such V2X data may enable autonomous driving and improve road safety and traffic efficiency. For example, the exchanged V2X data may be utilized by a V2X connected vehicle 202 and / or vehicle 204 to provide in-vehicle collision warnings, road hazard warnings, approaching emergency vehicle warnings, pre- / post-crash warnings and information, emergency brake warnings, traffic jam ahead warnings, lane change warnings, intelligent navigation services, and other similar information. In addition, V2X data received by a V2X connected mobile device of a pedestrian / cyclist (e.g., pedestrian 208) may be utilized to trigger a warning sound, vibration, flashing light, etc., in case of imminent danger.
[0073] The sidelink communication between vehicle 202, vehicle 204, roadside unit 206, a UE of pedestrian 208, UE 214, and / or UE 216, may occur over a sidelink 212 utilizing a proximity service (ProSe) PC5 interface. In various aspects of the disclosure, the PC5 interface may further be utilized to support D2D sidelink 212 communication in other proximity use cases (e.g., other than V2X) . Examples of other proximity use cases may include smart wearables, public safety, or commercial (e.g., entertainment, education, office, medical, and / or interactive) based proximity services.
[0074] According to various aspects of the present disclosure, server 218 (which is provided as an example of a map-sharing device) may share maps with one or more devices in wireless communication network 200 (e.g., vehicle 202, vehicle 204, roadside unit 206, pedestrian 208, UE 214, and / or UE 216) . The maps may include keypoints (which may be referred to alternatively as "visual features" or as "points of interest" ) . Server 218 may share the maps to enable the devices (which are provided as examples of localizing devices) to determine their respective locations. For example, server 218 may share a map including descriptors of keypoints. Vehicle 202 may determine its location based on the shared keypoints and descriptors.
[0075] FIG. 3 is a diagram illustrating an example of an image 300 including a keypoint p according to various aspects of the present disclosure. Keypoint p is surrounded by a window 302 of pixels 304 in the image 300. Keypoint p may be selected such that keypoint p can be matched between images. For example, Keypoint p may be visually distinct in image 300. Keypoint p may be, as an example, a corner point on an object. In the art a keypoint may be alternatively referred to as a visual feature, a point of interest or a key point. An example keypoint-detection method is described with regard to FIG. 3. In particular, FIG. 3 illustrates the Features from Accelerated Segment Test (FAST) technique (Machine Learning for High-Speed Corner Detection, Edward Rosten &Tom Drummond, ECCV 2006: Computer Vision –ECCV 2006 pp 430–443, Part of the Lecture Notes in Computer Science book series (LNIP, volume 3951) ) . In the FAST method, a pixel under test (e.g., pixel p) with intensity Ip may be identified as an interest point. A circle 306 of sixteen pixels (pixels 1–16) around the pixel under test p (e.g., a Bresenham circle of radius 3) may then be identified. The pixel under test p may be considered a keypoint if there exists a set of n contiguous pixels in circle 306 of sixteen pixels that are all brighter than Ip + t, or all darker than Ip -t, where t is a threshold value and n is configurable. In this example, n may be twelve. For example, the intensity of pixels 1, 5, 9, and 13 of the circle may be compared with Ip. If at least three of the four pixels do not satisfy the threshold criteria, the pixel p is not considered an interest point. As can be seen in FIG. 3, at least three of the four pixels satisfy the threshold criteria. Therefore, all sixteen pixels may be compared to pixel p to determine if twelve contiguous pixels meet the threshold criteria. This process may be repeated for each of pixels 304 in the image 300 to identify the corner points corresponding to keypoint p in image 300.
[0076] Although FIG. 3 illustrates a FAST keypoint-identifying method, it should be understood that the present disclosure is applicable to any keypoint-identifying method. Examples of keypoint-identifying methods include speeded-up robust features (SURF ) , scale-invariant feature transform (SIFT) , binary robust independent elementary feature (BRIEF) , oriented FAST and rotated BRIEF (ORB) , and Harris corner point.
[0077] As indicated above, a keypoint p represents a feature of an image 300 that may be matched between multiple images of a scene (e.g., captured from different viewing angles and / or with different intrinsic camera parameters) . For example, various cross-correlation or optical flow methods may match features (keypoints) across multiple images. In some examples, each feature may further include a feature descriptor that assists with the matching process. A feature descriptor may summarize, in vector format (e.g., of constant length) one or more characteristics of pixels 304 of window 302. For example, the feature descriptor may correspond to the intensity of pixels 304 of window 302. In general, feature descriptors are independent of the positions of keypoint p, robust against image transformations, lighting of the scene, and / or weather of the scene, and scale independently. Thus, keypoints with feature descriptors may be independently re-detected in each image frame and then subjected to a keypoint matching / tracking procedure. For example, the keypoints in two different images with matching descriptors and the smallest distance between them may be considered to be matching keypoints. Examples of feature-descriptor methods may include, but are not limited to, ORB, SURF, and BRIEF.
[0078] A relative pose of two cameras (or of one camera at two times) may be calculated based on the two-dimensional displacement of a plurality of keypoints in images from each of the cameras (or in images captured by the one camera at the two times) . For example, the pose may be determined by forming and factoring an essential matrix using eight keypoints or using Nister’s method with five keypoints. As another example, a Perspective-n-Point (PnP) algorithm with three keypoints may be used to determine the pose if keypoint depth is also being tracked. In some aspects, images captured by different cameras (e.g., of different devices or vehicles) that contain a minimum number of the same features (e.g., based on the pose determination method) may be used to determine the relative pose between the cameras.
[0079] FIG. 4 is a diagram illustrating an example of relative-pose determination using keypoints from an image at a camera C1 according to various aspects of the present disclosure. Camera C1 may be included in, for example, an extended reality (XR) device, a mobile device, a vehicle, or a roadside unit. Real points M1, M2, and M3 in three-dimensional space (x, y, z) may be projected onto a respective image plane I1 of cameras C1 to produce features (keypoints) m1, m2, and m3. If the location of the real points M1, M2, and M3 are known, a relative pose of C1 may be determined based on the positions of m1, m2, and m3 in I1. For example, the relative pose of C1 may be triangulated based m1, M1, m2, M2, m3, and M3.
[0080] Returning to FIG. 1, vehicle 102 may communicate with map server 106 via a network (e.g., wireless communication network 200 of FIG. 2) . Vehicle 102 may request that map server 106 provide map 112 to vehicle 102 and vehicle 102 may determine its location in environment 114 based on map 112.
[0081] FIG. 5 is a block diagram illustrating data of map 112 of FIG. 1, according to various aspects of the present disclosure. Map 112 includes map points 502. Each of map points 502 may be, or may include, a location 508. Each location 508 may be respective three-dimensional coordinates of the map point. Further, map 112 may include descriptors 506 for each of map points 502. For example, each map point may be identified as a keypoint in one or more images of the environment. Further, map 112 may include descriptors 506 based on the keypoint as the keypoint appears in the one or more images captured from a respective observer 504. The observers 504 may correspond to keyframes 510. For example, each map points 502 may include a list of indices of keyframes 510 for which descriptors 506 includes a descriptor.
[0082] For example, keypoint p of FIG. 3 may correspond to a map point (for example, keypoint p may be an image of a map point) . A descriptor based on the intensity of pixels 304 of FIG. 3 of window 302 of FIG. 3 may be determined. Map 112 may include the map point (e.g., the three-dimensional coordinates of the map point) and the descriptor of the map point.
[0083] Image 300 may have been captured of keypoint p from a pose within the environment. Keypoint p may have been captured in multiple images from multiple poses. Map 112 may include multiple descriptors of keypoint p based on the appearance of keypoint p in the multiple images. For example, for each of a number of keyframes (e.g., predefined poses in the environment) , map 112 may include a descriptor for each map point.
[0084] Returning to FIG. 1, vehicle 102 may capture an image of its environment, identify keypoints in the image (e.g., as described with regard to FIG. 3) , and determine descriptors of the identified keypoints. Further, vehicle 102 may compare the descriptors with descriptors received in map 112. At least some of the descriptors of the image captured by vehicle 102 may match with descriptors of map 112. For each of the matching descriptors, vehicle 102 may identify map points. The map points of map 112 may include the three-dimensional coordinates of the map points. For example, the map points may provide M1, M2, and M3 of FIG. 4. The locations of the keypoints in the captured image may provide m1, m2, and m3 of FIG. 4. Accordingly, vehicle 102 may be able to determine its relative position (e.g., relative to M1, M2, and M3) based on the image captured by vehicle 102 and map 112 as described with regard to FIG. 4. If map 112 includes global coordinates (e.g., latitude and longitude of the map points) , map 112 may determine its absolute location.
[0085] Map server 106 may store a highly-detailed map of environment 114. For example, map server 106 may store a map including thousands of map points. The map points may be captured in images from hundreds of keyframes (e.g., predetermined poses within the environment) . The map may include one descriptor for each of the keyframes for each of the map points. Thus, the map may be highly-detailed and very large (in data) .
[0086] In some cases, vehicle 102 may not efficiently use a highly-detailed map of environment 114. For example, ADAS 104 may not have sufficient memory to store and / or processing capacity to process a highly-detailed map. Additionally or alternatively, ADAS 104 may not need to use all the details of a highly-detailed map of environment 114. For example, in some cases, vehicle 102 may use relatively few (e.g., 1 out of 10 or fewer) of the descriptors of a highly-detailed map.
[0087] For example, vehicle 102 may need only a coarse localization solution (e.g., accurate to 1.5 meters) and may be able to determine the coarse localization solution using tens of map points rather than the thousands of map points of the highly-detailed map. As another example, many of the keyframes of the highly-detailed map may be relatively close to one another, for example, so close that descriptors of a map point may be substantially similar when viewed from two close keyframes. As such, vehicle 102 may be able to match map points of a received map with keypoints of images captured by vehicle 102 using fewer than all of the keyframes of the highly-detailed map.
[0088] To conserver computational resources of map server 106 (including, as examples, computing time, power, transmission power, and / or transmission time) and computational resources of vehicle 102 (including memory space, receiving time, processing time, and / or power) , vehicle 102 may request a constrained map rather than a highly-detailed map. For example, ADAS 104 may determine map constraint 110, and vehicle 102 may submit request 108 to map server 106 including or based on map constraint 110. Map server 106 may obtain map 112 based on map constraint 110 and provide map 112 to vehicle 102.
[0089] For example, map server 106 may make different qualities of camera NSL map available to potential consumer vehicles based on map constraints indicated by the vehicles. The different qualities can vary in terms of a total number of map points, a total number of descriptors per map point, a data size of descriptors, and / or a total number of keyframes and observers per map point. The different qualities may reduce map size from a baseline high quality version.
[0090] For vehicles to consume reduced-size maps, maps and / or map messages may be modified. For example, the maps may include, and / or be relative to, observations regions rather than keyframes. Observations regions may be sized and / or shaped to reduce a data size of maps (for example, by reducing the number of descriptors per map point by reducing the number of viewing angles from which descriptors of map points are stored) .
[0091] In generating a reduced-size map, each keyframe may be correlated to an impact region to be used in localization algorithm. After reducing keyframe size based on impact regions, a vehicle may have difficulty in identifying which impact region to use to determine descriptors to match. The keyframes may be replaced with observation regions. Further, map server 106 may provide meta data useful for localization as part of a onetime message from map server 106 to vehicle 102. Information regarding relationships between impact region and keyframes may be communicated (e.g., in a one-time communication) . Additionally or alternatively, such information may be communicated with regard to observation region.
[0092] Further a descriptor-distance threshold may be provided to vehicle 102 based on a reduction in a size or number of descriptors. For example, if the reduced-size map uses fewer bits to represent descriptors, map server 106 may provide an indication of a descriptor-distance threshold to change matching thresholds used by vehicle 102 when attempting to match keypoints in images captured by vehicle 102 with map points. Vehicle 102 may use the descriptor-distance threshold when matching keypoints with map points.
[0093] Additionally or alternatively, a codebook of descriptors maybe generated by map server 106 and communicated to vehicle 102. Succeeding map updates from map server 106 to vehicle 102 can include codeword indices instead of actual descriptors for each map point to reduce map size.
[0094] Returning to FIG. 5, Oi= {Oi, 1, …, Oi, N} can be used to represent a list of observers for map point i and Di= {Di, 1, …, Di, M} can represent a list of descriptors for map point i. Here, each Di, j can be a vector scalars (e.g., 50 or more scalars, each of the scalars can be stored as a double, a float or a uint8) for instance. In general, the length of Oi can be different than Di. However considering the densest map representation, every observer will accompany a descriptor in which case M=N, with N being a very large number (consider keyframes spaced at potentially as close as every meter or half a meter for instance) . Note that in such representation although a keyframe corresponds to a unique descriptor for a map point, across different map points the descriptors can be different.
[0095] An example map format is given below. Alternatives are possible without compromising information, for example, by representing the rotation matrices by Rodrigues angles or Euler angles.
[0096] The size of a map can be reduced by reducing the number of map points. For example, by considering the actual locations of the map points in a world frame and using grid average or random subsampling to reduce number of map points. Additionally or alternatively, map points far enough apart that have similar descriptors can be dropped.
[0097] Additionally or alternatively, the map size may be reduced without changing the number of map points but by reducing the number of descriptors. Reducing the number of descriptors may significantly reduce map size because descriptors may form a significant portion of camera NSL maps.
[0098] There can be certain map points which have more than 20 or 30 descriptors owing to perceived variability in pixel neighborhood as observed from different viewing angles, distances, weather conditions, etc. All that is required for localization is map points that to have enough distinctiveness compared to other map points which will have projection close to it to enable accurate localization. From this perspective, the number of descriptors per map point can be reduced without compromising localization. This will help not only reduce the size of the database but also less load on wireless exchange links between vehicle and mapping server, and lesser real time compute during camera NSL implementation.
[0099] One example technique for reducing the number of descriptors of a map includes combining nearby keyframes into sparser keyframes (for example, from every 1 meter (m) , to every 10m) and accordingly update the observer vector per map point Oi.
[0100] With this update now each new observer can have multiple descriptors for a map point. So if there are K descriptors for a map point observed by a particular keyframe and if K>T, then use k means or k medoid (this may be better than k means since it picks one of the sample points as cluster center) or dBscan or any other clustering mechanism to reduce number of descriptors to be at most T (can be set to 1 in which case there will still be one descriptor per sparser keyframe) .
[0101] As is commonly done for clustering on high dimensional data, the descriptors could also be reduced to a lower dimensional space using multidimensional scaling (MDS) algorithms followed by doing the clustering since there may not be enough descriptor samples to cluster in the high dimensional space.
[0102] The MDS algorithm may solve the following problem given A points in B dimensions, A < B, MDS attempts to arrive at a lower dimensional set of points in P dimensions such that pairwise distances in original points similar to pairwise distances in new points.
[0103] In the MDS algorithm, setting P = A can be a safe upper bound option.
[0104] Another thing is after dimensionality reduction algorithm is employed; a straight line can be fit on the scatter plot of the pairwise distances obtained after dimensionality reduction with original pairwise distances. If the straight line goes through close to origin and has close to 45 degree slope the dimensionality reduction is a success.
[0105] FIG. 6 includes an image 602 to illustrate dimensionality reduction and clustering therein according to various aspects of the present disclosure. In image 602, an example using descriptor dimensionality reduction using features rather than the map is provided as an example to show how low dimensions can showcase clustering of descriptor data. For instance, considering descriptors in a neighborhood of a selected point, descriptors of a certain resolution can be reduced to a smaller dimension. In one illustrative example, considering all descriptors in a 10-pixel neighborhood of a selected point, 48-dimensional descriptors can be reduced down to 3 dimensions. A 3D plot 604 of FIG. 6 illustrates an example showcasing 3 clusters.
[0106] All descriptors in a cluster can be combined as cluster center if a property P of the cluster is acceptable. For instance, property P can be average distance between cluster center and members of the cluster. Another example of property P can be max distance between cluster center and members of the cluster. In some cases, outliers during clustering are kept (e.g., not deleted) since they contain notably different information compared to other descriptors.
[0107] Another alternative to clustering is to simply look at pairwise descriptor distances. First sort all pairwise descriptor distance. Start with smallest one, and then check if 2 descriptors have distance within a threshold. If yes, decide to combine them into a single descriptor (centroid of the two) . Delete all other pairs including at least one of the combined descriptors. Then continue to find next smallest descriptor distance pair and redo the check to continue combining if applicable. This will at most reduce from a total of N descriptors for the map point down to T=ceil
[0108] Another example technique for reducing the number of descriptors of a map includes not changing the keyframes at all, but per map point if there are K>T total descriptors then use mechanisms mentioned above to reduce down to T. This way the density of observers is not changed, but closeness of descriptors is used to reduce map size. A mapping for which observer corresponds to which descriptor will be used to implement this example technique.
[0109] Another example technique for reducing the number of descriptors of a map includes usage of a descriptor codebook. Since the descriptor space is high dimensional, the descriptors for all map points that are relevant in a particular region in space are already sparse. A maximum number of unique descriptors (codewords) can be allowed to represent the space of descriptors. The larger the number of codewords in a codebook, the more accurate the representation of descriptors of each map point. One example method for codebook generation is a greedy algorithm.
[0110] The greedy algorithm may list all descriptors in map data into a candidate set. Next, the greedy algorithm may pick first codeword as descriptor that has maximum neighbors with distance < T, for T set as 0.1 say. The greedy algorithm may delete all descriptors that fall in this cluster from candidate set. Then, the greedy algorithm may pick second codeword as descriptor with most neighbors from remaining candidate set and continue this process to get a codebook. The smaller the value of T, the larger the size of codebook. There are other ways of determining a codebook. Such other ways are contemplated by this disclosure.
[0111] With such a codebook available at mapping server, the mapping server may communicate the codebook to vehicle (e.g., once) . Then the vehicle then may receive succeeding map data for each map tile, replacing descriptor data structure with just codeword index corresponding to the codebook. A codebook ID may be communicated to ensure the vehicle and server are talking about the same codebook and thus the codeword indices are correctly interpreted during localization.
[0112] There may be variations in map format based on implications for localization. For example, during localization, matching of 2D features generated online may be done with 3D map points using descriptors. However, if all descriptors per map point are used for matching the complexity will increase and thus keyframes help pruning which descriptors to consider during the matching process. For instance, if ego pose is within 5m radius of certain keyframes then the descriptors corresponding to observers having those keyframe IDs are picked for matching. Thus, every keyframe has an impact region and can be defined using configuration parameters like radius of circle (e.g., 5m) or field of view (e.g., within +-60 degrees of ego yaw direction of motion) . Such a solution may be suboptimal since not all keyframes will have same impact region in reality. However, appending such information in the map data is even more data intensive. Accordingly, a configurable impact region that is pre-agreed between mapping server and vehicle may be used. Note that during first handshake of camera NSL mapping server and the vehicle, the mapping server may communicate such information to the vehicle to accurately use the map data.
[0113] If keyframes are made sparser as per above mentioned algorithm, then the impact region of each sparse keyframe is larger than original impact region and can be treated as a union of original impact regions. In this case, it is not necessary to be able to efficiently represent the new impact region with simple parameterization like 5m radius and 60-degree field of view. In such a case, the map format can change by defining a region of interest instead of a keyframes. An exemplary region of interest can be a bounding box along with direction of travel information for instance but could be any 3D convex region defined and could be different for different observation region instances.
[0114] Alterations to communications and / or maps may have implications for map crowdsourcing and / or vehicle to everything (V2X) links. For example, the reduction of number of descriptors and / or map points can be done in mapping server or in vehicle. If done on mapping server, the vehicle sends a flag indicating the maximum camera NSL map size it can consume given its sensor stack. Depending on the flag the server uses one or more of the methods above to reduce camera NSL map size and send it to the vehicle. Vehicle’s requirement on camera NSL map size not only depends on its compute availability but also on the usage of camera NSL localization layer in its ADAS stack. Considering that there may be semantic layer localization using lane boundaries and traffic signs, or other non-semantic layers like radar or lidar, one may not want an accuracy of 10 centimeters (cm) or so from camera NSL layer. The vehicle identifies what is the role of camera NSL layer in its current operation, say it identifies only needed for multi-hypothesis convergence (so about 1.5m accuracy enough) . Then in this case, it can request for a reduced size of the map. However, if the vehicle identifies it needs camera NSL localization support with more “within lane” level granularity then it may request a camera NSL map with larger size.
[0115] There can also be situations where same vehicle A requests a different camera NSL map type at different times T1 and T2. For instance, the vehicle may identify that a certain geographical region in its planned path is where there have been recent reports on vehicles getting out of self-driving mode often. The vehicle may then request best quality camera NSL map from the mapping server pre-empting challenging situation. In normal situations, it may choose a different map type as default. Similarly, one may have a case where mapping server may have recommended defaults of the map size type for different geographical region in case vehicle just wants to go with map server recommendation instead of relying on its own decision making.
[0116] For example, FIG. 7 illustrates an example scenario in which a vehicle 702 sends a request 708 for a map while vehicle 702 is in a situation 716. Later, while vehicle 702 is in situation 726, vehicle 702 may send a request 718 for a map.
[0117] For example, at a time T1, while vehicle 702 is in situation 716, ADAS 704 of vehicle 702 may determine a map constraint 710 based on an intended use of the map. Vehicle 702 may transmit request 708, based on or including, map constraint 710, to map server 706. Map server 706 may obtain (e.g., generate or select) map 712 based on map constraint 710. Map server 706 may provide (e.g., transmit) map 712 to vehicle 702. ADAS 704 of vehicle 702 may determine a location of vehicle 702 based on map 712 (and based on an image captured by vehicle 702 at T1) .
[0118] At T2, vehicle 702 may be in situation 726 that may be different from situation 716. For example, at T1, vehicle 702 may be travelling at relatively high speed through a portion of environment 714 that includes relatively deviations in the road. As such, while in situation 716, ADAS 704 may determine that ADAS 704 needs only a coarse localization solution (e.g., based on the intended use of the map) . Later, at T2, vehicle 702 may enter a portion of environment 714 that includes more road deviations. Additionally or alternatively, vehicle 702 may enter a portion of environment 714 in which a number of other vehicles have reported vehicles experiencing difficulties navigating. Accordingly, ADAS 704 may determine that ADAS 704 would benefit from a more precise localization solution (e.g., based on situation 726 and / or an intended use of the map according to situation 726) . ADAS 704 may determine map constraint 720 that may be for a larger map than map constraint 710. Vehicle 702 may transmit request 718, including or based on map constraint 720, to map server 706. Map server 706 may obtain (e.g., generate or select) map 722 based on map constraint 720. Map server 706 may provide (e.g., transmit) map 722 to vehicle 702. ADAS 704 of vehicle 702 may determine a location of vehicle 702 based on map 722 (and based on an image captured by vehicle 702 at T2) .
[0119] There can be situations where vehicle A and vehicle B are requesting different camera NSL map type depending on their available sensor stack and compute available.
[0120] For example, FIG. 8 illustrates an example scenario in which a vehicle 802 sends a request 808 for a map and vehicle 822 sends a request 828 for a map. For example, ADAS 804 of vehicle 802 may determine a map constraint 810 based on ADAS 804 (e.g., based on computational resources of ADAS 804) . Vehicle 802 may transmit request 808, based on or including, map constraint 810, to map server 806. Map server 806 may obtain (e.g., generate or select) map 812 based on map constraint 810. Map server 806 may provide (e.g., transmit) map 812 to vehicle 802. ADAS 804 of vehicle 802 may determine a location of vehicle 802 based on map 812 (and based on an image captured by vehicle 802) .
[0121] The ADAS 824 of vehicle 822 may determine a map constraint 830 based on ADAS 824 (e.g., based on computational resources of ADAS 824) . Vehicle 822 may transmit request 828, based on or including, map constraint 830, to map server 806. Map server 806 may obtain (e.g., generate or select) map 832 based on map constraint 830. Map server 806 may provide (e.g., transmit) map 832 to vehicle 822. ADAS 824 of vehicle 822 may determine a location of vehicle 822 based on map 832 (and based on an image captured by vehicle 822) .
[0122] Map 832 may, or may not, be the same as map 812. For example, map 832 may be larger than (in data size) map 812 based on differences between map constraint 830 and map constraint 810. For example, ADAS 824 may have more memory and / or processing capacity than ADAS 804. Accordingly, vehicle 822 may request a larger map than is requested by vehicle 802.
[0123] In another embodiment, the camera NSL map transmitted from serve to vehicle without any map size optimization, and the vehicle dynamically decides whether it needs to trim map size down or not depending on availability of compute or real time determination of camera NSL importance for smooth ego vehicle operation. For instance, in a region wherein there is good HD semantic layer map coverage.
[0124] FIG. 9 is a flow diagram illustrating a process 900 for obtaining map information, in accordance with aspects of the present disclosure. One or more operations of process 900 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the process 900. The one or more operations of process 900 may be implemented as software components that are executed and run on one or more processors.
[0125] At block 902, a computing device (or one or more components thereof) may determine a map constraint related to a vehicle. For example, ADAS 104 of FIG. 1 of vehicle 102 of FIG. 1 may determine map constraint 110 of FIG. 1.
[0126] In some aspects, the map constraint may be determined based on a requirement of an advanced driver assistance system (ADAS) of the vehicle. For example, ADAS 104 may determine map constraint 110 based on a requirement of ADAS 104.
[0127] In some aspects, the requirement of the ADAS may be based on computational resources of the ADAS. For example, ADAS 104 may determine map constraint 110 based on computational resources (e.g., memory and / or processing bandwidth) of ADAS 104.
[0128] In some aspects, the map constraint includes at least one of: a map-size constraint related to a size of the map; a map-point constraint related to a number of map points of a point cloud of the map; a descriptor constraint related to a number of descriptors of the map or a data size of descriptors; or a keyframe constraint related to a number of keyframes of the map. For example, map constraint 110 may be, may include, or may be based on a map-size constraint related to a size of the map, a map-point constraint related to a number of map points of a point cloud of the map, a descriptor constraint related to a number of descriptors of the map or a data size of descriptors, a keyframe constraint related to a number of keyframes of the map, or any combination thereof.
[0129] At block 904, the computing device (or one or more components thereof) may cause at least one transmitter to transmit a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint. For example, ADAS 104 may cause a transmitter of vehicle 102 to transmit request 108 of FIG. 1 to map server 106 of FIG. 1. Request 108 may include an indication of map constraint 110.
[0130] In some aspects, the computing device (or one or more components thereof) may include or control the at least one transmitter. For example, vehicle 102 may include a transmitter. ADAS 804 may be capable of using the transmitter to transmit requests.
[0131] At block 906, the computing device (or one or more components thereof) may receive the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint. For example, ADAS 104 may receive map 112 of FIG. 1. Map 112 may be constrained according to map constraint 110. Map 112 may be generated (e.g., by map server 106) based on map constraint 110 or selected (e.g., by map server 106) from among a plurality of maps.
[0132] In some aspects, the map of the location may be, or may include, a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points. For example, map 112 may be, or may include, map points 502 of FIG. 5. Each map point of map points 502 may include a location 508, which may be, or may include, three-dimensional coordinates of the map point (e.g., in environment 114 of FIG. 1) . Map 112 may include descriptors 506 of FIG. 5. Each map point of map points 502 may be described by a plurality of descriptors 506.
[0133] In some aspects, the map of the location may be, or may include, a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes. For example, map 112 may include keyframes 510 of FIG. 5. Keyframes 510 may describe poses in environment 114. Map 112 may include a relationship between keyframes 510 and descriptors 506. For example, map 112 may include relationships between map points 502, keyframes 510, and descriptors 506. For instance map 112 may include relationships describing which map points are visible from which keyframes and which descriptors correspond a given map point and corresponding key frame.
[0134] In some aspects, the map constraint may be, or may include, a first map constraint. The request may be, or may include, a first request. The map of the location may be, or may include, a first map of the location. The computing device (or one or more components thereof) may determine a second map constraint related to the vehicle; cause the at least one transmitter to transmit a second request for a second map of the location to the map server, wherein the second request comprises an indication of the second map constraint; and receive the second map of the location, constrained according to the second map constraint, from the map server, wherein the second map of the location is either generated based on the second map constraint or selected from among a plurality of maps of the location based on the second map constraint. For example, ADAS 704 of FIG. 7 of vehicle 702 of FIG 7 may determine map constraint 710 of FIG. 7 according to situation 716 of FIG. 7. ADAS 704 may transmit request 708 of FIG. 7 to map server 706 of FIG. 7. Request 708 may include an indication of map constraint 710 of FIG. 7. Map server 706 may provide map 712 to vehicle 702. According to situation 726 of FIG. 7, ADAS 704 may determine map constraint 720 of FIG. 7. ADAS 704 may transmit request 718 of FIG. 7 to map server 706. Request 718 may include an indication of map constraint 720. Map server 706 may provide map 722 of FIG. 7 to vehicle 702 responsive to the request. Map 722 may be constrained according to map constraint 720. Map 722 may be generated (e.g., by map server 706) based on map constraint 720 or selected (e.g., by map server 706) from among a plurality of maps.
[0135] In some aspects, the first map constraint may be determined based on a first requirement of an advanced driver assistance system (ADAS) of the vehicle; and the second map constraint may be determined based on a second requirement of the ADAS. For example, according to situation 716, ADAS 704 may determine map constraint 710 based on a first requirement of ADAS 704. Further, according to situation 726, ADAS 704 may determine map constraint 720 based on a second requirement of ADAS 704. In some aspects, map 712 and map 722 may describe the same location (e.g., within environment 714) of FIG. 7) .
[0136] In some aspects, the request may be, or may include, a first request. The map of the location may be, or may include, a map of a first location. The computing device (or one or more components thereof) may cause the at least one transmitter to transmit a second request for a map of a second location to the map server, wherein the second request comprises an indication of the map constraint; and receive the map of the second location, constrained according to the map constraint, from the map server, wherein the map of the second location is either generated based on the second map constraint or selected from among a plurality of maps of the second location based on the second map constraint. For example, according to situation 716, ADAS 704 may transmit request 708 to map server 706. Request 708 may be based on, and / or include, map constraint 710 and an indication of a first location within environment 714. Map server 706 may transmit map 712 to vehicle 702. Map 712 may be, or may include, a map of the first location. According to situation 726, ADAS 704 may transmit request 718 to map server 706. Request 718 may be based on, and / or include, map constraint 710 and an indication of a second location within environment 714. Map server 706 may transmit map 722 to vehicle 702. Map 722 may be, or may include, a map of the second location. Both map 712 and map 722 may be generated based on map constraint 710. Additionally or alternatively, both map 712 and map 722 may be selected from among a plurality of maps (e.g., stored by map server 706) .
[0137] In some aspects, the map of the location received by the vehicle is different from another map of the location received by another vehicle from the map server. For example, ADAS 804 of FIG. 8 of vehicle 802 of FIG. 8 may generate map constraint 810 of FIG. 8. ADAS 804 may further transmit request 808 of FIG. 8, which may include an indication of map constraint 810 to map server 806 of FIG. 8. Map server 806 may transmit map 812 of FIG. 8 to vehicle 802. ADAS 824 of FIG. 8 of vehicle 822 of FIG. 8 may generate map constraint 830 of FIG. 8. ADAS 824 may further transmit request 828 of FIG. 8, which may include an indication of map constraint 830 to map server 806. Map server 806 may transmit map 832 of FIG. 8 to vehicle 822. Map 812 may be different that map 832. Map server 806 may generate or select map 812 based on map constraint 810 and map server 806 may generate or select map 832 based on map constraint 830.
[0138] In some aspects, the map constraint may be determined based on computational resources of an advanced driver assistance system (ADAS) of the vehicle; the other map received by the other vehicle may be constrained according to another map constraint related to the other vehicle; and the other map constraint may be determined based on computational resources of an ADAS of the other vehicle. For example, ADAS 804 may determine map constraint 810 based on computational resources of ADAS 804. ADAS 824 may determine map constraint 830 based on computational resources of ADAS 824.
[0139] FIG. 10 is a flow diagram illustrating a process 1000 for providing map information, in accordance with aspects of the present disclosure. One or more operations of process 1000 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc. ) of the computing device. The computing device may be a mobile device (e.g., a mobile phone) , a network-connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the process 1000. The one or more operations of process 1000 may be implemented as software components that are executed and run on one or more processors.
[0140] At block 1002, a computing device (or one or more components thereof) may receive a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint. For example, map server 106 of FIG. 1 may receive request 108 of FIG. 1. Request 108 may include an indication of map constraint 110 of FIG. 1.
[0141] At block 1004, the computing device (or one or more components thereof) may obtain a map of the location constrained according to the map constraint. For example, map server 106 may obtain map 112 of FIG. 1. Map 112 may be constrained according to map constraint 110.
[0142] In some aspects, the map of the location may be, or may include, a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points. For example, map 112 may be, or may include, map points 502 of FIG. 5. Each map point of map points 502 may include a location 508, which may be, or may include, three-dimensional coordinates of the map point (e.g., in environment 114 of FIG. 1) . Map 112 may include descriptors 506 of FIG. 5. Each map point of map points 502 may be described by a plurality of descriptors 506.
[0143] In some aspects, the map of the location may be, or may include, a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes. For example, map 112 may include keyframes 510 of FIG. 5. Keyframes 510 may describe poses in environment 114. Map 112 may include a relationship between keyframes 510 and descriptors 506. For example, map 112 may include relationships between map points 502, keyframes 510, and descriptors 506. For instance map 112 may include relationships describing which map points are visible from which keyframes and which descriptors correspond a given map point and corresponding key frame.
[0144] In some aspects, to obtain the map of the location, the computing device (or one or more components thereof) may generate the map of the location based on the map constraint. For example, map server 106 may generate map 112 based on map constraint 110.
[0145] In some aspects, to obtain the map of the location, the computing device (or one or more components thereof) may select, based on the map constraint, the map of the location from among a plurality of maps of the location. For example, map server 106 may store several maps of the location. Map server 106 may select map 112 from among the several maps based on map 112 being constrained according to map constraint 110.
[0146] In some aspects, the request is received at a map server. The map may be a first map of the location. The first map of the location may be transmitted from the map server. The computing device (or one or more components thereof) may store, at the map server, a second map of the location, wherein the second map of the location has a larger data size than the first map of the location. For example, map server 106 may store a highly-detailed map of environment 114. Map server 106 may generate map 112 based on the highly-detailed map. In some aspects, the first map of the location may include fewer map points than the second map of the location. For example, map server 106 may select map points of map 112 from among the map points of the highly-detailed map. In some aspects, the first map of the location may include fewer descriptors than the second map of the location. For example, map server 106 may select descriptors of map points of map 112 from among descriptors of the highly-detailed map. In some aspects, the first map of the location may include indices into a descriptor codebook. For example, map server 106 may generate map 112 to include indices into a codebook. For instance, rather than include descriptors, map 112 may include indices into a codebook that includes descriptors. In some aspects, the first map of the location may include indices to observation regions. For example, map 112 may include indices of observation regions. For instance, rather than including keyframes, map 112 may include indices of observation regions that describe regions of the environment from which keypoints are viewable.
[0147] At block 1006, the computing device (or one or more components thereof) may cause at least one transmitter to transmit the map of the location to the vehicle. For example, map server 106 may transmit map 112 to vehicle 102 of FIG. 1.
[0148] In some aspects, the computing device (or one or more components thereof) may include or control the at least one transmitter. For example, map server 106 may include a transmitter.
[0149] In some examples, as noted previously, the methods described herein (e.g., process 900 of FIG. 9, process 1000 of FIG. 10 and / or other methods described herein) can be performed, in whole or in part, by a computing device or apparatus. In one example, one or more of the methods can be performed by ADAS 104 of vehicle 102 of FIG. 1, map server 106 of FIG. 1, ADAS 704 of vehicle 702 of FIG. 7, map server 706 of FIG. 7, ADAS 804 of vehicle 802 of FIG. 8, map server 806 of FIG. 8, ADAS 824 of vehicle 822 of FIG. 8, or by another system or device. In another example, one or more of the methods (e.g., process 900 of FIG. 9, process 1000 of FIG. 10 and / or other methods described herein) can be performed, in whole or in part, by the computing-device architecture 1300 shown in FIG. 13. For instance, a computing device with the computing-device architecture 1300 shown in FIG. 13 can include, or be included in, the components of the ADAS 104, map server 106, ADAS 704, map server 706, ADAS 804, map server 806, and / or ADAS 824 and can implement the operations of process 900, process 1000, and / or other process described herein. In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component (s) that are configured to carry out the steps of processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component (s) . The network interface can be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.
[0150] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs) , digital signal processors (DSPs) , central processing units (CPUs) , and / or other suitable electronic circuits) , and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0151] Process 900, process 1000, and / or other process described herein are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.
[0152] Additionally, process 900, process 1000, and / or other process described herein can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0153] As noted above, various aspects of the present disclosure can use machine-learning models or systems.
[0154] FIG. 11 is an illustrative example of a neural network 1100 (e.g., a deep-learning neural network) that can be used to implement machine-learning based feature segmentation, implicit-neural-representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc. ) , feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, neural network 1100 can implement, feature generation, for instance to generate descriptors of map points based on images.
[0155] An input layer 1102 includes input data. In one illustrative example, input layer 1102 can include data representing images of an environment. Neural network 1100 includes multiple hidden layers, for example, hidden layers 1106a, 1106b, through 1106n. The hidden layers 1106a, 1106b, through hidden layer 1106n include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. Neural network 1100 further includes an output layer 1104 that provides an output resulting from the processing performed by the hidden layers 1106a, 1106b, through 1106n. In one illustrative example, output layer 1104 can provide descriptors of keypoints of images of the environment.
[0156] Neural network 1100 may be, or may include, a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, neural network 1100 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some cases, neural network 1100 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.
[0157] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of input layer 1102 can activate a set of nodes in the first hidden layer 1106a. For example, as shown, each of the input nodes of input layer 1102 is connected to each of the nodes of the first hidden layer 1106a. The nodes of first hidden layer 1106a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 1106b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 1106b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 1106n can activate one or more nodes of the output layer 1104, at which an output is provided. In some cases, while nodes (e.g., node 1108) in neural network 1100 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.
[0158] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of neural network 1100. Once neural network 1100 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset) , allowing neural network 1100 to be adaptive to inputs and able to learn as more and more data is processed.
[0159] Neural network 1100 may be pre-trained to process the features from the data in the input layer 1102 using the different hidden layers 1106a, 1106b, through 1106n in order to provide the output through the output layer 1104. In an example in which neural network 1100 is used to identify features in images, neural network 1100 can be trained using training data that includes both images and labels, as described above. For instance, training images can be input into the network, with each training image having a label indicating the features in the images (for the feature-segmentation machine-learning system) or a label indicating classes of an activity in each image. In one example using object classification for illustrative purposes, a training image can include an image of a number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0 0] .
[0160] In some cases, neural network 1100 can adjust the weights of the nodes using a training process called backpropagation. As noted above, a backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until neural network 1100 is trained well enough so that the weights of the layers are accurately tuned.
[0161] For the example of identifying objects in images, the forward pass can include passing a training image through neural network 1100. The weights are initially randomized before neural network 1100 is trained. As an illustrative example, an image can include an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like) .
[0162] As noted above, for a first training iteration for neural network 1100, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1) . With the initial weights, neural network 1100 is unable to determine low-level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a cross-entropy loss. Another example of a loss function includes the mean squared error (MSE) , defined as The loss can be set to be equal to the value of Etotal.
[0163] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. Neural network 1100 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as where w denotes a weight, wi denotes the initial weight, and η denotes a learning rate. The learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.
[0164] Neural network 1100 can include any suitable deep network. One example includes a convolutional neural network (CNN) , which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling) , and fully connected layers. Neural network 1100 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs) , a Recurrent Neural Networks (RNNs) , among others.
[0165] FIG. 12 is an illustrative example of a convolutional neural network (CNN) 1200. The input layer 1202 of the CNN 1200 includes data representing an image or frame. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like) . The image can be passed through a convolutional hidden layer 1204, an optional non-linear activation layer, a pooling hidden layer 1206, and fully connected layer 1208 (which fully connected layer 1208 can be hidden) to get an output at the output layer 1210. While only one of each hidden layer is shown in FIG. 12, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 1200. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.
[0166] The first layer of the CNN 1200 can be the convolutional hidden layer 1204. The convolutional hidden layer 1204 can analyze image data of the input layer 1202. Each node of the convolutional hidden layer 1204 is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 1204 can be considered as one or more filters (each filter corresponding to a different activation or feature map) , with each convolutional iteration of a filter being a node or neuron of the convolutional hidden layer 1204. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28×28 array, and each filter (and corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 1204. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptive field in the input image. Each node of the convolutional hidden layer 1204 will have the same weights and bias (called a shared weight and a shared bias) . For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for an image frame example (according to three color components of the input image) . An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.
[0167] The convolutional nature of the convolutional hidden layer 1204 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 1204 can begin in the top-left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 1204. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top-left corner of the input image array) . The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 1204. For example, a filter can be moved by a step amount (referred to as a stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 1204.
[0168] The mapping from the input layer to the convolutional hidden layer 1204 is referred to as an activation map (or feature map) . The activation map includes a value for each node representing the filter results at each location of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a stride of 1) of a 28 x 28 input image. The convolutional hidden layer 1204 can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 12 includes three activation maps. Using three activation maps, the convolutional hidden layer 1204 can detect three different kinds of features, with each feature being detectable across the entire image.
[0169] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 1204. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f (x) = max (0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the non-linear properties of the CNN 1200 without affecting the receptive fields of the convolutional hidden layer 1204.
[0170] The pooling hidden layer 1206 can be applied after the convolutional hidden layer 1204 (and after the non-linear hidden layer when used) . The pooling hidden layer 1206 is used to simplify the information in the output from the convolutional hidden layer 1204. For example, the pooling hidden layer 1206 can take each activation map output from the convolutional hidden layer 1204 and generates a condensed activation map (or feature map) using a pooling function. Max-pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 1206, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 1204. In the example shown in FIG. 12, three pooling filters are used for the three activation maps in the convolutional hidden layer 1204.
[0171] In some examples, max-pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a stride (e.g., equal to a dimension of the filter, such as a stride of 2) to an activation map output from the convolutional hidden layer 1204. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes in the previous layer (with each node being a value in the activation map) . For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max-pooling filter is applied to an activation filter from the convolutional hidden layer 1204 having a dimension of 24x24 nodes, the output from the pooling hidden layer 1206 will be an array of 12x12 nodes.
[0172] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2×2 region (or other suitable region) of an activation map (instead of computing the maximum values as is done in max-pooling) and using the computed values as an output.
[0173] The pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image and discards the exact positional information. This can be done without affecting results of the feature detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max-pooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 1200.
[0174] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 1206 to every one of the output nodes in the output layer 1210. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 1204 includes 3×24×24 hidden feature nodes based on application of a 5×5 local receptive field (for the filters) to three activation maps, and the pooling hidden layer 1206 includes a layer of 3×12×12 hidden feature nodes based on application of max-pooling filter to 2×2 regions across each of the three feature maps. Extending this example, the output layer 1210 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 1206 is connected to every node of the output layer 1210.
[0175] The fully connected layer 1208 can obtain the output of the previous pooling hidden layer 1206 (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 1208 can determine the high-level features that most strongly correlate to a particular class and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 1208 and the pooling hidden layer 1206 to obtain probabilities for the different classes. For example, if the CNN 1200 is being used to predict that an object in an image is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the object, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person) .
[0176] In some examples, the output from the output layer 1210 can include an M-dimensional vector (in the prior example, M=10) . M indicates the number of classes that the CNN 1200 has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the M-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0] , the vector indicates that there is a 5%probability that the image is the third class of object (e.g., a dog) , an 80%probability that the image is the fourth class of object (e.g., a human) , and a 15%probability that the image is the sixth class of object (e.g., a kangaroo) . The probability for a class can be considered a confidence level that the object is part of that class.
[0177] FIG. 13 illustrates an example computing-device architecture 1300 of an example computing device which can implement the various techniques described herein. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a personal computer, a laptop computer, a video server, a vehicle (or computing device of a vehicle) , or other device. For example, the computing-device architecture 1300 may include, implement, or be included in any or all of ADAS 104 of vehicle 102 of FIG. 1, map server 106 of FIG. 1, ADAS 704 of vehicle 702 of FIG. 7, map server 706 of FIG. 7, ADAS 804 of vehicle 802 of FIG. 8, map server 806 of FIG. 8, ADAS 824 of vehicle 822 of FIG. 8. Additionally or alternatively, computing-device architecture 1300 may be configured to perform process 900, and / or other process described herein.
[0178] The components of computing-device architecture 1300 are shown in electrical communication with each other using connection 1312, such as a bus. The example computing-device architecture 1300 includes a processing unit (CPU or processor) 1302 and computing device connection 1312 that couples various computing device components including computing device memory 1310, such as read only memory (ROM) 1308 and random-access memory (RAM) 1306, to processor 1302.
[0179] Computing-device architecture 1300 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1302. Computing-device architecture 1300 can copy data from memory 1310 and / or the storage device 1314 to cache 1304 for quick access by processor 1302. In this way, the cache can provide a performance boost that avoids processor 1302 delays while waiting for data. These and other modules can control or be configured to control processor 1302 to perform various actions. Other computing device memory 1310 may be available for use as well. Memory 1310 can include multiple different types of memory with different performance characteristics. Processor 1302 can include any general-purpose processor and a hardware or software service, such as service 1 1316, service 2 1318, and service 3 1320 stored in storage device 1314, configured to control processor 1302 as well as a special-purpose processor where software instructions are incorporated into the processor design. Processor 1302 may be a self-contained system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0180] To enable user interaction with the computing-device architecture 1300, input device 1322 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. Output device 1324 can also be one or more of a number of output mechanisms known to those of skill in the art, such as a display, projector, television, speaker device, etc. In some instances, multimodal computing devices can enable a user to provide multiple types of input to communicate with computing-device architecture 1300. Communication interface 1326 can generally govern and manage the user input and computing device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0181] Storage device 1314 is a non-volatile memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random-access memories (RAMs) 1306, read only memory (ROM) 1308, and hybrids thereof. Storage device 1314 can include services 1316, 1318, and 1320 for controlling processor 1302. Other hardware or software modules are contemplated. Storage device 1314 can be connected to the computing device connection 1312. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1302, connection 1312, output device 1324, and so forth, to carry out the function.
[0182] The term “substantially, ” in reference to a given parameter, property, or condition, may refer to a degree that one of ordinary skill in the art would understand that the given parameter, property, or condition is met with a small degree of variance, such as, for example, within acceptable manufacturing tolerances. By way of example, depending on the particular parameter, property, or condition that is substantially met, the parameter, property, or condition may be at least 90%met, at least 95%met, or even at least 99%met.
[0183] Aspects of the present disclosure are applicable to any suitable electronic device (such as security systems, smartphones, tablets, laptop computers, vehicles, drones, or other devices) including or coupled to one or more active depth sensing systems. While described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to specific devices.
[0184] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, one processing system and so on) . As used herein, a device may be any electronic device with one or more parts that may implement at least some portions of this disclosure. While the below description and examples use the term “device” to describe various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. While the below description and examples use the term “system” to describe various aspects of this disclosure, the term “system” is not limited to a specific configuration, type, or number of objects.
[0185] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by one of ordinary skill in the art that the aspects may be practiced without these specific details. For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blocks including devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.
[0186] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0187] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.
[0188] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction (s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD) , flash memory, magnetic or optical disks, USB devices provided with non-volatile memory, networked storage devices, any suitable combination thereof, among others. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0189] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0190] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor (s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0191] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0192] In the foregoing description, aspects of the application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0193] One of ordinary skill will appreciate that the less than ( “<” ) and greater than ( “>” ) symbols or terminology used herein can be replaced with less than or equal to ( “≤” ) and greater than or equal to ( “≥” ) symbols, respectively, without departing from the scope of this description.
[0194] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0195] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0196] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on) , or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0197] Claim language or other language reciting “at least one processor configured to, ” “at least one processor being configured to, ” “one or more processors configured to, ” “one or more processors being configured to, ” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation (s) . For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0198] Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0199] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method) , the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function) .
[0200] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0201] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general-purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium including program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random-access memory (RAM) such as synchronous dynamic random-access memory (SDRAM) , read-only memory (ROM) , non-volatile random-access memory (NVRAM) , electrically erasable programmable read-only memory (EEPROM) , flash memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0202] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs) , general-purpose microprocessors, an application specific integrated circuits (ASICs) , field programmable logic arrays (FPGAs) , or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor, ” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0203] Illustrative aspects of the disclosure include:
[0204] Aspect 1. An apparatus for obtaining map information, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: determine a map constraint related to a vehicle; cause at least one transmitter to transmit a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and receive the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0205] Aspect 2. The apparatus of aspect 1, wherein the map of the location comprises: a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.
[0206] Aspect 3. The apparatus of aspect 2, wherein the map of the location further comprises: a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.
[0207] Aspect 4. The apparatus of any one of aspects 1 to 3, wherein the map constraint is determined based on a requirement of an advanced driver assistance system (ADAS) of the vehicle.
[0208] Aspect 5. The apparatus of aspect 4, wherein the requirement of the ADAS is based on computational resources of the ADAS.
[0209] Aspect 6. The apparatus of any one of aspects 1 to 5, wherein, the map constraint comprises a first map constraint, wherein the request comprises a first request, and wherein the map of the location comprises a first map of the location, wherein the at least one processor is configured to: determine a second map constraint related to the vehicle; cause the at least one transmitter to transmit a second request for a second map of the location to the map server, wherein the second request comprises an indication of the second map constraint; and receive the second map of the location, constrained according to the second map constraint, from the map server, wherein the second map of the location is either generated based on the second map constraint or selected from among a plurality of maps of the location based on the second map constraint.
[0210] Aspect 7. The apparatus of aspect 6, wherein: the first map constraint is determined based on a first requirement of an advanced driver assistance system (ADAS) of the vehicle; and the second map constraint is determined based on a second requirement of the ADAS.
[0211] Aspect 8. The apparatus of any one of aspects 1 to 7, wherein the request comprises a first request and wherein the map of the location comprises a map of a first location, wherein the at least one processor is configured to: cause the at least one transmitter to transmit a second request for a map of a second location to the map server, wherein the second request comprises an indication of the map constraint; and receive the map of the second location, constrained according to the map constraint, from the map server, wherein the map of the second location is either generated based on the second map constraint or selected from among a plurality of maps of the second location based on the second map constraint.
[0212] Aspect 9. The apparatus of any one of aspects 1 to 8, wherein the map of the location received by the vehicle is different from another map of the location received by another vehicle from the map server.
[0213] Aspect 10. The apparatus of aspect 9, wherein: the map constraint is determined based on computational resources of an advanced driver assistance system (ADAS) of the vehicle; the other map received by the other vehicle is constrained according to another map constraint related to the other vehicle; and the other map constraint is determined based on computational resources of an ADAS of the other vehicle.
[0214] Aspect 11. The apparatus of any one of aspects 1 to 10, wherein the map constraint comprises at least one of: a map-size constraint related to a size of the map; a map-point constraint related to a number of map points of a point cloud of the map; a descriptor constraint related to a number of descriptors of the map or a data size of descriptors; or a keyframe constraint related to a number of keyframes of the map.
[0215] Aspect 12. The apparatus of any one of aspects 1 to 11, further comprising the at least one transmitter.
[0216] Aspect 13. A method for obtaining map information, the method comprising: determining a map constraint related to a vehicle; transmitting a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; and receiving the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.
[0217] Aspect 14. The method of aspect 13, wherein the map of the location comprises: a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.
[0218] Aspect 15. The method of aspect 14, wherein the map of the location further comprises: a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.
[0219] Aspect 16. The method of any one of aspects 13 to 15, wherein the map constraint is determined based on a requirement of an advanced driver assistance system (ADAS) of the vehicle.
[0220] Aspect 17. The method of aspect 16, wherein the requirement of the ADAS is based on computational resources of the ADAS.
[0221] Aspect 18. The method of any one of aspects 13 to 17, wherein, the map constraint comprises a first map constraint, wherein the request comprises a first request, and wherein the map of the location comprises a first map of the location, the method further comprising: determining a second map constraint related to the vehicle; transmitting a second request for a second map of the location to the map server, wherein the second request comprises an indication of the second map constraint; and receiving the second map of the location, constrained according to the second map constraint, from the map server, wherein the second map of the location is either generated based on the second map constraint or selected from among a plurality of maps of the location based on the second map constraint.
[0222] Aspect 19. The method of aspect 18, wherein: the first map constraint is determined based on a first requirement of an advanced driver assistance system (ADAS) of the vehicle; and the second map constraint is determined based on a second requirement of the ADAS.
[0223] Aspect 20. The method of any one of aspects 13 to 19, wherein the request comprises a first request and wherein the map of the location comprises a map of a first location, the method further comprising: transmitting a second request for a map of a second location to the map server, wherein the second request comprises an indication of the map constraint; and receiving the map of the second location, constrained according to the map constraint, from the map server, wherein the map of the second location is either generated based on the second map constraint or selected from among a plurality of maps of the second location based on the second map constraint.
[0224] Aspect 21. The method of any one of aspects 13 to 20, wherein the map of the location received by the vehicle is different from another map of the location received by another vehicle from the map server.
[0225] Aspect 22. The method of aspect 21, wherein: the map constraint is determined based on computational resources of an advanced driver assistance system (ADAS) of the vehicle; the other map received by the other vehicle is constrained according to another map constraint related to the other vehcile; and the other map constraint is determined based on computational resources of an ADAS of the other vehicle.
[0226] Aspect 23. The method of any one of aspects 13 to 22, wherein the map constraint comprises at least one of: a map-size constraint related to a size of the map; a map-point constraint related to a number of map points of a point cloud of the map; a descriptor constraint related to a number of descriptors of the map or a data size of descriptors; or a keyframe constraint related to a number of keyframes of the map.
[0227] Aspect 24. An apparatus for providing map information, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; obtain a map of the location constrained according to the map constraint; and cause at least one transmitter to transmit the map of the location to the vehicle.
[0228] Aspect 25. The apparatus of aspect 24, wherein the map comprises: a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.
[0229] Aspect 26. The apparatus of aspect 25, wherein the map further comprises: a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.
[0230] Aspect 27. The apparatus of any one of aspects 24 to 26, wherein, to obtain the map of the location, the at least one processor is configured to generate the map of the location based on the map constraint.
[0231] Aspect 28. The apparatus of any one of aspects 24 to 27, wherein, to obtain the map of the location, the at least one processor is configured to select, based on the map constraint, the map of the location from among a plurality of maps of the location.
[0232] Aspect 29. The apparatus of any one of aspects 24 to 28, wherein the request is received at a map server, the map is a first map of the location, and the first map of the location is transmitted from the map server, wherein the at least one processor is configured to: store, at the map server, a second map of the location, wherein the second map of the location has a larger data size than the first map of the location.
[0233] Aspect 30. The apparatus of aspect 29, wherein the first map of the location comprises fewer map points than the second map of the location.
[0234] Aspect 31. The apparatus of any one of aspects 29 or 30, wherein the first map of the location comprises fewer descriptors than the second map of the location.
[0235] Aspect 32. The apparatus of any one of aspects 29 to 31, wherein the first map of the location comprises indices into a descriptor codebook.
[0236] Aspect 33. The apparatus of any one of aspects 29 to 32, wherein the first map of the location comprises indices to observation regions.
[0237] Aspect 34. The apparatus of any one of aspects 24 to 33, further comprising the at least one transmitter.
[0238] Aspect 35. A method for providing map information, the method comprising: receiving a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint; obtaining a map of the location constrained according to the map constraint; and transmitting the map of the location to the vehicle.
[0239] Aspect 36. The method of aspect 35, wherein the map comprises: a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; and a plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.
[0240] Aspect 37. The method of aspect 36, wherein the map further comprises: a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; and a relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.
[0241] Aspect 38. The method of any one of aspects 35 to 37, wherein obtaining the map of the location comprises generating the map of the location based on the map constraint.
[0242] Aspect 39. The method of any one of aspects 35 to 38, wherein obtaining the map of the location comprises selecting, based on the map constraint, the map of the location from among a plurality of maps of the location.
[0243] Aspect 40. The method of any one of aspects 35 to 39, wherein the request is received at a map server, the map is a first map of the location, and the first map of the location is transmitted from the map server, the method further comprising: storing, at the map server, a second map of the location, wherein the second map of the location has a larger data size than the first map of the location.
[0244] Aspect 41. The method of aspect 40, wherein the first map of the location comprises fewer map points than the second map of the location.
[0245] Aspect 42. The method of any one of aspects 40 or 41, wherein the first map of the location comprises fewer descriptors than the second map of the location.
[0246] Aspect 43. The method of any one of aspects 40 to 42, wherein the first map of the location comprises indices into a descriptor codebook.
[0247] Aspect 44. The method of any one of aspects 40 to 43, wherein the first map of the location comprises indices to observation regions.
[0248] Aspect 45. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations according to any of aspects 13 to 23 or 35 to 44.
[0249] Aspect 46. An apparatus for providing virtual content for display, the apparatus comprising one or more means for perform operations according to any of aspects 13 to 23 or 35 to 44.
Claims
1.An apparatus for obtaining map information, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:determine a map constraint related to a vehicle;cause at least one transmitter to transmit a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; andreceive the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.2.The apparatus of claim 1, wherein the map of the location comprises:a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; anda plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.3.The apparatus of claim 2, wherein the map of the location further comprises:a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; anda relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.4.The apparatus of claim 1, wherein the map constraint is determined based on a requirement of an advanced driver assistance system (ADAS) of the vehicle.5.The apparatus of claim 4, wherein the requirement of the ADAS is based on computational resources of the ADAS.6.The apparatus of claim 1, wherein, the map constraint comprises a first map constraint, wherein the request comprises a first request, and wherein the map of the location comprises a first map of the location, wherein the at least one processor is configured to:determine a second map constraint related to the vehicle;cause the at least one transmitter to transmit a second request for a second map of the location to the map server, wherein the second request comprises an indication of the second map constraint; andreceive the second map of the location, constrained according to the second map constraint, from the map server, wherein the second map of the location is either generated based on the second map constraint or selected from among a plurality of maps of the location based on the second map constraint.7.The apparatus of claim 6, wherein:the first map constraint is determined based on a first requirement of an advanced driver assistance system (ADAS) of the vehicle; andthe second map constraint is determined based on a second requirement of the ADAS.8.The apparatus of claim 1, wherein the request comprises a first request and wherein the map of the location comprises a map of a first location, wherein the at least one processor is configured to:cause the at least one transmitter to transmit a second request for a map of a second location to the map server, wherein the second request comprises an indication of the map constraint; andreceive the map of the second location, constrained according to the map constraint, from the map server, wherein the map of the second location is either generated based on the second map constraint or selected from among a plurality of maps of the second location based on the second map constraint.9.The apparatus of claim 1, wherein the map of the location received by the vehicle is different from another map of the location received by another vehicle from the map server.10.The apparatus of claim 9, wherein:the map constraint is determined based on computational resources of an advanced driver assistance system (ADAS) of the vehicle;the other map received by the other vehicle is constrained according to another map constraint related to the other vehicle; andthe other map constraint is determined based on computational resources of an ADAS of the other vehicle.11.The apparatus of claim 1, wherein the map constraint comprises at least one of:a map-size constraint related to a size of the map;a map-point constraint related to a number of map points of a point cloud of the map;a descriptor constraint related to a number of descriptors of the map or a data size of descriptors; ora keyframe constraint related to a number of keyframes of the map.12.The apparatus of claim 1, further comprising the at least one transmitter.13.A method for obtaining map information, the method comprising:determining a map constraint related to a vehicle;transmitting a request for a map of a location to a map server, wherein the request comprises an indication of the map constraint; andreceiving the map of the location, constrained according to the map constraint, from the map server, wherein the map of the location is either generated based on the map constraint or selected from among a plurality of maps of the location based on the map constraint.14.The method of claim 13, wherein the map of the location comprises:a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; anda plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.15.The method of claim 14, wherein the map of the location further comprises:a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; anda relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.16.The method of claim 13, wherein the map constraint is determined based on a requirement of an advanced driver assistance system (ADAS) of the vehicle.17.The method of claim 16, wherein the requirement of the ADAS is based on computational resources of the ADAS.18.The method of claim 13, wherein, the map constraint comprises a first map constraint, wherein the request comprises a first request, and wherein the map of the location comprises a first map of the location, the method further comprising:determining a second map constraint related to the vehicle;transmitting a second request for a second map of the location to the map server, wherein the second request comprises an indication of the second map constraint; andreceiving the second map of the location, constrained according to the second map constraint, from the map server, wherein the second map of the location is either generated based on the second map constraint or selected from among a plurality of maps of the location based on the second map constraint.19.The method of claim 18, wherein:the first map constraint is determined based on a first requirement of an advanced driver assistance system (ADAS) of the vehicle; andthe second map constraint is determined based on a second requirement of the ADAS.20.The method of claim 13, wherein the request comprises a first request and wherein the map of the location comprises a map of a first location, the method further comprising:transmitting a second request for a map of a second location to the map server, wherein the second request comprises an indication of the map constraint; andreceiving the map of the second location, constrained according to the map constraint, from the map server, wherein the map of the second location is either generated based on the second map constraint or selected from among a plurality of maps of the second location based on the second map constraint.21.The method of claim 13, wherein the map of the location received by the vehicle is different from another map of the location received by another vehicle from the map server.22.The method of claim 21, wherein:the map constraint is determined based on computational resources of an advanced driver assistance system (ADAS) of the vehicle;the other map received by the other vehicle is constrained according to another map constraint related to the other vehcile; andthe other map constraint is determined based on computational resources of an ADAS of the other vehicle.23.The method of claim 13, wherein the map constraint comprises at least one of:a map-size constraint related to a size of the map;a map-point constraint related to a number of map points of a point cloud of the map;a descriptor constraint related to a number of descriptors of the map or a data size of descriptors; ora keyframe constraint related to a number of keyframes of the map.24.An apparatus for providing map information, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:receive a request for a map of a location from a vehicle, wherein the request comprises an indication of a map constraint;obtain a map of the location constrained according to the map constraint; andcause at least one transmitter to transmit the map of the location to the vehicle.25.The apparatus of claim 24, wherein the map comprises:a point-cloud map of the location comprising a plurality of map points, wherein each map point of the plurality of map points comprises respective three-dimensional coordinates; anda plurality of descriptors, wherein each descriptor of the plurality of descriptors is descriptive of a respective map point of the plurality of map points.26.The apparatus of claim 25, wherein the map further comprises:a plurality of keyframes, wherein each keyframe of the plurality of keyframes describes a pose relative to the location; anda relationship between descriptors of the plurality of descriptors and keyframes of the plurality of keyframes.27.The apparatus of claim 24, wherein, to obtain the map of the location, the at least one processor is configured to generate the map of the location based on the map constraint.28.The apparatus of claim 24, wherein, to obtain the map of the location, the at least one processor is configured to select, based on the map constraint, the map of the location from among a plurality of maps of the location.29.The apparatus of claim 24, wherein the request is received at a map server, the map is a first map of the location, and the first map of the location is transmitted from the map server, wherein the at least one processor is configured to:store, at the map server, a second map of the location, wherein the second map of the location has a larger data size than the first map of the location.30.The apparatus of claim 29, wherein the first map of the location comprises fewer map points than the second map of the location.
Citation Information
Patent Citations
Use of map data difference tiles to iteratively provide map data to a client device
CN105359189A
Map updating system and method for automatic driving
CN109783593A
Scalable 3D mapping system
CN111767358A
Electronic device and vehicle control method of electronic device, server and method for providing precise map data of server
CN112740134A
Progressive map maintenance at a mobile navigation unit
US20170254665A1