Road point diagram generation using semantic map information for route planning
Through the road map generation technology based on semantic maps, the problem of inability to accurately reflect the cost of navigable surfaces in the prior art is solved, and accurate route planning is generated in multiple surface environments and the optimal path is identified.
Patent Information
- Application Number
- CN202510103303.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-29
AI Technical Summary
Existing route planning techniques, when dealing with many different types of navigable surfaces, cannot accurately reflect the cost of robots moving in the environment, resulting in the generation of suboptimal route planning.
By generating a path map based on semantic maps, using semantic information to determine the cost of the road map edge, considering the types and weights of different navigable surfaces, an accurate route plan is generated.
It realizes the generation of accurate route planning in a variety of navigable surface environments, identifying the best routes, and improving the accuracy and efficiency of route planning.
Smart Images

Figure CN120385328A_ABST
Abstract
Description
Background Art
[0001] To traverse an environment, many autonomous or semi-autonomous mobile robots cross the environment to a given destination according to a route determined by a route planner. The route planner may use a waypoint graph to determine the route, where the waypoint graph represents regions in the physical environment that the robot can cross. The waypoint graph also represents regions that the robot cannot cross, such as physical obstacles. The physical environment may be a building, such as a warehouse, and the waypoint graph may be, for example, a map of the building floor area. As another example, the waypoint graph may represent a road network for vehicle traffic, and the given destination may be a street address.
[0002] In existing methods, the route planner uses an occupancy map to generate the waypoint graph. The occupancy map represents the physical environment as a grid of cells and indicates whether each cell is occupied by an obstacle. The route planner uses the waypoint graph to determine a route through the environment. The waypoint graph includes a set of vertices representing positions in the environment and a set of edges connecting the vertices. The edges represent navigable surfaces that the robot can physically occupy. The edges are associated with costs, which may correspond to the travel time between the vertices. However, there can be different types of navigable surfaces in the environment, and different types of navigable surfaces can have different effects on route planning. In existing methods, the waypoint graph does not include information about the characteristics of the navigable surfaces related to how the robot moves in the environment. Thus, for example, the edge cost cannot accurately represent the travel time on a surface (such as a ramp) where the robot decelerates. Using inaccurate edge costs may cause the route planner to generate suboptimal route plans. For example, the robot moves at a reduced speed on a particular ramp, but the edge cost of the ramp in the prior art is not adjusted to reflect the reduced speed. Therefore, the actual navigation cost of the ramp is greater than the cost used by the route planner, and the actual travel time of the route through the ramp is greater than the expected travel time used by the route planner. A route plan generated based on inaccurate costs may not be optimal because such a route plan is likely to be selected in place of other route plans with lower true costs.
[0003] Therefore, more effective techniques are needed to generate route plans in an environment with multiple different types of navigable surfaces that have different effects on route planning in an autonomous or semi-autonomous system. Summary of the Invention
[0004] Embodiments of the present disclosure relate to generating a route plan based on a waypoint graph, where the edge costs in the waypoint graph are determined by a semantic graph. The techniques described herein include generating a route plan by at least receiving, at a computing device, a semantic graph representing a physical environment. The techniques may also include generating a roadmap based on the semantic graph, where the roadmap includes one or more roadmap edges, and each roadmap edge has an associated location in one or more graph regions of the semantic graph and an associated cost. The techniques may also include determining the cost of a particular roadmap edge based on a region type, where the region type is determined based on the graph region in which the particular roadmap edge is located. The techniques may also include using the roadmap to generate a route plan for a mobile robot to move from a given starting position on the semantic graph to a given target position on the semantic graph.
[0005] One technical advantage of the disclosed techniques over prior methods is the ability to generate a route plan based on semantic information associated with waypoint graph edges, such that the route plan accurately reflects the navigation cost of each region represented by the waypoint graph edges. For example, the cost of each edge is based on the type of navigable surface represented by the edge. Using edge costs based on the type of navigable surface allows the disclosed graph generator to accurately identify the best path between given locations. These technical advantages represent one or more technical improvements over prior methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present system and method for generating a route plan based on a waypoint graph for a robotic system and application are described in detail below with reference to the accompanying drawings, in which:
[0007] Figure 1 A computing device configured to implement one or more aspects of the various embodiments is shown;
[0008] Figure 2 is a more detailed illustration of a Figure 1 waypoint graph generator according to the various embodiments;
[0009] Figure 3A An example Voronoi diagram on an example semantic graph according to the various embodiments is shown;
[0010] Figure 3B An example Voronoi diagram and a corresponding example waypoint graph according to the various embodiments are shown;
[0011] Figure 4 A flowchart of a method for generating a route plan using a waypoint graph having edge costs determined based on region information from a semantic graph according to the various embodiments is shown;
[0012] Figure 5A An illustration of an example autonomous vehicle according to some embodiments of the present disclosure;
[0013] Figure 5B Examples of camera positions and fields of view of exemplary autonomous vehicles in accordance with some embodiments of the present disclosure Figure 5A ;
[0014] Figure 5C Examples of exemplary system architectures of exemplary autonomous vehicles in accordance with some embodiments of the present disclosure Figure 5A ;
[0015] Figure 5D System diagrams of communications between a cloud-based server and an exemplary autonomous vehicle in accordance with some embodiments of the present disclosure Figure 5A ;
[0016] Figure 6 Block diagrams of exemplary computing devices suitable for implementing some embodiments of the present disclosure; and
[0017] Figure 7 Block diagrams of exemplary data centers suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Systems and methods for generating a waypoint graph with edge costs determined based on region information provided by a semantic map are disclosed. Although the present disclosure may be described with respect to exemplary autonomous or semi-autonomous vehicles or machines 500 (alternatively referred to herein as "vehicle 500" or "self-machine 500"), examples of which are described with respect to Figures 5A - 5D , this is not limiting. For example, the systems and methods described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, vessels, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other vehicle types. Additionally, although the present disclosure may be described with respect to monitoring sensor performance in autonomous and / or semi-autonomous vehicles, this is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field in which sensor monitoring may be used.
[0019] Figure 1Computing device 100 configured to implement one or more aspects of various embodiments is shown. In at least one embodiment, computing device 100 includes a desktop computer, a laptop computer, a smartphone, a personal digital assistant (PDA), a tablet computer, a server, one or more virtual machines, an embedded system, an embedded hardware module including a system-on-chip, a system-on-chip, a computing system of an autonomous, semi-autonomous or non-autonomous machine, and / or any other type of computing device configured to receive input, process data, and optionally display images, and suitable for practicing one or more embodiments. Computing device 100 includes a memory 116, one or more processors 102, an interconnect 112, a storage device 114, an input / output (I / O) device interface 104, and a network interface 106. Computing device 100 also includes one or more I / O devices 108 and / or communicates with one or more I / O devices 108. I / O device 108 communicates with interconnect 112 via I / O device interface 104. Memory 116 may be volatile random access memory or other suitable type of memory.
[0020] In a particular embodiment, waypoint map generator 122, waypoint map 124, and route planner 126 are stored in memory 116. Waypoint map generator 122 generates waypoint map 124, which has waypoint map edges representing navigable regions of a semantic map. Waypoint map 124 also includes a set of waypoint map vertices that are connected by waypoint map edges. Thus, waypoint map edges represent path segments that a robot can follow to reach each navigable region.
[0021] Waypoint map generator 122 determines waypoint map edge costs for the waypoint map edges. Each waypoint map edge cost (“waypoint edge cost”) represents the cost of navigating through the corresponding navigable region represented by the waypoint map edge. For example, the waypoint edge cost can be proportional to the expected amount of time required for a robot to traverse the corresponding navigable region. The waypoint edge cost is determined based on semantic information associated with the navigable regions of the semantic map. The semantic information may include a region type characterizing the region. The region type may indicate that the region is a generally navigable surface with a relatively low cost, or a surface with a higher cost, such as a room, a narrow corridor, a ramp, or a room entrance. The waypoint edge cost for a particular edge can be determined by multiplying the length of the corresponding navigable region by a weight value associated with the region type of the corresponding navigable region. If a particular waypoint map edge traverses multiple different types of navigable regions, the waypoint edge cost can be determined as a weighted sum of the lengths of the navigable regions, where the lengths are weighted by the respective weight values associated with the respective region types of the navigable regions. Waypoint map 124 (including edge costs) is provided to route planner 126, which generates a route between a given start waypoint vertex and a destination waypoint vertex based on the waypoint map edges of waypoint map 124 and the waypoint edge costs associated with the waypoint map vertices.
[0022] It should be noted that the computing device described herein is illustrative, and any other technically feasible configuration falls within the scope of the present disclosure. For example, multiple instances of the waypoint map generator 122 and / or the route planner 126 can be executed on a set of nodes in a distributed and / or cloud computing system to implement the functions of the computing device 100. Alternatively, the computing device 100 can be implemented similar to the computing device of the example autonomous or semi-autonomous machine 500 described at least with reference to Figures 5A to 5D described example autonomous or semi-autonomous machine 500.
[0023] In at least one embodiment, the computing device 100 includes, but is not limited to, an interconnect (bus) 112 connecting one or more processors 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage device 114, and / or a network interface 106. The processor 102 can include any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, a deep learning accelerator (DLA), a parallel processing unit (PPU), a data processing unit (DPU), a vector or vision processing unit (VPU), a programmable vision accelerator (PVA), any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in cooperation with a GPU. Generally, the processor 102 can include any technically feasible hardware unit capable of processing data and / or executing software applications. Additionally, in the context of the present disclosure, the computing elements shown in the computing device 100 can correspond to a physical computing system (e.g., a system in a data center or a machine) and / or can correspond to a virtual computing instance executed in a computing cloud.
[0024] In at least one embodiment, the I / O device 108 includes devices capable of receiving input (e.g., a keyboard, a mouse, a touchpad, VR / MR / AR headsets, a gesture recognition system, a steering wheel, mechanical, digital, or touch-sensitive buttons or input components, and / or a microphone), and devices capable of providing output (e.g., a display device, a haptic device, and / or a speaker). Additionally, the I / O device 108 can include devices capable of both receiving input and providing output, such as a touch screen, a universal serial bus (USB) port, etc. The I / O device 108 can be configured to receive various types of input from an end user (e.g., a designer) of the computing device 100, and also provide various types of output to the end user of the computing device 100, such as a displayed digital image or digital video or text. In some embodiments, one or more I / O devices 108 are configured to couple the computing device 100 to a network 110.
[0025] In at least one embodiment, network 110 is a communication network of any technically feasible type that allows for the exchange of data between computing device 100 and internal, local, remote, or external entities or devices (such as a web server or another networked computing device). For example, network 110 may include a wide area network (WAN), a local area network (LAN), a wireless (e.g., WiFi) network, a cellular network, and / or the Internet, among others.
[0026] In at least one embodiment, storage device 114 includes non-volatile storage for applications and data and may include fixed or removable disk drives, flash devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. Waypoint map generator 122 and / or route planner 126 may be stored in storage device 114 and loaded into memory 116 when executed by processor 102.
[0027] In one embodiment, memory 116 includes random access memory (RAM) modules, flash memory cells, and / or any other type of memory cells or combinations thereof. Processor 102, I / O device interface 104, and network interface 106 may be configured to read data from and write data to memory 116. Memory 116 may include various software programs or more generally software code executable by processor 102 and application data associated with the software programs, including waypoint map generator 122 and / or route planner 126.
[0028] Figure 2 is in accordance with various embodiments Figure 1 A more detailed illustration of waypoint map generator 122. As shown, semantic map 202 is provided as an input to waypoint map generator 122. Semantic map 202 may be received from another component, such as a mapping system (not shown). Semantic map 202 may correspond to map information 594, an example of which is referenced Figure 5DA description is given. The semantic map 202 includes information about one or more map regions 204. For example, each region may represent a space that can be occupied by a robot. The semantic map 202 includes two or more map regions 204. Map region 204A includes a region boundary 206A and a region type 210A. The region boundary 206A identifies the location and shape of map region 204A on the semantic map. The region boundary 206A is specified by a set of boundary coordinates 208A, which may be the coordinates of the points of a polygon representing the shape of map region 204A. The term "point" as used herein refers to a location in a discrete representation of the environment. A point can be identified by coordinates, such as x and y or latitude and longitude coordinates in two dimensions, or x, y, z or latitude, longitude and altitude coordinates in three dimensions. In one example, a point corresponds to a cell in a grid of cells representing the environment.
[0029] Map region 204A may include other semantic information in addition to or in place of region type 210A. For example, map region 204A may include one or more time ranges that indicate the times when navigation in map region 204A is allowed. As another example, map region 204A may include a set of time ranges, and each time range may be associated with a weight value to be used within the corresponding time range. For example, a larger weight value (corresponding to a higher cost) may be associated with a time when map region 204A is expected to have more people and / or robot traffic. As another example, a smaller weight value (corresponding to a lower cost) may be associated with a time range when map region 204A is expected to have less traffic.
[0030] The semantic map 202 also includes map region 204N, which represents another region of the semantic map 202. Map region 204N includes a region boundary 206N, boundary coordinates 208N, and a region type 210N, which are similar to the corresponding region boundary 206A, boundary coordinates 208A, and region type 210A described herein.
[0031] The waypoint map generator 122 includes a merged surface generator 220, a Voronoi map generator 226, and a waypoint map creator 234. In some embodiments, the merged surface generator 220 receives the semantic map 202 and generates a merged surface 222 representing the map regions 204 of the semantic map 202. The merged surface 222 can be a polygon having edges for each region boundary 206 of the semantic map 202. For example, the merged surface generator 220 can generate the merged surface 222 by merging line segments and / or polygons specified by the boundary coordinates 208 of the semantic map 202 into one or more combined polygons (e.g., a single polygon). The combined polygon is represented as a set of polygon coordinates 224, which are the coordinates of points on the boundary of the combined polygon. In some embodiments, the merged surface 222 is suitable for use as an input to the Voronoi map generator 226, which requires one or more combined polygons as input. In other embodiments, the Voronoi map generator 226 accepts inputs in formats other than the merged surface 222. Thus, in other embodiments, the merged surface generator 220 can be omitted, and the boundary coordinates 208 of the map regions 204 of the semantic map 202 can be provided as an input to the Voronoi map generator 226, and the Voronoi map generator 226 can use the boundary coordinates 208 of the map regions 204 to generate a Voronoi map 228. The Voronoi map 228 includes Voronoi edges 320 that form paths in the map regions 204 along which a robot can reach a large portion of each map region. Although in the examples described herein, the paths through the map regions are generated as Voronoi edges of a Voronoi map, in other examples, any suitable technique can be used to generate paths through the map regions. For example, a search can be performed using any suitable pathfinding technique to identify map regions that can be occupied by the robot, and a graph representing the identified map regions can be generated.
[0032] The Voronoi diagram generator 226 generates a Voronoi diagram 228 based on the merged surface 222 (or other suitable representation of the graph regions of the semantic graph 202, such as the set of boundary coordinates 208). The Voronoi diagram 228 includes a set of Voronoi edges 230 that connect Voronoi vertices 232 and are located between the boundaries (e.g., walls) of the graph region. The Voronoi diagram 228 is a representation of a path that a robot can follow through the navigable region, so the route between a given start position and a destination position corresponds to a portion of the Voronoi diagram 228 that connects the start position and the destination position. To construct the Voronoi diagram, the Voronoi diagram generator 226 identifies the edges of the merged polygons specified by the polygon coordinates 224 of the merged surface 222. The edges of the merged polygons represent the shape and location of the navigable region (e.g., a room, a corridor, a room entrance, or generally a navigable surface). A Voronoi diagram generator algorithm that accepts line segments as input is used to generate a set of Voronoi edges 230 based on the input line segments corresponding to the edges of the merged polygons. Voronoi edges 240 correspond to line segments that connect vertices referred to herein as Voronoi vertices 232. Each point on a Voronoi edge 230 is equidistant from two or more boundaries of the regions in the merged polygons. In this way, the Voronoi edges 230 represent paths through the navigable regions of the semantic graph, and in the sense that the points on the paths are equidistant from two or more boundaries of the graph region, the paths pass through the middle of the graph region.
[0033] The waypoint graph generator 122 provides the Voronoi diagram 228 as input to the waypoint graph creator 234. In some embodiments, the waypoint graph creator 234 generates a waypoint graph 124 that has the same graph structure as the Voronoi diagram 228. The waypoint graph creator 234 can generate waypoint graph edges 240 corresponding to the Voronoi edges 230 and waypoint graph vertices 242 corresponding to the Voronoi vertices 232. In other embodiments, the waypoint graph creator 234 uses the Voronoi diagram 228 as the waypoint graph 124, in which case the Voronoi edges 230 are used as the waypoint graph edges 240 and the Voronoi vertices 232 are used as the waypoint graph vertices 242, and the waypoint graph edges 240 and waypoint graph vertices 242 refer to the Voronoi edges 230 and Voronoi vertices 232, respectively.
[0034] The waypoint graph creator 234 includes an edge cost determiner 236 that determines the waypoint graph edge cost 244 of the waypoint graph edge 240. The graph edge cost 244 is determined based on semantic information associated with regions of the semantic graph. The semantic information includes a region type associated with each graph region. The region type can indicate that the associated graph region is a generally navigable surface that can be navigated without changing speed, or a surface type suitable for changing speed, such as a narrow corridor, a room entrance, a room, or a ramp. The waypoint graph generator determines the cost of each graph edge based on the region type of one or more regions through which the graph edge passes and the length of the portion of the graph edge passing through that region.
[0035] Each region type is associated with a numerical weight. The numerical weight can be a predetermined value related to the speed at which a robot can move through a region of that region type. Example numerical weights are 1.0 for a generally navigable surface, 1.3 for a room, 1.5 for a narrow corridor or ramp, and 2.0 for a room entrance. The length of the graph edge corresponds to the length of the path through the region represented by the edge. As an example, the cost of a graph edge can be determined by multiplying the numerical weight of the edge type by the length of the edge. If a graph edge passes through multiple regions of different types, the cost of the graph edge can be calculated as the weighted average of the lengths of the different portions of the graph edge passing through the different regions. In the average calculation, each length is weighted by the numerical weight of the region type associated with the corresponding region.
[0036] The waypoint graph generator 122 provides the waypoint graph 124 to the route planner 126, which generates one or more route plans 246. The route planner 126 receives a start waypoint vertex 250 and a destination waypoint vertex 252 as input and uses the waypoint graph 124 (including edge costs) to determine a route plan 246 from the start waypoint vertex 250 to the destination waypoint vertex 252. The route plan 246 can be a sequence of waypoint graph vertices 242 connected by one or more waypoint graph edges 240 such that the sum of the costs of the waypoint graph edges 240 is minimized.
[0037] Figure 3AShows an example Voronoi diagram 312 on an example semantic map 300 according to various embodiments. The example semantic map includes a set of map regions. Each map region is a room 302, a generally navigable surface 304, a narrow corridor 306, an entrance 308, or a ramp 310. The Voronoi diagram 312 is generated by a Voronoi diagram generator 226 based on a merged surface that includes the boundaries of the map regions. The Voronoi diagram 312 includes Voronoi edges 320 and Voronoi vertices 322. Obstacles (such as walls) are shown as non-shaded regions outside the boundaries of the map regions. Each point on each Voronoi edge 320 has a determined distance from two or more boundaries of the map region in which the Voronoi edge 320 lies. The determined distance can be determined such that the Voronoi edge 320 is equidistant from two or more boundaries of the map region in which the Voronoi edge 320 lies. For example, each point on the edge 320 shown in room 302D is equidistant from the upper and lower boundaries of room 302D.
[0038] Room 302 includes room 302A, which includes four Voronoi edges 320 and four Voronoi vertices 322. Entrance 308A (e.g., a doorway) enables movement between room 302A and narrow corridor 306. The Voronoi edges 320 located in room 302A, entrance 308A, and narrow corridor 306 represent segments of paths along which a robot can move between room 302A and narrow corridor 306. The Voronoi edges 320 connect the Voronoi vertices 322.
[0039] Other rooms 302 of the semantic map 300 similarly include Voronoi edges 320 and Voronoi vertices 322 connected by the Voronoi edges 320. The Voronoi edges 320 form paths in the rooms along which a robot can reach most of each room. Room 302 does not include Voronoi edges 320 that are less than a threshold minimum distance from the region boundary (such as a room wall). The threshold minimum distance can be the minimum distance between the position of the mobile robot and the position of the boundary. For example, the side of the robot can prevent the robot from moving closer to the region boundary than the threshold minimum distance. If the Voronoi diagram generator 226 generates a Voronoi edge 320 that is less than the threshold distance from the region boundary, the Voronoi diagram generator 226 or other components (e.g., the waypoint map creator 234) can remove the Voronoi edge 320 from the Voronoi diagram 228.
[0040] Each room 302 has one or more entrances 308 that enable movement between the room 302 and the narrow corridor 306 or the generally navigable surface 304. A ramp 310 is located between the narrow corridor 306 and the generally navigable surface 304 to enable movement between the narrow corridor 306 and the generally navigable surface 304. Room 302 has two entrances 308B and includes a set of Voronoi edges 320 and a set of Voronoi vertices 322. Room 302B has two entrances, including entrance 302B, a set of nine Voronoi edges 320, and a set of eight Voronoi vertices 322. Room 302C has entrance 308, four Voronoi edges 320, and four Voronoi vertices 322. Room 302D has two entrances 308 and includes a set of Voronoi edges 320 and a set of Voronoi vertices 322. Each of rooms 302F, 302G, 302H, 302I, and the generally navigable surface 304 includes at least one entrance, a set of Voronoi edges 320, and a set of Voronoi vertices 322.
[0041] Figure 3B An example Voronoi diagram 312 and a corresponding example waypoint diagram 330 in accordance with various embodiments are shown. Figure 3B The portion of the Voronoi diagram 312 shown in includes Voronoi edges 320 and Voronoi vertices 322 located in three diagram regions that are part of room 302A, entrance 308A, and narrow corridor 306. Voronoi edges 320A, 320B, 320E, and 320F are located in room 302A. Voronoi edge 320C is located in entrance 308A, and Voronoi edge 320D is located in narrow corridor 306. Each Voronoi edge 320 connects two Voronoi vertices 322. Voronoi edge 320A connects Voronoi vertices 322A and 322B. Voronoi edge 320B connects Voronoi vertices 322B and 322C. Voronoi edge 320C connects Voronoi vertices 322C and 322D. Voronoi edge 320D connects Voronoi vertices 322D and 322E.
[0042] Each graph region 204 of the semantic graph 300 is associated with a corresponding region type 210. Each Voronoi edge 320 can also be associated with a corresponding region type 210, which is the region type 210 of the graph region 204 where the Voronoi edge 320 is located. For example, the room 302 has the region type "room", the entrance 308A has the region type "entrance", and the narrow corridor 306 has the region type "narrow corridor". Since the Voronoi edges 320A, 320B, 320E, and 320F are located in the room 302, each of these Voronoi edges 320 is associated with the "room" region type. In addition, the Voronoi edge 320C in the entrance 308A is associated with the "entrance" region type, and the Voronoi edge 320D in the narrow corridor 306 is associated with the "narrow corridor" region type. Each region type can be associated with a numerical weight. Example weights are 1.3 for the room, 2.0 for the entrance, and 1.5 for the narrow corridor. In addition, each Voronoi edge 320 has an associated length, which is a measure of the length represented by the Voronoi edge 320 in the graph region 204 that contains the Voronoi edge 320. Example side lengths are 5 meters for the Voronoi edge 320A, 20 meters for the Voronoi edge 320B, 3 meters for the Voronoi edge 320C, and 4 meters for the Voronoi edge 320D.
[0043] The example waypoint graph 330 is generated by the waypoint graph creator 234 by adding waypoint graph edges 340 to the waypoint graph 330 for each Voronoi edge 320 and adding waypoint graph vertices 342 to the waypoint graph 330 for each Voronoi vertex 322. The waypoint graph edges 340 and the waypoint graph vertices 342 can be added to the waypoint graph 330 such that the waypoint graph 330 has the same graph structure as the Voronoi graph 312. The waypoint graph 330 generated from the Voronoi graph 312 includes waypoint graph edges 340A, 340B, 340C, 340D, 340E, and 340F, which are generated from the Voronoi edges 320A, 320B, 320C, 320D, 320E, and 320F respectively and correspond to the Voronoi edges 320A, 320B, 320C, 320D, 320E, and 320F respectively. The waypoint graph 330 also includes waypoint graph vertices 342A, 342B, 342C, 342D, and 342E, which are generated by the Voronoi vertices 322A, 322B, 322C, 322D, and 322E respectively and correspond to the Voronoi vertices 322A, 322B, 322C, 322D, and 322E respectively. In addition, the region type and side length associated with the Voronoi edge 320 can be added to the waypoint graph 330.
[0044] The edge cost determiner 236 determines the edge cost of each waypoint graph edge 340 based on the region type of the area where the waypoint graph edge 340 is located and the length associated with the waypoint graph edge 340. The edge cost determiner 236 can determine the edge cost of individual waypoint graph edges 340 and / or a path including two or more waypoint graph edges 340. To determine the edge cost of an individual waypoint graph edge 340, the edge cost determiner 236 multiplies the length of the waypoint graph edge 340 by the region type of the waypoint graph edge 340. Alternatively, the edge cost and region type of the corresponding Voronoi edge 320 can be accessed.
[0045] For example, the edge cost of waypoint graph edge 340A is the length of waypoint graph edge 340A (which is 5 meters) multiplied by the weight associated with the room region type (which is 1.3) (5 x 1.3 = 6.5). As another example, the edge cost of waypoint graph edge 340B is the length of waypoint graph edge 340B (which is 20 meters) multiplied by the weight associated with the room region type (which is 1.3) (20 x 1.3 = 26). As yet another example, the edge cost of waypoint graph edge 340C is the length of waypoint graph edge 340C (which is 3 meters) multiplied by the weight associated with the entrance region type (which is 2.0) (3 x 2.0 = 6). As a further example, the edge cost of waypoint graph edge 340D is the length of waypoint graph edge 340D (which is 4 meters) multiplied by the weight associated with the narrow corridor region type (which is 1.5) (4 x 1.5 = 6). The edge cost determiner 236 can similarly determine the edge cost of each edge 320 of the waypoint graph 330.
[0046] To determine the edge cost of a path including two or more waypoint graph edges 340, the edge cost determiner 236 can calculate the weighted sum of the lengths of the waypoint graph edges 340, where the waypoint graph edges 340 are weighted by the weights associated with the respective region types. For example, to determine the edge cost of the path between waypoint graph vertex 342A and waypoint graph vertex 342E, the edge cost determiner 236 can calculate the weighted sum as (5 meters x 1.3) + (20 meters x 1.5) + (3 meters x 2.0) + (4 meters x 1.5) = 6.5 + 30 + 6 + 6 = 48.5. The edge cost determiner 236 can similarly determine the edge cost of any path of the waypoint graph 330 including two or more waypoint graph edges 340.
[0047] Figure 4FIG. 400 is a flow chart showing a method for generating a route plan using a waypoint graph with edge costs determined based on region information from a semantic graph according to various embodiments. Each block of method 400 described herein includes a computing process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service), or a plug-in of another product, to name a few. Further, method 400 is described by way of example with respect to Figures 1 - 2 the system of. However, the method can alternatively or additionally be performed by any one system or any combination of systems, including but not limited to the systems described herein. Further, operations in method 400 can be omitted, repeated, and / or performed in any order without departing from the scope of the present disclosure.
[0048] As Figure 4 shown, method 400 begins at operation 402, where waypoint graph generator 122 generates merged polygons based on one or more regions 204 of semantic graph 202. The merged polygons can be generated by merged surface generator 220 and have edges of each region boundary 206 of semantic graph 202. For example, merged surface generator 220 can generate merged surface 222 by merging line segments and / or polygons specified by boundary coordinates 208 of semantic graph 202 into one or more combined polygons (e.g., a single polygon).
[0049] In operation 404, waypoint graph generator 122 generates Voronoi diagram 228 based on one or more line segments of the merged polygons. The line segments can be boundaries of regions 204 specified by semantic graph 202. Voronoi diagram 228 includes one or more Voronoi edges 230 and one or more Voronoi vertices 232. In a particular embodiment, each of a plurality of points on Voronoi edge 230 is located at a respective distance determined from each of two or more boundary lines of the graph region in which Voronoi edge 230 lies. The determined respective distances can be equal distances from each of two or more boundary lines of the graph region in which Voronoi edge 230 lies.
[0050] In operation 406, the waypoint graph generator 122 generates a waypoint graph 124 based on one or more Voronoi edges 230 and one or more Voronoi vertices 232 of the Voronoi graph 228, wherein each waypoint graph edge 240 of the waypoint graph 124 is associated with a corresponding Voronoi edge 230, and the corresponding Voronoi edge 230 is located in a corresponding graph region 204 of the semantic graph 202. The waypoint graph 124 has the same graph structure as the Voronoi graph 228. The waypoint graph 124 may include a plurality of waypoint graph edges and a plurality of waypoint graph vertices, each waypoint graph edge corresponding to a corresponding roadmap edge among one or more roadmap edges, and each waypoint graph vertex corresponding to a corresponding roadmap vertex. In some embodiments, the waypoint graph 124 may be generated by copying the Voronoi edges 230 and Voronoi vertices 232 to the waypoint graph 124. In other embodiments, the Voronoi graph 228 may be used as the waypoint graph 124.
[0051] In operation 408, the waypoint graph generator 122 determines one or more edge costs for one or more waypoint graph edges 240, wherein for each waypoint graph edge 240, the corresponding edge cost is based on the region type 210 of the graph region 204 that contains the Voronoi edge 230 associated with the waypoint graph edge 240. The region type may correspond to a numerical value, and the cost of the waypoint graph edge 240 may be determined based on the numerical value.
[0052] The cost of the waypoint graph edge 240 may be further determined based on the product of the numerical value corresponding to the region type and the length associated with the waypoint graph edge 240. The length associated with the waypoint graph edge 240 may correspond to the length represented by the waypoint graph edge 240 in the graph region where the waypoint graph edge 240 is located. In one example, the numerical value may be proportional to the speed limit associated with the region type.
[0053] In operation 410, the waypoint graph generator 122 uses the waypoint graph 124 to generate a route plan 246 from a starting waypoint vertex 250 to a destination waypoint vertex 252.
[0054] It should be understood that the arrangements described herein and other arrangements are listed only by way of example. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) may be used in addition to or in place of the shown arrangements and elements, and certain elements may be omitted altogether. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be executed by hardware, firmware, and / or software. For example, the various functions may be executed by a processor that executes instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may use components, features, and / or functions similar to those of the example autonomous vehicle 500 of Figures 5A - 5D and the example computing device 600 of Figure 6 and / or components, features, and / or functions similar to those of the example data center 700 of Figure 7 to perform.
[0055] In summary, the disclosed waypoint map generator generates a waypoint map that represents the navigable regions of a semantic map and has waypoint map edge costs based on semantic information associated with the navigable regions of the semantic map. The waypoint map (including the waypoint map edge costs) is provided to a route planner that generates a route between given vertices based on the edges connecting the given vertices and the corresponding waypoint map edge costs. To generate the waypoint map, the waypoint map generator constructs a Voronoi diagram based on the navigable regions of the semantic map. The Voronoi diagram includes a set of line segments, herein referred to as "Voronoi edges," that represent the path segments that a robot can follow to reach each navigable region. The Voronoi diagram also includes a set of vertices, herein referred to as "Voronoi vertices," that are connected by the Voronoi edges. Each Voronoi edge lies between the boundaries (e.g., walls) of at least one navigable region of the semantic map. The Voronoi edges and Voronoi vertices are used to generate the waypoint map. For example, in certain embodiments, the Voronoi edges and Voronoi vertices may be used as the waypoint map edges and waypoint map vertices. In other embodiments, the waypoint map edges and waypoint map vertices may be generated from the Voronoi edges and Voronoi vertices such that the waypoint map has the same structure as the Voronoi diagram.
[0056] The waypoint graph generator determines the waypoint graph edge cost for each waypoint graph edge based on the semantic information of the region where the waypoint graph edge provided by the semantic graph is located. The semantic information includes the region type associated with each navigable region. The region type can indicate that the associated navigable region is a general navigable surface that can be navigated without changing speed, or a surface type suitable for changing speed, such as a narrow corridor, a room entrance, a room, or a ramp. Each waypoint graph edge cost is determined based on the region type of one or more regions passed by the corresponding waypoint graph edge and the length of the portion of the waypoint graph edge passing through the region. The waypoint graph (including the cost associated with the waypoint graph edge) is provided to the route planner. The route planner generates a route between a given start waypoint graph vertex and a destination waypoint graph vertex based on the waypoint graph edges of the waypoint graph and the waypoint edge costs associated with the waypoint graph edges.
[0057] One technical advantage of the disclosed technology over prior methods is the ability to generate route planning based on the semantic information associated with waypoint graph edges, such that the route planning accurately reflects the navigation cost in each region represented by the waypoint graph edges. For example, the cost of each edge is based on the type of navigable surface represented by the edge. Using edge costs based on the type of navigable surface enables the disclosed graph generator to accurately identify the optimal path between given locations. These technical advantages represent one or more technical improvements over prior methods.
[0058] The systems and methods described herein can be used by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, submarines, drones, and / or other vehicle types, but are not limited thereto. Additionally, the systems and methods described herein can be used for a variety of purposes, such as but not limited to machine control, machine movement, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation, and / or digital twins, data center processing, conversational AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.
[0059] Embodiments of the present disclosure are included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models such as one or more large language models (LLMs) that can process text, audio, and / or image data, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0060] Example autonomous vehicle
[0061] Figure 5AFIG. is an illustration of an exemplary autonomous vehicle 500 in accordance with some embodiments of the present disclosure. The autonomous vehicle 500 (alternatively, referred to herein as “vehicle 500”) can include, but is not limited to, passenger vehicles such as cars, trucks, buses, emergency vehicles, shuttle vehicles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, submarines, robotic vehicles, drones, airplanes, vehicles coupled to trailers (semi-tractor-trailer trucks for hauling goods) and / or other types of vehicles (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are generally described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, issued June 15, 2018, Standard No. J3016-201609, issued September 30, 2016, and previous and future versions of the standard). The vehicle 500 may be capable of implementing one or more functions corresponding to levels 3-5 of the autonomous driving levels. The vehicle 500 may be capable of implementing one or more functions corresponding to levels 3-5 of the autonomous driving levels. For example, depending on the embodiment, the vehicle 500 may be capable of implementing driver assistance (level 1), semi-automation (level 2), conditional automation (level 3), high automation (level 4), and / or full automation (level 5). The term “autonomous” as used herein may include any and / or all types of autonomy of the vehicle 500 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistive autonomy, semi-autonomous, primarily autonomous, or other designations.
[0062] The vehicle 500 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. The propulsion system 550 may be connected to the driveline of the vehicle 500, which may include a transmission, to effect the propulsion of the vehicle 500. The propulsion system 550 may be controlled in response to receiving a signal from the throttle / accelerator 552.
[0063] A steering system 554, which may include a steering wheel, can be used to steer the vehicle 500 (e.g., along a desired path or route) while the propulsion system 550 is operating (e.g., while the vehicle is in motion). The steering system 554 can receive signals from a steering actuator 556. For fully autonomous (Level 5) functionality, the steering wheel can be optional.
[0064] The brake sensor system 546 can be used to operate the vehicle brakes in response to receiving signals from a brake actuator 548 and / or a brake sensor.
[0065] One or more controllers 536, which may include one or more system-on-chips (SoCs) 504 ( Figure 5C ) and / or one or more GPUs, can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 500. For example, one or more controllers can send signals to operate the vehicle brakes via one or more brake actuators 548, to operate the steering system 554 via one or more steering actuators 556, and to operate the propulsion system 550 via one or more throttles / accelerators 552. One or more controllers 536 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 500. One or more controllers 536 can include a first controller 536 for autonomous driving functionality, a second controller 536 for functional safety functionality, a third controller 536 for artificial intelligence functionality (e.g., computer vision), a fourth controller 536 for infotainment functionality, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 536 can handle two or more of the above functions, two or more controllers 536 can handle a single function, and / or any combination thereof.
[0066] One or more controllers 536 may provide signals for controlling one or more components and / or systems of vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data may be received from, for example and without limitation, a Global Navigation Satellite System sensor (“GNSS”) 558 (e.g., a Global Positioning System sensor), a RADAR sensor 560, an ultrasonic sensor 562, a LiDAR sensor 564, an Inertial Measurement Unit (IMU) sensor 566 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 596, a stereo camera 568, a wide-angle camera 570 (e.g., a fisheye camera), an infrared camera 572, a surround camera 574 (e.g., a 360-degree camera), a long-range and / or mid-range camera 598, a speed sensor 544 (e.g., for measuring the speed of vehicle 500), a vibration sensor 542, a steering sensor 540, a brake sensor (e.g., as part of a brake sensor system 546), and / or other sensor types. One or more of the (one or more) controllers 536 may include one or more instances of a waypoint map generator 122 and / or a route planner 126 to monitor sensor performance based on the corresponding sensor data.
[0067] One or more of the controllers 536 may receive inputs (e.g., represented by input data) from the instrument cluster 532 of vehicle 500 and provide outputs (e.g., represented by output data, display data, etc.) via a Human Machine Interface (HMI) display 534, an audible annunciator, a speaker, and / or via other components of vehicle 500. These outputs may include messages such as vehicle speed, rate, time, map data (e.g., Figure 5C a high-definition (“HD”) map 522), location data (e.g., the location of vehicle 500 on a map, for example), direction, the locations of other vehicles (e.g., occupancy grids), information about objects and object states as perceived by the controller 536, and so on. For example, the HMI display 534 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).
[0068] Vehicle 500 further includes a network interface 524, which may communicate over one or more networks using one or more wireless antennas 526 and / or a modem. For example, network interface 524 may be capable of communicating via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. One or more wireless antennas 526 may also be used to enable communication between objects (such as vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc. and / or one or more low power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.
[0069] Figure 5B For an example autonomous vehicle 500 according to some embodiments of the present disclosure for Figure 5A Examples of camera positions and fields of view of the example autonomous vehicle 500. The cameras and respective fields of view are one example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on vehicle 500.
[0070] The type of camera used for the cameras may include, but is not limited to, digital cameras that may be adapted to be used with components and / or systems of vehicle 500. The cameras may operate under Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as cameras having RCCC, RCCB, and / or RBGC color filter arrays may be used in an effort to increase light sensitivity.
[0071] In some examples, one or more of the cameras may be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera may be installed to provide functions including lane departure warning, traffic sign assistance, and smart headlight control. One or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0072] One or more of the cameras may be mounted in a mounting assembly such as a custom-designed (three-dimensional "3D" printed) component to cut off stray light and reflections from within the vehicle (such as reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the camera. Regarding the wing mirror mounting assembly, the wing mirror assembly may be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side vision cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.
[0073] A camera (e.g., a front camera) having a field of view that includes an environmental portion in front of the vehicle 500 may be used for surround view to help identify forward paths and obstacles and to assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 536 and / or a control SoC. The front camera may be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera may also be used for ADAS functions and systems, including lane departure warning ("LDW"), adaptive cruise control ("ACC"), and / or other functions such as traffic sign recognition.
[0074] A variety of cameras may be used in a front-facing configuration, including, for example, a monocular camera platform that includes a complementary metal oxide semiconductor ("CMOS") color imager. Another example may be a wide-angle camera 570, which may be used to sense objects (such as pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 5B only one wide-angle camera is illustrated, any number (including zero) of wide-angle cameras 570 may be present on the vehicle 500. Additionally, any number of long-range cameras 598 (such as a long-view stereo camera pair) may be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range cameras 598 may also be used for object detection and classification and basic object tracking.
[0075] One or more stereo cameras 568 may also be included in a front-facing configuration. In at least one embodiment, one or more of the stereo cameras 568 may include an integrated control unit that includes a scalable processing unit that may provide a multi-core microprocessor and programmable logic ("FPGA") with an integrated controller area network ("CAN") or Ethernet interface on a single chip. Such a unit may be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternatively, the stereo camera 568 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that may measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (such as metadata). Other types of stereo cameras 568 may be used in addition to or in place of those described herein.
[0076] Cameras having a field of view that includes an environmental portion of the side of the vehicle 500 (e.g., side-view cameras) may be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, surround cameras 574 (e.g., four surround cameras 574 as shown in Figure 5B may be disposed on the vehicle 500. The surround cameras 574 may include wide-angle cameras 570, fisheye cameras, 360-degree cameras, and / or the like. By way of example, four fisheye cameras may be disposed on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 574 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a forward camera) as a fourth surround camera.
[0077] Cameras having a field of view that includes an environmental portion of the rear of the vehicle 500 (e.g., rear-view cameras) may be used for assisting in parking, surround view, rear collision warnings, and creating and updating an occupancy grid. A variety of cameras may be used, including but not limited to cameras that may also be suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range cameras 598, stereo cameras 568, infrared cameras 572, etc.).
[0078] Figure 5C For use in accordance with some embodiments of the present disclosure Figure 5ABlock diagram of an example system architecture of an example autonomous vehicle 500. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.
[0079] Figure 5C Each of the components, features, and systems in vehicle 500 is illustrated as being connected via bus 502. Bus 502 may include a Controller Area Network (CAN) data interface (alternatively referred to herein as the “CAN bus”). The CAN may be a network within vehicle 500 that is used to assist in controlling various features and functions of vehicle 500, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, and so on. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle state indicators. The CAN bus may be ASIL B compliant.
[0080] Although bus 502 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or instead of the CAN bus, FlexRay and / or Ethernet may be used. Further, although a single line is used to represent bus 502, this is not intended to be limiting. For example, any number of buses 502 may be present, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 502 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 502 may be used for collision avoidance functions, and a second bus 502 may be used for drive control. In any example, each bus 502 may communicate with any component of vehicle 500, and two or more buses 502 may communicate with the same component. In some examples, each SoC 504, each controller 536, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 500) and may be connected to a common bus such as the CAN bus.
[0081] Vehicle 500 may include one or more controllers 536, such as those described herein with respect to Figure 5A the controllers described. The controller 536 may be used for a variety of functions. The controller 536 may be coupled to any other different components and systems of the vehicle 500 and may be used for the control of the vehicle 500, the artificial intelligence of the vehicle 500, the infotainment for the vehicle 500, and / or the like.
[0082] Vehicle 500 may include one or more system-on-chips (SoCs) 504. The SoC 504 may include a CPU 506, a GPU 508, a processor 510, a cache 512, an accelerator 514, a data store 516, and / or other components and features not shown. In a variety of platforms and systems, the SoC 504 may be used to control the vehicle 500. For example, one or more SoCs 504 may be combined with an HD map 522 in a system (such as a system of the vehicle 500), and the HD map may obtain map refreshes and / or updates from one or more servers (such as Figure 5D one or more servers 578) via a network interface 524.
[0083] The CPU 506 may include a CPU cluster or a CPU complex (alternatively, referred to herein as "CCPLEX"). The CPU 506 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 506 may include four dual-core clusters, each of which has a dedicated L2 cache (such as a 2MB L2 cache). The CPU 506 (such as CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 506 can be active at any given time.
[0084] The CPU 506 may implement power management capabilities including one or more of the following features: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 506 may further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this work is offloaded to the microcode.
[0085] The GPU 508 may include an integrated GPU (alternatively referred to herein as "iGPU"). The GPU 508 may be programmable and efficient for parallel workloads. In some examples, the GPU 508 may use an enhanced tensor instruction set. The GPU 508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 508 may include at least eight streaming microprocessors. The GPU 508 may use a compute application programming interface (API). Additionally, the GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0086] In the case of automotive and embedded use cases, the GPU 508 can be power-optimized for optimal performance. For example, the GPU 508 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be limiting, and the GPU 508 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to enable efficient execution of workloads by leveraging a mix of computing and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and collaboration between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0087] The GPU 508 may include a high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem that provides a peak memory bandwidth of approximately 900GB / s in some examples. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), may be used.
[0088] The GPU 508 may include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory range shared among the processors. In some examples, address translation service (ATS) support may be used to allow the GPU 508 to directly access the CPU 506 page tables. In such examples, when the GPU 508 memory management unit (MMU) experiences a miss, an address translation request may be transmitted to the CPU 506. In response, the CPU 506 may look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 508. In this way, the unified memory technology may allow a single unified virtual address space for the memories of both the CPU 506 and the GPU 508, thereby simplifying GPU 508 programming and porting applications to the GPU 508.
[0089] In addition, the GPU 508 may include access counters that may track how frequently the GPU 508 accesses the memories of other processors. The access counters may help ensure that memory pages are moved to the physical memories of the processors that most frequently access those pages.
[0090] The SoC 504 may include any number of caches 512, including those described herein. For example, the cache 512 may include an L3 cache (e.g., which is connected to both the CPU 506 and the GPU 508) that is available to both the CPU 506 and the GPU 508. The cache 512 may include a write-back cache that may track the state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.
[0091] The SoC 504 may include one or more arithmetic logic units (ALUs) that may be used to perform processing for any of the various tasks or operations regarding the vehicle 500, such as processing a DNN. In addition, the SoC 504 may include a floating point unit (FPU) or other math co-processor or digital co-processor type for performing mathematical operations within the system. For example, the SoC 504 may include one or more FPUs integrated as execution units within the CPU 506 and / or the GPU 508.
[0092] The SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memories. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to supplement the GPU 508 and offload some of the tasks of the GPU 508 (e.g., freeing up more cycles of the GPU 508 for performing other tasks). As an example, the accelerator 514 may be used for targeted workloads that are stable enough to be easily accelerated (e.g., perception, convolutional neural network (CNN), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0093] The accelerator 514 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0094] The DLA may quickly and efficiently execute neural networks, especially CNNs, for any of a variety of functions on processed or unprocessed data, such as, but not limited to: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner recognition using data from a camera sensor; and / or CNNs for security and / or safety-related events.
[0095] The DLA may perform any function of the GPU 508, and by using inference accelerators, for example, the designer may target the DLA or the GPU 508 for any function. For example, the designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 508 and / or other accelerators 514.
[0096] The accelerator 514 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0097] The RISC cores may interact with an image sensor (e.g., the image sensor of any camera described herein), an image signal processor, and / or the like. Each of these RISC cores may include any number of memories. Depending on the embodiment, the RISC cores may use any of several protocols. In some examples, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores may include an instruction cache and / or tightly coupled RAM.
[0098] The DMA may enable the components of the PVA to access system memory independently of the CPU 506. The DMA may support any number of features used to provide optimizations to the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA may support addressing up to six or more dimensions, which may include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0099] The vector processor may be a programmable processor that may be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW may enhance throughput and rate.
[0100] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. Additionally, a PVA may include additional error correction code (ECC) memory to enhance overall system security.
[0101] The accelerator 514 (e.g., the hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 514. In some examples, the on-chip memory may include at least 4MB of SRAM composed of, for example, and without limitation, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include an on-chip computer vision network (e.g., using APB) that interconnects the PVA and the DLA to the memory.
[0102] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transmission. This type of interface may conform to the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.
[0103] In some examples, SoC 504 can include a real-time ray tracing hardware accelerator as described, for example, in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.
[0104] Accelerator 514 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving uses. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.
[0105] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0106] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.
[0107] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 566 related to the vehicle 500 orientation, distance, a 3D position estimate of an object obtained from the neural network and / or other sensors (such as a LiDAR sensor 564 or a RADAR sensor 560), etc.
[0108] The SoC 504 can include one or more data stores 516 (e.g., memories). The data store 516 can be an on-chip memory of the SoC 504, which can store a neural network to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 516 can be large enough in capacity to store multiple instances of the neural network. The data store 512 can include an L2 or L3 cache 512. References to the data store 516 can include references to memories associated with the PVA, DLA, and / or other accelerators 514 as described herein.
[0109] The SoC 504 may include one or more processors 510 (e.g., embedded processors). The processor 510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 504 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 504 heat and temperature sensor management, and / or SoC 504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 504 may use the ring oscillator to detect the temperature of the CPU 506, GPU 508, and / or accelerator 514. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 504 in a lower power state and / or place the vehicle 500 in a driver safety stop mode (e.g., safely stop the vehicle 500).
[0110] The processor 510 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0111] The processor 510 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0112] The processor 510 may further include a security cluster engine that includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.
[0113] The processor 510 may further include a real-time camera engine that may include a dedicated processor subsystem for handling real-time camera management.
[0114] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0115] The processor 510 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 570, the surround camera 574, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.
[0116] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by adjacent frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.
[0117] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 508 does not need to continuously render new surfaces, the video image compositor may further be used for user interface composition. Even when the GPU 508 is powered on and active for 3D rendering, the video image compositor may be used to relieve the burden on the GPU 508 to improve performance and responsiveness.
[0118] The SoC 504 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 504 may further include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.
[0119] The SoC 504 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC 504 can be used to process data from cameras (connected via Gigabit Multimedia Serial Link and Ethernet), sensors (such as LiDAR sensor 564, RADAR sensor 560, etc. that can be connected via Ethernet), data from bus 502 (such as the speed of vehicle 500, steering wheel position, etc.), and data from GNSS sensor 558 (connected via Ethernet or CAN bus). The SoC 504 may further include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and can be used to free the CPU 506 from routine data management tasks.
[0120] The SoC 504 can be an end-to-end platform with a flexible architecture that spans levels 3 - 5 of automation, thus providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. The SoC 504 can be faster, more reliable, and even more energy - efficient and space - efficient than conventional systems. For example, when combined with the CPU 506, GPU 508, and data storage 516, the accelerator 514 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.
[0121] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using high - level programming languages such as the C programming language to perform various processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in - vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.
[0122] In contrast to conventional systems, the techniques described herein allow multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3-5 autonomous driving capabilities. For example, a CNN executed on a DLA or a dGPU (e.g., GPU 520) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the sign, and passing that semantic understanding to a path planning module running on the CPU complex. The DLA can also utilize metrics related to sensor performance as inputs to one or more neural networks.
[0123] As another example, as required for level 3, 4, or 5 driving, multiple neural networks can run simultaneously. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with the lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network deployed over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on GPU 508.
[0124] In some examples, a CNN for face recognition and owner recognition can use data from a camera sensor to recognize the presence of an authorized driver and / or owner of vehicle 500. A processing engine always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a security mode, to disable the vehicle when the owner leaves the vehicle. In this way, SoC 504 provides security against theft and / or carjacking.
[0125] In another example, a CNN for emergency vehicle detection and identification can use data from microphone 596 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, SoC 504 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle operates as identified by GNSS sensor 558. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 562, a control program can be used to execute emergency vehicle safety routines to slow the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.
[0126] The vehicle can include a CPU 518 (e.g., a discrete CPU or dCPU) that can be coupled to SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 can include, for example, an X86 processor. The CPU 518 can be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and SoC 504, and / or monitoring the status and health of controller 536 and / or infotainment SoC 530.
[0127] The vehicle 500 can include a GPU 520 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 520 can provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of vehicle 500.
[0128] Vehicle 500 may further include a network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 524 can be used to enable wireless connections via the Internet to the cloud (e.g., to server 578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the passenger's client device). To communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and via the Internet). The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 500 with information about vehicles approaching vehicle 500 (e.g., vehicles in front of, to the side of, and / or behind vehicle 500). This function can be part of the cooperative adaptive cruise control function of vehicle 500.
[0129] The network interface 524 may include a SoC that provides modulation and demodulation functions and enables the controller 536 to communicate via a wireless network. The network interface 524 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0130] Vehicle 500 may further include a data storage 528 that may include off-chip (e.g., outside of SoC 504) storage. The data storage 528 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.
[0131] Vehicle 500 may further include a GNSS sensor 558. The GNSS sensor 558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 558 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.
[0132] Vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 can be used by vehicle 500 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level can be ASIL B. The RADAR sensor 560 can use CAN and / or bus 502 (e.g., to transmit data generated by the RADAR sensor 560) for control as well as access to object tracking data and, in some examples, access Ethernet to access raw data. A variety of RADAR sensor types can be used. For example and without limitation, the RADAR sensor 560 can be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.
[0133] The RADAR sensor 560 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 560 can help distinguish between static and moving objects and can be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor can include a single-station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas can create a focused beam pattern designed to record the surroundings of vehicle 500 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of vehicle 500.
[0134] As an example, a mid-range RADAR system can include a range of up to 560m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). The short-range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots beside the vehicle.
[0135] The short-range RADAR system can be used in the ADAS system for blind spot detection and / or lane change assistance.
[0136] Vehicle 500 may further include ultrasonic sensors 562. Ultrasonic sensors 562 that may be disposed in front of, behind, and / or on the sides of vehicle 500 may be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 562 may be used, and different ultrasonic sensors 562 may be used for different detection ranges (e.g., 2.5 m, 4 m). Ultrasonic sensors 562 may operate at ASIL B for functional safety levels.
[0137] Vehicle 500 may include a LiDAR sensor 564. The LiDAR sensor 564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 564 may be at ASIL B for functional safety levels. In some examples, vehicle 500 may include multiple LiDAR sensors 564 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a gigabit Ethernet switch).
[0138] In some examples, the LiDAR sensor 564 may be capable of providing a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 564 may have, for example, an advertised range of approximately 500 m, an accuracy of 2 cm - 3 cm, and support for a 500 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 564 may be used. In such examples, the LiDAR sensor 564 may be implemented as a small device that may be embedded in the front, behind, on the sides, and / or corners of vehicle 500. In such examples, the LiDAR sensor 564 may provide a field of view of up to 120 degrees horizontal and 35 degrees vertical even for low-reflectivity objects, with a range of 200 m. The front-mounted LiDAR sensor 564 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0139] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses the flash of a laser as the emission source to illuminate the vehicle's surrounding environment up to about 200m. The flash LiDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR can allow for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle 500. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) that have no moving parts other than a fan. The flash LiDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 564 can be less susceptible to motion blur, vibration, and / or shock.
[0140] The vehicle can further include an IMU sensor 566. In some examples, the IMU sensor 566 can be located at the center of the rear axle of the vehicle 500. The IMU sensor 566 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 566 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 can include an accelerometer, a gyroscope, and a magnetometer.
[0141] In some embodiments, the IMU sensor 566 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 566 can enable the vehicle 500 to estimate the heading without input from a magnetic sensor by directly observing the change in velocity from the GPS to the IMU sensor 566 and correlating them. In some examples, the IMU sensor 566 and the GNSS sensor 558 can be integrated into a single unit.
[0142] The vehicle can include a microphone 596 placed in and / or around the vehicle 500. Among other things, the microphone 596 can be used for emergency vehicle detection and identification.
[0143] The vehicle may further include any number of camera types, including a stereo camera 568, a wide-angle camera 570, an infrared camera 572, a surround camera 574, a long-range and / or mid-range camera 598, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 500. The camera types used depend on the embodiment and the requirements of the vehicle 500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 500. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 5A and Figure 5B described in more detail.
[0144] The vehicle 500 may further include a vibration sensor 542. The vibration sensor 542 can measure the vibration of components of the vehicle such as the axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 542 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-rotating axle).
[0145] The vehicle 500 may include an ADAS system 538. In some examples, the ADAS system 538 may include a SoC. The ADAS system 538 may include autonomous / adaptive / automated cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0146] The ACC system can use RADAR sensors 560, LiDAR sensors 564, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and recommends that the vehicle 500 change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0147] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a network interface 524 and / or a wireless antenna 526 via a wireless link or through a network connection (e.g., via the Internet). A direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicle immediately in front (e.g., a vehicle immediately in front of vehicle 500 and in the same lane), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given information about the vehicle in front of vehicle 500, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0148] The FCW system is designed to alert the driver of a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of, for example, audible, visual warnings, vibrations, and / or rapid braking pulses.
[0149] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.
[0150] The LDW system provides visual, audible, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 500 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating a turn signal. The LDW system can use a front-side facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0151] The LKA system is a variant of the LDW system. If the vehicle 500 starts to leave the lane, then the LKA system provides steering input or braking to correct the vehicle 500.
[0152] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0153] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 500 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0154] Conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide if a safe condition truly exists and act accordingly. However, in an autonomous vehicle 500, in the case of conflicting results, the vehicle 500 itself must decide whether to heed the results from the main computer or an auxiliary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 can be a backup and / or auxiliary computer for providing perception information to a backup computer sanity module. The backup computer sanity monitor can run redundant diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 538 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0155] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the host computer's direction regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine an appropriate result.
[0156] The supervisory MCU may be configured to run a neural network that is trained and configured to determine, based on outputs from the host computer and the secondary computer, conditions under which the secondary computer provides false alarms. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with an associated memory. In a preferred embodiment, the supervisory MCU may include components of and / or be included as components of the SoC 504.
[0157] In other examples, the ADAS system 538 may include a secondary computer that performs ADAS functions using traditional computer vision rules. As such, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, the diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for faults caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0158] In some examples, the output of the ADAS system 538 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 538 indicates a forward collision warning due to an object being immediately in front, the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network that is trained and thus reduces the risk of false positives as described herein.
[0159] The vehicle 500 can further include an infotainment SoC 530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 530 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, radio data systems, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 500. For example, the infotainment SoC 530 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 530 can further be used to provide information (e.g., visual and / or auditory) to the user of the vehicle, such as information from the ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0160] The infotainment SoC 530 can include GPU functionality. The infotainment SoC 530 can communicate with other devices, systems, and / or components of the vehicle 500 via a bus 502 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 530 can be coupled to the supervisory MCU such that in the event of a failure of the main controller 536 (e.g., the main and / or standby computer of the vehicle 500), the GPU of the infotainment system can perform some autonomous driving functions. In such examples, the infotainment SoC 530 can place the vehicle 500 in the driver safe parking mode as described herein.
[0161] Vehicle 500 may further include an instrument cluster 532 (such as a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 532 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, and so on. In some examples, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.
[0162] Figure 5D A system schematic diagram of communication between a cloud-based server and Figure 5A an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure. System 576 may include a server 578, a network 590, and vehicles including vehicle 500. Server 578 may include multiple GPUs 584(A)-584(H) (collectively referred to herein as GPUs 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switches 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPUs 580). The GPUs 584, CPUs 580, and PCIe switches may be interconnected by high-speed interconnects such as, for example and without limitation, an NVLink interface 588 developed by NVIDIA and / or PCIe connections 586. In some examples, the GPUs 584 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 584 and the PCIe switches 582 are connected via a PCIe interconnect. Although eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, each of the servers 578 may include eight, sixteen, thirty-two, and / or more GPUs 584.
[0163] Server 578 can receive image data from the vehicle via network 590, the image data representing an image showing an unexpected or changed road condition such as a recently started roadwork. Server 578 can transmit neural network 592, updated neural network 592, and / or map information 594, including information about traffic and road conditions, to the vehicle via network 590. Updates to the map information 594 can include updates to the HD map 522, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 592, updated neural network 592, and / or map information 594 can be represented and / or generated based on data received from new training and / or data from any number of vehicles in the environment and / or experience from training performed at the data center (e.g., using server 578 and / or other servers).
[0164] Server 578 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by the vehicle and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 590), and / or the machine learning model can be used by server 578 to remotely monitor the vehicle.
[0165] In some examples, server 578 can receive data from the vehicle and apply the data to the latest real-time neural network for real-time intelligent inference. Server 578 can include a deep learning supercomputer powered by GPU 584 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 578 can include a deep learning infrastructure of a data center powered only by a CPU.
[0166] The deep learning infrastructure of server 578 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 500. For example, the deep learning infrastructure may receive periodic updates from vehicle 500, such as an image sequence and / or objects located in the image sequence that vehicle 500 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify the objects and compare them with the objects identified by vehicle 500. If the results do not match and the infrastructure concludes that the AI in vehicle 500 has malfunctioned, then server 578 may transmit a signal to vehicle 500 instructing the fail-safe computer in vehicle 500 to take control, notify the passengers, and complete a safe parking operation.
[0167] For inference, server 578 may include a GPU 584 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration may enable real-time response. In other examples, such as when performance is less critical, a CPU, FPGA, and other processor-powered servers may be used for inference.
[0168] Example computing device
[0169] Figure 6 FIG. is a block diagram of an example computing device 600 suitable for implementing some embodiments of the present disclosure. Computing device 600 may include an interconnect system 602 that directly or indirectly couples the following devices: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., a display), and one or more logic units 620. In at least one embodiment, computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 608 may include one or more vGPUs, one or more CPUs 606 may include one or more vCPUs, and / or one or more logic units 620 may include one or more virtual logic units. Thus, computing device 600 may include discrete components (e.g., a complete GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof.
[0170] Although Figure 6The various blocks are shown as being connected via an interconnection system 602 having circuitry, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a rendering component 618, such as a display device, may be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, the CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device in addition to the memory of the GPU 608, CPU 606, and / or other components). In other words, Figure 6 the computing devices are merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types because all of these are considered within the Figure 6 scope of the computing devices.
[0171] The interconnection system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 602 may include one or more types of links or buses, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 606 may be directly connected to the memory 604. Additionally, the CPU 606 may be directly connected to the GPU 608. In cases where there are direct or point-to-point connections between components, the interconnection system 602 may include a PCIe link to perform the connection. In these examples, a PCI bus need not be included in the computing device 600.
[0172] The memory 604 may include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 600. The computer-readable media may include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, the computer-readable media may include computer storage media and communication media.
[0173] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 can store computer-readable instructions (e.g., which represent programs and / or program elements such as operating systems). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic tape cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 600. As used herein, computer storage media does not include the signals themselves.
[0174] Computer storage media can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, computer storage media can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.
[0175] CPU 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. CPU 606 can include any type of processor and can include different types of processors depending on the type of computing device 600 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as a math coprocessor, computing device 600 can also include one or more CPUs 606.
[0176] In addition to or instead of the CPU 606, the GPU 608 may also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. One or more GPUs 608 may be an integrated GPU (e.g., having one or more CPUs 606) and / or one or more GPUs 608 may be a discrete GPU. In an embodiment, one or more GPUs 608 may be a coprocessor of one or more CPUs 606. The computing device 600 may use the GPU 608 to render graphics (e.g., 3D graphics) or perform general computing. For example, the GPU 608 may be used for general-purpose computing on the GPU (GPGPU). The GPU 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 may generate pixel data for an output image in response to a rendering command (e.g., a rendering command from the CPU 606 received via the host interface). The GPU 608 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory may be included as part of the memory 604. The GPU 608 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 608 may generate pixel data or GPGPU data for different portions of the output or for different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.
[0177] In addition to or instead of the CPU 606 and / or the GPU 608, the logic unit 620 may be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. In an embodiment, the CPU 606, the GPU 608, and / or the logic unit 620 may perform any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 620 may be part of and / or integrated in one or more CPUs 606 and / or one or more GPUs 608, and / or one or more logic units 620 may be discrete components of the CPU 606 and / or the GPU 608 or otherwise external thereto. In an embodiment, one or more logic units 620 may be a processor of one or more CPUs 606 and / or one or more GPUs 608.
[0178] Examples of the logic unit 620 include one or more processing cores and / or their components, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel vision cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, etc.
[0179] In various embodiments, one or more CPUs 606, GPUs 608, and / or logic units 620 are configured to execute one or more instances of the waypoint map generator 122 and / or the route planner 126. Then, the route planner 126 can use the waypoint map 124 generated by the waypoint map generator 122 to generate a route from the starting waypoint vertex 250 to the destination waypoint vertex 252.
[0180] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 610 may include components and functions that enable communication through any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) for directly delivering data received through the network and / or the interconnect system 602 to one or more GPUs 608 (e.g., the memory of the GPU 608).
[0181] The I / O port 612 can enable the computing device 600 to be logically coupled to other devices including I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 614 can provide a natural user interface (NUI) for processing user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 600 (described in more detail below). The computing device 600 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touch screen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 600 can include an accelerometer or gyroscope enabling motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 600 to render immersive augmented reality or virtual reality.
[0182] The power supply 616 can include hardwired power, battery power, or a combination thereof. The power supply 616 can power the computing device 600 to enable the components of the computing device 600 to operate.
[0183] The presentation component 618 can include a display (e.g., a monitor, touch screen, television screen, head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 can receive data from other components (e.g., GPU 608, CPU 606, DPU, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0184] Example data center
[0185] Figure 7 An example data center 700 is shown, which can be used in at least one embodiment of the present disclosure. The data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0186] As Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1)-716(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units (“CPU”) or other processors (including DPU, accelerator, field programmable gate array (FPGA), graphics processor or graphics processing unit (GPU), etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, etc. In some embodiments, one or more of the node C.R. 716(1)-716(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R. 716(1)-716(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the node C.R. 716(1)-716(N) may correspond to a virtual machine (VM).
[0187] In at least one embodiment, the grouped computing resources 714 may include separate groupings (not shown) of the node C.R. 716 housed within one or more racks, or numerous racks (also not shown) within data centers located in various geographical locations. The separate groupings of the node C.R. 716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several of the node C.R. 716 including CPUs, GPUs, DPUs, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0188] The resource coordinator 712 may configure or otherwise control one or more of the node C.R. 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource coordinator 712 may include a software design infrastructure (“SDI”) management entity for the data center 700. The resource coordinator 712 may include hardware, software, or some combination thereof.
[0189] In at least one embodiment, as Figure 7As shown, the framework layer 720 may include a job scheduler 733, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 may include a framework that supports the software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application 742 may respectively include web-based service software or applications, such as the service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 733 may include a Spark driver for facilitating the scheduling of workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. The resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 733. In at least one embodiment, the cluster or grouped computing resources may include the grouped computing resources 714 at the data center infrastructure layer 710. The resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0190] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus browsing software, database software, and streaming video content software.
[0191] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least portions of nodes C.R. 716(1)-716(N), grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.) and / or other machine learning applications used in conjunction with one or more embodiments.
[0192] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilized and / or poorly performing portions of the data center.
[0193] The data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 700. In at least one embodiment, by using the weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 700 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks, such as, but not limited to, those described herein.
[0194] In at least one embodiment, the data center 700 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. In addition, one or more of the above software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0195] Example Network Environment
[0196] The network environment suitable for implementing the embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the Figure 6 computing device 600 - for example, each device may include similar components, features, and / or functions of the computing device 600. Additionally, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 700, an example of which is described in more detail herein with respect to Figure 7 more detail.
[0197] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or networks within multiple networks. For example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide wireless connectivity.
[0198] A compatible network environment may include one or more peer - to - peer network environments (in which case servers may not be included in the network environment), and one or more client - server network environments (in which case one or more servers may be included in the network environment). In a peer - to - peer network environment, the functions described herein with respect to servers may be implemented on any number of client devices.
[0199] In at least one embodiment, the network environment may include one or more cloud - based network environments, distributed computing environments, combinations thereof, etc. A cloud - based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting one or more applications of a software layer and / or an application layer. The software or application may respectively include network - based service software or applications. In an embodiment, one or more client devices may use network - based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open - source software web application framework, such as may be used for large - scale data processing (e.g., “big data”) using a distributed file system.
[0200] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any one of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, the globe, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. The cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0201] The client device can include at least some of the components, features, and functions of the example computing device 600 described herein. By way of example and not limitation, the client device can be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these described devices, or any other suitable device. Figure 6 1. In some embodiments, a method includes: receiving, at a computing device, a semantic map representing a physical environment; generating a roadmap at least based on the semantic map, the roadmap including one or more roadmap edges, each roadmap edge including an associated location in one or more graph regions of the semantic map and an associated cost; determining the cost of a particular roadmap edge at least based on a region type of the particular roadmap edge, the region type being determined at least based on the graph region in which the particular roadmap edge is located; and using the roadmap to generate a route plan for moving a mobile robot from a first position on the semantic map to a second position on the semantic map.
[0202] 2. The method of clause 1, wherein each point among a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
[0203]
[0204] 3. The method according to clause 1 or 2, wherein each of the determined respective distances is equidistant from each of the two or more boundary lines of the graph region in which the specific roadmap edge is located.
[0205] 4. The method according to any one of clauses 1 - 3, wherein the region type of the specific roadmap edge corresponds to a numerical value, and the cost of the specific roadmap edge is determined based at least on this numerical value.
[0206] 5. The method according to any one of clauses 1 - 4, wherein the cost of the specific roadmap edge is further determined based at least on the product of the numerical value of the region type and the length associated with the specific roadmap edge.
[0207] 6. The method according to any one of clauses 1 - 5, wherein the length associated with the specific roadmap edge corresponds to the length represented by the specific roadmap edge in the graph region in which the specific roadmap edge is located.
[0208] 7. The method according to any one of clauses 1 - 6, wherein the numerical value is proportional to the speed limit associated with the region type.
[0209] 8. The method according to any one of clauses 1 - 7, further comprising determining the cost of a path including a plurality of roadmap edges by: determining a weighted sum of a plurality of side lengths, wherein each respective side length is the length of a respective roadmap edge among the plurality of roadmap edges, and each respective side length is weighted in the weighted sum by a respective numerical value determined based at least on the respective region type associated with the respective roadmap edge.
[0210] 9. The method according to any one of clauses 1 - 8, wherein the roadmap represents one or more routes through the graph region of the semantic graph.
[0211] 10. The method according to any one of clauses 1 - 9, wherein each roadmap edge is at least a threshold distance away from each boundary of each graph region of the semantic graph, and the threshold distance corresponds to the minimum distance between the position of the mobile robot and the position of the boundary.
[0212] 11. The method according to any one of clauses 1 - 10, wherein each graph region represents a part of the environment that the mobile robot can reach.
[0213] 12. The method according to any one of clauses 1 - 11, wherein the roadmap further includes a plurality of roadmap vertices, wherein each roadmap edge connects a respective first roadmap vertex located at a respective first endpoint of the respective roadmap edge to a respective second roadmap vertex located at a respective second endpoint of the respective roadmap edge.
[0214] 13. The method according to any one of clauses 1-12 further includes: generating a waypoint graph including a plurality of waypoint graph edges and a plurality of waypoint graph vertices, each waypoint graph edge corresponding to a respective roadmap edge among the one or more roadmap edges, and each waypoint graph vertex corresponding to a respective roadmap vertex.
[0215] 14. In some embodiments, one or more processors include: processing circuitry configured to perform operations including: receiving a semantic map representing a physical environment; generating at least based on the semantic map a roadmap including one or more roadmap edges, the one or more roadmap edges including associated locations and associated costs in one or more graph regions of the semantic map; determining the cost of a particular roadmap edge among the roadmap edges at least based on a region type of the particular roadmap edge, the region type being determined at least based on the graph region in which the particular roadmap edge is located; using the roadmap to generate a route plan for moving a mobile robot from a first position on the semantic map to a second position on the semantic map; and causing the mobile robot to navigate through at least a portion of the physical environment at least based on the route plan.
[0216] 15. The one or more processors according to clause 14, wherein each of a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
[0217] 16. The one or more processors according to clause 14 or 15, wherein the determined respective distances are equal distances from each of the two or more boundary lines of the graph region in which the particular roadmap edge is located.
[0218] 17. One or more processors according to any one of clauses 14-16, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing 3D asset collaborative content creation; a system for performing deep learning operations; a system implemented using edge devices; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using robots; a system for performing conversational AI operations; a system for performing one or more generative AI operations; a system implementing one or more large language models (LLMs); a system for generating synthetic data; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0219] 18. In some embodiments, a system includes: one or more processors for performing operations including: receiving a semantic map representing a physical environment; generating a roadmap based at least on the semantic map, the roadmap including one or more roadmap edges, each roadmap edge including an associated location in one or more graph regions of the semantic map and an associated cost; determining the cost of a particular roadmap edge based at least on a region type of the particular roadmap edge, the region type being determined at least based on the graph region in which the particular roadmap edge is located; and generating a route plan using the roadmap to move a mobile robot from a first position on the semantic map to a second position on the semantic map.
[0220] 19. The system according to clause 18, wherein each of a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
[0221] 20. A system according to clause 18 or 19, wherein the system is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a sensing system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing 3D asset collaborative content creation; a system for performing deep learning operations; a system implemented using edge devices; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using robots; a system for performing conversational AI operations; a system for performing one or more generative AI operations; a system implementing one or more large language models (LLMs); a system for generating synthetic data; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0222] The present disclosure may be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0223] As used herein, the recitation of "and / or" with respect to two or more elements should be construed to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0224] This disclosure describes the subject matter of the present disclosure in detail to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from those described herein or combinations of similar steps in conjunction with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is explicitly described.
Claims
1. A method, comprising: Receiving, at a computing device, a semantic map representing a physical environment; Generating a roadmap based at least on the semantic map, the roadmap including one or more roadmap edges, each roadmap edge including an associated location in one or more graph regions of the semantic map and an associated cost; Determining the cost of a particular roadmap edge among the roadmap edges based at least on a region type of the particular roadmap edge, the region type being determined based at least on the graph region in which the particular roadmap edge is located; And Using the roadmap to generate a route plan for moving a mobile robot from a first position on the semantic map to a second position on the semantic map.
2. The method according to claim 1, wherein Each point among a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
3. The method according to claim 2, wherein The determined respective distances are equidistant from each of the two or more boundary lines of the graph region in which the particular roadmap edge is located.
4. The method according to claim 1, wherein, The region type of the particular roadmap edge corresponds to a numerical value, and the cost of the particular roadmap edge is determined based at least on the numerical value.
5. The method according to claim 4, wherein, The cost of the particular roadmap edge is further determined based at least on a product of the numerical value of the region type and a length associated with the particular roadmap edge.
6. The method according to claim 5, wherein The length associated with the particular roadmap edge corresponds to the length represented by the particular roadmap edge in the graph region in which the particular roadmap edge is located.
7. The method according to claim 4, wherein, The numerical value is proportional to a speed limit associated with the region type.
8. The method according to claim 1, further comprising determining the cost of a path including a plurality of roadmap edges by: Determine the weighted sum of multiple side lengths, where, Each respective side length is the length of a respective roadmap edge among the plurality of roadmap edges, and each respective side length is weighted in the weighted sum by a respective numerical value determined based at least on a respective region type associated with the respective roadmap edge.
9. The method according to claim 1, wherein The roadmap represents one or more routes through the graph regions of the semantic map.
10. The method according to claim 1, wherein Each roadmap edge is at least a threshold distance from each boundary of each graph region of the semantic map, and the threshold distance corresponds to a minimum distance between the position of the mobile robot and the position of the boundary.
11. The method according to claim 1, wherein, Each graph region represents a portion of the environment that the mobile robot can reach.
12. The method according to claim 1, wherein, The roadmap further includes a plurality of roadmap vertices, wherein each roadmap edge connects a respective first roadmap vertex located at a respective first endpoint of the respective roadmap edge to a respective second roadmap vertex located at a respective second endpoint of the respective roadmap edge.
13. The method according to claim 1, further comprising: Generating a waypoint map including a plurality of waypoint map edges and a plurality of waypoint map vertices, each waypoint map edge corresponding to a respective roadmap edge among the one or more roadmap edges, and each waypoint map vertex corresponding to a respective roadmap vertex.
14. One or more processors, comprising: Processing circuitry configured to perform operations including: Receiving a semantic map representing a physical environment; Generate a roadmap based at least on the semantic map, the roadmap including one or more roadmap edges, the one or more roadmap edges including associated positions in one or more graph regions of the semantic map, and associated costs; Determine the cost of a particular roadmap edge among the roadmap edges based at least on the region type of the particular roadmap edge, the region type being determined based at least on the graph region in which the particular roadmap edge is located; Use the roadmap to generate a route plan for a mobile robot to move from a first position on the semantic map to a second position on the semantic map; and Cause the mobile robot to navigate through at least a portion of the physical environment based at least on the route plan.
15. One or more processors according to claim 14, wherein, Each of a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
16. One or more processors according to claim 15, wherein, The determined respective distances are equidistant from each of the two or more boundary lines of the graph region in which the particular roadmap edge is located.
17. One or more processors according to claim 14, wherein, The one or more processors are included in at least one of the following: A control system for an autonomous or semi-autonomous machine; A perception system for an autonomous or semi-autonomous machine; A system for performing simulation operations; A system for performing digital twin operations; A system for performing optical transmission simulation; A system for performing 3D asset collaborative content creation; A system for performing deep learning operations; A system implemented using edge devices; A system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system implemented using robots; A system for performing conversational AI operations; A system for performing one or more generative AI operations; A system implementing one or more large language models (LLMs); A system for generating synthetic data; A system including one or more virtual machines (VMs); A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.
18. A system, comprising: One or more processors, the one or more processors configured to perform operations including: Receiving a semantic map representing a physical environment; Generating a roadmap based at least on the semantic map, the roadmap including one or more roadmap edges, each roadmap edge including associated positions in one or more graph regions of the semantic map, and associated costs; Determining the cost of a particular roadmap edge among the roadmap edges based at least on the region type of the particular roadmap edge, the region type being determined based at least on the graph region in which the particular roadmap edge is located; and Using the roadmap to generate a route plan for a mobile robot to move from a first position on the semantic map to a second position on the semantic map.
19. The system according to claim 18, wherein Each of a plurality of points on the particular roadmap edge is located at a respective distance determined from each of two or more boundary lines of the graph region in which the particular roadmap edge is located.
20. The system according to claim 18, wherein, The system is included in at least one of the following: A control system for an autonomous or semi-autonomous machine; Perception systems for autonomous or semi-autonomous machines; Systems for performing simulation operations; Systems for performing digital twin operations; Systems for performing optical transmission simulations; Systems for performing 3D asset collaborative content creation; Systems for performing deep learning operations; Systems implemented using edge devices; Systems for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; Systems implemented using robots; Systems for performing conversational AI operations; Systems for performing one or more generative AI operations; Systems implementing one or more large language models (LLMs); Systems for generating synthetic data; Systems containing one or more virtual machines (VMs); Systems implemented at least in part in a data center; or Systems implemented at least in part using cloud computing resources.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2