Generating lane segments for autonomous vehicle navigation using embedding
By using trained artificial intelligence or machine learning models to process autonomous vehicle and robot sensor data, potential lane segment marking sets in the environment are generated, solving the problem of inefficient navigation in the prior art and achieving more efficient and accurate autonomous navigation.
Patent Information
- Application Number
- CN202380076624.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is inefficient in analyzing and processing autonomous vehicle and robot sensor data, making it difficult to effectively navigate through complex and dynamic environments.
Using trained artificial intelligence (AI) or machine learning (ML) models, by encoding sensor data and map data into tensors, the machine learning model is applied to generate a set of markers, define potential lane segments in the environment, and update the graphs to support autonomous navigation.
Improves the efficiency and accuracy of autonomous navigation, can process data in complex and dynamic environments more effectively, and supports autonomous vehicles and robots to navigate safely and reliably in multiple environments.
Smart Images

Figure CN120152893A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 377,954, filed Sep. 30, 2022, which is hereby incorporated by reference in its entirety for all purposes. Technical Field
[0003] The present disclosure generally relates to artificial intelligence-based modeling techniques for analyzing image data and predicting occupancy attributes of an environment for an ego body. Background Art
[0004] Due to the rapid development of computer technology, autonomous navigation technologies for autonomous vehicles and robots (collectively referred to as ego bodies) have become ubiquitous. These advancements allow for safer and more reliable autonomous navigation of ego bodies. Ego bodies typically need to navigate in complex and dynamic environments and terrains, which can include vehicles, traffic, pedestrians, cyclists, and various other static or dynamic obstacles. To navigate through such complex and dynamic environments, computational systems on the ego body can perform path planning through the environment using sensor data on the surrounding environment. Path planning can identify a trajectory or lane along which the ego body is to pass through the environment. Understanding the ego body's surrounding environment is necessary for informed and capable decision-making to avoid collisions and successfully execute path planning. However, techniques for analyzing and processing sensor data can be cumbersome and too slow to effectively navigate the ego body through the environment. Summary of the Invention
[0005] To facilitate autonomous navigation of an ego body (e.g., a vehicle or a robot) through an environment, a computational system of the ego body can be configured with a trained artificial intelligence (AI) or machine learning (ML) model. The ML model can obtain inputs from an embedding set that encodes low-dimensional representations of sensor data (e.g., video from the surrounding environment) and map data (e.g., a navigation map of the environment). Using the inputs, the ML model can be used to generate a graph with a set of labels that define a linguistic representation of potential lane segments within the environment in which the ego body can navigate. The graph can correspond to a sparse set of lane segments and their connectivity specified by their usage coefficients (e.g., spline coefficients).
[0006] Aspects of the present disclosure relate to systems, methods, devices, apparatuses, and non-transitory computer-readable media for generating a path for autonomous navigation through an environment. One or more processors may identify a tensor that includes a plurality of encodings derived from sensor data from an ego and map data defining an environmental topology around the ego. By applying at least a first portion of the plurality of encodings to a machine learning (ML) model, one or more processors may determine a first index value that defines a point within a first plurality of points of a first grid defined on the environment. By applying at least a second portion of the plurality of encodings and the first index value to the ML model, one or more processors may determine a second index value that defines a point within a second plurality of points of a second grid within a subset of the first plurality of points of the first grid. Based on the first index value and the second index value for the point, one or more processors may generate a marker for at least one of a plurality of paths through the environment. One or more processors may store a graph to include the marker to be used for autonomous navigation of the ego through the environment via one or more of the plurality of paths.
[0007] In one embodiment, by applying at least a third portion of the plurality of encodings and a third index value for a second point to the ML model, one or more processors may classify the second point as a continuation topology type that depends on the first point. In response to the classification of the second point as a continuation topology type, one or more processors may determine a plurality of spline coefficients that define a path of a plurality of paths through the environment between the first point and the second point. One or more processors may generate a second marker based on the third index value for the second point and the continuation topology type. One or more processors may update the graph to include the second marker and the plurality of spline coefficients to be used for autonomous navigation of the ego through the environment.
[0008] In another embodiment, by applying at least a third portion of the plurality of encodings and a third index value for a second point to the ML model, one or more processors may classify the second point as a bifurcation topology type from the point with respect to a third point. In response to the classification of the second point as a terminal topology type, one or more processors may determine a path through the environment defined by the first point and the second point. One or more processors may generate a second marker based on the third index value for the second point and the terminal topology type. One or more processors may update the graph to include the second marker to be used for autonomous navigation of the ego through the environment.
[0009] In another embodiment, one or more processors may identify a second figure including a plurality of markers to be used to autonomously navigate a second self through an environment via one or more of a second plurality of passages. One or more processors may use the figure and the second figure to determine that at least one first passage of the plurality of passages for the self intersects at least one second passage of the second plurality of passages for the second self. In response to determining that at least one first passage intersects the second passage, one or more processors may perform an action on at least one of the self or the second self.
[0010] In another embodiment, one or more processors may use sensor data from the self to identify the presence of a second stationary self in the environment. One or more processors may use the figure to determine that at least one first passage of the plurality of passages for the self intersects the stationary second self. In another embodiment, by applying at least a third portion of the plurality of encodings and a second index value to an ML model, one or more processors may classify the point as a topological type indicating the start of at least one passage of the plurality of passages. One or more processors may generate a marker for at least one passage of the plurality of passages through the environment based on the topological type.
[0011] In another embodiment, one or more processors may use a plurality of markers of the figure to determine a trajectory that defines the navigation of the self via passages of the plurality of passages through the environment. In another embodiment, one or more processors may present the figure via a graphical user interface (GUI) that defines the plurality of passages relative to the topology of the environment surrounding the self. In another embodiment, one or more processors may generate a marker using (i) a first embedding generated from a first index value, (ii) a second embedding generated from a second index value, and (iii) one or more embeddings associated with the point. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Non-limiting embodiments of the present disclosure are described by way of example in connection with the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as representing the background art, the drawings represent various aspects of the present disclosure.
[0013] Figure 1A Illustrates components of an AI-enabled visual data analysis system according to an illustrative embodiment.
[0014] Figure 1B Illustrates various sensors associated with a self according to an illustrative embodiment.
[0015] Figure 1C Illustrates components of a vehicle according to an illustrative embodiment.
[0016] Figure 2The block diagram of a system for generating a tensor representation of a path for autonomous navigation through an environment according to an illustrative embodiment is shown.
[0017] Figures 3A to 3H The block diagrams of processes for generating a path for autonomous navigation through an environment according to illustrative embodiments are each shown.
[0018] Figure 4 The diagram of a scenario according to an illustrative embodiment is shown, in which a first self uses a graph representing a path to autonomously navigate through an environment.
[0019] Figure 5 The diagram of a scenario is shown, in which a first self uses a graph to detect that the path the first self is passing through intersects with the path a second self is about to pass through.
[0020] Figure 6 The diagram of a scenario according to an illustrative embodiment is shown, in which a first self uses a graph to detect the presence of a second self that is stationary in the environment.
[0021] Figure 7 The flowchart of a method for generating a path for autonomous navigation through an environment according to an illustrative embodiment is shown. Detailed Description
[0022] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, and specific language will be used herein to describe them. However, it is to be understood that no limitation of the scope of the claims or the present disclosure is thereby intended. Changes and further modifications of the features described herein and additional applications of the principles of the subject matter described herein that are contemplated by those of ordinary skill in the relevant art and having the present disclosure will be considered to be within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the detailed description are not meant to limit the subject matter presented.
[0023] The ego vehicle can be an autonomous vehicle (e.g., car, truck, bus, motorcycle, all-terrain vehicle, cart), a robot, or other automated device. The ego vehicle can use one or more artificial intelligence (AI) algorithms or machine learning (ML) models to autonomously navigate the ego vehicle through the environment. To facilitate autonomous navigation, when the ego vehicle traverses, the ego vehicle can use a lane detection algorithm to identify lane segments on the environmental road. For example, the ego vehicle can acquire sensor data (e.g., light detection and ranging (LiDAR) and optical images), and apply an image segmentation model to detect lines along the road surface to identify lane segments on the road. However, this method may be limited to detecting lane segments from several different types of geometries, such as a single lane and its adjacent lanes along the road, with minimal ability to detect bifurcations and merges. As a result, the image segmentation model can confine the autonomous navigation of the ego vehicle to highly structured environments, such as highways with fewer lanes. Additionally, it may be difficult (if not impossible) for the ego vehicle to rely on such a model to autonomously navigate through more complex environments, such as intersections on local roads.
[0024] To address these and other technical constraints, the computing system on the ego vehicle can be configured with a collection of AI algorithms and ML models to generate a graph defining a set of lane segments (sometimes referred to herein as paths) and connectivity. To this end, the computing system on the ego vehicle can acquire sensor data (e.g., optical camera images and LiDAR) of the environment around the ego vehicle as well as map data (e.g., navigation map). Using a first set of encoders, the computing system can generate a set of tensors (or embeddings) as a reduced-dimensional representation of the sensor data. The computing system on the ego vehicle can enhance the set of tensors by applying a second set of encoders to the map data to add tensors to embed values related to topology and road layout. The resulting set of tensors can be a rich, dense representation of the environment around the ego vehicle.
[0025] Using the output set of tensors, the computing system on the ego vehicle can apply a third set of encoders (e.g., an autoregressive decoder) to generate a set of tokens for the graph. Each token can define various characteristics of points that form one or more of the lane segments in the lane segments passing through the environment. When applying the encoders to the set of tensors, the computing system can determine a first index value to define the position of points defined within a coarse grid. Using the first index value, the computing system can calculate a second index value to define the position of points within a fine grid. Leveraging the definition of points within the grid, the computing system can classify the points by topological type, such as start point, continuation, bifurcation, or end point, etc.
[0026] Continuing, when the point is linked to another point (e.g., if classified as a continuation or a fork), the computing system can also identify the index value of the reference point. The computing system can also calculate spline coefficients that define the connectivity between the two points to define the corresponding lane segment through the environment. Using a third set of encoders, the computing system can generate embeddings to represent the index value and the topological type, and then combine the embeddings to form tokens to insert into the graph. By repeating this process, the computing system on the ego can generate a set of tokens for the graph to define a linguistic representation of potential lane segments within the environment that the ego can navigate through. Using the lane segments defined by the graph, the ego can autonomously navigate through the environment.
[0027] Figure 1A are non-limiting examples of components of a system in which the methods and systems discussed herein can be implemented. For example, an analytics server can train an AI model and use the trained AI model to generate occupancy datasets and / or maps for one or more egos. Figure 1A Illustrates components of an AI-enabled visual data analysis system 100. System 100 can include an analytics server 110a, a system database 110b, an administrator computing device 120, egos 140a to 140b (collectively referred to as (multiple) egos 140), ego computing devices 141a to 141c (collectively referred to as ego computing devices 141), and a server 160. System 100 is not limited to the components described herein and can include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.
[0028] The components mentioned herein can be connected via a network 130. Examples of network 130 can include, but are not limited to, private or public LANs, WLANs, MANs, WANs, and the Internet. Network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.
[0029] Communication over network 130 can be performed according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 can also include communication over a cellular network, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.
[0030] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models such as (multiple) AI models 110c. Specifically, as Figure 1AAs depicted and described herein, the analytics server 110a can use the methods discussed herein to train the AI model(s) 110c using data retrieved from the ego 140 (e.g., by using data streams 172 and 174). When the AI model(s) 110c have been trained, each ego in the ego 140 can access and execute the trained AI model(s) 110c. For example, a vehicle 140a having an ego computing device 141a can transmit its camera feed to the trained AI model(s) 110c and can generate a graphic defining lane segments in the environment (e.g., data stream 174). Also, data ingested and / or predicted by the AI model(s) 110c with respect to the ego 140 (at inference time) can also be used to improve the AI model(s) 110c. Thus, the system 100 depicts a continuous loop that can periodically improve the accuracy of the AI model(s) 110c. Also, the system 100 depicts a loop in which, in addition to the inference phase, data received by the ego 140 can also be used in the training phase.
[0031] The analytics server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the ego 140. The collected data can then be processed and prepared into a training dataset. The training dataset can then be used to train one or more AI models, such as the AI model 110c. The analytics server 110a can also be configured to collect visual data from the ego 140. Using the AI model 110c (trained using the methods and systems discussed herein), the analytics server 110a can generate a dataset and / or occupancy map for the ego 140. The analytics server 110a can display the occupancy map on the ego 140 and / or transmit the occupancy map / dataset to the ego computing device 141, the administrator computing device 120, and / or the server 160.
[0032] In Figure 1A , the AI model 110c is illustrated as a component of the system database 110b, but the AI model 110c can be stored in a different or separate component, such as a cloud storage device or any other data repository accessible to the analytics server 110a.
[0033] The analytics server 110a can also be configured to display an electronic platform that illustrates various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 such that an analyst can monitor the training of the AI model 110c. An example of an electronic platform generated and hosted by the analytics server 110a can be a web-based application or website configured to display the training dataset collected from the ego 140 and / or the training status / metrics of the AI model 110c.
[0034] The analysis server 110a can be any computing device including a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analysis server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment (such as a cloud environment).
[0035] The ego 140 can represent various electronic data sources that transfer data associated with its previous or current navigation session to the analysis server 110a. The ego 140 can be any device configured for navigation, such as a vehicle 140a and / or a truck 140c. The ego 140 is not limited to being a vehicle and can also include robotic devices. For example, the ego 140 can include a robot 140b, which can represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b can be equipped with software for achieving balance, navigation, perception, or interacting with the physical world. The robot 140b can also include various cameras configured to transfer visual data to the analysis server 110a.
[0036] Although referred to herein as an "ego", the ego 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the ego 140 can be controlled by a human operator or a remote processor. The ego 140 can include various sensors, such as Figure 1B the sensors depicted in. The sensors can be configured to collect data as the ego 140 navigates various terrains (such as roads). The analysis server 110a can collect the data provided by the ego 140. For example, the analysis server 110a can obtain navigation session and / or road / terrain data (such as an image of the ego 140 navigating on a road) from various sensors, such that the collected data is ultimately used by the AI model 110c for training purposes.
[0037] As used herein, a navigation session corresponds to the journey of the ego 140's travel route, regardless of whether the journey is autonomous or controlled by a human. In some embodiments, the navigation session can be used for data collection and model training purposes. However, in some other embodiments, the ego 140 can refer to a vehicle purchased by a consumer, and the purpose of the journey can be classified as daily use. A navigation session can start when the ego 140 moves more than a threshold distance (such as 0.1 mile, 100 feet) or more than a threshold rate (such as more than 0 mph, more than 1 mph, more than 5 mph) from a non-moving position. A navigation session can end when the ego 140 returns to a non-moving position and / or shuts down (such as when the driver leaves the vehicle).
[0038] The self 140 can represent a collection of selves monitored by the analytics server 110a for training the AI model(s) 110c. For example, a driver of a vehicle 140a can authorize the analytics server 110a to monitor data associated with their corresponding vehicle. As a result, the analytics server 110a can utilize the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analytics server 110a can then apply the trained AI model(s) 110c to analyze data associated with the self 140 and predict an occupancy map for the self 140. Moreover, additional / ongoing data associated with the self 140 can also be processed and added to the training dataset so that the analytics server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a cycle in which navigation data received from the self 140 can be used to train the AI model(s) 110c. The self 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the self 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the self 140 represents a self that can be used for training, executing / using, and recalibrating the AI model(s) 110c. In a non-limiting example, the self 140 represents vehicles purchased by customers that can use the AI model(s) 110c for autonomous navigation while improving the AI model(s) 110c.
[0039] The self 140 can be equipped with various technologies that allow the self to collect data from its surrounding environment and (possibly) navigate autonomously. For example, the self 140 can be equipped with an inference chip to run autonomous driving software.
[0040] Various sensors for each self 140 can monitor data collected associated with different navigation sessions and transmit the data to the analytics server 110a. Figures 1B to 1C A block diagram of sensors integrated within the self 140 according to an embodiment is illustrated. The number and location of each sensor discussed with respect to Figures 1B to 1C can depend on the type of self discussed in Figure 1A . For example, the robot 140b can include different sensors than the vehicle 140a or the truck 140c. For example, the robot 140b may not include an airbag activation sensor 170q. Moreover, the sensors of the vehicle 140a and the truck 140c can be positioned differently than illustrated in Figure 1C .
[0041] As discussed herein, the various sensors integrated within each vehicle body 140 can be configured to measure various data associated with each navigation session. The analysis server 110a can periodically collect the data monitored and collected by these sensors, where the data is processed according to the methods described herein and is used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.
[0042] The vehicle body 140 can include a user interface 170a. The user interface 170a can refer to the user interface of a vehicle body computing device (such as Figure 1A the vehicle body computing device 141 in). The user interface 170a can be implemented as a display screen integrated with or coupled to the interior of the vehicle, a head-up display, a touch screen, etc. The user interface 170a can include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a can be adapted to provide user input (such as a signal and / or sensor information) to other devices or sensors (such as Figure 1B the illustrated sensors) of the vehicle body 140, such as the controller 170c.
[0043] The user interface 170a can also be implemented with one or more logic devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, transmit and / or receive communications (such as sensor signals, control signals, sensor information, user input, and / or other information) or execute various other processes and / or methods. In another example, the driver can use the user interface 170a to control the temperature of the vehicle body 140 or activate its features (such as the autonomous driving or steering system 170o). Thus, the user interface 170a can combine with other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analysis server 110a and / or the AI model 110c.
[0044] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of other digital or analog devices capable of measuring the orientation of the vehicle body 140 (such as the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations, such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide heading measurements to the vehicle body 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rates to the vehicle body 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a specific coordinate system of the vehicle body 140.
[0045] The controller 170c can be implemented as any suitable logic device (e.g., a processing device, a microcontroller, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of the self 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via the user interface 170a), querying the device for operating parameters, selecting operating parameters for the device, or performing any of the various operations described herein.
[0046] The communication module 170e can be implemented as any wired and / or wireless interface configured to transmit sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (e.g., the analytics server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are implemented within Figure 1B one or more of the elements shown and within the sensors. In some embodiments, the communication module 170e can delay the transmission of sensor data. For example, when the self 140 does not have network connectivity, the communication module 170e can store the sensor data in a temporary data storage device and transmit the sensor data when the self 140 is identified as having appropriate network connectivity.
[0047] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind rate sensor (e.g., direction and magnitude), and / or other devices capable of measuring or determining the linear speed of the self 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the self 140) and providing such measurement as a sensor signal (which can be transmitted to various devices).
[0048] The gyroscope / accelerometer 170f can be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration of the self 140 (e.g., direction and magnitude) and providing such measurement as a sensor signal (which can be transmitted to various devices, such as the analytics server 110a). The gyroscope / accelerometer 170f can be positioned and / or adapted to make such measurements relative to a particular coordinate system of the self 140. In various embodiments, the gyroscope / accelerometer 170f can be associated with Figure 1BThe other elements depicted are implemented in a common housing and / or module to ensure a common reference frame or known transformations between reference frames.
[0049] The Global Navigation Satellite System (GNSS) 170h can be implemented as a Global Positioning Satellite receiver and / or another device capable of determining the absolute and / or relative position of the vehicle body 140 based on wireless signals received, for example, from space and / or ground sources and capable of providing measurements such as sensor signals (which can be transmitted to various devices). In some embodiments, the GNSS 170h can be adapted to determine the rate, speed, and / or yaw rate of the vehicle body 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute rate and / or angular rate of the vehicle body 140.
[0050] The temperature sensor 170i can be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other devices capable of measuring the temperature associated with the vehicle body 140 and providing such measurement as a sensor signal. The temperature sensor 170i can be configured to measure the ambient temperature associated with the vehicle body 140, such as the cockpit or dashboard temperature, and for example, this ambient temperature can be used to estimate the temperature of one or more elements of the vehicle body 140.
[0051] The humidity sensor 170j can be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the vehicle body 140 and providing such measurement as a sensor signal.
[0052] The steering sensor 170g can be adapted to physically adjust the heading of the vehicle body 140 according to one or more control signals provided by a logic device (such as the controller 170c) and / or user input. The steering sensor 170g can include one or more actuators and control surfaces (such as a rudder or other types of steering or adjustment mechanisms) of the vehicle body 140, and can be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g can also be adapted to sense the current steering angle / position of such a steering mechanism and provide such measurement.
[0053] The propulsion system 170k can be implemented as a propeller, a turbine, or other thrust-based propulsion systems, mechanical wheeled and / or tracked propulsion systems, wind / sail-based propulsion systems, and / or other types of propulsion systems that can be used to provide power to the vehicle body 140. The propulsion system 170k can also monitor the direction of the power and / or thrust of the vehicle body 140 relative to the reference coordinate system of the vehicle body 140. In some embodiments, the propulsion system 170k can be coupled to the sensor 170g and / or integrated with the steering sensor 170g.
[0054] The passenger restraint sensor 170l can monitor seat belt detection and locking / unlocking assemblies, as well as other passenger restraint subsystems. The passenger restraint sensor 170l can include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the vehicle 140. For example, the passenger restraint sensor 170l can be configured to receive motion and / or status data from Figure 1B the other sensors depicted. The passenger restraint sensor 170l can determine whether a safety measure (e.g., seat belt) is being used.
[0055] As Figure 1C depicted, the camera 170m can refer to one or more cameras integrated within the vehicle 140, and can include multiple cameras integrated (or retrofitted) into the vehicle 140. The camera 170m can be an interior- or exterior-facing camera of the vehicle 140. For example, as Figure 1C depicted, the vehicle 140 can include one or more interior-facing cameras that can monitor and collect footage of the passengers of the vehicle 140. The vehicle 140 can include eight exterior-facing cameras. For example, the vehicle 140 can include a front camera 170m-1, front side cameras 170m-2, 170m-3, rear side cameras 170m-4 on each front fender, cameras 170m-5 on each side (e.g., integrated within the B-pillar), and a rear camera 170m-6.
[0056] Referring to Figure 1B , the radar 170n and ultrasonic sensors 170p can be configured to monitor the distance of the vehicle 140 to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). The vehicle 140 can also include an autonomous driving or steering system 170o configured to autonomously navigate the vehicle 140 using data collected via various sensors, such as the radar 170n, speed sensor 170d, and / or ultrasonic sensors 170p.
[0057] Thus, the autonomous driving or steering system 170o can analyze the various data collected by one or more of the sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to the various features discussed herein, such as an analysis server.
[0058] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can transmit data regarding airbag deployment, including data associated with the event that caused the deployment.
[0059] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analytics server 110a (such as various analytics metrics and risk scores), where the system administrator can monitor the various models utilized by the analytics server 110a, review feedback, and / or facilitate the training of the AI model(s) 110c maintained by the analytics server 110a.
[0060] (The) vehicle(s) 140 can be any device configured to navigate various routes, such as a vehicle 140a or a robot 140b. As discussed with respect to Figures 1B to 1C , the vehicle(s) 140 can include various telemetry sensors. The vehicle(s) 140 can also include a vehicle computing device 141. Specifically, each vehicle can have its own vehicle computing device 141. For example, a truck 140c can have a vehicle computing device 141c. For the sake of brevity, the vehicle computing devices are collectively referred to as the vehicle computing device(s) 141. The vehicle computing device 141 can control content presentation on the infotainment system of the vehicle 140, process commands associated with the infotainment system, aggregate sensor data, manage communication of data to electronic data sources, receive updates, and / or transmit messages. In one configuration, the vehicle computing device 141 communicates with an electronic control unit. In another configuration, the vehicle computing device 141 is an electronic control unit. The vehicle computing device 141 can include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the AI model(s) 110c described herein can be stored and executed (or directly accessed) by the vehicle computing device 141. Non-limiting examples of the vehicle computing device 141 can include vehicle multimedia and / or display systems.
[0061] In one example of training the AI model 110c, the analytics server 110a can collect data from the vehicle 140 to train the AI model(s) 110c. Before executing the AI model(s) 110c to generate or predict a graphic defining a lane segment, the analytics server 110a can use various methods to train the AI model(s) 110c. Training allows the AI model(s) 110c to ingest data from one or more cameras of one or more vehicles 140 (without receiving radar data) and predict occupancy data for the surrounding environment of the vehicle. The operations described in this example can be performed by the Figure 1A and 1Bexecuted by any number of computing devices (e.g., processors of the vehicle 140) operating in the distributed computing system described.
[0062] To train the AI model(s) 110c, the analytics server 110a can first employ one or more vehicles in the vehicle 140 to drive a specific route. While driving, the vehicle 140 can use one or more of its sensors (including one or more cameras) to generate navigation session data. For example, one or more vehicles 140 equipped with various sensors can navigate a designated route. As one or more vehicles 140 in the vehicle 140 traverse the terrain, their sensors can capture continuous (or periodic) data of their surrounding environment. The sensors can indicate the occupancy status of the surrounding environment of one or more vehicles 140. For example, the sensor data can indicate various objects with mass in the surrounding environment of one or more vehicles 140 as they navigate their route.
[0063] In operation, as one or more vehicles 140 navigate, their sensors collect data and transmit the data to the analytics server 110a, as depicted by the data stream 172. In some embodiments, one or more vehicles 140 can include one or more high-resolution cameras that capture a continuous visual data stream from the surrounding environment of one or more vehicles 140 as one or more vehicles 140 navigate through the route. The analytics server 110a can then use the camera feed to generate a second data set that includes visual elements / depictions of different voxels of the surrounding environment of one or more vehicles 140. In operation, as one or more vehicles 140 navigate, their cameras collect data and transmit the data to the analytics server 110a, as depicted by the data stream 172. For example, the vehicle computing device 141 can use the data stream 172 to transmit image data to the analytics server 110a.
[0064] The analytics server 110a can use the data collected from the vehicle 140 (e.g., the camera feed received from the vehicle 140) to generate a training data set. The training data set can identify or include a set of examples. Each example can identify or include input data and expected output data from the input data. In each example, the input can include the data collected, such as sensor data (e.g., video or images from one or more cameras) and map data from the vehicle 140 (e.g., navigation map). The output can include environmental features (e.g., attributes collected from sensor data), map features (e.g., attributes in the navigation map, such as topological features and road layout), classifications (e.g., topological types), and output markers (e.g., a combination of environmental features, map features, and classifications) to be included in the graphics defining the lane segments, etc. In some embodiments, the output can be created by a human reviewer who examines the input data.
[0065] Using the training data set, the analysis server 110a can feed a series of training data sets into the AI model(s) 110c and obtain a set of predicted outputs (such as environmental features, map features, classifications, and output tags). The analysis server 110a can then compare the predicted data with the ground truth data to determine the differences and train the AI model(s) 110c by adjusting the internal weights and parameters of the AI model(s) 110c proportional to the determined differences according to a loss function. The analysis server 110a can train the AI model(s) 110c in a similar manner until the predictions of the trained AI model(s) 110c are accurate to a certain threshold (such as recall or precision).
[0066] In some embodiments, the analysis server 110a can use a supervised training method. For example, using the ground truth and the received visual data, the AI model(s) 110c can train itself such that it can predict an output. As a result, when being trained, the AI model(s) 110c can receive sensor data and map data, analyze the received data, and generate tags. In some embodiments, the analysis server 110a can use an unsupervised method in which the training data set is not labeled. Since labeling the data within the training data set can be time-consuming and may require excessive computing power, the analysis server 110a can utilize unsupervised training techniques to train the AI model 110c.
[0067] With the AI model 110c established, the analysis server 110a can transmit, send, or otherwise distribute the weights of the AI model 110c to each of the autonomous computing devices 141a to 141c. Once received, the autonomous computing devices 141a to 141c can store and maintain the AI model 110c on a local storage device. Once stored and loaded, the autonomous computing devices 141a to 141c can use it when processing newly acquired data (such as sensor and map data) to create a graph to define lane segments to autonomously navigate the corresponding autonomous entities 140a to 140c through the environment. The analysis server 110a can transmit, send, or otherwise distribute updated weights of the AI model 110c from time to time to update the instances of the AI model 110c on the autonomous computing devices 141a to 141c.
[0068] Now refer to Figure 2, which depicts a block diagram of the architecture of a system 200 for generating a tensor representation of a pathway for autonomous navigation through an environment. System 200 may be implemented using any of the components described herein, such as the self-computing devices 141a to 141c and the AI model(s) 110c. System 200 may include the components and steps described herein. However, other embodiments may include additional or alternative components and steps, or one or more components and steps may be omitted. System 200 and its architecture may be implemented by an analytics server (e.g., a computer similar to analytics server 110a) or a self-computing device (e.g., self-computing devices 141a to 141c), or across multiple computing systems. However, one or more steps of method 200 may be implemented by any number of computing devices (e.g., processors of self 140 and / or self-computing devices 141) operating in the Figures 1A to 1C distributed computing system described herein. For example, one or more computing devices of the self may locally implement Figure 2 some or all of the components and steps described herein.
[0069] System 200 may include at least one vision component 205 to process sensor data. The vision component 205 may include a set of cameras, such as at least one main camera 202A, at least one left or right pillar camera 202B, and at least one backup camera 202C, etc. The set of cameras 202A to 202C (collectively referred to herein as camera 202) may be an instance of camera 170m and may be multiple cameras integrated (or retrofitted) into self 140 as discussed herein. Each camera 202 may acquire and collect images of the environment surrounding the self (e.g., optical vision images). The images may be in the form of a set of frames of a video acquired by camera 202. In some embodiments, the vision component 205 may include other sensors, such as radar 170n and ultrasonic sensors 170p, as discussed herein. In some embodiments, the vision component 205 may rely solely on optical images captured by the optical cameras 202.
[0070] The vision component 205 may include artificial intelligence (AI) algorithms or machine learning (ML) models to process sensor data from the set of cameras and other sensors. The ML models may be trained and established as discussed herein. Generally, the ML models of the vision component 205 may generate a set of embeddings corresponding to a reduced-dimensional representation of the sensor data. The ML models in the vision component 205 may include, for example, a set of self-regulating networks (RegNet) 210A to 210C (collectively referred to hereinafter as residual networks 210), a set of feature pyramid networks (FPN) 215A to 215C (collectively referred to hereinafter as feature pyramid networks 215), at least one transformer 220, and at least one video module 225, etc.
[0071] Each self-regulating network 210 may include a set of weights arranged according to a collection of convolutional recurrent neural networks (RNNs) (e.g., a collection including long short-term memory networks (LSTMs) or gated recurrent units (GRUs)) and operators (e.g., concatenators and activation functions). Each self-regulating network 210 may retrieve, identify, or otherwise receive sensor data from a corresponding camera 202 (or other sensor). Once received, the self-regulating network 210 may process the sensor data according to the set of weights to generate a set of embeddings (e.g., feature maps). The set of embeddings may be a reduced-dimensional representation of the sensor data, particularly having spatio-temporal features. The self-regulating network 210 may feed the output set of embeddings forward to the corresponding feature pyramid network 215.
[0072] Each feature pyramid network 215 may include a set of weights arranged at multiple scales according to a collection of convolutional neural networks (CNNs) to detect features within the input data. Each feature pyramid network 215 may retrieve, identify, or otherwise receive the output set of embeddings from a corresponding self-regulating network 210. With the receipt, the feature pyramid network 215 may use the set of weights to process the set of embeddings from the self-regulating network 210 to generate another set of embeddings. This set of embeddings may be a further reduced-dimensional representation of the sensor data. Additionally, the transformer 220 may include a set of weights arranged according to a transformer architecture with a multi-head attention mechanism. The feature pyramid network 215 may feed the output set of embeddings forward to the transformer 220.
[0073] The transformer 220 may receive, collect, or otherwise aggregate the set of embeddings generated by the feature pyramid network 215 and by the extended self-regulating network 210 from the sensor data. The transformer 220 may process the set of embeddings according to the set of weights to generate another set of embeddings (or output tokens). The video module 225 may receive, collect, or otherwise aggregate the set of embeddings output by the transformer 220. Once received, the video module 225 may combine the set of embeddings corresponding to the sensor data acquired within a defined time period. With this combination, the video module 225 may generate an aggregated set of embeddings for feed-forward. The resulting set of embeddings may represent or define a reduced-dimensional representation of the sensor data acquired via the camera 202 (or other sensor).
[0074] System 200 may include at least one map component 230 to process map data, such as navigation map 235. Navigation map 235 may include or identify data defining a map or topology of the environment surrounding the ego, such as terrain type, elevation, or semantic labels of roads (e.g., navigable routes, unpaved paths, streets, avenues, highways, bus lanes, lane counts, and ramps) or other features (e.g., geometry, buildings, signage, and flora) by coordinates in the environment (e.g., Global Positioning System (GPS) coordinates). Map component 230 may retrieve, obtain, or otherwise acquire navigation map 235 relative to the position of the ego in the environment. For example, map component 230 may acquire navigation map 235 defining a size (e.g., 2km by 2km) along the direction of travel of the ego.
[0075] Map component 230 may include at least one lane guidance module 240 to enhance the embedding set from vision component 205 with an embedding set derived from navigation map 235. Lane guidance module 240 may include a set of weights arranged according to an encoder (e.g., a Convolutional Neural Network (CNN) ensemble). Lane guidance module 240 may retrieve, obtain, or otherwise identify navigation map 235 defining the topology surrounding the ego. Using this identification, lane guidance module 240 may use the set of weights to process the data of navigation map 235 to generate an embedding set. The embedding set may be a low-dimensional representation of relevant features in navigation map 235. In some embodiments, lane guidance module 240 may utilize the embedding set from lane component 205 to process the data from navigation map 235. Lane guidance module 240 may combine the embedding set derived from sensor data with the embedding set derived from navigation map 235 to produce, output, or otherwise generate at least one tensor 250. Tensor 250 may include an aggregated embedding set derived from sensor data and map data.
[0076] System 200 may include at least one lane language component 255 to process tensor 250 from map component 230. The lane language component 255 may include at least one lane encoder 260. The lane encoder 260 may include a set of weights arranged according to an encoder or decoder (such as an autoregressive decoder, etc.). The autoregressive decoder of the lane encoder 260 may include a model sequence to process a portion of the input embedding from tensor 250 to generate an output embedding that depends on another output embedding derived from a previous portion of the input embedding. From processing tensor 250, the lane encoder 260 may produce, output, or otherwise generate a set of lane instances 265 and at least one adjacency matrix 270. The lane instances 265 may define one or more lane segments through which an ego may potentially navigate the environment. The adjacency matrix 270 may define the connectivity or relationships between the lane segments defined in lane segments 265. The lane instances 265 and the adjacency matrix 270 may be collectively referred to as a lane graph or lane language. Additional details regarding the operation of the lane language component 255 and the lane encoder 260 are provided herein in connection with Figures 3A to 3H Additional details regarding the operation of the lane language component 255 and the lane encoder 260 are provided.
[0077] Now referring to Figures 3A to 3H , a block diagram of a process 300 for generating a path for autonomous navigation through an environment is depicted. The process 300 may be implemented using any of the components described herein, such as ego computing devices 141a to 141c and (a) plurality of AI models 110c. The process 300 may include the steps described herein. However, other embodiments may include additional or alternative steps, or one or more steps may be omitted. The process 300 may be executed by an analytics server (e.g., a computer similar to analytics server 110a) or an ego computing device (e.g., ego computing devices 141a to 141c). However, one or more steps of the process 300 may be executed by any number of computing devices (e.g., processors of ego 140 and / or ego computing devices 141 or a centralized service such as analytics server 110a) operating in a Figures 1A to 1C distributed computing system described herein. For example, one or more computing devices of the ego may execute locally Figures 3A to 3H some or all of the steps described herein.
[0078] From Figure 3ABeginning, under process 300, a computing system can retrieve, receive, or otherwise identify at least one vector space encoding 302 (such as tensor 250) to be used to form or generate at least one graphic 304. The computing system performing process 300 can correspond to an on-ego computing device 141 on the on-ego 140 or a centralized service such as analytics server 110a. The vector space encoding 302 (sometimes referred to herein as a tensor) can identify or include a set of encodings (sometimes referred to herein as a set of embeddings or feature maps). The set of encodings of the vector space encoding 302 can be generated, determined, or otherwise derived from sensor data acquired by the on-ego (such as from cameras and other sensors) and map data defining the environmental topology surrounding the on-ego (such as navigation map 235). The encodings can be a low-dimensional representation of latent features derived from the sensor data and the map data. The graphic 304 can be used to define a set of pathways (sometimes referred to herein as lane segments) along which the on-ego can autonomously navigate through the environment.
[0079] Using this identification, the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a first grid 308 defined on the environment represented by map 310. The point 306A can correspond to a starting point from which the on-ego navigates through the environment as defined by map 310. The grid 308 can specify or define the set of grid points (or coordinates) at a coarser or lower resolution than the original resolution of the grid points defined by map 310. The map 310 can correspond to the topological layout surrounding the on-ego and can be acquired as part of the map data (such as navigation map 235).
[0080] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to a first point predictor unit 310. The portion of the vector space encoding 302 input into the first point predictor unit 310 can correspond to the point 306A defined within the set of grid points in the first grid 308 in the environment. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to be input based on the point 306A. The first point predictor unit 310 can correspond to a portion of a machine learning (ML) model (such as lane encoder 260 discussed herein) and can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the first point predictor unit 310 to compute, determine, or otherwise generate at least one coordinate 330A corresponding to the point 306A. Using this point, the computing system can use the first point predictor unit 310 to compute, produce, or otherwise determine a first index value 330’A for the point 306A.
[0081] Move toFigure 3B , the computing system can find, determine, or otherwise identify at least one point 306A within a set of grid points in a second grid 308' defined on an environment represented by map 312. The computing system can correspond to an on-vehicle computing device 141 on vehicle 140 or a centralized service such as analytics server 110a. Point 306A can correspond to a starting point from which the vehicle navigates through the environment as defined by map 312. The second grid 308' can specify or define the set of grid points (or coordinates) with a finer or higher resolution than the resolution of the first grid 308. The second grid 308' can correspond to a subset of the first grid 308 around or surrounding the vehicle. Map 312 can correspond to a topological layout around the vehicle and can be obtained as part of map data (e.g., navigation map 235).
[0082] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 and the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’A) to the second point predictor unit 312. The portion of the vector space encoding 302 can correspond to a point 306A defined within a set of grid points in the first grid 308' in the environment. The portion of the vector space encoding 302 input to the second point prediction unit 312 can be the same as or different from the portion of the vector space encoding 302 input to the first point predictor unit 310. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to be input based on point 306A.
[0083] The second point predictor unit 312 can correspond to a part of a machine learning (ML) model (e.g., the lane encoder 260 discussed herein) and can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332A corresponding to point 306A. Using this point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332’A for point 306A.
[0084] Continue to Figure 3C, the computing system may input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the topology predictor unit 314. The computing system may correspond to an on-vehicle computing device 141 on the vehicle 140 or a centralized service such as the analytics server 110a. Additionally, the computing system may apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’A) or the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’A) to the topology predictor unit 314. The portion of the vector space encoding 302 input to the topology predictor unit 314 may be the same as or different from the portion of the vector space encoding 302 input to the first point prediction unit 310 or the second point prediction unit 312. The topology predictor unit 314c may correspond to a portion of a machine learning (ML) model (e.g., the lane encoder 260 discussed herein) and may include a set of weights arranged according to cross-attention, self-attention, and transformers, etc.
[0085] By applying the topology predictor unit 314, the computing system may determine, classify, or otherwise categorize the point 306A as at least one topology type 334A. The topology type 334A may semantically specify, identify, or otherwise define the point 306A as a function of the path along which the vehicle navigates through the environment as represented by the map 310. The topology types may include, for example, start types, continuation types, fork types, or terminal types, etc. In the depicted example, the topology type 334A may be a start type to indicate the start of at least one path in the path.
[0086] Next to Figure 3D , the computing system may input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the point attribute predictor 316. The computing system may correspond to an on-vehicle computing device 141 on the vehicle 140 or a centralized service such as the analytics server 110a. Additionally, the computing system may apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’A), the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’A), or the output (e.g., an embedding derived from the topology type 334’A) to the point attribute predictor 316. The portion of the vector space encoding 302 input to the point attribute predictor 316 may be the same as or different from the portion of the vector space encoding 302 input to the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314.
[0087] The point attribute predictor 316 may correspond to a part of a machine learning (ML) model (such as the lane encoder 260 discussed herein), and may include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. By applying the point attribute predictor 316, the computing system may generate or determine a set of point attributes 336A for the point 306A. The set of point attributes 336A may identify or include, for example: at least one bifurcation point that identifies the index value of a previous point from which the current point 306A forms a bifurcation; at least one merge point that identifies the index value of a previous point from which the current point 306A forms a merge; and a set of spline coefficients 338A, etc. In the depicted example, since the current point 306A is a starting point, there are no other points, and thus the set of point attributes 336A may be empty.
[0088] Using the generated output, the computing system may use predictors in the ML model to generate respective embeddings to form or output an embedding set 340A for the point 306A. Using the first point predictor 310, the computing system may determine or generate at least one corresponding embedding 330”A from the first index value 330’A. Using the second point predictor 312, the computing system may determine or generate at least one corresponding embedding 332”A from the second index value 332’A. Using the second point predictor 312, the computing system may determine or generate at least one corresponding embedding 332”A from the second index value 332’A. Using the topology type predictor 314, the computing system may determine or generate at least one corresponding embedding 334’A from the topology type 334A. Similarly, using the point attribute predictor 316, the computing system may determine or generate an embedding set 336’A. The embedding set 336’A may lack spline coefficients.
[0089] When generating respective embeddings (such as embeddings 330”A, 332”A, 334’A, and 336’A), the computing system may write, produce, or otherwise generate an embedding set 340A. Based on the embedding set 340A, the computing system may produce, output, or otherwise generate at least one marker 342A for at least one path in the passage through the environment. In some embodiments, the computing system may generate the marker 342A based on the first index value 330A, the second index value 332A, the topology type 334A, or the set of point attributes 336A, or any combination thereof. With this generation, the computing system may add, insert, or otherwise include the marker 342A in the graph 304. The computing system may store and maintain the graph 304 including the marker 342A to be used to navigate itself through the environment via the path. The graph 304 may be maintained as one or more data structures (such as an array, linked list, matrix, tree, hash, heap, table, or graph) on a data storage device.
[0090] Now refer to Figure 3E, the computing system may repeat some of the functions described herein for the second point 306B. The computing system may correspond to an on-body computing device 141 on the body 140 or a centralized service such as the analytics server 110a. Using the identification of the second point 306B, the computing system may find, determine, or otherwise identify at least one point 306B within a set of grid points in the first grid 308 defined on the environment represented by the map 310. The point 306B may correspond to an intermediate point that the body will pass through when navigating through the environment defined by the map 310. The grid 308 may specify or define the set of grid points (or coordinates) at a resolution coarser or lower than the original resolution of the grid points defined by the map 310. The map 310 may correspond to the topological layout around the body and may be obtained as part of the map data (e.g., the navigation map 235).
[0091] The computing system may input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the first point predictor unit 310. The portion of the vector space encoding 302 input into the first point predictor unit 310 may correspond to the point 306B defined within the set of grid points in the first grid 308 in the environment. In some embodiments, the computing system may identify or select the portion of the vector space encoding 302 to input based on the point 306B. In the feeding, the computing system may process the portion of the vector space encoding 302 according to the weights of the first point predictor unit 310 to calculate, determine, or otherwise generate at least one coordinate 330B corresponding to the point 306B. Using this point, the computing system may use the first point predictor unit 310 to calculate, produce, or otherwise determine a first index value 330’B for the point 306B.
[0092] The computing system may find, determine, or otherwise identify at least one point 306B within a set of grid points in the second grid 308’ defined on the environment represented by the map 312. The point 306B may correspond to a starting point from which the body will navigate through the environment defined by the map 310. The second grid 308’ may specify or define the set of grid points (or coordinates) at a resolution finer or higher than that of the first grid 308. The second grid 308’ may correspond to a subset of the first grid 308 around or surrounding the body. The map 312 may correspond to the topological layout around the body and may be obtained as part of the map data (e.g., the navigation map 235).
[0093] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 and the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’B) to the second point predictor unit 312. The portion of the vector space encoding 302 can correspond to the point 306B defined within the set of grid points in the first grid 308’ in the environment. The portion of the vector space encoding 302 input into the second point prediction unit 312 can be the same as or different from the portion of the vector space encoding 302 input into the first point predictor unit 310. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to be input based on the point 306B. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332B corresponding to the point 306B. Using this point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332’B for the point 306B.
[0094] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the topology predictor unit 314. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’B) or the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’B) to the topology predictor unit 314. The portion of the vector space encoding 302 input into the topology predictor unit 314 can be the same as or different from the portion of the vector space encoding 302 input into the first point predictor unit 310 or the second point predictor unit 312.
[0095] By applying the topology predictor unit 314, the computing system can determine, classify, or otherwise categorize the point 306B as at least one topology type 334B. The topology type 334B can semantically specify, identify, or otherwise define the function of the point 306B relative to a path along which an agent navigates through the environment as represented by the map 310. The topology type can include, for example, a start type, a continuation type, a fork type, or a terminal type, etc. In the depicted example, the topology type 334B can be a continuation type to indicate the continuation of at least one path in the path between the point 306A and the point 306B.
[0096] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the point property predictor 316. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’B), the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’B), or an output (e.g., an embedding derived from the topology type 334’B) into the point property predictor 316. The portion of the vector space encoding 302 that is input into the point property predictor 316 can be the same as or different from the portion of the vector space encoding 302 that is input into the first point prediction unit 310, the second point prediction unit 312, or the topology prediction unit 314.
[0097] By applying the point property predictor 316, the computing system can generate or determine a point property set 336B for the point 306B. The point property set 336B can identify or include, for example: at least one bifurcation point that identifies an index value of a previous point from which the current point 306B forms a bifurcation; at least one merge point that identifies an index value of a previous point from which the current point 306B forms a merge; and a set of spline coefficients 338B, etc. In the depicted example, since the current point 306B is a continuation point that depends on the point 306A rather than a bifurcation or a merge, the index can be empty. The set of spline coefficients 338B can include a set of values that define a spline curve 338’B between the points 306A and 306B. The set of spline coefficients 338B can define a path through the environment between the points 306A and 306B.
[0098] Using the generated output, the computing system can use predictors in the ML model to generate respective embeddings to form or output an embedding set 340A for the point 306B. Using the first point predictor 310, the computing system can determine or generate at least one corresponding embedding 330”B from the first index value 330’B. Using the second point predictor 312, the computing system can determine or generate at least one corresponding embedding 332”B from the second index value 332’B. Using the second point predictor 312, the computing system can determine or generate at least one corresponding embedding 332”B from the second index value 332’B. Using the topology type predictor 314, the computing system can determine or generate at least one corresponding embedding 334’B from the topology type 334B. Similarly, using the point property predictor 316, the computing system can determine or generate an embedding set 336’B.
[0099] When generating each embedding (such as embeddings 330”B, 332”B, 334’B, and 336’B), the computing system can write, produce, or otherwise generate an embedding set 340A. Based on the embedding set 340A, the computing system can produce, output, or otherwise generate at least one marker 342B for at least one path in the paths through the environment. The computing system can also apply the self-attention unit 318 when determining the marker 342B. In some embodiments, the computing system can generate the marker 342B based on the first index value 330B, the second index value 332B, the topology type 334B, or the set of point attributes 336B, or any combination thereof. With this generation, the computing system can add, insert, or otherwise include the marker 342B in the graph 304. The computing system can update the graph 304 to include the marker 342B for use in autonomously navigating the self through the environment via the paths.
[0100] Proceeding to Figure 3F , the computing system can repeat some of the functions described herein for the third point 306C. The computing system can correspond to an on-self computing device 141 on the self 140 or a centralized service such as the analytics server 110a. Using the identification of the third point 306C, the computing system can find, determine, or otherwise identify at least one point 306C within a set of grid points in the first grid 308 defined on the environment represented by the map 310. The point 306C can correspond to an end point near the boundary of the map 310 and can be the point to which the self is to navigate through the environment as defined by the map 310. The grid 308 can specify or define the set of grid points (or coordinates) at a resolution coarser or lower than the original resolution of the grid points defined by the map 310. The map 310 can correspond to the topological layout around the self and can be obtained as part of the map data (such as the navigation map 235).
[0101] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the first point predictor unit 310. The portion of the vector space encoding 302 input to the first point predictor unit 310 can correspond to the point 306C defined within the set of grid points in the first grid 308 in the environment. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to input based on the point 306C. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the first point predictor unit 310 to calculate, determine, or otherwise generate at least one coordinate 330C corresponding to the point 306C. Using this point, the computing system can use the first point predictor unit 310 to calculate, produce, or otherwise determine a first index value 330’C for the point 306C.
[0102] The computing system can find, determine, or otherwise identify at least one point 306C within a set of grid points in a second grid 308' defined on an environment represented by map 312. The point 306C can correspond to a starting point from which the self is to navigate through an environment defined by, for example, map 310. The second grid 308' can specify or define a set of grid points (or coordinates) with a finer or higher resolution than the resolution of the first grid 308. The second grid 308' can correspond to a subset of the first grid 308 around or surrounding the self. The map 312 can correspond to a topological layout around the self and can be obtained as part of map data (e.g., navigation map 235).
[0103] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 and the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330'C) to the second point predictor unit 312. The portion of the vector space encoding 302 can correspond to a point 306C defined within a set of grid points in the first grid 308' in the environment. The portion of the vector space encoding 302 input into the second point prediction unit 312 can be the same as or different from the portion of the vector space encoding 302 input into the first point predictor unit 310. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to be input based on the point 306C. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332C corresponding to the point 306C. Using this point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332'C for the point 306C.
[0104] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the topological predictor unit 314. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330'C) or the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332'C) to the topological predictor unit 314. The portion of the vector space encoding 302 input into the topological predictor unit 314 can be the same as or different from the portion of the vector space encoding 302 input into the first point prediction unit 310 or the second point prediction unit 312.
[0105] By applying the topology predictor unit 314, the computing system can identify, classify, or otherwise categorize the point 306C as at least one topology type 334C. The topology type 334C can semantically specify, identify, or otherwise define the function of the point 306C relative to a path along which the ego navigates through the environment as represented by the map 310. The topology types can include, for example, a start type, a continuation type, a fork type, or a terminal type, etc. In the depicted example, the topology type 334C can be a continuation type to indicate the continuation of at least one path in the path between the point 306A and the point 306C.
[0106] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the point attribute predictor 316. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’C), the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’C), or the output (e.g., an embedding derived from the topology type 334’C) into the point attribute predictor 316. The portion of the vector space encoding 302 input into the point attribute predictor 316 can be the same as or different from the portion of the vector space encoding 302 input into the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314.
[0107] By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336C for the point 306C. The set of point attributes 336C can identify or include, for example: at least one fork point that identifies the index value of the previous point from which the current point 306C forms a fork; at least one merge point that identifies the index value of the previous point from which the current point 306C forms a merge; and a set of spline coefficients 338C, etc. In the depicted example, since the current point 306C is a continuation point that depends on the point 306A rather than a fork or a merge, the index can be empty. The set of spline coefficients 338C can include a set of values that define the spline curve 338’C between the points 306A and 306C. The set of spline coefficients 338C can define the path through the environment between the points 306A and 306C.
[0108] In some embodiments, the computing system may also determine whether point 306C is an end point towards the boundary of map 310 based on coordinates from the first grid 308 or the second grid 308'. When the coordinates of the point are within the margin of the boundary of map 310 as defined by the first grid 308 or the second grid 308', it can be determined that the point is terminal. The end point may correspond to the end of a path (or line segment) opposite to the initial starting point (e.g., point 306A). To determine, the computing system may identify the coordinates of point 306C defined by the first grid 308 or the second grid 308'. Using this identification, the computing system may compare the coordinates of point 306C with the coordinates corresponding to the boundary of the acquired map 310. When the coordinates of point 306C are within the margin (e.g., equivalent to 1m to 10m) of the boundary of the acquired map 310, the computing system may determine that point 306C is an end point (e.g., as depicted). Additionally, the computing system may determine that a path is defined between point 306A and point 306C (e.g., via point 306B). Otherwise, when the coordinates of point 306C are outside the margin (e.g., equivalent to 1m to 10m) of the boundary of the acquired map 310, the computing system may determine that point 306C is not an end point (e.g., as depicted).
[0109] Using the generated output, the computing system may use predictors in the ML model to generate respective embeddings to form or output an embedding set 340A for point 306C. Using the first point predictor 310, the computing system may determine or generate at least one corresponding embedding 330”C from the first index value 330’C. Using the second point predictor 312, the computing system may determine or generate at least one corresponding embedding 332”C from the second index value 332’C. Using the second point predictor 312, the computing system may determine or generate at least one corresponding embedding 332”C from the second index value 332’C. Using the topology type predictor 314, the computing system may determine or generate at least one corresponding embedding 334’C from the topology type 334C. Similarly, using the point attribute predictor 316, the computing system may determine or generate an embedding set 336’C.
[0110] When generating each embedding (e.g., embeddings 330°C, 332°C, 334°C, and 336°C), the computing system can write, generate, or otherwise produce an embedding set 340A. Based on the embedding set 340A, the computing system can generate, output, or otherwise produce at least one marker 342C for at least one path in the passage through the environment. The computing system can also apply the self-attention unit 318 when determining the marker 342C. In some embodiments, the computing system can generate the marker 342C based on the first index value 330C, the second index value 332C, the topology type 334C, or the set of point attributes 336C, or any combination thereof. With this generation, the computing system can add, insert, or otherwise include the marker 342C in the graph 304. The computing system can update the graph 304 to include the marker 342C for use in autonomously navigating the self through the environment via the passage.
[0111] Move to Figure 3G , the computing system can repeat some of the functions described herein for the fourth point 306D. The computing system can correspond to a self-computing device 141 on the self 140 or a centralized service such as the analysis server 110a. Using the identification of the fourth point 306D, the computing system can find, determine, or otherwise identify at least one point 306D within a set of grid points in the first grid 308 defined on the environment represented by the map 310. The point 306D can correspond to an intermediate point that the self is to cross when navigating through the environment defined by the map 310. The grid 308 can specify or define a set of grid points (or coordinates) at a resolution that is coarser or lower than the original resolution of the grid points defined by the map 310. The map 310 can correspond to the topological layout around the self and can be obtained as part of the map data (e.g., the navigation map 235).
[0112] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the first point predictor unit 310. The portion of the vector space encoding 302 input to the first point predictor unit 310 can correspond to a point 306D defined within a set of grid points in the first grid 308 in the environment. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to input based on the point 306D. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the first point predictor unit 310 to calculate, determine, or otherwise generate at least one coordinate 330D corresponding to the point 306D. Using this point, the computing system can use the first point predictor unit 310 to calculate, produce, or otherwise determine a first index value 330'D for the point 306D.
[0113] The computing system can find, determine, or otherwise identify at least one point 306D within a set of grid points in a second grid 308' defined on an environment represented by a map 312. The point 306D can correspond to a starting point from which the self is to navigate through the environment defined by a map 310. The second grid 308' can specify or define the set of grid points (or coordinates) with a finer or higher resolution than the resolution of the first grid 308. The second grid 308' can correspond to a subset of the first grid 308 around or surrounding the self. The map 312 can correspond to a topological layout around the self and can be obtained as part of map data (e.g., navigation map 235).
[0114] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 and the output of the first point predictor unit 310 (e.g., an embedding derived from a first index value 330'D) to a second point predictor unit 312. The portion of the vector space encoding 302 can correspond to a point 306D defined within a set of grid points in the first grid 308' in the environment. The portion of the vector space encoding 302 input to the second point prediction unit 312 can be the same as or different from the portion of the vector space encoding 302 input to the first point prediction unit 310. In some embodiments, the computing system can identify or select the portion of the vector space encoding 302 to be input based on the point 306D. In the feed, the computing system can process the portion of the vector space encoding 302 according to the weights of the second point predictor unit 312 to calculate, determine, or otherwise generate at least one coordinate 332D corresponding to the point 306D. Using this point, the computing system can use the second point predictor unit 312 to calculate, produce, or otherwise determine a first index value 332'D for the point 306D.
[0115] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to a topology predictor unit 314. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from a first index value 330'D) or the output of the second point predictor unit 312 (e.g., an embedding derived from a second index value 332'D) to the topology predictor unit 314. The portion of the vector space encoding 302 input to the topology predictor unit 314 can be the same as or different from the portion of the vector space encoding 302 input to the first point prediction unit 310 or the second point prediction unit 312.
[0116] By applying the topology predictor unit 314, the computing system can determine, classify, or otherwise categorize the point 306D as at least one topology type 334D. The topology type 334D can semantically specify, identify, or otherwise define the function of the point 306D relative to a path along which the ego navigates through an environment as represented by the map 310. The topology types can include, for example, start type, continuation type, fork type, or terminal type, etc. In the depicted example, the topology type 334D can be a fork type to indicate the continuation of at least one path in the path between the point 306A and the point 306D and a fork relative to the path defined between the point 306A and the point 306B and in turn between the point 306C.
[0117] The computing system can input, feed, or otherwise apply at least a portion of the vector space encoding 302 to the point attribute predictor 316. Additionally, the computing system can apply the output of the first point predictor unit 310 (e.g., an embedding derived from the first index value 330’D), the output of the second point predictor unit 312 (e.g., an embedding derived from the second index value 332’D), or the output (e.g., an embedding derived from the topology type 334’D) into the point attribute predictor 316. The portion of the vector space encoding 302 input into the point attribute predictor 316 can be the same as or different from the portion of the vector space encoding 302 input into the first point prediction unit 310, the second point prediction unit 312, or the topology predictor unit 314.
[0118] By applying the point attribute predictor 316, the computing system can generate or determine a set of point attributes 336D for the point 306D. The set of point attributes 336D can identify or include, for example: at least one fork point identifying an index value of a previous point from which the current point 306D forks; at least one merge point identifying an index value of a previous point from which the current point 306D merges; and a set of spline coefficients 338D, etc. In the depicted example, since the current point 306D is a fork point dependent on the point 306A, the index for the fork can refer to the index value of the point 306A. The set of spline coefficients 338D can include a set of values defining the spline curve 338’D between the points 306A and 306D. The set of spline coefficients 338D can define the path through the environment between the points 306A and 306D.
[0119] Using the generation output, the computing system can use predictors in the ML model to generate respective embeddings to form or output an embedding set 340A for point 306D. Using the first point predictor 310, the computing system can determine or generate at least one corresponding embedding 330”D from the first index value 330’D. Using the second point predictor 312, the computing system can determine or generate at least one corresponding embedding 332”D from the second index value 332’D. Using the second point predictor 312, the computing system can determine or generate at least one corresponding embedding 332”D from the second index value 332’D. Using the topology type predictor 314, the computing system can determine or generate at least one corresponding embedding 334’D from the topology type 334D. Similarly, using the point attribute predictor 316, the computing system can determine or generate an embedding set 336’D.
[0120] When generating respective embeddings (e.g., embeddings 330”D, 332”D, 334’D, and 336’D), the computing system can write, produce, or otherwise generate the embedding set 340A. Based on the embedding set 340A, the computing system can produce, output, or otherwise generate at least one marker 342D for at least one of the passages in the passage through the environment. The computing system can also apply the self-attention unit 318 when determining the marker 342D. In some embodiments, the computing system can generate the marker 342D based on the first index value 330D, the second index value 332D, the topology type 334D, or the point attribute set 336D, or any combination thereof. Using this generation, the computing system can add, insert, or otherwise include the marker 342D in the graph 304. The computing system can update the graph 304 to include the marker 342D to be used to navigate the self through the environment via the passage.
[0121] Continue to Figure 3H, the computing system can repeat the above functions any number of times to reach the end point 306N. The computing system can correspond to an on - self computing device 141 on the self - body 140 or a centralized service such as the analysis server 110a. The computing system can repeat some of the functions described herein for point 306N. With the identification of point 306N, the computing system may have determined an index value based on the grids 308 and 308' described herein and may have determined the corresponding embedding for indexing numerical values. The computing system can classify point 306N as at least one topological type that identifies the end - of - sentence (or end - of - graph) topological point. The end - of - sentence can correspond to the completion of a graph derived from a map 310 obtained from the surrounding environment of the self - body. In the depicted example, the computing system can determine point 306N as the end - of - sentence topological type. The computing system can complete the generation of the graph 304 in a manner similar to that discussed herein by generating a label for point 306N and inserting it into the graph 304. When the self - body navigates through the environment, the graph 304 can be continuously generated and updated.
[0122] Using generation, the computing system on the self - body can use the graph 304 to autonomously navigate the self - body through the environment along one of the paths defined by the set of labels in the graph 304. The computing system can use the set of labels of the graph 304 to generate, calculate, or otherwise determine at least one trajectory. The trajectory can identify, specify, or otherwise define the navigation of the self - body through the environment via the corresponding path in the set of paths defined by the graph 304. Upon determination, the computing system can monitor the position and movement of the self - body in the environment and select one path from the set of paths based on the position and movement (e.g., by proximity). Using selection, the computing system can project or determine a trajectory using the path from the position of the self - body within the environment. The use of these trajectories is detailed below with respect to Figures 4 to 6 Details.
[0123] In some embodiments, the computing system can display, render, or otherwise present on a graphical user interface (GUI) the graph 304 that defines the set of paths through the environment. The set of paths can represent potential lanes, routes, or trajectories along which the self - body can navigate through the environment. The GUI can be presented on a display communicatively coupled to the computing system, such as a touch - screen display within an autonomous vehicle (e.g., the self - body). The set of paths can be presented on the GUI as defined by the graph 304 with respect to the environmental topology around the self - body defined by the map 310 (e.g., overlapping as typically depicted along the right side). Each path can be presented with one or more nodes and the edges between the nodes. Each node can correspond to one of the points (e.g., points 306A to 306N), and each edge can correspond to a line or curve between the corresponding pair of nodes using a set of spline coefficients.
[0124] Now refer to Figure 4 , which depicts a diagram of a scenario for environment 400, where a first self - vehicle uses a graphic representing a path to autonomously navigate through the environment. In environment 400, the self - vehicle 405 can navigate through an intersection of roads. As the self - vehicle 405 traverses environment 400, cameras (and other sensors) on the self - vehicle 405 can acquire a set of videos, such as: a first video 415A from the left - front camera; a second video 415B from the central front - facing camera; and a third video 415C from the right - front camera. The computing device on the self - vehicle 405 can also acquire map data that defines various information about the topology of environment 400. Using the acquired data, the computing device can generate graphics as detailed herein to define a set of paths 420A to 420C along the road. Using sensor data, the computing device can detect the presence of other self - vehicles 425 and human bystanders 435 in environment 400. Based on this and other data, the computing device on the self - vehicle 405 can determine to autonomously navigate along path 420B.
[0125] Now refer to Figure 5 , which depicts a diagram of a scenario 500 for an environment, where a first self - vehicle uses a graphic to detect that the path the first self - vehicle is traversing intersects with the path that a second self - vehicle is about to traverse. In environment 505, the first self - vehicle 505 and the second self - vehicle 510 (or a non - autonomous vehicle) can navigate through an intersection of roads. As the self - vehicle 505 traverses environment 500, cameras (and other sensors) on the self - vehicle 505 can acquire at least one video 515. From the video 515, the computing device on the self - vehicle 505 can detect the presence of other self - vehicles (or vehicles on the road). Additionally, the computing device on the self - vehicle 505 can generate graphics as detailed herein to define a set of paths 520A and 520B along the road through the environment discussed herein. In a similar manner, the computing device on the self - vehicle 510 can generate graphics as detailed herein to define path 525. In some embodiments, the computing device on the self - vehicle 505 can generate multiple graphics for the self - vehicles detected in the environment, including a graphic that defines a set of potential paths for the self - vehicle 510.
[0126] Using generation, a computing device on the ego body 505 can identify a graph generated for its own navigation through the environment. Additionally, the computing device on the ego body 505 can use its own set of markers to identify a graph for another ego body 510 to define autonomous navigation for the ego body 510 through the environment via one or more paths. Once identified, the computing device can use these two graphs to determine whether the path 520A of the ego body 505 intersects with the path 525 of another ego body 510. In making this determination, the computing device can determine a predicted trajectory for each ego body 510. When the paths do not intersect, the computing device on the ego body 505 can continue to autonomously navigate along the path (e.g., path 520A). On the other hand, if the paths intersect, the computing device can detect or determine that there is a potential collision between the ego bodies 505 and 510. In response to this determination, the computing device on the ego body 505 (or the ego body 510) can perform an action on the ego body 505 (or the ego body 515) to avoid the potential collision. For example, the computing device on the ego body 505 can stop the propulsion of the ego body 505 or change the trajectory of the ego body 505 away from the ego body 515.
[0127] Now referring to Figure 6 , depicted is a diagram of a scene 600 for an environment, where a first ego body uses a graph to detect the presence of a stationary second ego body in the environment. In the environment 600, the ego body 605 can navigate along a road. As the ego body 605 traverses the environment 600, cameras (and other sensors) on the ego body 605 can acquire a set of videos, such as: a first video 615A from a left front camera; a second video 615B from a central front camera; and a third video 615C from a right front camera. The computing device on the ego body 605 can also acquire map data that defines various information about the topology of the environment 600. Using the acquired data, the computing device can generate a graph as detailed herein to define at least one path 620.
[0128] Combined, based on sensor data, the computing device on the ego body 605 can determine or detect the presence of other ego bodies 620 and 625 in the environment 600. From the sensor data, the computing device can determine or identify the presence of a stationary ego body 620 in the environment 600. Using the detection, the computing device on the ego body 605 can determine at least one path defined by the graph for the ego body 605 that intersects with the ego body 620 identified as stationary. In response to detecting the intersection, the computing device on the ego body 605 can perform an action on the ego body 605 to avoid the stationary ego body 620. For example, the computing device on the ego body 605 can use the sensor data and the map data to calculate a new trajectory to bypass the ego body 620 on the road. Using the calculation, the computing device can control the ego body 605 to take the bypass trajectory.
[0129] Now referring to Figure 7, which depicts a flowchart of a method 700 for generating a path for autonomous navigation through an environment. Method 700 can be implemented using any of the components described herein, such as the autonomous computing devices 141a to 141c and the AI model(s) 110c. Method 700 can include the steps described herein. However, other embodiments can include additional or alternative steps, or one or more steps can be omitted. Method 700 can be executed by an analysis server (e.g., a computer similar to analysis server 110a) or an autonomous computing device (e.g., autonomous computing devices 141a to 141c). However, one or more steps of process 300 can be executed by any number of computing devices (e.g., processors of autonomous 140 and / or autonomous computing devices 141 or a centralized service such as analysis server 110a) operating in the Figures 1A to 1C distributed computing system described in. For example, one or more computing devices of the autonomous can execute locally Figure 7 some or all of the steps described in.
[0130] Under method 700, at step 705, a computing system (e.g., a processor of autonomous 140 and / or autonomous computing devices 141 or a centralized service such as analysis server 110a, etc.) can obtain, retrieve, or otherwise identify a tensor (e.g., vector space encoding 302). The sensor can include a set of encodings derived from sensor data obtained via one or more cameras and map data defining the environmental topology around the autonomous. The encoding can be a low-dimensional representation of features in the sensor data and map data that are used to define lane segments in the environment.
[0131] At step 710, the computing system can determine, select, or otherwise identify a point in the environment (e.g., points 306A to 306N). The point can be a location, place, or point along which the autonomous will navigate through the environment as defined by the map data. The point can be a starting point, an intermediate (or continuation) point, or an end point of one or more potential lane segments defined in the potential lane segments in the environment. The map data can correspond to the topological layout around the autonomous and can be used to identify points within the environment.
[0132] At step 715, the computing system can calculate, determine, or otherwise generate a first index value for a point in a first grid (e.g., first grid 308) defined on the environment. The first grid can correspond to a set of grid points at a coarser or lower resolution than the original resolution of the map. To generate, the computing system can apply a portion of the tensor corresponding to the point to a first predictor unit. The first predictor unit can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. From the application, the computing system can determine a set of coordinates corresponding to the point within the first grid and use the set of coordinates to determine the first index value for the point.
[0133] At step 720, the computing system can compute, determine, or otherwise generate a second index value for a point in a second grid (e.g., second grid 308’) defined on the environment. The second grid can correspond to a set of grid points at a finer or higher resolution than the first grid. To generate, the computing system can apply a portion of the tensor corresponding to the point and the first index value (or an embedding derived therefrom) to a second predictor unit. The second predictor unit can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. From the application, the computing system can determine a set of coordinates corresponding to the point within the second grid and use the set of coordinates to determine the second index value for the point.
[0134] At step 725, the computing system can classify, assign, or otherwise categorize the topological type of the point. The topological type can semantically define the point as a function of the lane segment and the self to navigate through the environment. The topological type can include, for example, start type, continuation type, fork type, or terminal type, etc. To classify, the computing system can apply a portion of the tensor corresponding to the point and the first index value or the second index value (or derivative embedding) to a topology predictor unit. The second predictor unit can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. From the application, the computing system can classify the point as one of the topological types.
[0135] At step 730, the computing system can determine, generate, or otherwise identify point attributes. The point attributes can be at least one fork point that identifies the index value of a previous point from which the current point forms a fork; at least one merge point that identifies the index value of a previous point from which the current point forms a merge; and a set of spline coefficients from another point, etc. To identify, the computing system can apply a portion of the tensor along with other outputs (e.g., an embedding derived from the index value and the topological type) to a point attribute predictor. The point attribute predictor can include a set of weights arranged according to cross-attention, self-attention, and transformers, etc. From the application, the computing system can identify a set of point attributes.
[0136] At step 735, the computing system can use the first index value, the second index value, the topological type, and the set of point attributes to produce, determine, or otherwise generate a set of embeddings. Each embedding can correspond to a respective output from a machine learning model and can be a reduced-dimensional representation of the output. The embeddings can be generated from the weights of the applied model. At step 740, the computing system can use the set of embeddings to write, generate, or otherwise create tokens. To create tokens, the computing system can combine the set of embeddings. Each token can correspond to a point that defines a portion of one or more lane segments through the environment.
[0137] At step 745, the computing system may add, include, or otherwise insert markers into a map. The map may semantically define a set of potential lane segments along which an ego vehicle may navigate through an environment. At step 750, the computing system may determine whether there are additional points in the environment. This determination may be based on a topological type. If the topological type indicates the end of a sentence, the computing system may determine that there are no points in the environment to evaluate. On the other hand, if the topological type indicates another type other than the end of a sentence, the computing system may repeat method 700 starting from step 710. At step 755, the computing system may store and maintain the map. The map may be used by the ego vehicle to perform autonomous navigation through the environment.
[0138] Additionally or alternatively, the analytics server may transmit the generated map to a downstream software application or another server. The prediction results may be further analyzed and used in various models and / or algorithms to perform various actions. For example, a software model or a processor associated with the ego vehicle's autonomous navigation system may receive occupancy data predicted by the trained AI model, and navigation decisions may be made based on this data.
[0139] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.
[0140] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. Code segments or machine-executable instructions may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment may be coupled to another code segment or a hardware circuit. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, tag passing, network transmission, etc.
[0141] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or this disclosure. Accordingly, the operation and behavior of the systems and methods are described without reference to the specific software code, it being understood that the software and control hardware can be designed to implement the systems and methods based on the description herein.
[0142] When implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be implemented in a processor-executable software module that may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer or a processor. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), Blu-ray disc, and floppy disk where a “disk” typically reproduces data magnetically and a “disc” optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as code and / or instructions on a non-transitory processor-readable medium and / or a computer-readable medium that may be incorporated into a computer program product.
[0143] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein and their variations. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Accordingly, this disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0144] Although various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for illustrative purposes and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
1. A method of generating a path for autonomous navigation through an environment, comprising: identifying, by one or more processors, a tensor comprising a plurality of encodings derived from sensor data from the self and map data defining a topology of the environment surrounding the self; determining, by the one or more processors, a first index value by applying at least a first portion of the plurality of encodings to a machine learning (ML) model, the first index value defining a point within a first plurality of points of a first grid defined on the environment; determining, by the one or more processors, a second index value by applying at least a second portion of the plurality of encodings and the first index value to the ML model, the second index value defining the point within a second plurality of points of a second grid within a subset of the first plurality of points of the first grid; generating, by the one or more processors, a marker for at least one of a plurality of paths through the environment based on the first index value and the second index value for the point; and storing, by the one or more processors, a graph comprising the marker to be used for autonomous navigation of the self through the environment via one or more of the plurality of paths.
2. The method of claim 1, further comprising classifying, by the one or more processors, the second point as a continuation topology type dependent on the first point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; determining, in response to the classification of the second point as the continuation topology type, by the one or more processors, a plurality of spline coefficients that define a path of the plurality of paths through the environment between the first point and the second point; generating, by the one or more processors, a second marker based on the third index value for the second point and the continuation topology type; and updating, by the one or more processors, the graph to include the second marker and the plurality of spline coefficients to be used for autonomous navigation of the self through the environment.
3. The method of claim 1, further comprising: classifying, by the one or more processors, the second point as a bifurcation topology type from the point with respect to a third point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; determining, in response to classifying the second point as the bifurcation topology type, by the one or more processors, a fourth index value for the point referencing the marker; generating, by the one or more processors, a second marker for a first path different from a second path associated with the third point based on the third index value, the fourth index value, and the bifurcation topology type; and updating, by the one or more processors, the graph to include the second marker to be used for autonomous navigation of the self through the environment.
4. The method of claim 1, further comprising: By the one or more processors, classifying the second point as a terminal topology type dependent on the first point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; In response to the classification of the second point as the terminal topology type, by the one or more processors, determining a passage through the environment defined by the first point and the second point; By the one or more processors, generating a second marker based on the third index point for the second point and the terminal topology type; And By the one or more processors, updating the graph to include the second marker for use in autonomously navigating the self through the environment.
5. The method according to claim 1, further comprising By the one or more processors, identifying a second graph including a plurality of markers for use in autonomously navigating a second self through the environment via one or more of a second plurality of passages; By the one or more processors, using the graph and the second graph to determine that at least one first passage of the plurality of passages for the self intersects at least one second passage of the second plurality of passages for the second self; And In response to determining that at least one first passage intersects the second passage, by the one or more processors, performing an action on at least one of the self or the second self.
6. The method according to claim 1, further comprising: By the one or more processors, using the sensor data from the self to identify the presence of a second self that is stationary in the environment; And By the one or more processors, using the graph to determine that at least one first passage of the plurality of passages for the self intersects the stationary second self.
7. The method according to claim 1, further comprising: By the one or more processors, classifying the point as a topology type indicating the start of at least one of the plurality of passages by applying at least a third portion of the plurality of encodings and the second index value to the ML model, and wherein generating the marker further comprises generating the marker for at least one of the plurality of passages through the environment based on the topology type.
8. The method according to claim 1, further comprising: By the one or more processors, using the plurality of markers of the graph to determine a trajectory that defines the navigation of the self via passages of the plurality of passages through the environment.
9. The method according to claim 1, further comprising: By the one or more processors, presenting the graph via a graphical user interface (GUI), the graph defining the plurality of passages with respect to the topology of the environment around the self.
10. The method according to claim 1, wherein generating the marker further comprises generating the marker using (i) a first embedding generated from the first index value, (ii) a second embedding generated from the second index value, and (iii) one or more embeddings associated with the point.
11. A system for generating a path for autonomous navigation through an environment, comprising: one or more processors coupled to a memory and configured to: identify a tensor that includes a plurality of encodings derived from sensor data from the self and map data defining a topology of the environment surrounding the self; determine a first index value by applying at least a first portion of the plurality of encodings to a machine learning (ML) model, the first index value defining a point within a first plurality of points of a first grid defined on the environment; determine a second index value by applying at least a second portion of the plurality of encodings and the first index value to the ML model, the second index value defining the point within a second plurality of points of a second grid within a subset of the first plurality of points of the first grid; generate a marker for at least one of a plurality of paths through the environment based on the first index value and the second index value for the point; and store a graph to include the marker for use in autonomously navigating the self through the environment via one or more of the plurality of paths.
12. The system of claim 11, wherein the one or more processors are further configured to classify a second point as a continuation topology type dependent on the first point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; determine a plurality of spline coefficients in response to the classification of the second point as the continuation topology type, the plurality of spline coefficients defining a path of the plurality of paths through the environment between the first point and the second point; generate a second marker based on the third index value for the second point and the continuation topology type; and update the graph to include the second marker and the plurality of spline coefficients for use in autonomously navigating the self through the environment.
13. The system of claim 11, wherein the one or more processors are further configured to: classify a second point as a bifurcation topology type from the point with respect to a third point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; determine a fourth index value for the point referencing the marker in response to classifying the second point as the bifurcation topology type; generate a second marker for a first path different from a second path associated with the third point based on the third index value, the fourth index value, and the bifurcation topology type; and update the graph to include the second marker for use in autonomously navigating the self through the environment.
14. The system of claim 11, wherein the one or more processors are further configured to: classify a second point as a terminal topology type dependent on the first point by applying at least a third portion of the plurality of encodings and a third index value for the second point to the ML model; determine a path through the environment defined by the first point and the second point in response to the classification of the second point as the terminal topology type; Generate a second marker based on the third index value for the second point and the terminal topology type; and Update the graph to include the second marker for use in autonomously navigating the self through the environment.
15. The system of claim 11, wherein the one or more processors are further configured to: Identify a second graph including a plurality of markers for use in autonomously navigating a second self through the environment via one or more of a second plurality of paths; Use the graph and the second graph to determine that at least one first path of the plurality of paths for the self intersects at least one second path of the second plurality of paths for the second self; and In response to determining that at least one first path intersects the second path, perform an action on at least one of the self or the second self.
16. The system of claim 11, wherein the one or more processors are further configured to: Use the sensor data from the self to identify the presence of a second stationary self in the environment; and Use the graph to determine that at least one first path of the plurality of paths for the self intersects the stationary second self.
17. The system of claim 11, wherein the one or more processors are further configured to: Classify the point as a topology type indicating the start of at least one of the plurality of paths by applying at least a third portion of the plurality of encodings and the second index value to the ML model, and Generate the marker for at least one of the plurality of paths through the environment based on the topology type.
18. The system of claim 11, wherein the one or more processors are further configured to: Use the plurality of markers of the graph to determine a trajectory that defines the navigation of the self via a path among the plurality of paths through the environment.
19. The system of claim 11, wherein the one or more processors are further configured to: Present the graph via a graphical user interface (GUI), the graph defining the plurality of paths with respect to the topology of the environment around the self.
20. The system of claim 11, wherein the one or more processors are further configured to: Generate the marker using (i) a first embedding generated from the first index value, (ii) a second embedding generated from the second index value, and (iii) one or more embeddings associated with the point.