AI inference compiler and runtime tool chain

By decomposing the neural network architecture into multiple sub-neural networks and assigning them to multiple processing units, and executing multiple compilers on each sub-neural network, the problem of inefficient compilation and deployment of autonomous vehicle neural network architectures in the prior art is solved, and more efficient data processing and computing efficiency is achieved.

CN120112913APending Publication Date: 2025-06-06TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072102.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently compile and deploy neural network architectures for autonomous vehicles, especially when processing large amounts of data and complex environments, with increased resource requirements resulting in insufficient computing efficiency of hardware and software components.

Method used

By designing a method, the method includes obtaining functional software programming of multiple sub-neural networks, assigning these sub-neural networks to multiple processing units of the autonomous vehicle, and generating execution instructions for the multiple processing units by executing multiple compilers on multiple functions of each sub-neural network.

Benefits of technology

It improves the performance of autonomous vehicles in processing sensor data, realizes more efficient compilation and deployment of neural network architectures, and meets the requirements for computing efficiency in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112913A_ABST
    Figure CN120112913A_ABST
Patent Text Reader

Abstract

Embodiments include systems and methods for processing sensor data and generating operational instructions for hardware of an autologous (e.g., an autonomous vehicle, robot). The self-body includes any number of machine learning architectures, typically neural network architectures, for processing sensor data and identifying the environment around the self-body and making decisions on the behavior of the self-body. The autologous neural network architecture ingests sensor data and uses the sensor data to perform any number of operations related to a particular domain or task, such as object recognition or path planning. A graph zoner is trained to assign functions in software of the neural network and sensor data to certain hardware processing units. Several compilers are used to generate instructions based on the assigned types of the processing units.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 377,954, filed on September 30, 2022, which is incorporated herein by reference in its entirety for all purposes. Technical Field

[0003] The present application relates generally to implementing a neural network architecture for autonomous vehicles or other autonomous electronic equipment, and more particularly to systems and methods for remotely and efficiently compiling and deploying such a neural network architecture. Background Art

[0004] Autonomous navigation technology for autonomous vehicles and robots (sometimes called "agents") has become ubiquitous due to rapid advances in computer technology. These advances allow for safer and more reliable autonomous navigation of the agents. Agents are often required to navigate complex and dynamic environments and terrains that can include vehicles, traffic, pedestrians, cyclists, and a variety of other static or dynamic obstacles. Understanding the agent's surroundings is necessary for intelligent and competent decision making to avoid conflicts. This involves developing and deploying complex neural network architectures on the agents.

[0005] The increase in data volume and feature sophistication naturally raises resource demand issues, requiring solutions for improving the computational efficiency of both the hardware and software components of the self. This requires sophisticated mechanisms for compiling and deploying neural network architectures on the self. In some cases, the challenge is enhanced by, for example, remote deployment from a software development source to a remote self and / or deploying software updates of the neural network architecture to predefined and fixed execution hardware of the self. Summary of the invention

[0006] The embodiments described herein include systems and methods that address various shortcomings in the art, and may also provide various additional or alternative benefits. Embodiments include hardware and software configurations that improve the performance of processing sensor data through software components and computing hardware of an autonomous vehicle, such as an autonomous vehicle, a robot. The autonomous vehicle includes any number of machine learning architectures, typically neural network architectures, for processing sensor data and identifying the environment around the autonomous vehicle, and making decisions about the behavior of the autonomous vehicle. The autonomous vehicle's neural network architecture ingests sensor data and uses the sensor data to perform any number of operations related to a specific domain or task, such as object recognition or path planning. Any number of compilers transform the software functions and sensor data of the neural network architecture into machine code execution instructions for execution by hardware components.

[0007] Embodiments may include a method comprising: obtaining, by a computer, software programming comprising a plurality of functions of a plurality of sub-neural networks of a neural network architecture; assigning, by the computer, one or more sub-neural networks to a plurality of processing units of the self, wherein for each sub-neural network, the computer assigns a processing unit to compute the plurality of functions of the sub-neural network; and generating, by the computer, a plurality of execution instructions for the plurality of processing units by executing a plurality of compilers on the plurality of functions of each sub-network. For each execution instruction, the computer uses a compiler from the plurality of compilers according to the processing unit of the self assigned to the plurality of functions of the sub-neural network.

[0008] The method may include: generating, by a computer, a computer file including a plurality of execution instructions for a plurality of processing units to execute a plurality of sub-neural networks.

[0009] The method may include: sending, by the computer, the computer file to the self.

[0010] At least one execution instruction may cause the circuit of the body including the plurality of chips to operate in an extended computing mode for parallel execution of the plurality of execution instructions.

[0011] At least one execution instruction may instruct a circuit of the body to operate in a redundant mode for mainly executing the plurality of execution instructions by a master chip among the plurality of chips, the circuit of the body comprising a plurality of chips including a plurality of processing units.

[0012] The method may include: applying, by a computer, a schedule optimizer engine to the execution instructions to generate an execution schedule for the execution instructions, the schedule optimizer engine comprising a neural network layer trained to generate an execution schedule for minimizing latency.

[0013] The one or more processing units may include at least one of a GPU, a CPU, or an accelerator device.

[0014] One or more processing units may be heterogeneous, including at least two types of processing units.

[0015] Based on one or more diagrams representing a circuit architecture of an entity having multiple processing units, a computer can assign processing units among the multiple processing units of the entity to apply a sub-neural network.

[0016] The method may include: training, by a computer, one or more sub-neural networks for quantization-aware training based on one or more processing units by applying each sub-neural network to a training data set including data having expected quantization characteristics.

[0017] Embodiments may include a system comprising: a computer including a processor configured to: obtain software programming including a plurality of functions of a plurality of sub-neural networks of a neural network architecture; assign one or more sub-neural networks to a plurality of processing units of the self, wherein for each sub-neural network, the computer assigns a processing unit to calculate the plurality of functions of the sub-neural network; and generate a plurality of execution instructions for the plurality of processing units of the self by executing a plurality of compilers on the plurality of functions of each sub-neural network. For each execution instruction, the computer uses a compiler from the plurality of compilers according to the processing unit of the self assigned to the plurality of functions of the sub-neural network.

[0018] The computer may also be configured to generate a computer file comprising a plurality of execution instructions for a plurality of processing units to execute a plurality of sub-neural networks.

[0019] The computer may also be configured to send the computer file to itself.

[0020] At least one execution instruction may cause the plurality of chips of the entity to operate in an extended computing mode for parallel execution of the plurality of execution instructions.

[0021] At least one execution instruction may cause the plurality of chips of the self to operate in a redundant mode for primarily executing the plurality of execution instructions by a master chip among the plurality of chips.

[0022] The computer may also be configured to apply a schedule optimizer engine to the execution instructions to generate an execution schedule for executing the instructions. The schedule optimizer engine includes a neural network layer trained to generate instructions for minimizing latency.

[0023] The plurality of processing units may include at least one of a GPU, a CPU, or an accelerator device.

[0024] The plurality of processing units may be heterogeneous, including at least two types of processing units.

[0025] Based on one or more diagrams representing a circuit architecture of an entity having multiple processing units, a computer can assign processing units from the multiple processing units of the entity to sub-neural networks.

[0026] The computer may also be configured to train one or more sub-neural networks for quantization-aware training based on the one or more processing units by applying each sub-neural network to a training data set stored in a database, the training data set including data having expected quantization characteristics.

[0027] An embodiment may include a body comprising: a circuit board comprising multiple system-on-chip (SOC) devices and multiple microcontrollers corresponding to the SOC devices; and a processor of the SOC device, the processor of the SOC device being configured to: send an initial timing message to a first microcontroller at an initial time based on a core clock of the processor; receive a response message from the first microcontroller indicating a response time at a first controller clock of the first microcontroller; determine a completion time based on the core clock of the processor in response to receiving the response message from the first microcontroller; calculate an error rate of the core clock representing the difference between the core clock of the processor and the controller clock of the first microcontroller based on the initial clock time, the completion time, and the response time; and adjust the frequency of the core clock based on the error rate to reduce the error rate between the core clock and the first controller clock.

[0028] The processor of the SOC device is configured to send an initial timing message in response to receiving a boot instruction from the first microcontroller.

[0029] A first microcontroller coupled to a processor may be configured to: calculate a second error rate representing a difference between a first controller clock of the first microcontroller and a second controller clock of the second microcontroller based on an initial clock time between the first microcontroller and the second microcontroller, a completion time between the first microcontroller and the second microcontroller, and a response time between the first microcontroller and the second microcontroller; and adjust a second frequency of the second controller clock based on the second error rate to reduce a second error rate between the first controller clock and the second controller clock.

[0030] The first microcontroller is configured to determine whether a difference between a first controller clock and a second controller clock of the first microcontroller satisfies a threshold difference.

[0031] The first microcontroller is configured to send an initial timing message in response to executing a boot function of the first microcontroller.

[0032] The second microcontroller is configured to perform a reboot function, and wherein a second controller clock of the second microcontroller is updated to match a first controller clock of the first microcontroller indicated in a timing message received at the second microcontroller from the first microcontroller for the reboot function.

[0033] An embodiment may include a method comprising: sending, by a processor of a system on chip (SOC), an initial timing message to a first microcontroller at an initial time based on a core clock of the processor; receiving, by the processor, a response message from the first microcontroller, the response message indicating a response time at a first controller clock of the first microcontroller; determining, by the processor in response to receiving the response message, a completion time based on the core clock of the processor; calculating, by the processor, an error rate of the core clock representing the difference between the core clock of the processor and the first controller clock of the first microcontroller based on the initial clock time, the completion time, and the response time; and adjusting, by the processor, the frequency of the core clock based on the error rate to reduce the error rate between the core clock and the first controller clock.

[0034] The method may include receiving, by a processor, a boot function from the first microcontroller, wherein the processor of the SOC device is configured to send an initial timing message in response to receiving the boot function.

[0035] The processor and the first microcontroller exchange one or more timing messages at boot time according to a boot loader function of the first microcontroller.

[0036] The method may include: calculating, by the first microcontroller, a second error rate representing the difference between a first controller clock of the first microcontroller and a second controller clock of the second microcontroller based on an initial clock time between the first microcontroller and the second microcontroller, a completion time between the first microcontroller and the second microcontroller, and a response time between the first microcontroller and the second microcontroller; and, at the second microcontroller, adjusting a second frequency of the second controller clock based on the second error rate to reduce a second error rate between the first controller clock and the second controller clock.

[0037] An embodiment may include a body comprising: a circuit board comprising multiple system-on-chip (SOC) devices and multiple microcontrollers corresponding to the SOC devices; and a first microcontroller configured to: send an initial timing message to a second microcontroller at an initial time based on a first controller clock of the first microcontroller; receive a response message from the second microcontroller, the response message indicating a response time at a second controller clock of the second microcontroller; determine a completion time based on the first controller clock of the first microcontroller in response to receiving the response message from the second microcontroller; calculate an error rate representing the difference between the first controller clock of the first microcontroller and the second controller clock of the second microcontroller based on the initial clock time, the completion time and the response time; and adjust the frequency of the second controller clock based on the error rate to reduce the error rate between the first controller clock and the second controller clock.

[0038] The first microcontroller may be configured to determine whether a difference between a first controller clock and a second controller clock of the first microcontroller satisfies a threshold difference.

[0039] The first microcontroller may also be configured to send a boot signal to a processor core of the SOC coupled to the first microcontroller.

[0040] The processor of the SOC can be configured to: calculate a second error rate representing the difference between a core clock of a core of the processor and a first controller clock of the first microcontroller based on an initial clock time between the SOC and the first microcontroller, a completion time between the SOC and the first microcontroller, and a response time between the SOC and the first microcontroller; and adjust the core frequency of the core clock based on the error rate to reduce the error rate between the core clock and the first controller clock.

[0041] The second microcontroller may be configured to send a second guidance signal to a second core of a second processor of a second SOC coupled to the second microcontroller.

[0042] The second microcontroller may perform a reboot function.The second controller clock is updated to match the first controller clock indicated in a timing message received at the second microcontroller from the first microcontroller.

[0043] An embodiment may include a method comprising: sending, by a first microcontroller coupled to a first SOC, an initial timing message to a second microcontroller coupled to a second SOC at an initial time according to a first controller clock of the first microcontroller; receiving, by the first microcontroller, a response message from the second microcontroller, the response message indicating a response time at a second controller clock of the second microcontroller; determining, by the first microcontroller, a completion time according to the first controller clock of the first microcontroller in response to receiving the response message from the second microcontroller; calculating, by the first microcontroller, an error rate representing a difference between the first controller clock of the first microcontroller and the second controller clock of the second microcontroller based on the initial clock time, the completion time, and the response time; and adjusting, by the first microcontroller, the frequency of the second controller clock based on the error rate to reduce the error rate between the first controller clock and the second controller clock.

[0044] The first microcontroller may be configured to determine whether a difference between a first controller clock and a second controller clock of the first microcontroller satisfies a threshold difference.

[0045] The method may include sending, by the first microcontroller, a boot signal to a processor core of a SOC coupled to the first microcontroller.

[0046] The method may include: calculating, by a processor of the first SOC, a second error rate representing the difference between a core clock of a processor core and a first controller clock of a first microcontroller based on an initial clock time between the SOC and the first microcontroller, a completion time between the SOC and the first microcontroller, and a response time between the SOC and the first microcontroller; and adjusting, by the processor of the first SOC, a core frequency of the core clock based on the error rate to reduce the error rate between the core clock and the first controller clock.

[0047] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Non-limiting embodiments of the present disclosure are described by way of example in relation to the accompanying drawings, which are schematic and are not intended to be drawn to scale. Unless indicated as representing background art, the drawings represent various aspects of the present disclosure.

[0049] Figure 1A Components of an autonomic AI-enabled visual data analysis system according to an embodiment are illustrated.

[0050] Figure 1B Various sensors associated with a vehicle (or other type of entity) are illustrated in accordance with an embodiment.

[0051] Figure 1C Components of a body according to an embodiment are illustrated.

[0052] Figure 1D Certain hardware and software components of an autonomous vehicle for performing full or partial autonomous driving (SD) operations according to an embodiment are illustrated.

[0053] Figure 1E Certain hardware and software components of a clock synchronization module according to an embodiment are illustrated.

[0054] Figures 2A to 2B The data flow between hardware and software computing components of a native computing system is illustrated according to an embodiment.

[0055] Figure 2C Execution instructions are illustrated being arranged into an execution schedule for execution by IC hardware components of a system according to an embodiment. DETAILED DESCRIPTION

[0056] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, and specific language will be used herein to describe them. However, it is to be understood that it is not intended to limit the scope of the claims or the disclosure. Changes and further modifications to the inventive features described herein that may be conceived by a technician in the relevant field and in possession of the disclosure, as well as additional applications of the subject principles described herein, will be considered to be within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the disclosure. The illustrative embodiments described in the specific embodiments are not meant to limit the subject matter presented.

[0057] The embodiments described herein include systems and methods that address various shortcomings in the art, and may also provide various additional or alternative benefits. Embodiments include hardware and software configurations that improve the performance of processing sensor data through software components and computing hardware of an entity (e.g., an autonomous vehicle, a robot). The entity includes any number of machine learning architectures, typically neural network architectures, for processing sensor data and recognizing the environment around the entity, and making decisions about the behavior of the entity.

[0058] At the software level, a local or remote computer processing device of the host can execute various software routines of a neural network architecture (or other machine learning architecture). The software defines the layers and functions of the neural network architecture, or defines a hierarchical parent neural network architecture and one or more hierarchical child neural network architectures (sometimes referred to as child nets or subnets). The host neural network architecture ingests sensor data and uses the sensor data to perform any number of operations related to a specific domain or task, such as object recognition or path planning.

[0059] At the hardware level, the self includes various types of computing hardware resources, including various integrated circuit (IC) components and related firmware or controllers, etc. The computing hardware components implement and execute the software functions of the neural network architecture using sensor data. Any number of compilers transform the software functions of the neural network architecture and sensor data into machine code execution instructions for execution by the hardware components.

[0060] Embodiments include various software routines for improving the performance and efficiency of processing sensor data within computing hardware by partitioning the sensor data into multiple data partitions or portions and partitioning the neural network architecture structure in subnets. The software routines may include a machine learning architecture that includes a neural network layer that defines a graph partitioner. The graph partitioner is configured and trained to assign sensor data portions to certain functions of the subnet, and then assign one or more hardware processing units (e.g., GPUs, CPUs, dedicated hardware AI accelerator devices) to perform functions using the sensor data portions. The neural network of the graph partitioner is trained to identify and assign functions to be applied to the sensor data based on, for example, the type of sensor data or the type of features of the sensor. In some embodiments, the graph partitioner may include a hard-coded or pre-configured mapping between the type of sensor data or features and the software functions of the neural network architecture.

[0061] A graph partitioner can be configured and trained to identify functional sensor data portions and assign them to specific hardware processing units. For example, a graph partitioner can be trained to select a hardware processing unit for performing a function based on a desired performance behavior or result, such as optimizing the efficiency of the computing hardware, maximizing a performance metric of the computing hardware, or minimizing a performance metric of the computing hardware. For example, a graph partitioner can be trained to assign certain sophisticated functions to specially designed AI accelerator devices.

[0062] Embodiments may include a heterogeneous collection of hardware processing units. A graph partitioner may be trained to identify and assign a compiler that is capable of generating execution instructions having machine code compatible with the assigned processing units. A graph partitioner is trained, for example, to assign processing units to perform functions to optimize functionality of computing hardware, maximize certain performance behaviors of computing hardware, or minimize certain performance behaviors. By training a graph partitioner to dynamically assign heterogeneous processing units to perform specified functions of a neural network architecture, the graph partitioner optimizes the execution of neural network functions within the hardware. This results in improved accuracy and performance of the neural network architecture when analyzing sensor data.

[0063] Embodiments include a scheduling optimizer that is preconfigured or trained to identify the relationship between different partitions of sensor data and a neural network architecture. The scheduling optimizer is used to combine multiple compiled code segments (e.g., execution instructions) into one or more executable files. In some cases, the link generates an ordering or arrangement for executing instructions by hardware components. In these cases, by applying the trained neural network architecture of the link engine to determine, for example, dependencies between data, functions to be performed by processing units, and processing units assigned to perform these functions, the scheduling optimizer can generate an execution schedule that minimizes latency. The link engine can then identify an execution schedule to optimize performance relative to dependent constraints. In this way, the scheduling optimizer ensures that execution instructions are organized in a manner that reduces delays and latency, thereby allowing improved real-time processing of sensor data.

[0064] The self's downstream hardware and software can ingest the outputs generated by the SD circuits executing the neural network architecture (such as trajectory, velocity, and other navigation determinations of the path planning network) to operate or manipulate the self within the environment.

[0065] The hardware of the self includes an SD circuit (e.g., an integrated circuit (IC) board) having two (or more) system-on-chip (SOC) chips or similar types of IC devices. The SOC chip can perform the functions of certain neural network architectures by executing execution instructions compiled from the source code of the specific neural network architecture into an execution library. The server of the development system trains the neural network architecture on historical and / or current sensor data and the output of other neural network architectures. The server then applies the compiler tool chain described in this article to the source code and data of the trained neural network architecture to generate an execution library. The neural network architecture of the training graph partitioner is assigned to portions of the execution library to corresponding processing units of the SOC chip. In some cases, the first neural network architecture is programmed to require or expect data input from the second neural network architecture. When the first SOC chip loads and executes execution instructions for the first neural network architecture and the second SOC chip loads and executes execution instructions for the second neural network architecture, then the second SOC chip will require or expect data input from the first SOC chip. The SD circuit includes a bus that allows the SOC chips to transmit data signals.

[0066] Embodiments include hardware and / or software components within the circuit assembly of the self that allow the SOC chip of the SD circuit board to operate as if the SOC chip were running on a single synchronous clock. In some embodiments, the SD circuit includes two (or more) SOC chips, each of which is coupled to a SOC microcontroller, which can be any microcontroller device or a similar processing circuit unit that maintains a chip clock for the corresponding SOC chip. The SOC chips are guided by the self-computer processor at the same time so that the SOC clocks of the SOCs are as close to each other as possible. Each SOC chip includes an operating system (OS) kernel that manages certain operations of the corresponding SOC chip. At preconfigured intervals, each SOC chip exchanges time messages with the microcontroller of the corresponding SOC to confirm that the kernel clock of the SOC is within a threshold distance from the controller clock of the SOC, and corrects or adjusts the microcontroller of the SOC as needed. In this way, each SOC chip maintains clock synchronization between the OS kernel of the SOC and the microcontroller of the SOC. Furthermore, at preconfigured intervals, the SOC microcontrollers exchange time messages to confirm that a first controller clock of a first SOC microcontroller is within a threshold distance from a second SOC controller clock of a second SOC microcontroller. If the threshold distance is exceeded, the second microcontroller may be corrected or adjusted to reduce the distance between the controller clocks. In this manner, the microcontrollers maintain clock synchronization between the SOC microcontrollers and, by extension, between the SOC chips.

[0067] In some embodiments, the SD circuit includes a controller or other processing unit that maintains clock synchronization between SOC chips based on interpreting timestamps of sensor inputs (e.g., timestamps of camera inputs) and translating timestamps between SOC chips based on differences between current chip clocks of each SOC chip.

[0068] Figure 1A are non-limiting examples of components of the system 100 in which the methods and systems discussed herein may be implemented. For example, the analysis server may train an AI model and use the trained AI model to generate an occupancy dataset and / or map for one or more self-bodies. Figure 1A Components of an AI-enabled visual data analysis system 100 are illustrated. The system 100 may include an analysis server 110a, a system database 110b, an administrator computing device 120, autons 140a-140b (collectively referred to as (multiple) autons 140), auton computing devices 141a-141c (collectively referred to as auton computing devices 141), and a server 160. The system 100 is not limited to the components described herein, and may include additional or other components that are not shown for the sake of brevity, which are to be considered within the scope of the embodiments described herein.

[0069] The above components may be connected via a network 130. Examples of network 130 may include, but are not limited to, private or public LANs, WLANs, MANs, WANs, and the Internet. Network 130 may include wired and / or wireless communications according to one or more standards and / or via one or more transmission media.

[0070] Communications on network 130 may be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 may include wireless communications according to the Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 may also include communications over a cellular network, including, for example, a GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) network.

[0071] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models, such as (multiple) AI models 110c. Figure 1A As depicted and described herein, the analysis server 110a can use the methods discussed herein to train (multiple) AI models 110c using data retrieved from the self 140 (e.g., by using data streams 172 and 176). When the (multiple) AI model 110c has been trained, each of the self 140 can access and execute the (multiple) trained AI model 110c. For example, a vehicle 140a having an self computing device 141a can send its camera feed to the (multiple) trained AI model 110c and can determine the occupancy state of its surrounding environment (e.g., data stream 174). Moreover, the data ingested and / or predicted by the (multiple) AI model 110c relative to the self 140 (at inference) can also be used to improve the (multiple) AI model 110c. Therefore, the system 100 depicts a continuous cycle that can periodically improve the accuracy of the (multiple) AI model 110c. Furthermore, system 100 depicts a loop in which data received by self 140 may be used in a training phase in addition to an inference phase.

[0072] The analysis server 110a can be configured to collect, process and analyze navigation data (e.g., images captured while navigating) and various sensor data collected from the self 140. The collected data can then be processed and prepared into a training data set. The training data set can then be used to train one or more AI models, such as the AI ​​model 110c. The analysis server 110a can also be configured to collect visual data from the self 140. Using the AI ​​model 110c (trained using the methods and systems discussed herein), the analysis server 110a can generate a data set and / or occupancy map for the self 140. The analysis server 110a can display an occupancy map on the self 140 and / or send the occupancy map / data set to the self computing device 141, the administrator computing device 120 and / or the server 160.

[0073] exist Figure 1A , the AI ​​model 110 c is illustrated as a component of the system database 110 b , but the AI ​​model 110 c may be stored in a different or separate component, such as a cloud storage device or any other data repository accessible to the analysis server 110 a .

[0074] The analysis server 110a may also be configured to display an electronic platform that illustrates various training attributes used to train the AI ​​model 110c. The electronic platform may be displayed on the administrator computing device 120 so that an analyst can monitor the training of the AI ​​model 110c. An example of an electronic platform generated and hosted by the analysis server 110a may be a web-based application or website configured to display training data sets collected from the self 140 and / or training status / metrics of the AI ​​model 110c.

[0075] The analysis server 110a may be any computing device including a processor and non-transitory machine-readable storage capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices may include a workstation computer, a laptop computer, a server computer, etc. Although the system 100 includes a single analysis server 110a, the system 100 may include any number of computing devices operating in a distributed computing environment, such as a cloud environment.

[0076] Self 140 may represent various electronic data sources that send data associated with its previous or current navigation session to analysis server 110a. Self 140 may be any device configured for navigation, such as vehicle 140a and / or truck 140c. Self 140 is not limited to being a vehicle, and may also include robotic devices. For example, self 140 may include robot 140b, which may represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. Robot 140b may be equipped with software that enables balance, navigation, perception, or interaction with the physical world. Robot 140b may also include various cameras configured to send visual data to analysis server 110a.

[0077] Although referred to herein as an "autonomous device," the autonomous device 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the autonomous device 140 may be controlled by a human operator or a remote processor. The autonomous device 140 may include various sensors, such as Figure 1B The sensors depicted in . The sensors may be configured to collect data as the self 140 navigates various terrains (e.g., roads). The analysis server 110a may collect data provided by the self 140. For example, the analysis server 110a may obtain navigation sessions and / or road / terrain data (e.g., images of the self 140 navigating on the road) from various sensors, so that the collected data is ultimately used by the AI ​​model 110c for training purposes.

[0078] As used herein, a navigation session corresponds to a journey of a route taken by the vehicle 140, regardless of whether the journey is autonomous or controlled by a human. In some embodiments, a navigation session may be used for data collection and model training purposes. However, in some other embodiments, the vehicle 140 may refer to a vehicle purchased by a consumer, and the purpose of the journey may be classified as daily use. A navigation session may begin when the vehicle 140 moves beyond a threshold distance (e.g., 0.1 miles, 100 feet) or exceeds a threshold speed (e.g., greater than 0mph, greater than 1mph, greater than 5mph) from a non-moving position. The navigation session may end when the vehicle 140 returns to a non-moving position and / or is turned off (e.g., when the driver leaves the vehicle).

[0079] The self 140 may represent a collection of selfs monitored by the analysis server 110a to train (multiple) AI models 110c. For example, a driver for a vehicle 140a may authorize the analysis server 110a to monitor data associated with its corresponding vehicle. As a result, the analysis server 110a may collect sensor / camera data using various methods discussed herein and generate a training data set to train (multiple) AI models 110c accordingly. The analysis server 110a may then apply (multiple) trained AI models 110c to analyze the data associated with the self 140 and predict an occupancy map for the self 140. Moreover, additional / ongoing data associated with the self 140 may also be processed and added to the training data set so that the analysis server 110a recalibrates (multiple) AI models 110c accordingly. Therefore, the system 100 depicts a loop in which navigation data received from the self 140 may be used to train (multiple) AI models 110c. The self 140 may include a processor that executes (multiple) trained AI models 110c for navigation purposes. While navigating, the self 140 may collect additional data about its navigation session, and this additional data may be used to calibrate the AI ​​model(s) 110c. That is, the self 140 represents a self that may be used to train, execute / use, and recalibrate the AI ​​model(s) 110c. In a non-limiting example, the self 140 represents vehicles purchased by customers that may autonomously navigate using the AI ​​model(s) 110c while improving the AI ​​model(s) 110c.

[0080] The self 140 may be equipped with various technologies that allow the self to collect data from its surroundings and (potentially) navigate autonomously. For example, the self 140 may be equipped with an inference chip to run autonomous driving software.

[0081] Various sensors for each of the ego bodies 140 may monitor collected data associated with different navigation sessions and send the data to the analysis server 110a. Figures 1B to 1C FIG. 1 is a block diagram of a sensor integrated within the body 140 according to an embodiment. Figures 1B to 1C The number and location of each sensor discussed may depend on Figure 1A For example, robot 140b may include different sensors than vehicle 140a or truck 140c. For example, robot 140b may not include air bag activation sensor 170q. Also, the sensors of vehicle 140a and truck 140c may be different than those of vehicle 140a and truck 140c. Figure 1C The illustrated ones are positioned differently.

[0082] As discussed herein, various sensors integrated within each self 140 can be configured to measure various data associated with each navigation session. The analysis server 110a can periodically collect data monitored and collected by these sensors, where the data is processed according to the methods described herein and used to train the AI ​​model 110c and / or execute the AI ​​model 110c to generate an occupancy map.

[0083] The host 140 may include a user interface 170a. The user interface 170a may refer to a host computing device (e.g. Figure 1A The user interface 170a may be implemented as a display screen, a head-up display, a touch screen, etc. that is integrated with or coupled to the interior of the vehicle. The user interface 170a may include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a may be adapted to provide input to other devices or sensors (e.g., Figure 1B The illustrated sensor (such as controller 170c) provides user input (eg, as a signal and / or sensor information).

[0084] The user interface 170a can also be implemented with one or more logic devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, send and / or receive communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information), or perform various other processes and / or methods. In another example, the driver can use the user interface 170a to control the temperature of the self 140 or activate its features (e.g., autonomous driving or steering system 170o). Thus, the user interface 170a can be combined with other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analysis server 110a and / or the AI ​​model 110c.

[0085] The orientation sensor 170b may be implemented as one or more of a compass, a float, an accelerometer, and / or other digital or analog device capable of measuring the orientation of the self-body 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations (such as gravity and / or magnetic north)). The orientation sensor 170b may be adapted to provide a heading measurement to the self-body 140. In other embodiments, the orientation sensor 170b may be adapted to provide a roll, pitch, and / or yaw rate to the self-body 140 using a time series of orientation measurements. The orientation sensor 170b may be positioned and / or adapted to make orientation measurements relative to a particular coordinate system of the self-body 140.

[0086] The controller 170c may be implemented as any appropriate logic device (e.g., a processing device, a microcontroller, a processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that may be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement a control loop for controlling various operations of the self 140. Such software instructions may also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via the user interface 170a), querying the device for operating parameters, selecting operating parameters for the device, or performing any of the various operations described herein.

[0087] The communication module 170e may be implemented as any wired and / or wireless interface configured to transmit sensor data, configuration data, parameters and / or other data and / or signals to the Figure 1A As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are Figure 1B One or more elements and sensors shown are implemented in. In some embodiments, the communication module 170e can delay the transmission of sensor data. For example, when the self 140 does not have network connectivity, the communication module 170e can store the sensor data in a temporary data storage device and send the sensor data when the self 140 is identified as having appropriate network connectivity.

[0088] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind velocity sensor (e.g., direction and magnitude), and / or other device capable of measuring or determining the linear velocity of the self 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the self 140) and providing such measurement as a sensor signal (which can be transmitted to various devices).

[0089] The gyroscope / accelerometer 170f may be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring angular velocity / acceleration and / or linear acceleration (e.g., direction and magnitude) of the self-body 140 and providing such measurements as sensor signals (which may be transmitted to various devices, such as the analysis server 110a). The gyroscope / accelerometer 170f may be positioned and / or adapted to make such measurements relative to a particular coordinate system of the self-body 140. In various embodiments, the gyroscope / accelerometer 170f may be coupled to a computer program product 110a that is configured to provide a plurality of gyroscopes / accelerometers. Figure 1BThe other elements depicted are implemented in a common housing and / or module to ensure a common reference frame or known transformations between reference frames.

[0090] The global navigation satellite system (GNSS) 170h may be implemented as a global positioning satellite receiver and / or another device capable of determining the absolute and / or relative position of the self-body 140 based on, for example, wireless signals received from space and / or ground sources and capable of providing measurements such as sensor signals (which may be transmitted to various devices). In some embodiments, the GNSS 170h may be adapted to determine the velocity, speed, and / or yaw rate of the self-body 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute velocity and / or angular velocity of the self-body 140.

[0091] The temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other device capable of measuring a temperature associated with the self 140 and providing such a measurement as a sensor signal. The temperature sensor 170i may be configured to measure an ambient temperature associated with the self 140, such as a cockpit or dashboard temperature, for example, which may be used to estimate the temperature of one or more elements of the self 140.

[0092] Humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with body 140 and providing such a measurement as a sensor signal.

[0093] The steering sensor 170g may be adapted to physically adjust the heading of the self 140 according to one or more control signals and / or user input provided by a logic device (such as the controller 170c). The steering sensor 170g may include one or more actuators and control surfaces (e.g., a rudder or other type of steering or adjustment mechanism) of the self 140, and may be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g may also be adapted to sense the current steering angle / position of such a steering mechanism and provide such a measurement.

[0094] The propulsion system 170k may be implemented as a propeller, a turbine or other thrust-based propulsion system, a mechanical wheeled and / or tracked propulsion system, a wind / sail-based propulsion system, and / or other types of propulsion systems that may be used to provide power to the self-body 140. The propulsion system 170k may also monitor the direction of the power and / or thrust of the self-body 140 relative to the reference coordinate system of the self-body 140. In some embodiments, the propulsion system 170k may be coupled to the sensor 170g and / or integrated with the steering sensor 170g.

[0095] The passenger restraint sensor 1701 may monitor seat belt detection and locking / unlocking assemblies and other passenger restraint subsystems. The passenger restraint sensor 1701 may include various environmental and / or state sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the vehicle 140. For example, the passenger restraint sensor 1701 may be configured to detect a vehicle seat belt or a vehicle lock / unlock assembly. Figure 1B Other sensors depicted receive motion and / or state data.Occupant restraint sensors 1701 may determine whether safety measures (eg, seat belts) are being used.

[0096] like Figure 1C As depicted, camera 170m may refer to one or more cameras integrated within body 140, and may include multiple cameras integrated (or retrofitted) into body 140. Camera 170m may be a camera facing the inside or outside of body 140. For example, Figure 1C As depicted, vehicle 140 may include one or more interior-facing cameras that may monitor and collect footage of passengers of vehicle 140. Vehicle 140 may include eight exterior-facing cameras. For example, vehicle 140 may include a front-facing camera 170m-1, a front-looking side camera 170m-2, a front-looking side camera 170m-3, a rear-looking side camera 170m-4 on each front fender, a camera 170m-5 on each side (e.g., integrated into a B-pillar), and a rear-facing camera 170m-6.

[0097] Reference Figure 1B , radar 170n and ultrasonic sensor 170p can be configured to monitor the distance of the self 140 to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). The self 140 can also include an autonomous driving or steering system 170o configured to self-navigate the self 140 using data collected via various sensors (e.g., radar 170n, speed sensor 170d, and / or ultrasonic sensor 170p).

[0098] Thus, the autonomous driving or steering system 170o can analyze various data collected by one or more sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can send the analyzed data to various features discussed herein, such as an analysis server.

[0099] The airbag activation sensor 170q may predict or detect a collision and cause activation or deployment of one or more airbags. The airbag activation sensor 170q may transmit data regarding airbag deployment, including data associated with the event that caused the deployment.

[0100] Refer again Figure 1A , the administrator computing device 120 may represent a computing device operated by a system administrator. The administrator computing device 120 may be configured to display data retrieved or generated by the analysis server 110a (e.g., various analysis metrics and risk scores), wherein the system administrator may monitor various models utilized by the analysis server 110a, review feedback, and / or facilitate the training of the (multiple) AI models 110c maintained by the analysis server 110a.

[0101] The (multiple) agents 140 may be any device configured to navigate various routes, such as a vehicle 140a or a robot 140b. Figures 1B to 1C As discussed, the self 140 may include various telemetry sensors. The self 140 may also include a self computing device 141. Specifically, each self may have its own self computing device 141. For example, a truck 140c may have a self computing device 141c. For simplicity, the self computing devices are collectively referred to as (multiple) self computing devices 141. The self computing device 141 can control the content presentation on the infotainment system of the self 140, process commands associated with the infotainment system, aggregate sensor data, manage the communication of data to electronic data sources, receive updates and / or send messages. In one configuration, the self computing device 141 communicates with an electronic control unit. In another configuration, the self computing device 141 is an electronic control unit. The self computing device 141 may include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the (multiple) AI models 110c described herein may be stored and executed (or directly accessed) by the self computing device 141. Non-limiting examples of on-board computing devices 141 may include vehicle multimedia and / or display systems.

[0102] In one example of how the AI ​​model(s) 110c may be trained, the analysis server 110a may collect data from the self 140 to train the AI ​​model(s) 110c. Prior to executing the AI ​​model(s) 110c to generate / predict occupancy data sets, the analysis server 110a may train the AI ​​model(s) 110c using various methods. The training allows the AI ​​model(s) 110c to ingest data from one or more cameras of one or more self 140 (without receiving radar data) and predict occupancy data for the self's surroundings. The operations described in this example may be performed by the Figures 1A to 1DExecuted by any number of computing devices (eg, processors of self 140) operating in the distributed computing system described in .

[0103] The analysis server 110a can use the sensors of the self-body 140 to generate a first data set, which has a first set of data points, wherein each data point in the first set of data points corresponds to the position and sensor properties of at least one voxel in the space around the self-body 140, and the sensor properties indicate whether the at least one voxel is occupied by an object having mass.

[0104] To train (multiple) AI models 110c, the analysis server 110a may first employ one or more of the agents 140 to drive a particular route. While driving, the agent 140 may generate navigation session data using one or more of its sensors (including one or more cameras). For example, one or more agents 140 equipped with various sensors may navigate a specified route. As one or more of the agents 140 traverse the terrain, its sensors may capture continuous (or periodic) data of its surroundings. The sensors may indicate the occupancy state of the surroundings of the one or more agents 140. For example, the sensor data may indicate various objects having mass in the surroundings of one or more of the agents 140 as they navigate their routes.

[0105] The analysis server 110a can generate a first data set using sensor data received from one or more of the selfs 140. The first data set can indicate the occupancy state of different voxels within the surroundings of one or more of the selfs 140. As used herein in some embodiments, a voxel is a three-dimensional pixel that forms a building block of the surroundings of one or more of the selfs 140. In the first data set, each voxel can encapsulate sensor data that indicates whether a mass is identified for that particular voxel. As used herein, a mass can indicate or represent any object identified using a sensor. For example, in some embodiments, the self 140 can be equipped with a LiDAR that identifies the mass by emitting laser pulses and measuring the time required for these pulses to travel to an object (with mass) and return. The LiDAR sensor system can operate based on the principle of measuring the distance between the LiDAR sensor and an object in its field of view. Combined with other sensor data, this information can be analyzed to identify and characterize different masses or objects within the surroundings of one or more of the selfs 140.

[0106] Various additional data may be used to indicate whether voxels of the surroundings of one or more ego bodies 140 are occupied by objects having mass. For example, in some embodiments, a digital map of the surroundings of one or more ego bodies 140 (e.g., a digital map of the route the ego is traversing) may be used to determine the occupancy state of each voxel.

[0107] In operation, as one or more of the agents 140 navigate, their sensors collect data and send the data to the analysis server 110a, as depicted by data stream 176. For example, agent 140 computing device 141 may use data stream 176 to send sensor data to analysis server 110a.

[0108] The analysis server 110 a may generate a second data set having a second set of data points using a camera of the object 140 , wherein each data point in the second set of data points corresponds to a position and an image attribute of at least one voxel of the space around the object 140 .

[0109] The analysis server 110a may receive camera feeds from one or more agents 140 navigating the same route as in the first step. In some embodiments, the analysis server 110a may perform the first step and the second step simultaneously (or at the same time). Alternatively, two (or more) different agents 140 may navigate the same route, with one agent sending its sensor data and the second agent 140 sending its camera feed.

[0110] The one or more bodies 140 may include one or more high-resolution cameras that capture a continuous stream of visual data from the surroundings of the one or more bodies 140 as the one or more bodies 140 navigate the route. The analysis server 110a may then use the camera feeds to generate a second data set, wherein visual elements / depictations of different voxels of the surroundings of the one or more bodies 140 are included in the second data set.

[0111] In operation, as one or more autonomous entities 140 navigate, their cameras collect data and send the data to analysis server 110a, as depicted by data stream 172. For example, autonomous computing device 141 may use data stream 172 to send image data to analysis server 110a.

[0112] The analysis server 110a can use the first data set and the second data set to train the AI ​​model, whereby the AI ​​model 110c uses the corresponding position of each data point to train itself, associating each data point in the first data point set with the corresponding data point in the second data point set, wherein after being trained, the AI ​​model 110c is configured to receive a camera feed from the new self 140 and predict the occupancy state of at least one voxel of the camera feed.

[0113] Using the first data set and the second data set, the analysis server 110a can train (multiple) AI models 110c so that (multiple) AI models 110c can associate different visual properties of a voxel (within the camera feed within the second data set) with the occupancy state of the voxel (within the first data set). In this way, after being trained, the AI ​​model (multiple) 110c can receive a camera feed (e.g., from the new self 140) without receiving sensor data, and then determine the occupancy state of each voxel for the new self 140.

[0114] The analysis server 110a may generate a training data set including a first data set and a second data set. The analysis server 110a may use the first data set as ground truth. For example, the first data set may indicate different locations of voxels and their occupancy states. The second data set may include a visual (e.g., camera feed) illustration of the same voxels. Using the first data set, the analysis server 110a may label the data so that the (multiple) data records associated with each voxel corresponding to the object are indicated as having a positive occupancy state.

[0115] The labeling of the occupancy states of different voxels may be performed automatically and / or manually. For example, in some embodiments, the analysis server 110a may use a human reviewer to label the data. For example, as discussed herein, camera feeds from one or more cameras of a vehicle may be shown to a human reviewer on an electronic platform for labeling. Additionally or alternatively, the (multiple) AI model 110c may ingest the entire data, wherein the (multiple) AI model 110c identifies the corresponding voxels, analyzes the first digital map, and associates the (multiple) images of each voxel with its corresponding occupancy state.

[0116] Using the ground truth, the AI ​​model(s) 110c can be trained to analyze the visual elements of each voxel and associate it with whether the voxel is occupied by a mass body. Thus, the AI ​​model 110c can retrieve the occupancy state of each voxel (using the first dataset) and use that information as ground truth. The AI ​​model(s) 110c can also retrieve the visual attributes of the same voxel using the second dataset.

[0117] In some embodiments, the analysis server 110a may use a supervised training approach. For example, using the ground truth and received visual data, the AI ​​model(s) 110c may train itself so that it can predict the occupancy state for a voxel using only the image of that voxel. As a result, when trained, the AI ​​model(s) 110c may receive a camera feed, analyze the camera feed, and determine the occupancy state for each voxel within the camera feed (without the use of radar).

[0118] The analysis server 110a can feed a series of training data sets to the (multiple) AI model 110c and obtain a set of predicted outputs (e.g., predicted occupancy states). The analysis server 110a can then compare the predicted data with the ground truth data to determine the difference, and train the (multiple) AI model 110c by adjusting the internal weights and parameters of the AI ​​model 110c proportional to the determined difference according to the loss function. The analysis server 110a can train the (multiple) AI model 110c in a similar manner until the predictions of the trained AI model 110c are accurate to a certain threshold (e.g., recall or precision).

[0119] Additionally or alternatively, the analysis server 110a may use an unsupervised approach in which the training data set is not labeled. Because labeling the data within the training data set may be time consuming and may require excessive computing power, the analysis server 110a may utilize unsupervised training techniques to train the AI ​​model 110c.

[0120] After the AI ​​model 110c is trained, the self 140 can use it to predict occupancy data of the surrounding environment of one or more selfs 140. For example, (multiple) AI models 110c can divide the surrounding environment of the self into different voxels and predict the occupancy state for each voxel. In some embodiments, (multiple) AI models 110c (or analysis servers 110a using data predicted by using AI models 110c) can generate an occupancy map or occupancy network representing the surrounding environment of one or more selfs 140 at any given time.

[0121] In another example of how the (multiple) AI model 110c may be used, after training the (multiple) AI model 110c, the analysis server 110a (or a local chip of the self 140) may collect data from the self (e.g., one or more of the self 140) to predict an occupancy data set for the one or more selfs 140. This example describes how the (multiple) AI model 110c may be used to predict occupancy data for one or more selfs 140 in real time or near real time. The configuration may have a processor, such as the analysis server 110a, that executes the AI ​​model. However, one or more actions may be performed locally, for example, via a chip located within one or more selfs 140. In operation, the (multiple) AI model 110c may be executed locally via the self 140 so that the results may be used for autonomous navigation itself.

[0122] The processor may input image data of the space around the self object 140 into the AI ​​model 110c using the camera of the self object 140. The processor may collect and / or analyze data received from various cameras (e.g., externally facing cameras) of one or more self objects 140. In another example, the processor may collect and aggregate footage recorded by one or more cameras of the self object 140. The processor may then send the footage to (multiple) AI models 110c trained using the methods discussed herein.

[0123] The processor may predict occupancy properties of the plurality of voxels by executing the AI ​​model 110c. The AI ​​model(s) 110c may use the received image data to predict occupancy states for different voxels surrounding the one or more ego bodies 140 using the methods discussed herein.

[0124] The processor may generate a data set based on the plurality of voxels and their corresponding occupancy attributes. The analysis server 110a may generate a data set including the occupancy status of different voxels according to their corresponding coordinate values. The data set may be a queryable data set that may be used to send the predicted occupancy status to different software modules.

[0125] In operation, one or more of the selves 140 may collect image data from its camera and send the image data to a processor (located locally on the one or more of the selves 140) and / or the analysis server 110a, as depicted by data stream 172. The processor may then execute (multiple) AI models 110c to predict occupancy data for the one or more of the selves 140. If the prediction is performed by the analysis server 110a, the occupancy data may be sent to the one or more of the selves 140 using data stream 174. If the processor is placed locally within the one or more of the selves 140, the occupancy data is sent to the selves computing device 141 ( Figure 1A not shown).

[0126] Using the methods discussed herein, training of the AI ​​model(s) 110c may be performed such that execution of the AI ​​model(s) 110c may be performed locally (at inference time) on any of the agents 140. Collected data (e.g., navigation data collected during navigation of the agent 140, such as image data of a journey) may then be fed back into the AI ​​model(s) 110c such that the additional data may improve the AI ​​model(s) 110c.

[0127] Figure 1DCertain hardware and software components of the self 140 for performing full or partial autonomous driving (SD) operations according to an embodiment are shown. The self 140 includes an SD circuit 150 and a self computing device 141, which may include components that are the same as or different from the SD circuit 150. The SD circuit 150 includes SD chips 152a to 152b (commonly referred to as SD chips 152), such as a system-on-chip (SoC) integrated circuit chip. Each SD chip 152 includes a non-transitory machine-readable memory, such as DRAM 190a to 190b (commonly referred to as DRAM 190) and SRAM. The SD chip 152 also includes various types of processing units, including GPU 191, CPU 193a to 193c (commonly referred to as CPU 193), and specially designed AI accelerator devices 192a to 192b (commonly referred to as AI accelerator devices 192). SD chip 152 includes inter-chip interfaces 194 a to 194 b (generally referred to as chip interface 194 ) such as Peripheral Component Interconnect (PCI) or PCI Express (PCIe). SD chip 152 transmits signals via inter-chip bus 199 according to the protocol and programming of chip interface 194 .

[0128] As compared to Figure 1AAs mentioned, the analysis server 110 (or other computing device) can compile a compiled executable binary file of software for the neural network architecture and download it to the self 140 and / or the self computing device 141. The self computing device 141 can generate and / or execute various software programming operations and executable binary files for managing the operation of the SD circuit 150 (or other hardware), which can include execution instructions for applying the neural network architecture to sensor data types from sensors of the self 140. The executable instructions received, generated and / or executed by the self computing device 141 may include executable commands for managing the operation of components of the SD circuit 150. For example, the compiled executable binary file may include instructions, such as indicating a destination SD chip 152 for transmitting data signals between SD chips 152 via the bus 199, or indicating an execution SD chip 152 for performing the functions of certain neural network architectures. For example, instructions generated and compiled for a neural network architecture for identifying traffic sign objects using data from camera 170m may be loaded into and executed by components of the first SD chip 152a, where instructions compiled for a neural network architecture for path planning may be loaded and executed by components of the second SD chip 152b. If the path planning neural network architecture is trained to use the data output produced by the traffic sign neural network architecture, the instructions for the first SD chip 152a instruct the first SD chip 152a to pass the output data signal to the second SD chip 152b via bus 199; and the instructions for the second SD chip 152b instruct the second SD chip 152b to use such data signal received from the first SD chip 152a.

[0129] In an example embodiment, the SD circuit 150 includes two SD chips 152a-152b. In many cases, the SD chips 152 operate in a redundant mode or failover mode of operation, where the first SD chip 152a acts as a primary chip and the second SD chip 152b acts as a secondary chip. For example, the first SD chip 152a is prioritized to execute most executable instructions, and the second SD chip 152b is called to operate as a failover or redundancy in the event of a problem with the first SD chip 152a.

[0130] SD circuit 150 may operate in an extended computing mode that balances execution instruction pipelines among SD chips 152. As an example, the native computing device 141 executes a software routine for compiling execution instructions to be executed by processing units 191 to 193 of SD chip 152 and distributing the execution instructions to optimal hardware components of SD circuit 150.

[0131] SD chip 152 includes inter-chip memory sequencers 195a to 195b (commonly referred to as inter-chip memory sequencers 195). Inter-chip memory sequencers 195 include hardware IC devices used to coordinate signal communications between systems on a chip (SoCs) such as SD chip 152. In some implementations, inter-chip memory sequencers 195 may include non-transitory storage locations that provide shared memory space accessible by SD chip 152. In some implementations, inter-chip memory sequencers 195 perform operations to coordinate data signal transfer between SD chips 152 by, for example, generating various control signals. Inter-chip memory sequencers 195 may implement one or more inter-chip communication protocols, such as PCIe, SPI, or I2C.

[0132] The hardware and software components of the runtime system of the self 140 (eg, the self computing device 141, the SD board 150, the controller 180) receive the compiler schedule (eg, Figures 2A to 2C 152) of the execution schedule 217, which may be in the form of an execution binary file 216, and runs the execution instructions on various types of heterogeneous cores (e.g., processing units 190 to 193 of each chip 152) and on various chips 152 across the circuit 150. The runtime system (represented as being executed by the controller 180) includes software components for an inter-chip compute scheduler, a heterogeneous hardware scheduler (e.g., a CPU accelerator, a GPU accelerator, an AI accelerator), an inter-chip memory sequencer 195 for scheduling and managing inter-chip signals via a chip interface 194 and a bus 199, and a clock synchronizer (e.g., an OS kernel). In some implementations, runtime system programming supports model parallelism across multiple SD chips 152. In this way, the self 140 can review and delve into the clock synchronizer.

[0133] In some embodiments, the self 140 includes a controller 180, which performs various operations for managing the SD circuit 150. The controller 180 can perform various functions according to, for example, instructions from the self computing device 141 (or other components of the self 140) or configuration inputs from a management user. For example, the controller 180 switches, configures, or otherwise instructs the SD circuit 150 to operate in various operating modes. In some cases, for example, the controller 180 instructs the SD circuit 150 to operate in an extended computing mode, in which the first SD chip 152a executes a first instruction partition of execution instructions, and the second SD chip 152b executes a second instruction partition. As another example, in some cases, the controller 180 instructs the SD circuit 150 to operate in a failover mode, in which when the first SD chip 152a fails, the second SD chip 152b executes the execution instruction.

[0134] SD chip 152 includes one or more DRAM 190 or other types of non-transitory memory for storing data input for SD chip 152. Data input can be stored in DRAM 190 for reference by the processing unit for various calculations. In some configurations, AI accelerator device 192 includes SRAM, so that SD chip 152 moves data from DRAM 190 for storage in SRAM of AI accelerator device 192. AI accelerator device 192 performs calculations according to the execution instructions and moves data back to DRAM 190 or other destinations of SD circuit 150.

[0135] SD chip 152 includes various types of processing units, which may include any hardware integrated circuit (IC) processor device capable of performing the various processes and tasks described herein. Non-limiting examples of processing unit types include GPU 191, CPU 193, AI accelerator device 192, microcontroller, ALU, ASIC, and FPGA, etc. The processing unit can perform the computing functions of the programming layer that defines the neural network architecture or sub-architecture. The compiler outputs execution instructions representing the operation of the neural network architecture executed by the self-computing device 141 (or other components of the self 140).

[0136] AI accelerator devices 192 are hardware accelerators specifically designed for neural network operations, with a beneficial focus on, for example, optimizing power and performance (e.g., low latency) improvements. AI accelerator devices 192 include hardware IC devices (e.g., microcontrollers, ALUs, ASICs, FPGAs, processor devices) that are designed for fast operation when processing neural network architectures. For example, as transformer neural network architectures (e.g., GPT) and other types of neural network modeling techniques become more popular, other types of processing units (e.g., CPUs 193, GPUs 191) may be slower due to design theories intended for a wider range of implementation use cases. For example, neural network architectures, sub-neural networks (e.g., Figure 2BThe moving object network 206b) or child neural network performs computer vision or object recognition by implementing various GPTs (or other types of transformers) on image sensor data, thereby beneficially replacing previous techniques for post-processing of visual neural networks. The AI ​​accelerator device 192 is specifically designed for neural network operations that allow GPT transformers to run locally in the computing components of the self 140, so that the AI ​​accelerator device 192 provides faster and more efficient processing than a traditional GPU 191 or CPU 193 that performs similar GPT transforms. In this way, the AI ​​accelerator device 192 reduces or eliminates latency and improves overall efficiency, contributing to the ability of the self 140 to make real-time decisions. Moreover, when executing more sophisticated and complex functions of the neural network architecture, such as transformer networks (e.g., transformers), the structural design and design theory of the AI ​​accelerator device 192 draw relatively less power than a traditional GPU 191 or CPU 193.

[0137] In some embodiments, a transformer (e.g., a GPT) may be adapted for execution in the self 140, thereby improving the overall performance of a computing component for an autonomous or semi-autonomous self 140. For example, typical transformers are often resource-intensive, consume a lot of power, and / or cause a lot of latency when processing outputs, and thus hinder the overall performance of the self 140. As such, the transformer is a powerful neural network architecture that is not typically deployed in the self 140. To address this problem, the transformer of the self 140 described herein may be deployed without the attention module softmax typically found in conventional transformers. The embodiments described herein may include a transformer with softmax free attention, and may implement ReLU activation. By replacing softmax with ReLU in such a transformer, the transformer architecture may now be deployed in an autonomous or semi-autonomous self 140.

[0138] In some embodiments, the SD circuit includes a controller or other processing unit that maintains clock synchronization between SOC chips based on interpreting timestamps of sensor inputs (e.g., timestamps of camera inputs) and translating timestamps between SOC chips based on differences between current chip clocks of each SOC chip.

[0139] Figure 1ECertain hardware and software components of the self 140 for maintaining clock synchronization between SD chips 152 according to an embodiment are shown. In such an embodiment, the SD circuit 150 includes two (or more) SD chips 152 and two (or more) microcontrollers 155a to 155b (generally referred to as microcontrollers 155) coupled to corresponding microcontrollers 155. For example, the first SD chip 152a is coupled to the first microcontroller 155a, and the second SD chip 152b is coupled to the second microcontroller 155b. Each SD chip 152 includes an OS kernel 153 and executes the OS kernel 153, which manages the operation of the SD chip 152, including executing various execution instructions of the neural network architecture and managing the operation of the SD chip 152.

[0140] The OS kernel 153 may include any type of OS capable of performing the various processes and tasks described herein, including managing the execution of execution instructions and maintaining a software-based OS clock. The OS of the OS kernel 153 includes, for example, Linux, Unix, and the like.

[0141] The microcontroller 155 includes any type of processing circuit unit capable of performing the various processes and tasks described herein. Non-limiting examples include microcontrollers, controllers, ALUs, FPGAs, ASICs, etc. The SOC chip 152 includes one or more processing units that execute an OS kernel 153 (e.g., Linux) and one or more smaller microcontrollers 155. The microcontroller 155 includes low-level programming for performing various low-level functions. For example, the functions of the microcontroller 155 include boot loader functions, in which the microcontroller 155 guides various components of the SD chip 152, including guiding the processing unit that executes the OS kernel 153. In some embodiments, the microcontroller 155 or other devices of the SD circuit 150 may include or be coupled to a clock oscillator or counter that oscillates or increments a monotonic clock at a given frequency.

[0142] The microcontroller 155 can communicate with the processing unit executing the OS kernel 153 via a given interface (e.g., mailbox, Ethernet, UART, PCIe) to exchange data signals, such as time messages or correction instructions. Similarly, the first microcontroller 155a can communicate with the second microcontroller 155b via another interface (e.g., mailbox, Ethernet, UART, PCIe) to exchange time messages or correction instructions.

[0143] Typically, the microcontroller 155 performs a boot loader function to simultaneously boot the microcontroller 155 and the OS kernel 153 at a relatively early time after booting the body 140. At boot time, the microcontroller 155 and the OS kernel 153 communicate various timing messages to synchronize with each other. In this way, the boot and synchronization functions of the microcontroller 155 logically form a common monotonic time clock for the SD circuit 150 before the processor unit of the OS kernel 153 has a chance to begin execution.

[0144] When the programming of the OS kernel 153 (e.g., Linux) begins booting on the processor, early in the booting of the OS kernel 153, the OS kernel 153 resets the monotonic kernel clock of the OS kernel 153. The OS kernel 153 may reset the kernel clock to match the controller clock of the corresponding microcontroller 155. This reset may occur before anything or before a lot of operations have a chance to start on the SD chip 152, or before too much operations have started in the OS kernel 153. When the OS kernel 153 boots, the OS kernel 153 synchronizes the kernel clock to the corresponding microcontroller 155.

[0145] In some implementations, at preconfigured resynchronization stages or threshold times, the SD chip 152 and microcontroller 155 actively maintain synchronized kernel and controller clocks through a control loop operation in which the SD chip 152 and microcontroller 155 exchange timing messages. For example, at a synchronization interval (e.g., once per second), the OS kernel 153 will measure the time error relative to the corresponding microcontroller 155, and in some cases, the OS kernel 153 instructs the connected microcontroller 155 to initiate a small adjustment to correct the controller clock.

[0146] In some embodiments, the clock synchronization operation includes a redundant fallback function. In some cases, the SD chip 152 (e.g., the second SD chip 152b) may suffer a fatal error and must be rebooted, while another SD chip 152 (e.g., the first SD chip 152a) can continue to operate until the rebooted SD chip 152 (e.g., the second SD chip 152b) is restored. In this case, the operable SD chip 152 (e.g., the first SD chip 152a) continues to maintain the kernel clock of the OS kernel 153 (e.g., the first OS kernel 153a) and the controller clock of the corresponding microcontroller 155 (e.g., the first microcontroller 155a), and therefore, by extension, maintains the overall logical synchronization clock for the SD circuit 150. In this way, when the rebooted SD chip 152b is restored, the overall synchronization can be continued by, for example, a synchronization message between the first microcontroller 155a and the second microcontroller 155b. The rebooted SD chip 152b does not need to restart the new kernel clock and the new controller clock at zero or some other initialization time. The microcontroller 155b and the OS kernel 153b may start the kernel clock and the controller clock at the current monotonic time of the entire synchronous clock for the SD circuit 150. The microcontroller 155b may perform a recovery, reboot, or boot process, including a synchronization process with the operating microcontroller 155a, wherein the microcontroller 155 exchanges time messages to indicate the current time of the controller clock of the operating microcontroller 155a, which reflects the entire synchronous clock for the SD circuit 150. The recovered microcontroller 155b and the recovered OS kernel 153b may then exchange time messages indicating the current time of the controller clock of the microcontroller 155, which reflects the entire synchronous clock for the SD circuit 150. In this manner, the components of the SD circuit 150 do not need to calculate, distribute, or translate the time differences between the discontinuous clocks of the SD chip 152. After rebooting, the application software executing in the recovered SD chip 152b and the recovered OS kernel 153b can begin execution and participation almost immediately, with limited delay to reconstruct the state before the failure, and / or no continuous delay due to the continuous calculation for translating what would be a discontinuous clock. This beneficially improves fault tolerance and supports failover redundancy for the self 140.

[0147] The OS kernel 153 and / or microcontroller 155 may perform an error correction function that adjusts the frequency of the controller clock by a relatively small amount, causing the microcontroller 155 or OS kernel 153 to increase or decrease the frequency and clock time (eg, controller clock, kernel clock) by some amount.

[0148] As mentioned, each SD chip 152 has multiple types of synchronization operations, including kernel-to-controller synchronization operations between the OS kernel 153 and the corresponding microcontroller 155; and SOC synchronization or controller-to-controller synchronization operations between the microcontrollers 155 of the SD chip 152.

[0149] In a first type of synchronization operation (e.g., synchronizing the OS kernel 153a with the corresponding microcontroller 155a), the OS kernel 153 sends a timing message to the microcontroller 155. The OS kernel 153 and the microcontroller 155 of the SD chip 152 each include a communication interface (e.g., mailbox, Ethernet, UART, PCI) for exchanging timing messages (or other types of messages) via signal connections, wires, or buses according to the protocol of the particular interface.

[0150] At boot time and / or at preconfigured intervals, the OS kernel 153 sends a timing message to the microcontroller 155, and receives a return timing message from the microcontroller 155, and references the relevant clock time to determine whether one or more clocks have drifted beyond a threshold distance. The OS kernel 153 sends an initial timing message to the microcontroller 155 at a first time (T1). The OS kernel 153 retrieves the current time with reference to the kernel clock, and assigns the current time as the initial message time (T1) for the initial timing message. The microcontroller 155 receives the initial timing message, and references the current time of the controller clock of the microcontroller 155. The microcontroller 155 sends a response timing message to the OS kernel 153 at a response time (T2). The OS kernel 153 assigns a response time to the response timing message according to the current time of the controller clock. The OS kernel 153 receives a response timing message indicating a response time (T2), and references the kernel clock to retrieve the current time of the kernel clock. When the OS kernel 153 receives the response message, the OS kernel 153 assigns the current time of the kernel clock as the completion time (T3) to the response timing message. The OS kernel 153 calculates an average time based on an average value of kernel times including an initial message time (T1) and a completion time (T3). The OS kernel 153 then compares the average time with the response time (T2) received from the microcontroller 155 to calculate and output the difference between the average time and the response time. In an example configuration, the first OS kernel 153a may calculate an offset representing an estimated amount of time error between a first monotonic kernel clock of the first OS kernel 153a and a first monotonic controller clock of the first microcontroller 155a. In this example configuration, the offset is calculated by the difference between the average time and the response time.

[0151] In a second type of synchronization operation (e.g., synchronizing the first microcontroller 155a with the second microcontroller 155b), the OS kernel 153 sends timing messages to the microcontrollers 155. The microcontrollers 155 each include a communication interface (e.g., mailbox, Ethernet, UART, PCI) for exchanging timing messages (or other types of messages) with the other microcontrollers 155 via signal connections, wires, or buses according to the protocol of the particular interface.

[0152] At boot time and / or at preconfigured intervals, the microcontrollers 155 automatically begin exchanging timing messages with each other. The microcontrollers 155 do not need to establish a handshake, nor do they need to exchange any lookahead or assertion communications; the microcontrollers 155 may be sending timing messages. Each microcontroller 155 uses timing messages to capture and determine the message time (T1) and completion time (T3), and receive a response time (T2) back from the other microcontroller 155. Each microcontroller 155 can then calculate the offset as described above with respect to the OS kernel 153. Therefore, each microcontroller 155 can estimate the offset relative to the peer microcontroller 155. And they just do this automatically.

[0153] In some cases, a particular microcontroller 155 is rebooted and restored. In this case, the reboot or restore function of the restored microcontroller 155b can act as a follower of the operating microcontroller 155a, which acts as the master. The first microcontroller 155a and the second microcontroller 155b treat the controller clock of the first microcontroller 155a as the master controller clock. Upon reboot and restore, the first microcontroller 155a and the second microcontroller 155b exchange timing messages that directly assign the controller clock of the first microcontroller 155a to the controller clock of the second microcontroller 155b.

[0154] In some implementations, one or more offsets representing the difference between the two clocks are sent as error signals to, for example, a proportional integral (PI) controller or a phase-locked loop (PLL) controller. Using known algorithmic techniques, the PLL or PI controller can calculate an error rate based on the offset, where the input is an offset as a phase and the output is a ratio. The error rate is a ratio that the OS kernel 153 (or microcontroller 155) must correct. At each instance or interval (e.g., every second), a new measurement of the calculated ratio, the ratio difference (offset) between the two clocks, the OS kernel 153 (or microcontroller 155) can adjust the kernel clock (or controller clock) by applying the ratio to the kernel clock (or controller clock). As an example, if the calculated ratio indicates that the kernel clock of the first OS kernel 153a is two millionths faster than the first microcontroller 155a, then the first OS kernel 153a is adjusted to a new frequency that matches the two millionths slower frequency in the first OS kernel 153a to meet the frequency oscillation of the first microcontroller 155a.

[0155] The first microcontroller 155a sends a timing message to the second microcontroller 155b, and receives a return timing message from the second microcontroller 155b, and refers to the relevant clock time to determine whether one or more clocks have drifted beyond a threshold distance. The first microcontroller 155a sends an initial timing message to the second microcontroller 155b at a first time (T1). The OS kernel 153 retrieves the current time with reference to the kernel clock, and assigns the current time as the initial message time (T1) for the initial timing message. The microcontroller 155 receives the initial timing message, and refers to the current time of the controller clock of the microcontroller 155. The microcontroller 155 sends a response timing message to the OS kernel 153 at a response time (T2). The OS kernel 153 assigns the response time to the response timing message according to the current time of the controller clock. The OS kernel 153 receives a response timing message indicating the response time (T2), and refers to the kernel clock to retrieve the current time of the kernel clock. When the OS kernel 153 receives the response message, the OS kernel 153 assigns the current time of the kernel clock as the completion time (T3) to the response timing message. The OS kernel 153 calculates an average time based on an average value of kernel times including an initial message time (T1) and a completion time (T3). The OS kernel 153 then compares the average time with the response time (T2) received from the microcontroller 155 to calculate and output the difference between the average time and the response time. In an example configuration, the first OS kernel 153a may calculate an offset that represents an estimated amount of time error between a first monotonic kernel clock of the first OS kernel 153a and a first monotonic controller clock of the first microcontroller 155a. In this example configuration, the offset is calculated for the difference between the average time and the response time.

[0156] Figures 2A to 2B Depicted is the data flow between hardware and software computing components of system 200 for developing and compiling executable instructions 218a-218h (generally referred to as execution instructions 218) at development system 201 to be loaded into body 202 as execution binary file 216 according to an embodiment. Figure 2C Execution instructions 218 are illustrated as being generated and organized into an execution schedule 217 for execution by circuit hardware components of system 200 in accordance with an embodiment.

[0157] One or more computing devices (e.g., analysis server 110) of development system 201 may execute software programming that defines one or more neural network architectures 204, hardware model training engine 207, compiler 210, and execution scheduler 212, as well as other types of software programming routines. In addition, the computing device(s) of development system 201 may execute software programming for training, retraining, and tuning neural network architecture 204 or portions of neural network architecture 204 (e.g., parameters, hyperparameters, weights, layers, functions) on various forms of historical and / or current sensor data from any number of self-bodies 202 for prediction accuracy and consistency. Additionally or alternatively, the computing device(s) of development system 201 may execute software programming for training, retraining, and tuning neural network architecture 204 or portions of neural network architecture 204 (e.g., parameters, hyperparameters, weights, layers, functions) on various types of input data, output data, or prediction data of neural network architecture 204 that are relative to or scaled to, or optimized for, data sizes and formats implemented, for example, by hardware components of self-bodies 202.

[0158] During training or inference time, the computing device(s) of the development system 201 extract features or tensors from input data, such as historical or current sensor data acquired from sensors of the self 202 or retrieved from a database of the development system 201 containing historical data captured by the self 202. The computing device(s) of the development system 201 feeds the input data to the neural network architecture 204 or sub-architectures for various operations (e.g., computer vision, object recognition), and applies the neural network architecture 204 to the input data to generate predicted outputs, and adjusts or retrains portions of the neural network architecture 204 (e.g., parameters, hyperparameters, weights, layers) during training.

[0159] The (multiple) computing devices of the development system 201 apply a graph partitioner to the sensor data to generate data partitions or portions. The self computing device 141 applies a compiler set (not shown), which can logically form a compiler tool chain for the neural network architecture of the self 202 for compiling and debugging the code for executing the neural network architecture layer for sensor data interpretation. Each compiler is used to transform a high-level programming language into machine code including execution instructions executed by the hardware of the SD circuit 150. The compiler can be configured or optimized to compile programming code according to the specific architecture or type of the processing unit of the SD chip (e.g., CPU 193, GPU 191, or dedicated AI accelerator device 192 hardware). The scheduling optimizer of the execution scheduler can combine multiple compiled code segments (e.g., executable instructions) into one or more executable files or data streams (not shown) for execution scheduling.

[0160] The scheduling optimizer and the execution scheduler obtain the execution instruction set and map the execution instructions to the hardware components of the SD circuit (e.g., GPU 191, AI accelerator device 192, CPU 193) to execute specific execution instructions. In some implementations, the scheduling optimizer of the execution scheduler is trained to optimize the operations to be performed in the hardware components of the SD circuit. The scheduling optimizer is trained to determine or pre-configured with time or latency requirements for the hardware components to perform the operations of the execution instructions. This is usually possible because such performance timing or latency metrics are known, essentially static, quickly calculated, or pre-stored. In this way, the scheduling optimizer maps the execution instructions to the components of the SD circuit according to the minimized or optimized latency. Additionally or alternatively, the scheduling optimizer determines which hardware components of the SD circuit should execute which execution instructions based on the characteristics of the execution instructions (e.g., which compiler generates the machine code for the execution instructions). In this way, the scheduling optimizer maps the execution instructions to the processing unit based on the compiler that generates the specific execution instructions.

[0161] like Figure 2AAs shown, system 200 includes software programming for executing a neural network architecture 204 at a computing device (or devices) of a development system 201, including various domain-specific or task-specific subnetworks 206a to 206e (generally referred to as subnetworks 206), but other types of machine learning architectures may also be included. The source code of the software programming defines various aspects of the neural network architecture 204 (e.g., parameters, hyperparameters, weights, layers, functions), and includes source code defining any number of subnetworks 206, including a traffic signal network 206a, a moving object network 206b, a lane network 206c, an occupancy network 206d, and a path planning network 206e for performing operations for a specific domain or task, but embodiments may include additional or alternative types of subnetworks 206. The software components of the development system 201 may also include a compiler 210, a hardware model training engine 207, and an execution scheduler 212, which includes the function of defining a scheduling optimizer 214.

[0162] During training of the neural network architecture 204 and the subnet 206, the development system 201 executes the hardware model training engine 207 to train the subnet 206 on a model or representation data for the hardware components of the self 202. In this manner, the hardware model training engine 207 provides quantization-aware training (QAT) to the neural network architecture 204 so that the subnet 206 can be optimized for the hardware components and quantization resilience of the self 202. Beneficially, the QAT function of the hardware model training engine 207 trains the neural network architecture 204 and the subnet 206 to be more efficient and smaller in size. The QAT function of the hardware model training engine 207 trains the subnet 206 by applying the subnet 206 to various or desired quantized weights (e.g., 16-bit floating point values; 8-bit floating point values; 8-bit integers) and activations (e.g., 16-bit floating point values; 8-bit floating point values; 8-bit integers). The QAT function of the hardware model training engine 207 forces the neural network architecture 204 and / or each subnet 206 to learn to operate with lower precision numbers. For example, the hardware model training engine 207 may train the subnet 206 to ingest or produce 8-bit floating point or integer values, rather than, for example, ingesting or producing 16-bit floating point or integer values.

[0163] Beneficially, the hardware model training engine 207 can improve efficiency and reduce the demand on computational resources on the self 202 by using smaller quantized data sizes. Additionally, this can reduce the power consumption required by the hardware of the self 202. For example, by training the neural network architecture 204 to be quantization-aware or resilient, the neural network architecture 204 and the hardware executing the compiled neural network architecture 204 operate adequately on lower precision data values ​​or primitives. In this way, the hardware of the self 202 can run the neural network architecture 204 trained and compiled at lower precision and with relatively lower power consumption than would be used if the hardware of the self 202 ran the neural network architecture 204 trained and compiled at higher precision.

[0164] The neural network architecture 204 may include or be connected to a neural network of a graph partitioner 208. The graph partitioner 208 may partition the sensor data received via the ingestion layer 202 into subnets 206 for processing the sensor data. The neural network architecture 204 is logically partitioned into subnets 206. The neural network layers of the graph partitioner 208 are trained to parse the sensor data into data portions and then map the data portions to the subnets 206 to perform the functions of the specific subnets 206. The graph partitioner 208 maps the sensor data portions to the subnets 206 based on, for example, the type of data or features in the sensor data used by the subnets 206.

[0165] After assigning the sensor data and functions to the subnet 206, the graph partitioner 208 can then assign which hardware components of the SD circuit 201 should implement and execute the functions of the subnet 206. The neural network layer of the graph partitioner 208 is trained to assign the sensor data portion and the functions of a specific subnet 206 to a specific processing unit (e.g., CPU, GPU, dedicated hardware AI accelerator device) of the chips 203a to 203b (commonly referred to as chip 203) (e.g., SoC, SD chip) of the SD circuit 201. For example, the graph partitioner 208 is configured and trained to assign a relatively simple function of a subnet 206 to the CPU of the chip 203 using a sensor data portion, and to assign a relatively complex function of another subnet 206 to the AI ​​accelerator device of the chip 203 using another sensor data portion.

[0166] Reference Figure 2B, the domain-specific subnetwork 206 includes a traffic signal network 206a, a moving object network 206b, a lane network 206c, an occupancy network 206d, and a path planning network 206e. The subnetwork 206 performs various types of domain-specific or task-related functions for a given purpose (e.g., object recognition, path planning) according to the software programming of the neural network layer of the specific subnetwork 206. As an example, the function of the traffic sign network 206a includes recognizing certain object image data (or other types of sensor data), such as traffic control, stop signs, yield signs, speed signs, and topological signs. As another example, the function of the occupancy network 206d includes determining per-voxel occupancy, per-voxel rate, and 3D surface geometry semantics and other image-related metrics in the image data (or other types of sensor data). As another example, the function of the path planning network 206e includes using image data or other types of sensor data to generate a trajectory or path for navigating the self and adjusting the path for collision avoidance, etc.

[0167] The subnets 206 perform various operations or functions for computing sensor data and producing outputs for the particular subnet 206. In some cases, these functions include process computations or operations. In some cases, these functions include child neural networks for the particular subnet 206. Non-limiting examples of child network types for the subnet 206 include ingestion or ingestion layers (sometimes referred to as "head layers"), rectification layers, regularized neural network (RegNet) layers, transformer layers, and multi-layer perceptron (MLP) layers, among others.

[0168] In some cases, the graph partitioner 208 is also configured and trained to assign a portion of the input sensor data to a specific subnet 206 based on the type of child neural network in the subnet 206. Additionally or alternatively, the graph partitioner 208 is trained to assign functions and sensor data to the hardware of the SD circuit 201 based on the capabilities of the hardware. For example, the graph partitioner 208 is trained to optimize the efficiency of the hardware and / or reduce the latency of the hardware, or to implement any additional or alternative performance behavior of the hardware. As an example, the graph partitioner 208 assigns a series of interrelated or dependent functions of the child network within the subnet 206 to one or more CPUs of the same first chip 203a, which can improve efficiency. As another example, the graph partitioner 208 assigns the complex functions of the subnet 206 to the AI ​​accelerator device of the chip 203, which can improve the computing speed. The graph partitioner 208 can be trained to maximize or minimize the performance metric, or the graph partitioner can be trained to balance and optimize according to multiple performance metrics.

[0169] return Figure 2A, system 200 includes a compiler toolchain that includes a set of compilers 210a to 210c (generally referred to as compilers 210). Compiler 210 includes software programming that is configured to transform a high-level programming language of layers and functions of neural network architecture 204 and sensor data into machine code of execution instructions 218 that can be executed by hardware of SD circuit 201. System 200 includes a heterogeneous set of processing units and compilers 210, wherein compiler 210 is configured to transform from a given high-level programming language (e.g., functions and sensor data portions of subnet 206) into a given machine code (e.g., execution instructions 218) compatible with an assigned processing unit. In some embodiments, the software routines of compiler 210 can selectively compile machine code for more than one type of processing unit, so that compiler 210 can generate execution instructions 218 for more than one type of processing unit (e.g., CPU, GPU, AI accelerator device).

[0170] The graph partitioner 208 (or other component of the system 200) is configured and trained to identify which compiler 210 should be assigned to compile which functions of the subnet 206. For example, continuing with the previous example, the graph partitioner 208 is configured and trained to assign relatively simple functions of the subnet 206 to a first compiler 210a, which is programmed to generate execution instructions 218 for the CPU, and to assign relatively complex functions of another subnet 206 to a second compiler 210b, which is programmed to generate execution instructions 218 for the AI ​​accelerator device. The output of the compiler 210 is the execution instructions 218 compiled from the sensor data portion and the software programming for the functions of the subnet 206.

[0171] The system 200 includes an execution scheduler 212 neural network, which includes a layer defining a schedule optimizer 214 (sometimes referred to as a "linker"). The schedule optimizer 214 of the execution scheduler 212 can combine multiple compiled code segments (e.g., executable instructions 218) generated by the compiler 210 of the compiler tool chain into one or more execution binaries 216, including one or more executable files or data streams of execution instructions 218. The execution scheduler 212 can arrange or queue the execution instructions 218 for in-order execution by the hardware of the SD circuit 201.

[0172] The execution binary file 216 is downloaded from the software of the system 200 to the non-transitory memory of the hardware of the SD circuit 201. In the SD circuit 201, the software-based or firmware-based controller component of the SD circuit 201 parses the execution instructions 218 of the execution binary file 216 and loads the execution instructions 218 into one or more non-transitory memories (not shown) accessible to the assigned processing unit (or other hardware component) of the chip 203. The processing unit then executes the execution instructions 218 to perform the functions of the neural network architecture 204.

[0173] Reference Figure 2C In some cases, the execution instructions 218 generated by the compiler 210 can be arranged and logically represented as an execution schedule 217. The output of the compiler toolchain includes the execution instructions 218 for the functions of the subnet 206. For ease of understanding, Figure 2C The execution instructions 218 in FIG. 2 show the chip 203 assigned to perform the operation, the processing unit (e.g., GPU, CPU, AI accelerator device) assigned to perform the operation, and the function of which neural network architecture will be executed. However, the execution instructions 218 of potential embodiments may include additional or alternative types of information, such as input data sources or interfaces, output data destinations or interfaces, and computational instructions.

[0174] As an example, the first execution instruction 218a instructs the first GPU (gpu0) of the second chip 203b (SoC1) to be assigned to perform these functions. The first execution instruction 218a instructs the first GPU of the second chip 203b (i.e., running gpu0, soc1) to perform the functions of correcting the neural network architecture within the occupied network 206d.

[0175] The self's downstream hardware and software can ingest the outputs generated by the SD circuit 201 executing the neural network architecture 204, such as trajectory, velocity, and other navigational determinations of the path planning network 206e, to operate or manipulate the self within the environment.

[0176] Various illustrative logic blocks, modules, circuits and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits and steps have been generally described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Those skilled in the art can implement the described functionality in different ways for each specific application, but this implementation decision should not be interpreted as causing deviations from the scope of the present invention.

[0177] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description language or any combination thereof. Code segments or machine executable instructions can represent any combination of process, function, subroutine, program, routine, subroutine, module, software package, class, or instruction, data structure or program statement. Code segments can be coupled to another code segment or hardware circuit by transmitting and / or receiving information, data, independent variables, attributes or memory contents. Information, independent variables, attributes, data etc. can be transmitted, forwarded or sent via any suitable means (including memory sharing, message passing, token passing, network transmission etc.).

[0178] The actual software code or dedicated control hardware used to implement these systems and methods does not limit the present invention. Therefore, the operation and behavior of the systems and methods are described without reference to specific software code, and it is understood that software and control hardware can be designed to implement the systems and methods based on the description herein.

[0179] When implemented in software, functions can be stored as one or more instructions or codes on a non-transient computer-readable or processor-readable storage medium. The steps of the method or algorithm disclosed herein can be implemented in a processor-executable software module, which can reside on a computer-readable or processor-readable storage medium. Non-transient computer-readable or processor-readable media include both computer storage media and tangible storage media, and tangible storage media facilitate the transfer of computer programs from one place to another. Non-transient processor-readable storage media can be any available medium accessible by a computer. By way of example and not limitation, such non-transient processor-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other tangible storage media, which can be used to store desired program codes in the form of instructions or data structures and can be accessed by a computer or processor. Disks and optical disks used herein include optical disks (CDs), laser optical disks, optical disks, digital versatile disks (DVDs), floppy disks and blue optical disks, wherein disks usually reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above items should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0180] The foregoing description of the disclosed embodiments is provided to enable those skilled in the art to make or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be given the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0181] While various aspects and embodiments have been disclosed, other aspects and embodiments are also contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1. A method, include: Obtaining, by a computer, software programming of multiple functions of multiple sub-neural networks including a neural network architecture; assigning, by the computer, the one or more sub-neural networks to a plurality of processing units of the body, wherein for each sub-neural network, the computer assigns a processing unit to compute the plurality of functions of the sub-neural network; and The computer generates a plurality of execution instructions for the plurality of processing units by executing a plurality of compilers on the plurality of functions of each sub-network, wherein for each execution instruction, the computer uses a compiler from the plurality of compilers according to the processing unit of the body of the plurality of functions assigned to the sub-neural network.

2. The method according to claim 1, further comprising: include: A computer file including the plurality of execution instructions for the plurality of processing units is generated by the computer to execute the plurality of sub-neural networks.

3. The method according to claim 1, further comprising: include: By the computer, the computer file is sent to the self. 4 . The method of claim 1 , wherein at least one execution instruction causes the circuitry of the self to operate in an extended computing mode for parallel execution of the plurality of execution instructions, the circuitry of the self comprising the plurality of chips.

5. The method according to claim 1, wherein at least one execution instruction instructs the circuit of the self to operate in a redundant mode so that the multiple execution instructions are mainly executed by a main chip among the multiple chips, and the circuit of the self includes multiple chips, and the multiple chips include the multiple processing units.

6. The method according to claim 1, further comprising: include: Applying, by the computer, a schedule optimizer engine to the execution instructions to generate an execution schedule for the execution instructions, the schedule optimizer engine comprising a neural network layer trained to generate the execution schedule for minimizing latency. 7 . The method of claim 1 , wherein the one or more processing units include at least one of a GPU, a CPU, or an accelerator device.

8. The method of claim 1, wherein the one or more processing units are heterogeneous, comprising at least two types of processing units.

9. The method of claim 1 , wherein the computer assigns the processing unit among the plurality of processing units of the self to apply the sub-neural network based on one or more diagrams representing a circuit architecture of the self having the plurality of processing units.

10. The method according to claim 1, further comprising: include: The one or more sub-neural networks are trained, by the computer, based on the one or more processing units for quantization-aware training by applying each sub-neural network to a training data set including data having expected quantization characteristics.

11. A system, include: A computer including a processor configured to: Obtaining software programming for multiple functions of multiple sub-neural networks including a neural network architecture; Assigning the one or more sub-neural networks to a plurality of processing units of the body, wherein for each sub-neural network, the computer assigns a processing unit to compute the plurality of functions of the sub-neural network; as well as A plurality of execution instructions for the plurality of processing units of the self are generated by executing a plurality of compilers on the plurality of functions of each sub-neural network, wherein for each execution instruction, the computer uses a compiler from the plurality of compilers according to the processing unit of the self assigned to the plurality of functions of the sub-neural network.

12. The system of claim 11, wherein the computer is further configured to generate a computer file comprising the plurality of execution instructions for the plurality of processing units to execute the plurality of sub-neural networks.

13. The system of claim 11, wherein the computer is further configured to send the computer file to the self.

14. The system of claim 11, wherein at least one execution instruction causes the plurality of chips of the self to operate in an extended computing mode for parallel execution of the plurality of execution instructions.

15. The system of claim 11, wherein at least one execution instruction causes the plurality of chips of the self to operate in a redundant mode for primarily executing the plurality of execution instructions by a master chip of the plurality of chips.

16. The system of claim 11, wherein the computer is further configured to apply a schedule optimizer engine to the execution instructions to generate an execution schedule for the execution instructions, the schedule optimizer engine comprising a neural network layer trained to generate the instructions for minimizing latency.

17. The system of claim 11, wherein the plurality of processing units comprises at least one of a GPU, a CPU, or an accelerator device.

18. The system of claim 11, wherein the plurality of processing units are heterogeneous, comprising at least two types of processing units.

19. The system of claim 11, wherein the computer assigns the processing units of the plurality of processing units of the self to the sub-neural network based on one or more diagrams representing a circuit architecture of the self having the plurality of processing units.

20. The system of claim 11, wherein the computer is further configured to train the one or more sub-neural networks for quantization-aware training based on the one or more processing units by applying each sub-neural network to a training data set stored in a database, the training data set comprising data having expected quantization characteristics.