Door opening collision detection and avoidance
Patent Information
- Application Number
- CN202610288798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-17
- Filing Date
- 2026-03-10
- Publication Date
- 2026-09-18
Smart Images

Figure CN122773984A_ABST
Abstract
Description
[0001] Cross-reference to related applications This patent application claims priority to U.S. Provisional Patent Application No. 63 / 773,165, entitled “OPEN-DOOR DETECTION FOR SEMI-AUTONOMOUS AND AUTONOMOUS SYSTEMS AND APPLICATIONS,” filed on March 17, 2025, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0002] This disclosure generally relates to detection and / or prediction by autonomous or semi-autonomous vehicles, and in one aspect specifically relates to door opening collision detection and avoidance of such vehicles. Background Technology
[0003] The ability to detect and identify obstacles and objects during driving or navigation is a requirement for any semi-autonomous or autonomous system, such as vehicles, robots, warehouse machines, and / or other types of machines. For example, the ability to detect open doors of other vehicles or machines in the environment is crucial for ensuring the safety of the machine itself (or in the first person) and other vehicles or machines whose doors may be open. If an open door is not detected, the autonomous machine may be unable to navigate around it and / or slow down before reaching it, potentially reducing the overall safety of the system. Open door detection can focus on recognizing the vehicle or machine as a wider (e.g., a side door open) or longer (e.g., a trunk door open) conventional object detector and can update the occupancy grid or other representation methods to indicate the size of the space occupied by the vehicle or machine. This approach of incorporating open doors into the vehicle's dimensions can lead to inaccuracies in the vehicle's geometry. Furthermore, expanding the bounding box and perceived geometry of the vehicle or machine may overlook other contextual or useful information, such as changes in the probability of someone appearing before or after the door, or the possibility of someone leaving the vehicle through the open door. Given an increased probability of someone appearing now or soon, the autonomous machine may not be able to make appropriate corrections and take necessary precautions. Similarly, for an open rear door or trunk door, if the detected length of the vehicle or machine is greater than the actual length, the machine may not be able to infer the current situation or context, such as the vehicle or machine being parked and possibly remaining in that position for a period of time because its door is open (e.g., for loading or unloading cargo). Attached Figure Description
[0004] The present invention will be described in detail below with reference to the accompanying drawings, describing the system and method for detection and / or prediction of autonomous or semi-autonomous vehicles that support door opening collision detection and avoidance.
[0005] Figure 1A The illustrations depict systems for detection and / or prediction performed by autonomous or semi-autonomous vehicles to support door opening collision detection and avoidance in at least some embodiments.
[0006] Figure 1B The illustrations depict aspects of detection and / or prediction processes performed by autonomous or semi-autonomous vehicles to support door opening collision detection and avoidance in at least some embodiments.
[0007] Figure 2A The illustration shows bounding box details of detection and / or prediction performed by autonomous or semi-autonomous vehicles to support door opening collision detection and avoidance in at least some embodiments.
[0008] Figure 2B The illustration shows bounding box details of detection and / or prediction performed by autonomous or semi-autonomous vehicles to support door opening collision detection and avoidance in at least some embodiments.
[0009] Figure 3 The illustrations depict details related to artificial intelligence (AI) / machine learning (ML) for door opening collision detection and avoidance, according to at least some embodiments.
[0010] Figure 4 The illustration shows computer and processor aspects used in a system supporting door opening collision detection and avoidance according to at least some embodiments.
[0011] Figure 5 The illustration shows the process flow used by a system for door opening collision detection and avoidance according to at least some embodiments.
[0012] Figure 6 The illustration shows additional process flows used by a system for door opening collision detection and avoidance according to at least some embodiments.
[0013] Figure 7 The illustration shows the process flow of supporting door opening collision detection and avoidance using two-dimensional (2D) and three-dimensional (3D) information according to at least some embodiments.
[0014] Figure 8 The application is illustrated. Figure 1A-7 Example data center of at least one embodiment.
[0015] Figure 9 It is applicable Figure 1A-8 A block diagram of an example architecture of a computing system (e.g., a system-on-a-chip (SoC)) according to at least one embodiment of the present invention and according to at least some embodiments thereof.
[0016] Figure 10AThe illustrations depict example autonomous or semi-autonomous machines that benefit from door-opening collision detection and avoidance according to at least some embodiments of the present disclosure.
[0017] Figure 10B Examples of component and sensor locations on an autonomous or semi-autonomous vehicle according to at least some embodiments of the present disclosure are illustrated, the vehicle including door opening collision detection and avoidance.
[0018] Figure 10C This is a block diagram of an example system architecture for autonomous or semi-autonomous vehicles, robots, and / or other machine types, according to at least some embodiments of this disclosure.
[0019] Figure 10D This is a block diagram of an example architecture of a computing system (e.g., a system-on-a-chip (SoC)) according to at least some embodiments of the present disclosure.
[0020] Figure 10E This is a system diagram of communication between a cloud-based server and example autonomous or semi-autonomous vehicles, robots and / or other machine types, according to at least some embodiments of this disclosure.
[0021] Figure 11 This is a system diagram illustrating three computer ecosystems according to at least some embodiments of the present disclosure, including a computing system for generating or creating artificial intelligence (AI) (e.g., AI training and validation data), a computing system for training artificial intelligence, and a computing system for deploying AI at the edge.
[0022] Figure 12 This is a block diagram of an example computing system for generative artificial intelligence (AI) according to at least some embodiments of the present disclosure.
[0023] Figure 13 This is a block diagram of an example computing device according to at least some embodiments of the present disclosure. Detailed Implementation
[0024] To address door detection and obstacle avoidance issues in vehicles, proactive and explicit door opening detection (including trunk opening) can be implemented. This improves the system's overall safety relative to the autonomous machine and other vehicles within the autonomous machine's line of sight. It also allows the autonomous machine to more accurately reason about situations involving door opening detection and / or prediction. This detection and / or prediction can be performed by an autonomous or semi-autonomous vehicle acting as the autonomous machine. The detection and / or prediction can be performed on doors detected within the line of sight during the driving route and as part of a collision avoidance system.
[0025] In some examples, the detection and / or prediction can be performed by a dual machine learning (ML) system. The first ML model can be a predictor, and the second ML model can be a classifier. The first ML model can use images from a two-dimensional (2D) sensor (such as one of the vehicle's cameras) as part of the 2D information to predict the opening of the doors of at least one vehicle in the driving route. This prediction can be based in part on the dimensions of the bounding box associated with the vehicle captured by the 2D sensor. Therefore, the first ML model can include training using bounding boxes of vehicles with one or more doors closed or open. After the prediction is complete, a portion of the image can be applied to the second ML classifier to explicitly classify the door opening prediction for a specific type, such as door type (which could be a left door, right door, rear door, or trunk lid, etc.). In some examples, this portion of the image can have extended dimensions relative to the captured bounding box and relative to the trained bounding box. In some examples, this portion of the image can be located at a predetermined width from the edge of the captured bounding box. Furthermore, the second ML model can provide a label or annotation to this portion of the image after classification. This label or annotation can indicate the type of door based on the classification. This label or annotation can be used by the 2D-to-3D conversion subsystem to map a specific open door on a display representation of an autonomous or semi-autonomous vehicle with an open door. Furthermore, the driving subsystem of the autonomous or semi-autonomous vehicle can indicate a state based on this label or annotation, or perform or suggest a response to an open door. In the example, an open door might indicate that the vehicle has stopped and cannot move. The response could be an autonomous or semi-autonomous vehicle navigating around the vehicle, which is part of the open door collision detection and avoidance performed by the autonomous or semi-autonomous vehicle.
[0026] In some examples, 2D sensors may include, in addition to cameras, one or more LiDAR, RADAR, ultrasonic sensors, or similar devices. The 2D-3D conversion subsystem can provide a graphical representation of the vehicle's surroundings as a top-down or bird's-eye view (BEV) of the environment, and indicate the opening of a specific door based on its type. This type might be an explicit indication or signal that a door is open, allowing vehicles on the path to avoid guesswork based solely on geometry. Given the vehicle's known state, an autonomous or semi-autonomous vehicle can act as part of a driving subsystem, capable of planning and controlling actions or reactions such as maneuvering (as a door opening might indicate a long wait behind the vehicle—for unloading), decelerating, stopping, maintaining a greater safe distance or gap from other vehicles, anticipating potential personnel leaving the vehicle, or making slight adjustments to the left or right to safely maneuver around the vehicle.
[0027] In some examples, explicit detection of a target vehicle's door or trunk opening for semi-autonomous and autonomous systems and applications can include explicit prediction of the target vehicle's or machine's trunk or door opening status, for example, using 2D and / or 3D object detectors or sensors. These 2D and / or 3D object detectors or sensors can include cameras, LiDAR, RADAR, ultrasonic sensors, etc. Once the trunk or door opening of the target vehicle or machine is detected (e.g., using 2D sensors or object detectors, which may include a front-facing camera of the machine), a 2D-to-3D conversion can be performed. This 2D-to-3D conversion can be used to map the trunk or door opening information into a 3D space. In some examples, this 3D space can include a top-down (view-down) or BEV representation of the environment. Therefore, the systems and methods proposed herein can facilitate explicit determination of the "trunk open" or "door open" signal (as a status) for a specific vehicle or machine, without having to guess this status based solely on geometry. Once the state or situation is known, the autonomous machine's planning and control system (also referred to as the driving subsystem in this paper) can react more rationally to the target vehicle or machine and understand whether anyone might be present or about to appear. This understanding allows the autonomous machine to reason better, such as avoiding a vehicle unloading cargo from its trunk, slowing down or stopping when a side door opens onto the street to allow for pedestrians, making slight adjustments to the left or right to safely avoid vehicles or machines stopped for delivery or other purposes, and so on. This better reasoning enables the autonomous machine to plan and control more safely, resulting in a more reliable and powerful autonomous machine.
[0028] Figure 1AThe illustration depicts a system 100A for detection and / or prediction by an autonomous or semi-autonomous vehicle supporting door opening collision detection and avoidance, in at least some embodiments. System 100A may include at least one processor 102. Processor 102 may be located within or remotely from the autonomous or semi-autonomous vehicle 106 and provide input to it. Processor 102 may execute instructions from memory associated with the autonomous or semi-autonomous vehicle 106. The processor 102 executing the instructions may be assigned to perform a specific function. For example, processor 102 may be used to predict an open door 104A of vehicle 104. This prediction 112A may occur within the autonomous or semi-autonomous vehicle 106. The autonomous or semi-autonomous vehicle 106 may use prediction 112A to perform an action or reaction 108 (e.g., avoiding vehicle 104). While avoidance is shown in the figure, the action or reaction 108 may be deceleration, stopping, maintaining a greater safe distance or gap from the vehicle, considering the possibility of people disembarking, or making slight adjustments to the left or right to safely avoid the vehicle. The prediction may be based in part on the bounding box of the representation 110 applied to vehicle 104. For example, the representation 110 may be used in a first ML model 112, which may be trained using bounding boxes 110A for different vehicles with one or more closed or open doors.
[0029] Figure 1A As illustrated, in addition to prediction 112A, classification 114A can also be performed to classify the open door 104A into a specific type, based in part on at least a portion of representation 110, which is applied to a second ML model 114 trained using categories 114B for different types of open doors for different vehicles. Furthermore, processor 102 can be configured (e.g., by implementing or executing instructions) to provide representation 110 in the form of 2D information using 2D sensor 106A. Processor 102 can be configured to use a 2D-to-3D conversion subsystem (such as...) Figure 1A The 2D-3D association 116 shown in the diagram illustrates, in part, the open door 104A of vehicle 104 in 3D information form (e.g., on the dashboard of autonomous or semi-autonomous vehicle 106) based on the output of the second ML model 114. In some examples, the 2D-to-3D conversion subsystem may issue instructions or react 118 via the diagram 120 or via the driving subsystem 122. In some examples, the processor 102 may perform this reaction as part of the driving subsystem 122 of the autonomous or semi-autonomous vehicle 106 (e.g., located within, contained within, or supporting it), rather than as a supplement to the diagram 120. In some examples, the processor 102 may be configured to perform or suggest actions or reactions 108 to the open door 104A to the driver via the dashboard or other feedback mechanisms (in the case of a semi-autonomous vehicle).
[0030] Figure 1A The illustration also shows that processor 102 may be part of a 2D-to-3D conversion subsystem, and may be used to generate or support a top-down or BEV representation of illustration 120, which may be an open-door signal of an open door 104A of vehicle 104. In some examples, 2D sensor 106A is an onboard camera of autonomous or semi-autonomous vehicle 106. In some examples, the specific type supported by the classification 114A may be one of left passenger door open, right passenger door open, left driver's door open, right driver's door open, left door open, right door open, rear door open, sunroof or panoramic sunroof open, or hood open.
[0031] Therefore, in some examples, processor 102 can be one or more processors. One or more processors can be used to train an ML model (e.g., a second ML model 114) to determine a specific type of open door in vehicle 104 using different categories 114B associated with different types of open door situations for different vehicles. Furthermore, one or more processors can also be used to train an additional ML model (e.g., a first ML model 112) to predict the open door situation of vehicles, partly based on different bounding boxes of different vehicles that may have one or more closed or open doors, as combined... Figure 2A and Figure 2B As described. Specific types can include left passenger door opening, right passenger door opening, left driver's door opening, right driver's door opening, left door opening, right door opening, rear door opening, sunroof or panoramic sunroof opening, or hood opening. One or more processors can be used or configured to determine specific types of door opening situations for vehicle 104 based in part on an ML model (e.g., a second ML model 114), which is trained using different categories associated with different types of door opening for different vehicles and used after a prediction of potential door opening 104A for vehicle 104 is made using a first ML model.
[0032] Figure 1B The illustration depicts a processing aspect 100B for detection and / or prediction by an autonomous or semi-autonomous vehicle supporting door opening collision detection and avoidance, in at least some embodiments. As shown in processing aspect 100B, for the autonomous or semi-autonomous vehicle 106, there may be multiple 2D sensors 152A-152X. Each sensor may separately capture different views 150A-150X for use by the processor 102. Figure 1AAs shown, different views 150A-150X captured by multiple 2D sensors 152A-152X can be applied to the first ML model 112 and the second ML model 114. In some examples, views 150A-150X may have different zoom levels, come from different points in time, or come from different positions relative to vehicle 104, and are strictly from the point of view (POV) of the autonomous or semi-autonomous vehicle 106.
[0033] The first ML model 112 can provide predictions 112A of potential open doors 104A, and the second ML model 114 can provide classifications 114A of explicit open door situations, such as at least combining Figure 1A In some examples, a prediction 112A of a potential open door 104A can be output as an open door condition attribute, and a further classification 114A of the open door condition attribute can be output as an indication of trunk open 156A or left door open 156B, superimposed on one or more different views 150A-150X. In some examples, the trunk open 156A indication or the left door open 156B indication may be a definite open door condition signal and may not be superimposed or visible before performing 2D-to-3D association. In some examples, the trunk open 156A indication or the left door open 156B indication can be displayed in a 2D representation. Therefore, the trunk open 156A indication or the left door open 156B indication can be displayed for different 2D sensors 152A-152X and can be displayed in a perspective view.
[0034] A door opening signal can be assigned to vehicle 104 in the BEV via 2D-3D associations 116 established for each of the different views 150A-150X. Once the assignment is complete, processor 102 (e.g., downstream planners and controllers or driving subsystem 122) can trigger an action or reaction 108, such as a nudge or deceleration of the autonomous or semi-autonomous vehicle 106. In some examples, the action or reaction 108 can be supported or recorded via illustration 120A (e.g., via the dashboard of the autonomous or semi-autonomous vehicle 106) or within system 100A. Illustration 120A may include the driving route of the autonomous or semi-autonomous vehicle 106 and may include other vehicles and the environment in the route, such as... Figure 1B As shown.
[0035] Therefore, the systems and methods described herein can utilize one or more object detectors (e.g., ML models, neural networks, computer vision algorithms, etc. as described herein) that can be trained to predict open or closed door states or conditions, thereby identifying the opening status of left, right, rear, and / or other door / window locations of an agent vehicle or machine in the environment. To this end, the systems and methods described herein can predict open-open attributes and classify them more precisely; for example, in some examples, multiple cameras (e.g., 2D sensors 152A-152X) in perspective views 150A-150X can be used to identify “left door open” or “right door open”. Once a condition is detected, the predicted “door open” attribute can be assigned to a target vehicle or machine (e.g., vehicle 104) in a top-down or BEV representation (e.g., via 2D-3D association 116). Although the primary description focuses on converting 2D detection results to a 3D representation (e.g., BEV), this is not a limitation, and in other examples herein, open door condition signals in perspective views and / or other views 150A-150X can be used. Similarly, while 2D-to-3D detection has been discussed, this is not a limitation. In some examples, door opening conditions can be detected directly in 3D space using 3D information, for example, by using depth sensors (including LiDAR, RADAR, stereo cameras, etc.) and / or by training neural networks or ML models to directly predict 3D spatial information as output from 2D information as input. Once a specific vehicle or machine is assigned the "door opening condition" attribute or signal, downstream planners and controllers can trigger plans or controls to gently nudge, decelerate, evasive maneuvers, and / or otherwise safely respond to the situation.
[0036] Figure 2A The illustration shows bounding box details 200A for detection and / or prediction by an autonomous or semi-autonomous vehicle supporting door opening collision detection and avoidance in at least some embodiments. Bounding box detail 200A indicates that, in the autonomous or semi-autonomous vehicle 106 (or machine application), when a door is opened by the agent vehicle or machine (e.g., when a door is opened...), Figure 1A-2A As shown in vehicle 104, direct 3D detection may not accurately predict the door size, which can be used to combine with... Figure 1A and Figure 1B The system and method described herein can use 2D sensors 106A, 152A-152X (e.g., a camera that provides input to processor 102, such as...). Figure 1AAs shown, the detector explicitly predicts the "door open status" state. Specifically, in some examples, for a given agent vehicle 104 or machine, its slave machine (which may be autonomous or semi-autonomous vehicle 106) can predict or generate signals for the left door open attribute ("left door open"), right door open attribute ("right door open"), and rear door open attribute ("rear door open") to indicate different door open / close states of the agent vehicle 104 or machine.
[0037] In some examples, once a prediction 112A of a door-opening condition 104A of vehicle 104 is made in part based on the bounding box 202 of a representation 110 applied to vehicle 104 (which is applied to a first ML model 112), a classification 114A can be performed on the door-opening condition 104A to categorize it into a specific type. This may be in part based on at least one portion of the representation 110 applied to the second ML model 114. This at least one portion of the representation may be a partial bounding box in the representation 110 that has an extended dimension 206 relative to a bounding box 210 that can be used in the first ML model (e.g., a bounding box 210 used for training in the first ML model). In some examples, this at least one portion may be a predetermined portion of the representation 110 (e.g., a predetermined width 208) determined from the edge 204 of at least one dimension (e.g., width) of the bounding box 202. Therefore, once the first ML model has taken into account the entire bounding box 202, the second ML model can be directed to at least one portion of the representation 110 relative to the bounding box 202 (e.g., a portion of the bounding box 202 or a dimension derived therefrom) to determine the left-opening property.
[0038] In some examples, the predetermined portion may be a cropped portion of representation 110. In some examples, other image features besides the cropped portion (e.g., patterns from representation 110) may be applied to the second ML model. In some examples, using the first ML model first, followed by the second ML model, may represent a coarse prediction of representation 110 of vehicle 104 followed by fine-tuned classification. In some examples, boundary criteria may be used to automatically determine edges 204. For example, a third ML model trained using boundary images may be used to identify edges 204. The determination or recognition of edges in representation 110 of vehicle 104 can be used to determine the predetermined portion. For example, the determination of edge 204 can be used to expand and extract the predetermined portion and feed it to the second ML model for classification.
[0039] Figure 2B The illustration shows further bounding box details 200B for detection and / or prediction performed by an autonomous or semi-autonomous vehicle supporting door opening collision detection and avoidance in at least some embodiments. In some examples, vehicle 104 (also as Figure 1A and Figure 1BThe vehicle (as shown) may be parked in front of a self-driving machine (e.g., autonomous or semi-autonomous vehicle 106) with its trunk open, as shown in more bounding box details 200B. This can be a challenging situation because the self-driving machine, if using a traditional method that only adjusts the bounding box size for the open door, may not be able to determine whether it should wait behind the vehicle or go around it. By explicitly detecting the "rear door open" condition, the self-driving machine can understand the context of the situation and determine that the vehicle is parked, thus making it safe to go around the vehicle the best option.
[0040] and Figure 2A Similarly, once a prediction 112A of the door opening condition 104A of vehicle 104 is made in part based on the bounding box 252 of the representation 110A applied to vehicle 104 (which is used in conjunction with the first ML model 112), a classification 114A can be performed on the door opening condition 104A to categorize it into a specific type. This classification may be in part based on at least one portion of the representation 110 applied to the second ML model 114. The at least one portion of the representation may be a portion of the representation 110 that has an extended dimension 206 relative to the bounding box 210 that can be in the first ML model (e.g., the bounding box 210 used for training the first ML model). In some examples, the at least one portion may be a predetermined portion (e.g., a predetermined width 208) of the representation 110, which is determined by the edge 204 of at least one dimension (e.g., width) of the bounding box 202. Therefore, once the first ML model has taken into account the entire bounding box 202, the second ML model can be directed to at least one part of the representation 110, which can be obtained relative to the bounding box 202 (e.g., a portion of the bounding box or a dimension derived therefrom), to determine the rear door open attribute (“trunk open” or “rear door open”).
[0041] The described systems and methods ensure the safety of a self-operated machine (e.g., autonomous or semi-autonomous vehicle 106) when passing a vehicle or machine with its side doors, rear doors, and / or other types of doors 104A open. Detecting or predicting the opening state or condition of a door or trunk is crucial for determining the appropriate response of the self-operated machine. For example, when a vehicle's side door is open, a signal may be generated for the self-operated machine to encourage a gentle push to avoid a collision. Additionally or separately, this signal may also trigger a warning indicating that someone may be outside the vehicle or may soon be outside the vehicle when the determination (detection or prediction) of the door opening is made. In some cases, the self-operated machine may slow down to pass through vehicles with these doors open. Similarly, when a vehicle's trunk door is open, this may indicate that the vehicle will remain stationary for a period of time. When vehicle 104 obstructs the self-operated machine's path of travel, knowing that it is stationary and parked, the self-operated machine may attempt to go around vehicle 104—whereby simply increasing the vehicle's geometry (e.g., the bounding box in the representation) based on the trunk being open is insufficient to determine that vehicle 104 is stationary and parked. An open trunk can serve as a strong signal to trigger detours around such vehicles. Furthermore, the system and method described herein can also be used to identify target vehicles 104 parked in the middle of the street, which the machine may decide to detour based on the vehicle's parking purpose (e.g., unloading passengers or packages, if it is a delivery truck). Open doors (whether side or rear) strongly indicate that vehicle 104 is unlikely to move, and the system can utilize this information to improve the accuracy of "parked" status.
[0042] Figure 3 The illustration shows details 300 related to artificial intelligence (AI) / machine learning (ML) for door opening collision detection and avoidance in at least some embodiments. The AI / ML aspects may be supported or implemented in ML models 310A, 310B. ML models 310A, 310B may reside in ML subsystem 308, which may be independent of... Figure 1A The system 310A or system 310B can be located within the system 100A. ML models 310A and 310B can be trained using data from the data repository 302 and can be trained within the safe zone 306 of the ML subsystem 308, thereby making them independent of or isolated from the inference side of the ML models 310A and 310B. In some examples, the safe zone 306 also ensures that the ML models 310A and 310B are unbiased and uncontaminated / not poisoned.
[0043] ML models 310A and 310B can be trained using different data components. For example, the first ML model can be ML model 310A, which can be trained using first data 312A generated from the first bounding box 304A. The second ML model can be ML model 310B, which can be trained using second data 312B generated from the second bounding box 304B. The first bounding box 304A can surround one or more of the closed or open doors of different vehicles and can include historical or simulated data of vehicles with open doors. In this way, the first ML model can generate outputs, such as prediction output 314, to represent the prediction 112A of the vehicle's door opening status (e.g., Figure 1A (As shown). The second data 312B can be generated from the second bounding boxes 304B, which are different bounding boxes specifically set for different types of opening doors of different vehicles. The second bounding boxes 304B may also include historical or simulation data limited to the identified different types of opening doors. In this way, the second ML model can generate outputs, such as classification output 316, which represents the classification 114B of the opening door situation (e.g.). Figure 1A As shown), it is a specific type for vehicles from different types.
[0044] In some examples, a 2D sensor can be used to provide a representation of two-dimensional information 110 (e.g. Figure 1A (As shown). The depth sensor can be used in conjunction with 2D information representation to generate at least a portion of the representation in the form of 3D information. The 2D information can be provided as input to a second ML model, such as a 2D-to-3D information classifier 318. The 2D-to-3D information classifier 318 can be trained using both 2D and 3D information to classify open door situations in 3D space without requiring a 2D-to-3D assignment 316A from the classification output 316 (e.g., in...). Figure 1A (Used in 2D-3D association 116).
[0045] Figure 4 The illustration depicts computer and processor aspects used in a system supporting door opening collision detection and avoidance according to at least some embodiments. According to at least one embodiment, the computer and processor aspect 400 can be executed by one or more processors, including… Figure 9 The SoC or combination thereof shown may include an execution unit for executing instructions. One or more such processors may include a central processing unit (CPU), a data processing unit (DPU), and a graphics processing unit (GPU), and may be coupled with a processor 102 (such as...) Figure 1A (As shown) is associated with one or more processors that execute.
[0046] In at least one embodiment, the computer and processor aspect 400 may include, but is not limited to, computing components, such as processor 402, which employs an execution unit or processing unit, including logic for performing data processing algorithms, as described in this disclosure, such as in the embodiments described herein. In at least one embodiment, the computer and processor aspect 400 may include processors such as the Pentium® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors provided by Intel Corporation (Santa Clara, California), although other systems (including PCs, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, the computer and processor aspect 400 may run a version of the Windows® operating system provided by Microsoft Corporation (Redmond, Washington), although other operating systems (such as UNIX® and Linux®), embedded software, and / or graphical user interfaces may also be used.
[0047] These embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip, a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0048] In at least one embodiment, the computer and processor aspect 400 may include, but is not limited to, a processor 402, which may include, but is not limited to, one or more execution units 408, or in some examples, one or more processors 102, to be configured according to the description herein. Figure 1A-3 and Figure 5-9 The technologies described in at least one or more figures are used to perform the various aspects. In some examples, processor 402 may include multiple threads that act as individual processors. In some examples, there may be multiple processors 402, each acting as a processor. The computer and processor aspect 400 may be part of a host in a data center, and in some examples, may be located remotely from autonomous or semi-autonomous vehicle 106. In at least one embodiment, the computer and processor aspect 400 is a single-processor desktop or server system, but in another embodiment, the computer and processor aspect 400 may be a multi-processor system.
[0049] In at least one embodiment, processor 402 may include, but is not limited to: a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computer (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combinations, or any other processor device, such as, for example, a digital signal processor. In at least one embodiment, processor 402 may be coupled to processor bus 410, which enables the transmission of data signals between processor 402 and other components in the computer and processor aspect 400.
[0050] In at least one embodiment, processor 402 may include (but is not limited to) a level-one (“L1”) internal cache memory (“cache”) 404. In at least one embodiment, processor 402 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, cache 404 may be located external to processor 402. Other embodiments may also include a combination of internal and external caches, depending on specific implementation and requirements. In at least one embodiment, register file 406 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0051] In at least one embodiment, execution unit 408 (including, but not limited to, logic for performing integer and floating-point operations) is also located in processor 402. In at least one embodiment, processor 402 may further include microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, execution unit 408 may include logic for processing packaged instruction set 409.
[0052] In at least one embodiment, by including a packaged instruction set 409 and associated circuitry for executing instructions in the instruction set of a general-purpose processor, operations used by many multimedia applications can be performed using packaged data in processor 402. In at least one embodiment, by using the full width of the processor data bus to perform operations on packaged data, many multimedia applications can be accelerated and their execution efficiency improved, thereby avoiding the need to transfer smaller data units across the processor data bus to perform one or more operations on one data element at a time.
[0053] In at least one embodiment, execution unit 408 may also be used as a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, computer and processor aspect 400 may include, but is not limited to, memory 420. In at least one embodiment, memory 420 may be dynamic random access memory (“DRAM”), static random access memory (“SRAM”), flash memory, or other memory devices. In at least one embodiment, memory 420 may store instructions 419 and / or data 421 represented by data signals, which may be executed by processor 402.
[0054] In at least one embodiment, the system logic chip may be coupled to the processor bus 410 and the memory 420. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller center (“MCH”) 416, and the processor 402 may communicate with the MCH 416 via the processor bus 410. In at least one embodiment, the MCH 416 may provide a high-bandwidth memory path 418 to the memory 420 for storing instructions and data, as well as storing graphics commands, data, and textures. In at least one embodiment, the MCH 416 may direct data signals between the processor 402, the memory 420, and other components in the computer and processor aspect 400, and bridge data signals between the processor bus 410, the memory 420, and the system I / O interface 422. In at least one embodiment, the system logic chip may provide a graphics port to be coupled to a graphics controller. In at least one embodiment, the MCH 416 may be coupled to the memory 420 via the high-bandwidth memory path 418, and the graphics / video card 412 may be coupled to the MCH 416 via an Accelerated Graphics Port (“AGP”) interconnect 414.
[0055] In at least one embodiment, the computer and processor aspect 400 may use system I / O interface 422 as a dedicated hub interface bus to couple MCH 416 to I / O controller center (“ICH”) 430. In at least one embodiment, ICH 430 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 420, chipset, and processor 402. For example, it may include, but is not limited to, an audio controller 429, a firmware hub (“Flash BIOS”) 428, a wireless transceiver 426, a data storage device 424, a conventional I / O controller 423 including user input and keyboard interface 425, a serial expansion port 427 (e.g., a Universal Serial Bus (“USB”) port), and a network controller 434. In at least one embodiment, data storage device 424 may include a hard disk drive, floppy disk drive, CD-ROM device, flash memory device, or other mass storage device.
[0056] In at least one embodiment, Figure 4 The illustration depicts a computer and processor aspect 400, including interconnected hardware devices or "chips"; while in other embodiments, Figure 4 An exemplary SoC can be illustrated. In at least one embodiment, Figure 4 The devices shown can be interconnected via proprietary interconnects, standardized interconnects (such as PCIe®), or some combination thereof. In at least one embodiment, one or more components of the computer and processor aspect 400 are interconnected using a Compute Fast Link (CXL) interconnect.
[0057] According to at least some embodiments, Figure 5The illustration depicts a process flow or method 500 used by a system for door opening collision detection and avoidance. Method 500 may include the step of applying a bounding box 502 to a representation of a vehicle. Method 500 may include the step of providing a first machine learning (ML) model 504, which is trained using bounding boxes for different vehicles having one or more closed or open doors. Providing 504 may include allowing access to the first ML model (e.g., allowing input to the first ML model). Method 500 may include the step of predicting 506 a door opening condition of the vehicle, partially based on the bounding boxes applied to the representation of the vehicle (used in the first ML model). Method 500 may include the step of providing a second ML model 508, which has been trained using categories for different types of door opening conditions for different vehicles. Providing 508 may include allowing access to the second ML model (e.g., allowing input to the second ML model). Method 500 may include the step of classifying 510 a door opening condition into a specific type, partially based on at least one portion of the representation used for the second ML model.
[0058] Figure 6 The illustration shows a further process flow or method 600 used by a system for door opening collision detection and avoidance according to at least some embodiments. Method 600 can be used to support Figure 5 Method 500. For example, method 600 may include the step of generating 602 first data from bounding boxes surrounding one or more of the closed or open doors of different vehicles. Method 600 may include the step of training 604 a first ML model using the first data. The first ML model may generate output representing predictions of the open door situation of the vehicle to support step 506. Method 600 may include the step of generating 606 second data from bounding boxes specifically surrounding different types of open doors for different vehicles. Method 600 may include the step of training 608 a second ML model using the second data. The second ML model may generate output representing vehicle-specific classification of open doors into different types to support step 510.
[0059] Figure 7 The illustration depicts a process flow or method 700 that utilizes 2D and 3D information to support door opening collision detection and avoidance according to at least some embodiments. Method 700 may also support... Figure 6 Method 600 or Figure 5Method 500. For example, method 700 may include the step of providing a representation of 2D information to 702 using a 2D sensor. Method 700 may include using a 2D-to-3D conversion subsystem 704A, based in part on the output of a second ML model, to plot the vehicle's door opening status in the form of three-dimensional information. Method 700 may include using a depth sensor 704B in conjunction with 2D information to generate at least a portion of a representation in the form of 3D information, which serves as input to a second ML model trained using both 2D and 3D information to classify door opening status in 3D space. Steps 704A and 704B may be performed as two different methods, or as a verification of one method. Method 700 may include the step of generating or supporting a top-view or bird's-eye view (BEV) illustration of the vehicle's door opening status 706. Method 700 may include a step of performing or suggesting a response 708 to the door opening status.
[0060] Figure 8 The application is illustrated. Figure 1A-7 An example data center 800 is described in at least one embodiment. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840. The example data center 800 can be part of a multi-tenant environment, or support or allow a multi-tenant environment to support... Figure 1A-7 The data center 800 may use its included processors or their processing threads (e.g., CPUs, GPUs, DPUs, etc., also referred to as node computing resources (nodes CR816(1)-816(N)), which are described relative to the data center infrastructure layer 810) to support the system 100A. Furthermore, in at least one aspect, the data center 800 may include at least some computing components capable of performing the ML algorithms described herein.
[0061] In at least one embodiment, such as Figure 8As shown, the data center infrastructure layer 810 may include a resource coordinator 818, packet computing resources 814, and node computing resources (“nodes CR”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, nodes CR 816(1)-816(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (“FPGAs”), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 816(1)-816(N) may be servers having one or more of the aforementioned computing resources.
[0062] In at least one embodiment, the packet computing resource 814 may include multiple node CR packets housed within one or more racks (not shown), or multiple racks housed within data centers (not shown) in different geographical locations. Individual node CR packets within the packet computing resource 814 may include packet computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination thereof.
[0063] In at least one embodiment, resource coordinator 818 may be configured or otherwise control one or more nodes CR 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a software design infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource coordinator 812 may include hardware, software, or some combination thereof.
[0064] In at least one embodiment, such as Figure 8As shown, framework layer 820 includes job scheduler 822, configuration manager 824, resource manager 826, and distributed file system 828. In at least one embodiment, framework layer 820 may include a framework of software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. In at least one embodiment, software 832 or application 842 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 820 may be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize distributed file system 828 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 822 may include Spark drivers to facilitate scheduling workloads supported by the various layers of data center 800. In at least one embodiment, configuration manager 824 may be able to configure different layers, such as software layer 830 and framework layer 820, including Spark and distributed file system 828, to support large-scale data processing. In at least one embodiment, resource manager 826 may be able to manage cluster or group computing resources mapped to or allocated to support distributed file system 828 and job scheduler 822. In at least one embodiment, cluster or group computing resources may include group computing resources 814 at data center infrastructure layer 810. In at least one embodiment, resource manager 826 may coordinate with resource coordinator 818 to manage these mapped or allocated computing resources.
[0065] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes CR 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming media content software.
[0066] In at least one embodiment, the application 842 included in the application layer 840 may include one or more types of applications available for use by at least a portion of the nodes CR 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of applications may include, but are not limited to, genomics applications, cognitive computing applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0067] In at least one embodiment, any of the configuration manager 824, resource manager 826, and resource coordinator 818 can implement any number and type of self-modification operations based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification operations can enable the data center operator of data center 800 to avoid making potentially erroneous configuration decisions and may prevent underutilization and / or poor performance of parts of the data center.
[0068] In at least one embodiment, data center 800 may include tools, services, software, or other resources for training one or more machine learning models, or for using one or more machine learning models to predict or infer information, according to one or more embodiments described herein. For example, in at least one embodiment, the software and computing resources described above for data center 800 may be used to train machine learning models by calculating weight parameters based on a neural network architecture. In at least one embodiment, the resources described above for data center 800 may be used, and weight parameters calculated using one or more training techniques described herein may be used to infer or predict information using trained machine learning models corresponding to one or more neural networks.
[0069] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as services to allow users to train or perform information inference, such as data recognition, speech recognition, or other artificial intelligence services.
[0070] Figure 9 A computing system 900 according to at least some embodiments of the present disclosure (about Figure 1A-8A block diagram of an example architecture (a subset of the system described herein). Although the figure shows SoC 904, this is not a limitation, and the computing system may also include or alternatively include multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-package (SiP), heterogeneous integration (HI), single-board computers (SBCs), and / or other components and / or architectures without departing from the scope of this disclosure. SoC 904 may represent or support the processor 102 described herein.
[0071] In some embodiments, such as when the SoC 904 includes a GPU 908 with 2000 or more cores (e.g., 2048 cores), 60 or more tensor cores (e.g., 64 tensor cores) and a maximum GPU frequency exceeding 1 GHz (e.g., 1.3 GHz), a CPU 906 with 10 or more cores (e.g., 12 cores), 64-bit, 3 MB L2 and 6 MB L3 cache memory and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz), one or more deep learning accelerators (DLAs); deep learning accelerator clusters (XNNs), neural network accelerators (NNAs) or neural processing units (NPUs) 909 (e.g., 2 DLAs / XNNs / NNAs / NPUs 909s), and an accelerator (ACCLS) 907, a single SoC 904 may achieve AI performance of 275 trillion operations per second (TOPS). For example, NVIDIA's Jetson AGX Orin 64GB SoC meets these criteria and achieves this performance.
[0072] Similarly, in the following embodiments, the SoC 904 includes a GPU 908 with 1700 or more cores (e.g., 1792 cores), 50 or more tensor cores (e.g., 56 tensor cores) and a maximum GPU frequency exceeding 900MHz (e.g., 930MHz), a CPU 906 with 8 or more cores (e.g., 8 cores), 64-bit, 2MB L2 and 4MB L3 cache memory and a maximum frequency of 2GHz or more (e.g., 2.2GHz), one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs) or neural processing units (NPUs) 909 (e.g., 2 DLAs / XNNs / NNAs / NPUs 909), and an accelerator (e.g., a programmable accelerator 907). A single SoC 904 may have an AI performance of 200 trillion operations per second (TOPS). For example, NVIDIA's Jetson AGX Orin 32GB SoC meets these guidelines and achieves this performance.
[0073] In some embodiments, for example, when the SoC 904 includes a GPU 908 having 1,000 or more cores (e.g., 1,024 cores), 28 or more tensor cores (e.g., 32 tensor cores) and a maximum GPU frequency exceeding 900 MHz (e.g., 1,173 MHz), a CPU 906 having 8 or more cores (e.g., 8 cores), 64-bit, 2 MB L2 cache and 4 MB L3 cache memory and a maximum frequency of 2 GHz or higher (e.g., 2 GHz), one or more deep learning accelerators (DLA), deep learning accelerator clusters (XNN), neural network accelerators (NNA) or neural processing units (NPU) 909 (e.g., 1 DLA / XNN / NNA / NPU 909), and an accelerator (e.g., a programmable accelerator 907), a single SoC 904 may have an AI performance of 157 trillion operations per second (TOPS). For example, NVIDIA's Jetson AGX Orin NX16GB SoC meets these guidelines and achieves this performance.
[0074] In various embodiments, for example, in the case of a SoC 904 including a GPU 908 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 32 tensor cores) and a maximum GPU frequency exceeding 900MHz (e.g., 1020MHz), and a CPU 906 including 6 or more cores (e.g., 6 cores), 64-bit, 1.5MB L2 and 4MB L3 cache memory, and a maximum frequency of 1.5GHz or higher (e.g., 1.7GHz), a single SoC 904 can achieve AI performance of 67 trillion operations per second (TOPS). For example, NVIDIA's Jetson Orin Nano 8GB SoC meets these criteria and achieves this performance.
[0075] SoC 904 may include one or more CPUs 906. In various embodiments, CPU 906 may include CPU clusters or CPU complexes (hereinafter referred to as "CCPLEX"). CPU 906 may include multiple cores and / or (e.g., L2, L3) caches. For example, in some embodiments, CPU 906 may include twelve cores arranged in a coherent multiprocessor configuration. In some embodiments, CPU 906 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 3MB L2 cache). CPU 906 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, thereby allowing any combination of CPU 906 clusters to be active at any given time.
[0076] The SoC 904 may include any type and number of GPUs 908. For example, in some embodiments, an integrated GPU (referred to herein as an "iGPU") may be used. The GPU 908 may be programmable and capable of efficiently handling parallel workloads. In some examples, the GPU 908 may use an enhanced tensor instruction set. The GPU 908 may include one or more streaming microprocessors, each of which may include a cache (e.g., an L1 cache with a storage capacity of at least 96KB), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with a storage capacity of 512KB). In some embodiments, the GPU 908 may include at least eight streaming microprocessors. The GPU 908 may use a computing application programming interface (API). Furthermore, the GPU 908 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0077] The GPU 908 can be power-optimized for automotive, robotics, and / or other embedded applications to achieve optimal performance. For example, the GPU 908 can be manufactured using a FinFET process. However, this is not intended to limit it; the GPU 908 can also be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain several mixed-precision processing cores, which are divided into multiple modules. For example (but not limited to), 64 PF32 cores and 32 PF64 cores can be divided into four processing modules. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix operations, an instruction cache (e.g., L0), a thread bundle scheduler, a dispatch unit, and / or a register file (e.g., 64KB). Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to efficiently execute workloads that involve a mixture of computation and addressing operations. Streaming microprocessors can include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors may include combined L1 data caches and shared memory units to improve performance while simplifying programming.
[0078] The GPU 908 may include high-bandwidth memory (HBM) and / or (e.g., 16GB) HBM2 memory subsystems to provide peak memory bandwidth of approximately 900GB / s in some examples. In some examples, in addition to HBM memory, or as an alternative to HBM memory, synchronous graphics random access memory (SGRAM), such as graphics double data rate type 5 synchronous random access memory (GDDR5), may be used.
[0079] The GPU 908 may include unified memory technology, including access counters, to more accurately migrate memory pages to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support may be used, enabling the GPU 908 to directly access the CPU 906's page tables. In these examples, when a memory management unit (MMU) miss occurs in the GPU 908, an address translation request can be sent to the CPU 906. The CPU 906 responds to this request, looks up the virtual-to-physical mapping for that address in its page tables, and sends the translation result back to the GPU 908. Therefore, unified memory technology can provide a single, unified virtual address space for the memory of both the CPU 906 and the GPU 908, simplifying GPU 908 programming and application porting to the GPU 908.
[0080] SoC 904 may include any number of caches 912, including the caches described herein. For example, cache 912 may include L0 cache, L1 cache, L2 cache, L3 cache (e.g., a cache usable by both CPU 906 and GPU 908, e.g., a cache connected to both CPU 906 and GPU 908), etc. Cache 912 may include a write-back cache that can track the state of cache lines, for example, by using one or more cache coherence protocols (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, (e.g., L3) cache may contain 4MB or more, but smaller or larger cache sizes may also be used.
[0081] The SoC 904 may include one or more Arithmetic Logic Units (ALUs) 965, which can be used to perform processing for any of a variety of tasks or operations. Additionally, the SoC 904 may include one or more Floating Point Units (FPUs) 967, or other types of mathematical or numerical coprocessors, for performing mathematical operations within the system. For example, the SoC 904 may include one or more FPUs 967, which are integrated as execution units within the CPU 906 and / or GPU 908.
[0082] The SoC 904 may include one or more accelerators (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, the SoC 904 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. Large on-chip memory 915 (e.g., 4MB SRAM, 32GB and / or 64GB 256-bit LPDDR5 (at a rate of 204.8GB / s), 8GB and / or 16GB 128-bit LPDDR5 (at a rate of 102.4GB / s), and / or other types and sizes of memory) can enable the hardware acceleration cluster to accelerate neural network processing, converter processing, optical flow processing, data processing, and / or other computations or processing. The hardware acceleration cluster can be used to supplement the GPU 908 and offload some tasks from the GPU 908 (e.g., freeing up more cycles of the GPU 908 to perform other tasks). For example, accelerators can be used for target workloads (such as perceptual, convolutional neural network (CNN), deep neural network (DNN), language model (LLM, VLM, MMLM, VLA, etc.), transformer model, diffusion model, encoder-only model, encoder-decoder model, etc.), which must be stable enough to be suitable for acceleration.
[0083] Accelerators (such as hardware acceleration clusters) may include a Deep Learning Accelerator (DLA) 909 (also referred to herein as a “Deep Learning Accelerator Cluster (XNN) 909,” a “Neural Network Accelerator (NNA) 909,” or a “Neural Processing Unit (NPU) 909”). The DLA 909 may include one or more Tensor Processing Units (TPUs) 941, which may be configured to provide additional computing power, such as trillions of operations per second, for deep learning applications and inference. The TPUs 941 may be accelerators configured and optimized for performing processing functions, such as for CNNs, RCNNs, DNNs, etc. The DLA 909 may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperforms CPUs. The TPUs 941 can perform several functions, including single-instance convolution functions, for example, supporting INT8, INT16, and FP16 data types as features and weights, and post-processor functions. Although TPU941 is described as being included as part of DLA 909, this is not intended to be restrictive. TPU 941 may be included in additional or alternative accelerators and / or other components, and / or included as a discrete processing component.
[0084] The DLA 909 can quickly and efficiently execute neural networks on processed or unprocessed data to perform any of a variety of functions, including, but not limited to: feature recognition and detection using data from one or more sensor modalities; distance estimation using data from one or more sensor modalities; recognition and detection using data from microphones and / or data-based sensors; performing pick-and-place operations; performing manipulation operations; performing other operations using data from cameras and / or other sensor types; and / or for security and / or safety-related events, etc.
[0085] The DLA 909 can perform any function of the GPU 908. For example, by using an inference accelerator, designers can choose either the DLA 909 or the GPU 908 to perform any function as needed. For instance, designers can centralize DNN and floating-point operations on the DLA 909, leaving other functions to the GPU 908 and / or other accelerators. The DLA 909 can be used to run any type of network to enhance control and security; for example, it can run neural networks that output a confidence metric for each object detection.
[0086] Accelerators (e.g., hardware acceleration clusters) may include a programmable accelerator 907 (ACCLS), which may be alternatively referred to herein as a computer accelerator or generally as a data accelerator. Accelerator 907 may be designed and configured to accelerate computer algorithms for ML applications, security and surveillance applications, augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) applications, etc. Accelerator 907 can strike a balance between performance and flexibility. For example, each accelerator 907 may include (but is not limited to) any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) systems, Pixel Processing Engines (PPEs), Vector Processors or Vector Processing Units (VPUs), and / or other components. PVA engines may include advanced Very Long Instruction Word (VLIW) and Single Instruction Multiple Data (SIMD) digital signal processors. Accelerator 907 may be optimized for data processing and computer algorithm acceleration tasks. For example, the Accelerator 907 offers superior performance and extremely low power consumption, and can be used asynchronously or concurrently as part of a heterogeneous computing pipeline with the CPU 906, GPU 908, and / or other accelerators in the system (e.g., automation, etc.).
[0087] Accelerator 907 may include one or more (e.g., two) Vector Processing Subsystems (VPSs), each of which may include one or more Vector Processing Unit (VPU) cores, one or more Decoupled Lookup Units (DLUTs), one or more shared or vector memory (VMEMs), and one or more instruction caches (I-caches). The VPU core may be the main processing unit and may include a vector SIMD VLIW DSP 943 optimized for data processing. The VPU core can fetch instructions via the I-cache and access data via the VMEM. The DLUT may contain a dedicated hardware component to improve the efficiency of parallel lookup operations. For example, the DLUT allows parallel lookups using a single copy of the lookup table by performing these lookup operations in a decoupled pipeline independent of the main processor pipeline. In this way, the DLUT can minimize or reduce memory usage, increase throughput, and avoid data-related memory bank conflicts, ultimately improving overall system performance. The VPU VMEM can provide local data storage for the VPU, enabling efficient implementation of various data processing and computer algorithms. The VPU VMEM can support access from external hosts of the VPS, such as Direct Memory Access (DMA) and the CPU 906 (e.g., an ARM Cortex-R5 processor), thereby facilitating data exchange with the CPU 906 and other system-level components. The VPU instruction cache can provide instruction data to the VPU on request, request missing instruction data from system memory, and / or maintain temporary instruction storage for the VPU. For each VPU task, the CPU 906 can be configured with a DMA system to selectively prefetch VPU programs into the VPU I-cache and / or initiate each VPU-DMA pair to process the task. The accelerator 907 may also include L2 SRAM memory for sharing among one or more (e.g., two) VPSs and DMA sets. In some embodiments, one or more (e.g., two) DMA devices are used to move data between external memory, PVA L2 memory, VMEM (e.g., one in each VPS), CPU tightly coupled memory (TCM), DMA descriptor memory, and / or PVA-level configuration registers. In lightly loaded systems, the read / write bandwidth for two parallel DMA accesses to DRAM can reach 15 GB / s each; in heavily loaded systems, this bandwidth can reach 10 GB / s each. In terms of computational compactness, INT8 can achieve 2048 gigabyte multiply-accumulate operations per second (GMAC) or higher (excluding DLUT). FP32 can achieve 32 GMACs per PVA instance.
[0088] RISC cores can interact with sensors, data signal processors, and more. Each RISC core can contain memory of any capacity. RISC cores can use any of a variety of protocols, depending on the specific implementation. In some examples, a RISC core can run a real-time operating system (RTOS). A RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may contain an instruction cache and / or tightly coupled RAM.
[0089] The DMA system enables components of accelerator 907 to access system memory independently of CPU 906. DMA can support any number of functions to provide optimizations for accelerator 907, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0090] A vector processor, or VPU, can be a programmable processor designed to efficiently and flexibly execute computer algorithms and provide signal processing capabilities. In some examples, accelerator 907 may include a core and two vector processing subsystem partitions. The core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystems may operate as the main processing engine of accelerator 907 and may include one or more vector processing units (VPUs), one or more pixel processing engines (PPEs) (where PPEs may include a two-dimensional layout of interconnected processing elements (e.g., for communication between north, south, east, and west directions), one or more instruction caches, and / or one or more shared or vector memories (e.g., VMEMs). The VPU core may include a digital signal processor, such as, for example, a single-instruction, multiple-data (SIMD) or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.
[0091] In some embodiments, each vector processor may include an instruction cache and may be coupled to dedicated memory. Therefore, in some examples, each vector processor may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular accelerator 907 may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single accelerator 907 may execute the same computer algorithm but process different regions of data. In other examples, the vector processors included in a particular accelerator 907 may execute different computer algorithms simultaneously on the same data, or even execute different algorithms on sequential data or data segments. Furthermore, the hardware acceleration cluster may contain any number of PVAs, and each PVA may contain any number of vector processors. Additionally, the accelerator 907 may include additional error-correcting code (ECC) memory to enhance the overall security of the system.
[0092] Accelerators (e.g., hardware accelerator clusters) have wide applications in autonomous and semi-autonomous machine control. Accelerator 907 is likely a programmable accelerator used in critical processing stages such as perception, robot understanding, and reasoning. Accelerator 907's performance is well-suited for algorithmic domains requiring low power consumption, low latency, and predictable processing. In other words, Accelerator 907 performs well in semi-intensive or intensive rule computations even when processing small datasets, meeting the requirements for low latency, low power consumption, and predictable runtime. Therefore, Accelerator 907 is designed to run classic computer ML algorithms because they are highly efficient in object detection and integer arithmetic.
[0093] In some examples, accelerator 907 can be used to perform dense optical flow. According to this process, raw radar data (e.g., using a 4D Fast Fourier Transform) can be used to provide processed raw radar data. In other examples, accelerator 907 is used for time-of-flight depth processing, for example, to provide processed time-of-flight data by processing raw time-of-flight data.
[0094] While the VPU, DMA, RISC core, VMEM, and decoupled coprocessors (e.g., DLUT) are described as being included in the accelerator 907, this is not a limitation. In some embodiments, these components may be included in alternative or additional processing components and / or accelerators, and / or may be included as discrete components of the SoC 904 and / or other computing system architectures.
[0095] In some examples, SoC 904 may include a real-time ray tracing hardware accelerator (RTA) 951, which can be used to quickly and efficiently determine the location and extent of objects (e.g., in a world model), generate real-time or near-real-time visualization simulations, for radar signal interpretation, for acoustic propagation synthesis and / or analysis, for simulating sonar (SONAR), radar, lidar (LiDAR), camera and / or other sensor modes within the simulation, for general wave propagation simulation, for comparison with lidar data for localization, for generating realistic training data for training neural networks, and / or other functions and uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.
[0096] SoC 904 may include one or more Camera Serial Interfaces (CSI) 923s. For example, CSI 923 may include high-speed interfaces and / or data input modules that can be used for input functions. SoC 904 may also include one or more input / output controllers that can be software-controlled and can be used to receive I / O signals not delegated to a specific purpose. For example, CSI 923 may include MIPI CSI-2 connectors—e.g., a 16-channel MIPI CSI-2 connector, D-PHY 2.1 (up to 40Gbps) and C-PHY 2.0 (up to 164Gbps) for supporting 16 virtual channels and 6 or more cameras; an 8-channel MIPI CSI-2 connector, D-PHY 2.1 (up to 20Gbps) for supporting 8 virtual channels and 4 or more cameras; and / or a 2x MIPI CSI-2, 22-pin camera connector, depending on the embodiment and implementation.
[0097] Accelerators (e.g., hardware acceleration clusters) may include an on-chip computer network (CNOC) 963 and SRAM for providing high-bandwidth, low-latency SRAM to the accelerators. In some examples, on-chip memory may include at least 4 MB of SRAM, such as, but not limited to, eight field-configurable memory blocks accessible by accelerators 907, OFA 911, DLA 909, and / or other accelerators. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory 915 may be used. Accelerators 907, OFA 911, DLA 909, and / or other accelerators can access the memory via a backbone that provides high-speed memory access to the accelerators. This backbone may include an on-chip computer network for interconnecting the accelerators with the memory (e.g., using an APB).
[0098] CNOC 963 may include an interface that determines whether the accelerator provides a ready and valid signal before transmitting any control signals / addresses / data. Such an interface may provide separate stages and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. Such an interface may conform to ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.
[0099] SoC 904 may include a data repository 916 and / or memory 915. The data repository 916 may be on-chip memory 915 of the SoC 904 for storing neural networks and / or other algorithms to be executed on the CPU 906, GPU 908, and / or one or more accelerators. In some examples, the data repository 916 may be large enough to store multiple neural network instances for redundancy and security. For example, the data memory 916 may include L2 and / or L3 cache 912. The memory 915 may include SRAM, LPDDR5, and / or other types of memory. For example, the memory 915 may include 4MB SRAM, 32GB and / or 64GB 256-bit LPDDR5 (at a rate of 204.8GB / s), 8GB and / or 16GB 128-bit LPDDR5 (at a rate of 102.4GB / s), and / or other types and sizes of memory. References to data repository 916 may include references to memories associated with accelerator 907, OFA911, DLA 909 and / or other accelerators, as described herein.
[0100] Data repository 916 may include various storage types, such as eMMC, NVMe, etc. For example, SoC 904 may include storage in the form of an embedded multimedia card (eMMC) (e.g., 64GB eMMC 5.1) and / or an SD card slot, as well as external NVMe high-speed (NVMe) capabilities, such as via M.2 Key M. For example, data repository 916 and / or other storage may be accessed via, for example, NVMe using PCI Express (PCIe), RDMA, TCP, and / or other protocols.
[0101] SoC 904 may include one or more processors 402 (e.g., embedded processors). Processor 402 may include a boot and power management processor (BPMP) 953, which may be a dedicated processor and subsystem for handling boot power and management functions, as well as associated safety measures. BPMP 953 may be part of the SoC 904 boot sequence and may provide runtime power management services. BPMP 953 may provide clock and voltage programming, assist system low-power state transitions, manage the thermistors and temperature sensors of SoC 904, and / or manage the power state of SoC 904. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature. SoC 904 may use these ring oscillators to detect the temperature of CPU 906, GPU 908, accelerator, and / or other components. If the temperature is determined to exceed a threshold, BPMP 953 may enter a temperature fault routine and place SoC 904 into a low-power state.
[0102] Processor 402 may also include a set of embedded processors that can be used as the Audio Processing Engine (APE) 955. The APE 955 can be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces and offering a wide and flexible range of audio I / O interfaces. In some examples, the APE 955 is a dedicated processor core with a digital signal processor featuring dedicated RAM.
[0103] Processor 402 may also include an always-on processor engine (AOPE) 957, which provides the necessary hardware functionality to support low-power sensor management and wake-up use cases. AOPE 957 may include a processor core, tightly coupled RAM, peripheral support (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0104] Processor 402 may also include a security processor 913 (or alternatively referred to as "security island 913"), which may include a security cluster engine containing a dedicated processor or processor subsystem for handling security management for automotive, robotic, and / or other applications. The security processor 913 and / or the security cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and as a single core with comparison logic to detect any differences in their operation. In some embodiments, the security processor 913 may include one or more discrete processors to prevent failures of other system components from affecting the performance and availability of the security processor 913.
[0105] Processor 402 may also include a real-time or near-real-time sensor engine (SE) 959, which may include a dedicated processor subsystem for processing real-time or near-real-time data. Processor 402 may also include one or more signal processors (SP) 927, which may include high dynamic range signal processors and / or hardware engines as part of one or more sensor processing pipelines. Processor 402 may include a processing module 961 (e.g., implemented on a microprocessor) that implements post-processing functions. Processing module 961 includes enhanced temporal noise suppression for both spatial and temporal noise reduction.
[0106] The SoC 904 may also include a variety of peripheral interfaces 925 for input / output (I / O), such as for communicating with peripheral devices, audio codecs, power management, or other devices. The SoC 904 can be used to process data from cameras (e.g., via a gigabit multimedia serial link and / or Ethernet connection), sensors (via Ethernet connection), from the processor bus 410, and from GNSS sensors (e.g., via Ethernet or CAN bus connection). The SoC 904 may also include dedicated high-performance, high-capacity memory controllers, which may contain their own DMA engine and can be used to free the CPU 906 from routine data management tasks. In one embodiment, the SoC 904 I / O 925 may include headers (e.g., 40-pin headers or 40-pin extension headers) supporting Universal Asynchronous Receiver / Transmitter (UART), Serial Peripheral Interface (SPI), Inter-Integrated Circuit Audio (I2S), Inter-Integrated Circuit (I2C), Controller Area Network (CAN), Pulse Width Modulation (PWM), Digital Microphone Interface (DMIC), Digital Speaker Station (DSPK), General Purpose I / O (GPIO), etc.; automation headers (e.g., 12-pin automation headers); audio panel headers (e.g., 10-pin audio panel headers); Joint Test Action Group (JTAG) headers (e.g., 10-pin JTAG headers); fan headers (e.g., 4-pin fan headers); RTC battery backup connectors (e.g., 2-pin battery backup connectors); microSD card slot; DC power jack; power button; forced restart button; restore button; and reset button; and one or more display connectors (e.g., DisplayPort (DP), such as DP 1.4A (+MST), eDP 1.41, HDMI 2.1, and / or 4K30 multi-model DP). 1.2 (+MST) connector) and / or other I / O 925 elements, components or functions.
[0107] The SoC 904 may include intranetting capabilities using, for example, Ethernet (e.g., automotive Ethernet), SERDES, Controller Area Network (CAN), FlexRay, Local Interconnect Network (LIN), Low Voltage Differential Signaling (LVDS), Media-Oriented System Transport (MOST), other networking types, and / or combinations thereof. For example, the SoC 904 may include RJ45 connectors supporting up to 10GbE, 1GbE connectors, and / or other types of network connectors.
[0108] The SoC 904 may include one or more digital signal processors (DSPs) 943. For example, the DSP 943 may include a dedicated or custom-designed microprocessor chip optimized for digital signal processing, such as for audio signal processing, telecommunications, digital data processing and / or other sensor processing, speech recognition and / or other applications.
[0109] The SoC 904 may include one or more General Purpose Computing Acceleration Clusters (GCACs) 929. For example, the GCAC 929 may include various processor types that can be used to accelerate computing, such as one or more Vector Microcode Processors (VMPs) 933, one or more Multithreaded Processing Clusters (MPCs) 931, one or more Programmable Macroarrays (PMAs) 935, and / or one or more other types of processors. For example, the GCAC 929 may include a PMA 935, two VMPs 933, and two MPCs 931.
[0110] The SoC 904 may include one or more vector microcode processors (VMPs) 933. In some embodiments, the VMP 933 may include a wide vector (Very Long Instruction Word (VLIW) and Single Instruction Multiple Data (SIMD)) machine for performing various operations, such as short integer operations common in machine learning and deep learning algorithms, including those with at least... Figure 2B The operations related to the ML algorithm described in the document.
[0111] The SoC 904 may include one or more multi-threaded processing clusters (MPCs) 931. The MPC 931 may include processing clusters that, in some embodiments, are more general-purpose than GPUs and more efficient than CPUs. For example, the MPC 931 may include a multi-threaded processor that allows multiple threads to share resources and execute instructions concurrently.
[0112] The SoC 904 may include one or more programmable macro arrays (PMAs) 935. The PMA 935 may include a coarse-grained reconfigurable architecture (CGRA) dataflow machine, which has a unique architecture that enables powerful performance on intensive computer machine learning and deep learning algorithms that may not be achievable in traditional digital signal processing (DSP) architectures.
[0113] The SoC 904 may include one or more Display Processing Units (DPUs) 945 for performing hardware-accelerated processing. For example, the DPU 945 may retrieve pixel data from memory 915 and send it to a display peripheral via a standard interface. Thus, the DPU 945 can handle display processing and rendering for internal and / or external displays.
[0114] SoC 904 may include one or more Application Processing Units (APUs) 939. For example, APU 939 may include a quad-core or dual-core processor with 48KB / 32KB L1 cache with parity and ECC and 1MB L2 cache with ECC. APU 939 may support NEON instructions and single-precision and double-precision floating-point operations. SoC 904 may include one or more Real-Time Processing Units (RTPUs) 969. RTPU 969 may include a dual-core processor with 32KB / 32KB L1 cache and 256KBTCM with ECC. RTPU 969 may support single-precision and double-precision floating-point operations. SoC 904 may include one or more Built-in Self-Test (BIST) components 937. For example, BIST component 937 may include a Memory BIST (MBIST) for testing system memory and / or a Logic BIST (LBIST) for testing system logic. BIST component 937 may include embedded logic for directly testing system logic and / or memory.
[0115] SoC 904 may include one or more dynamically reconfigurable processors (DRPs) 971. For example, DRPs 971 can be used to accelerate various computational operations. For instance, in some embodiments, DRPs 971 may be used in combination with a MAC unit as an AI accelerator. In some embodiments, DRPs 971 may dynamically switch the circuit connection configuration of on-chip arithmetic units (e.g., ALUs) based on the content to be processed during each operating clock cycle, while simultaneously executing the application. Because only the necessary arithmetic circuitry is used, DRPs 971 may consume less power than CPU processing and may be faster. Furthermore, compared to CPUs where frequent external memory accesses due to cache misses or other reasons degrade performance, DRPs 971 can pre-build the necessary data paths in the hardware, thereby reducing performance degradation and speed fluctuations (jitter) caused by memory accesses. DRPs 971 may include dynamic loading functionality that switches circuit connection information each time an algorithm changes, enabling processing with limited hardware resources, even in robotics / automotive applications that require handling multiple algorithms.
[0116] In some embodiments, the accelerator may include an OpenCV accelerator for accelerating OpenCV processing. OpenCV is an open-source, industry-standard library for computer data processing. In some embodiments, one or more DRP 971 instances are deployed as AI accelerators and used in conjunction with OpenCV accelerators to enhance AI computations and other processing algorithms, enabling complex and computationally intensive operations such as visual simultaneous localization and mapping (SLAM).
[0117] Unlike traditional systems, the technology described herein allows multiple neural networks to be executed simultaneously (e.g., at least partially in parallel) and / or sequentially by providing CPU complexes, GPU complexes, and hardware acceleration clusters, and combining the results. Furthermore, since the SoC 904 can contain various computing engines (e.g., processor 402, CPU 906, GPU 908, accelerators, etc.), tasks can be distributed among and within these computing engines, and in some cases, common-cause failures are avoided due to the discrete layout of the computing engines. Moreover, since the SoC 904 can contain a dedicated security processor 913 (or security island 913), critical security or redundancy operations can be performed without causing noticeable failures in the main processing component or computing engine of the SoC 904.
[0118] In some examples, the machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged into microservices—such as inference microservices (e.g., NVIDIA NIM)—which may contain containers (e.g., operating system (OS) level virtualization packages) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine." For example, an inference microservice may contain the container itself as well as the model (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., with a small enough number of parameters), the model itself may be contained within a container. In other examples (e.g., when the model is large), the model may be hosted / stored in the cloud (e.g., in a data center), and / or may be hosted locally and / or at the edge (e.g., on a local server or compute component, but outside the container). In these embodiments, the model may be accessed via one or more APIs (e.g., REST APIs). Therefore, in some embodiments, the machine learning models described herein can be deployed as inference microservices to accelerate model deployment on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers (for simplified deployment), an optimized inference engine (e.g., execution software built using standardized AI model deployments, such as NVIDIA's Triton Inference Server), and / or one or more APIs for high-performance deep learning inference (which may include inference runtime and model optimizations for low latency and high throughput for production applications, such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The machine learning models described herein can be included as part of a microservice along with acceleration infrastructure that can be deployed with a single command and / or orchestrated and automatically scaled on the acceleration infrastructure using a container orchestration system (e.g., from a single device up to data center scale). Therefore, the inference microservice may include a machine learning model (e.g., a model optimized for high-performance inference), inference runtime software for executing the machine learning model and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, authentication, and / or other monitoring. In some embodiments, the inference microservice may include software for in-situ replacement and / or updating of the machine learning model. During replacement or updating, the software performing the replacement / update may maintain the user configurations of the inference runtime software and the enterprise management software.
[0119] While this article may describe examples of using machine learning models (such as neural networks), this is not limiting. For example (but not limited to), any machine learning model and / or neural network described herein may include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (KNN), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN)), perceptrons, long / short memory (LSTM) networks, multilayer perceptron (MLP) networks, deep stacked networks (DSN), generative pre-trained (GPT) models or networks, feedforward networks, radial basis function ANNs, self-organizing maps (SOM), Kohonen maps, Hopfield networks, Boltzmann machines, deep belief neural networks, deconvolutional neural networks, generative adversarial networks (GAN), liquid machines, etc. Modular neural networks, liquid machines, sequence-to-sequence models, networks using transformer architectures, state-space models (SSMs) (e.g., networks using Mamba architectures (e.g., Mamba-1, Mamba-2, etc.), networks using selective state-space models, networks using structured state-space sequence models, etc.), diffusion models (e.g., diffusion probability models, score-based generative models, etc.), neural radiation field (NeRF) models, Gaussian sputtering models, Kolmogorov-Arnold networks (KANs), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), multimodal language models (MMLMs), large action models (LAMs), and other machine learning models, and / or other types of machine learning models. All such machine learning aspects can be applied to at least... Figure 3 Related ML algorithms and discussions.
[0120] In some examples, the example architecture of computing system 900 can be used to enable at least one processor (labeled CPU906) to execute instructions to coordinate work among GPUs 908, thereby supporting streaming jobs. The CPU and GPU can execute or support scaling processes within a cluster of nodes in the data plane. The scaling process can retrieve progress information related to the streaming job from the data plane and use this progress information to request the manager subsystem of the control plane to allow or support the expansion of at least one cluster, thereby adding or removing nodes from available nodes to execute the streaming job.
[0121] In some examples, one or more CPUs or GPUs may be used to listen for data communication to nodes that are part of a streaming job. One or more CPUs or GPUs may be used to facilitate progress reports maintained in a file system. One or more CPUs or GPUs may be used to provide full or partial progress reports as progress information that can be used for explicit detection or prediction of open doors for collision detection and avoidance. One or more CPUs or GPUs may further execute instructions to configure one or more processors to determine the status associated with the streaming job. This status may be derived from the progress information. One or more CPUs or GPUs may be used to determine the number of nodes to add or remove as part of the expansion to complete the streaming job.
[0122] In some examples, one or more CPUs or GPUs may be used to scale to provide a first number of nodes, enabling processing of the streaming job at a rate greater than the input rate of the input stream, which is frequently processed as part of the streaming job. In some examples, one or more CPUs or GPUs may be used to scale to provide a second number of nodes, thereby keeping the size of the input stream within a predetermined hysteresis threshold. In some examples, one or more CPUs or GPUs may be used to provide scaling capabilities for individual input sources associated with different input sources in the streaming job, which may be contained in separate clusters within a cluster.
[0123] In some examples, one or more CPUs or GPUs can enable the execution of a streaming job, the processing rate of which includes a first predetermined value provided as part of the streaming job, which may represent a measure of the number of rows per second of records pending processing in the streaming job and is part of the current configuration of multiple nodes. The input rate may be a second predetermined value provided as part of the streaming job, which is based in part on the number of rows per second of records processed as part of the streaming job. The predetermined hysteresis threshold used may be a measure of the number of records remaining in the streaming job's input stream.
[0124] In some examples, one or more CPUs or GPUs can be used to train an ML model using data generated from one or more of the input rates, processing rates, or lag thresholds of different records processed as parts of different streaming jobs. In some examples, one or more CPUs or GPUs can use the ML model to generate outputs representing predictions or inferences about one or more of the input rates, processing rates, or latency thresholds, thereby allowing or supporting reactive or proactive applications to the scaling of at least one cluster in the cluster to add or remove nodes from the nodes that are part of the cluster for performing streaming jobs.
[0125] In some examples, one or more CPUs or GPUs can facilitate one or more APIs between the control plane and the data plane. These APIs can allow messaging between the management subsystem and the scaling processes, and support scaling of at least one cluster to add or remove nodes for executing streaming jobs. Furthermore, these APIs can perform O-auth authentication between the management subsystem and the scaling processes to support or initiate the messaging portion of communication between the scaling processes and the management subsystem.
[0126] In some examples, the at least one processor may be included in at least one of the following systems: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multimodal language models (MMLMs); a system for performing operations using one or more visual-language-action (VLA) models; a system for using or deploying one or more inference microservices; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0127] In some examples, the at least one processor may include (e.g., as part of a larger system) at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more LLMs; a system for performing operations using one or more VLMs; a system for performing operations using one or more MMLMs; a system for performing operations using one or more VLA models; a system for using or deploying one or more inference microservices; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system containing one or more virtual machines; a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0128] For example, the system described above may include at least one processor for explicitly detecting or predicting door openings in order to perform collision detection and avoidance. For example, the processor may be part of a system used for purposes such as (e.g., it may include or be included in another system): autonomous or semi-autonomous machines; analog operations; digital twin operations; optical transmission simulations; collaborative content creation for 3D assets; deep learning operations; edge devices; robots; generative AI operations; LLM; VLM; MMLM; VLA models; inference microservices; conversational AI operations; synthetic data; virtual reality content, augmented reality content, or mixed reality content; virtual machines; data centers; or cloud computing resources.
[0129] While this disclosure may refer to examples of autonomous or semi-autonomous vehicles, robots, and / or other machine types 1000 (which may instead be referred to herein as "vehicle 1000," "self-vehicle 1000," "machine 1000," "self-machine 1000," "robot 1000," and / or "self-robot 1000"), examples are found in... Figures 10A-10EThe description herein is provided in [the relevant section], but is not intended to limit the scope of this disclosure. For example, the systems and methods described herein may be used without limitation by: non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms (e.g., autonomous mobile robots (AMRs), humanoid robots, robotic arms and / or end effectors, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, water vehicles, shuttles (e.g., driverless taxis), emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, and underwater vehicles (e.g., manned or unmanned). Submarines), drones, and / or other types of vehicles, robots, or machines. Furthermore, while this disclosure may primarily describe the explicit detection or prediction of open doors in collision detection and avoidance, this is not intended to limit its application. The systems and methods described herein can also be used in augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security and surveillance (e.g., smart cities), autonomous or semi-autonomous machine applications, industrial manufacturing, simulation, and / or any other technical field where explicit detection or prediction of open doors can be used for collision detection and avoidance. In some embodiments, the systems, methods, and / or processes described herein can be used with… Figures 10A-10E The example machine 1000 shown Figure 11 The example shown calculates ecosystem 1100. Figure 12 The example generative language model system 1200 and / or shown Figure 13 The example computing device 1300 shown may be used to perform similar functions, features, and / or features to other components, features, and / or features.
[0130] In some embodiments, the systems and methods described herein can be performed using simulated data (e.g., simulated environment data of virtual or simulated vehicles, robots, or machines and simulated sensor data of simulated sensors) in a simulated environment (e.g., NVIDIA's DriveSIM, ISAAC Sim, ISAAC Gym, ISAAC Lab, etc.). For example, simulated input data (e.g., map data, perception data, self-motion data, haptic data, and / or any other data described herein) can be used for explicit detection or prediction of door openings in the collision detection and avoidance described herein, and can be used to perform virtual machine-related operations in a simulated environment. These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before deploying them to the real world. In some cases, simulation can be used to generate synthetic training data, such as sensor data from within the simulation. The synthetic training data (as an addition to or replacement of real-world data) can then be used or processed to perform explicit detection or prediction of door openings in the collision detection and avoidance described herein.
[0131] In any example, such as when using a simulated environment for testing, validation, training, etc., the simulated environment and / or associated training data may be rendered or otherwise generated using one or more optical transport simulation algorithms (e.g., one or more ray tracing and / or path tracing algorithms). When using optical transport simulation, the simulation system may employ one or more dedicated ray tracing hardware accelerators and / or processors (e.g., NVIDIA's RTX or other real-time ray tracing GPUs, such as GPUs containing one or more ray tracing (RT) cores) optimized to perform real-time or near-real-time optical transport simulation operations in conjunction with one or more other processors in the system (e.g., GPUs, CPUs, accelerators, etc.). In some embodiments, the simulated environment and / or one or more of its objects, features, or components may be generated or managed in a three-dimensional (3D) content collaboration platform (e.g., NVIDIA's OMNIVERSE), which may be optimized or suitable for industrial digitalization, generative physics, artificial intelligence, and / or other use cases, applications, and / or services. For example, a content collaboration platform or system may include systems for using or developing common scene descriptors (USD) (such as OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physical simulations (e.g., using NVIDIA's PhysX software development kit (SDK)) to simulate real physical phenomena and physical interactions with simulations hosted on the platform. The platform may integrate OpenUSD with ray tracing / path tracing / light transport simulations (such as NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, and / or testing AI systems, such as systems for testing, validating, and training (e.g., machine learning models, neural networks, etc.), and / or other tasks related to automobiles, robots, other machine types, and / or other systems and applications. In some examples, the simulated environment may include digital twins of real-world environments, such as specific road segments, warehouses, data centers, airports, geographic areas, ocean areas, and / or any other real-world environment that can be operated by autonomous or semi-autonomous vehicles or machines.
[0132] In some embodiments, a remote control or teleoperation system can be used to teleoperate or remotely control vehicles, robots, and / or other machines. For example, the systems and methods described herein can be used to explicitly detect or predict open doors, thereby enabling collision detection and avoidance. These systems and methods can be incorporated into the visualization or mapping of an environment to assist a remote operator in controlling an autonomous or semi-autonomous machine through the environment, or to provide waypoints or other control or navigation instructions for navigating the environment. Therefore, a remote operator can use visual, auditory, textual, and / or other cues or indicators generated using the systems and methods described herein to assist in navigating vehicles, robots, machines, etc., through real-world environments using a remote operating system.
[0133] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), deep learning accelerator cluster (XNN), neural processing unit (NPU), neural network accelerator (NNA), hardware-based programmable vision accelerator (PVA) – which may include one or more vector processing units (VPU), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.) and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models). This robotic system can utilize these processors to execute one or more machine learning models (e.g., language models, visual language models (VLM), large language models (LLM), visual-language-action (VLA) models, multimodal language models (MMLM), etc.) to enable it to autonomously or semi-autonomously perform complex tasks, such as interacting with and / or manipulating static and / or dynamic objects, or navigating in its environment using sensors such as cameras, LiDAR, RADAR, and ultrasonic sensors. The system can employ sensor fusion technology to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, radar, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server to perform computationally intensive tasks, such as 3D mapping or simultaneous localization and mapping (SLAM). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze the data and distribute optimized instructions to the entire fleet. In some embodiments, the machine learning models described herein (e.g., language models, VLM, VLA, LLM, MMLM, diffusion models, NeRF models, DNN, etc.) can be used to enable the robot to perceive and reason about its environment, and / or communicate with one or more other robots and / or people in the environment. In some embodiments, the robot may (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) communicate with one or more locally hosted servers / computing devices and / or with one or more remotely located servers / computing devices (e.g., located in one or more data centers).
[0134] In some embodiments, the systems and methods described herein can be deployed in in-vehicle infotainment (IVI) systems or in-cabin experience (IX) applications. For example, an infotainment system within a vehicle (e.g., a car, truck, drone, construction machinery, robot, semi-autonomous vehicle, or autonomous vehicle) may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), deep learning accelerator cluster (XNN), neural processing unit (NPU), neural network accelerator (NNA), hardware-based programmable vision accelerator (PVA)), which may include one or more vector processing units (VPU), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.) and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models), and memory and / or storage devices (e.g., for storing entertainment content, navigation data, and user preferences). The system can utilize these processors to execute one or more machine learning models (e.g., language models) to enable functions such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services via network connectivity. In-vehicle infotainment systems can also use natural language processing (NLP) models to enable voice-based interaction. One or more machine learning models can be stored locally or accessed via one or more APIs connected to cloud services, enabling the system to process requests in real-time or near real-time.
[0135] In some examples, the machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, visual-language-action (VLA) models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged into microservices—such as inference microservices (e.g., NVIDIA NIMs)—that may contain containers (e.g., operating system (OS) level virtualization packages) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model “engine.” For example, an inference microservice may contain the container itself as well as one or more models (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., with a small number of parameters), the model can be contained directly within the container itself. In other cases, such as when the model is large, the model can be hosted / stored in the cloud (e.g., in a data center), or hosted locally and / or at the edge (e.g., on a local server or computing device, but outside the container). In these embodiments, the model can be accessed via one or more APIs (e.g., REST APIs). Therefore, in some embodiments, the one or more machine learning models described herein can be deployed as inference microservices, thereby accelerating the deployment of one or more models on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., execution software built using standardized AI model deployments, such as NVIDIA's Triton Inference Server), and / or one or more simplified deployment APIs for high-performance deep learning inference, which may include inference runtimes and model optimizations that provide low latency and high throughput for production applications (e.g., NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The one or more machine learning models described herein may be included as part of a microservice along with acceleration infrastructure that can be deployed with a single command and / or orchestrated and automatically scaled on the acceleration infrastructure via a container orchestration system (e.g., from a single device to a data center scale). Therefore, an inference microservice may include one or more machine learning models (e.g., models optimized for high-performance inference), inference runtime software for executing one or more machine learning models and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, authentication, and / or other monitoring.In some embodiments, the inference microservice may include software for in-situ replacement and / or updating of one or more machine learning models. During replacement or updating, the software performing the replacement / update may maintain user configurations for the inference runtime software and enterprise management software.
[0136] While this article may describe examples of using machine learning models (such as neural networks), this is not limiting. For example (but not limited to), the various machine learning models and / or neural networks described herein can include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN)), perceptrons, long / short memory (LSTM) networks, multilayer perceptron (MLP) networks, deep stacked networks (DSN), generative pre-trained (GPT) models or networks, feedforward networks, radial basis function artificial neural networks, self-organizing maps (SOM), Kohonen maps, Hopfield networks, Boltzmann machines, deep belief neural networks, deconvolutional neural networks, generative adversarial networks (GAN), liquid machines, modular neural networks, and liquid... Machine learning models of various types, including machine learning models such as sequence-to-sequence models, networks using transformer architectures, state-space models (SSMs) (e.g., networks using Mamba architectures (e.g., Mamba-1, Mamba-2, etc.), networks using selective state-space models, networks using structured state-space sequence models, etc.), diffusion models (e.g., diffusion probability models, fraction-based generative models, etc.), neural radiation field (NeRF) models, Gaussian sputtering models, Kolmogorov-Arnold networks (KANs), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), visual language models (VLMs), multimodal language models (MMLMs), large action models (LAMs), visual-language-action (VLA) models, etc.), and / or one or more other types of machine learning models.
[0137] The systems and methods described herein can be used, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, water vehicles, shuttles (e.g., driverless taxis), emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles (e.g., manned or unmanned submarines), drones, and / or other types of vehicles. Furthermore, the systems and methods described herein can be used for a wide range of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets (e.g., NVIDIA's Omniverse), cloud computing, and / or any other suitable applications.
[0138] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines, etc.), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing language models (e.g., large language models (LLM), visual language models (VLM), visual-language-action (VLA) models, and / or multimodal language models), systems using or deploying one or more inference microservices, systems integrating and deploying one or more machine learning models in services or microservices and OS-level virtualization packages (e.g., containers), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI, systems for performing optical transmission simulation, systems for performing neural computing, systems for performing neural rendering, systems for performing 3D asset collaborative content creation, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0139] refer to Figure 1A-2B , Figure 4 and Figure 8-13This disclosure is based on several embodiments. It should be understood that the arrangements and other arrangements described herein are merely examples. Other arrangements, components, features, and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to the arrangements, components, features, and elements shown (e.g., machines, interfaces, functions, sequences, functional groups, etc.), and certain elements may be omitted entirely. Furthermore, many of the arrangements, components, features, elements, etc., described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be deployed in any suitable combination and at any suitable location (e.g., on a local device, vehicle, or machine at the edge; locally deployed—e.g., on a locally hosted server; remotely—e.g., on one or more computing or server devices in one or more data centers in the cloud; and / or other locations). The various functions performed by the entities described herein can be implemented using hardware, firmware, and / or software. For example, one or more processors (e.g., central processing unit (CPU), graphics processing unit (GPU)) can be used. Microprocessors, microcontrollers, embedded processors, digital signal processors (DSPs), image signal processors (ISPs), physical processing units (PPUs), field-programmable gate arrays (FPGAs), accelerators (e.g., deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs) and / or neural processing units (NPUs), programmable vision accelerators (PVAs), optical flow accelerators (OFAs), etc.), application-specific integrated circuits (ASICs), data processing units (DPUs), quantum processors, etc., execute instructions stored in memory to achieve various functions. In some embodiments, the systems, methods, and processes described herein can be used with... Figures 10A-10E Example machine 1000 Figure 11 Example calculation of ecosystem 1100, Figure 12 Example generative language model system 1200 and / or Figure 13 The example computing device 1300 uses components, characteristics, and / or functions similar to those of other components, characteristics, and / or functions to perform the task.
[0140] Furthermore, each module of the method described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, one or more processors (e.g., but not limited to the processors described herein) can be used to execute instructions stored in one or more memories or memory systems to perform various functions. In some embodiments, the computer process may also be embodied as computer-usable instructions stored on a computer storage medium. The method may be provided by a standalone application, service or managed service (alone or in combination with another managed service), application programming interface (API), and / or plug-in to another product, etc. Furthermore, the method may be provided as an example for… Figure 1A-4 and Figure 8-13 These methods can be executed by any system. Furthermore, these methods can also be executed, or alternatively, by any system or combination of any number of systems, including but not limited to the systems described herein.
[0141] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttle buses (e.g., robotaxis), emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles (e.g., manned or unmanned submarines), drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, including, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets (e.g., NVIDIA's Omniverse), cloud computing, and / or any other suitable application.
[0142] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines, etc.), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing language models (e.g., large language models (LLM), visual language models (VLM), and / or multimodal language models), systems using or deploying one or more inference microservices, systems containing one or more machine learning models deployed in services or microservices and OS-level virtualization packages (e.g., containers), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0143] Example autonomous or semi-autonomous machines Figure 10A Examples of sensor positions with corresponding fields of view or sensing fields of view of autonomous or semi-autonomous vehicles 1000A, autonomous mobile robots (AMRs) 1000B, and humanoid robots 1000C, according to some embodiments of this disclosure. While three types of machines 1000 are shown, this is not intended to be limiting, and the machines 1000 described herein may include vehicles, cars, trucks, buses, first-response vehicles, shuttle buses, electric or motorized bicycles, motorcycles, fire trucks, police or emergency vehicles, ambulances, boats, construction vehicles, underwater vehicles, robots (e.g., AMRs, humanoid robots, robotic arms, end effectors, forklifts, etc.), drones, aircraft, vehicles coupled to trailers (e.g., semi-trailer tractors for hauling goods), and / or another type of vehicle or machine (e.g., driverless and / or vehicles or machines accommodating one or more passengers). In some cases, vehicle 1000A, AMR 1000B, humanoid robot 1000C, and / or other machine types may be collectively referred to herein as machine 1000.
[0144] Regarding Vehicle 1000A, autonomous and semi-autonomous vehicles are typically described by automation levels, defined by the National Highway Traffic Safety Administration (NHTSA) (a division of the U.S. Department of Transportation) and the Society of Automotive Engineers (SAE) in their "Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 1000 may have one or more functions that meet Level 3 through Level 5 of autonomous driving. For example, Vehicle 1000 may be able to provide driver assistance (Level 1), partial automation (Level 2, Level 2+, Level 2++), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the specific implementation. The term “autonomy” as used herein can include any and / or all types of autonomy of Machine 1000 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, provision of auxiliary autonomy, semi-autonomy, primary autonomy or other specified autonomy.
[0145] about Figure 10A The sensors and their respective fields of view (not shown for clarity) or sensing fields (not shown for clarity) are an example embodiment and are not intended to be limiting. Although not shown, each sensor may have a corresponding field of view (e.g., a 360-degree field of view around camera 1068D, a 180-degree field of view around wide-angle camera 1068B, a 360-degree sensing field of view of LiDAR sensor 1064, etc.). For example, only a subset of the sensors shown may be included, additional sensors may be included, alternative sensors may be included, the number of each sensor mode may be different, the sensor modes may be different (e.g., LiDAR or RADAR may not be included, SONAR, thermal sensors, etc. may be included), and the sensor positions may differ from those shown on vehicles 1000A, AMR 1000B, and / or humanoid robots 1000C, etc. For example, for vehicle 1000A, the location, number, modality, and / or other sensor information may differ depending on the type (e.g., SUV, truck, car, robot, motorcycle, etc.), size (e.g., 18-wheeler, transport vehicle, small car, etc.), and related functions (e.g., L2 vs. L5). Similarly, for AMR 1000B and / or humanoid robot 1000C, shape, size, purpose, implementation, model, etc., can determine the number and type of sensors used.
[0146] like Figure 10AAs shown, autonomous or semi-autonomous vehicles 1000A, AMR 1000B, and humanoid robots 1000C may include different sensor types, numbers, and locations. As a non-limiting example, vehicle 1000A may include twelve cameras 1068, such as a front wide-angle camera (e.g., 120-degree field of view (FOV)), a front telephoto camera (e.g., 30-degree FOV), a side-rear left camera (e.g., 70-degree FOV), a side-rear right camera (e.g., 70-degree FOV), a front fisheye camera (e.g., 200-degree FOV), a rear fisheye camera (e.g., 200-degree FOV), a left fisheye camera (e.g., 200-degree FOV), a right fisheye camera (e.g., 200-degree FOV), a front telephoto satellite camera (e.g., 30-degree FOV), a rear telephoto camera (e.g., 30-degree FOV), a cross-left camera (e.g., 120-degree FOV), and a cross-right camera (e.g., 120-degree FOV). In one embodiment, the camera 1068 may use a gigabit multimedia serial link (GMSL) interface (e.g., GMSL2) as input / output (I / O).
[0147] In some embodiments, although Figure 10A As not shown, vehicle 1000A may include an in-cabin occupant and / or driver monitoring system, which may include various sensors. For example, in-cabin sensors may include various cameras 1068, such as a driver monitoring camera (e.g., located in front of the driver's seat at a 55-degree FOV facing the driver's seat), a front occupant monitoring camera (e.g., located in front of the front occupant seat at a 190-degree FOV facing the front occupant seat), and a rear occupant monitoring camera (e.g., located in front of the rear occupant seat at a 190-degree FOV facing the rear occupant seat). Similar to external cameras 1068, in embodiments, internal cameras 1068 may use a GMSL (e.g., GMSL2) interface for I / O.
[0148] As another non-limiting example, vehicle 1000A may also include nine RADAR sensors 1060. For example, vehicle 1000A may include a front center imaging RADAR sensor (e.g., 120-degree FOV or sensing field), a left front corner RADAR sensor (e.g., 160-degree FOV or sensing field), a right front corner RADAR sensor (e.g., 160-degree FOV or sensing field), a right rear corner RADAR sensor (e.g., 160-degree FOV or sensing field), a left RADAR sensor (e.g., 160-degree FOV or sensing field), a right RADAR sensor (e.g., 160-degree FOV or sensing field), a left rear RADAR sensor (e.g., 50-degree FOV or sensing field), and a right rear RADAR sensor (e.g., 50-degree FOV or sensing field). In embodiments, the RADAR sensors 1060 may use an Ethernet interface as I / O.
[0149] As a non-limiting example, the vehicle 1000A may also include twelve ultrasonic sensors 1062. Figure 10A As shown, the ultrasonic sensor can be placed along the front and rear bumpers of vehicle 1000A and along the sides of vehicle 1000A, and can be used to detect objects (static and dynamic) close to vehicle 1000A. In some embodiments, the ultrasonic sensor 1062 can use a DS13 interface as I / O.
[0150] As a non-limiting example, vehicle 1000A may also include a LiDAR sensor 1064, such as a front-center LiDAR sensor (e.g., a 120-degree horizontal FOV or sensing field and a 30-degree vertical FOV or sensing field). In some embodiments, such as when using additional or alternative LiDAR sensors, the LiDAR sensor may have different horizontal and vertical fields of view or sensing fields. For example, LiDAR sensor 1064 may include a 360-degree horizontal FOV or sensing field (e.g., located in a rotating LiDAR sensor) and a 90-degree vertical FOV or sensing field. In some embodiments, LiDAR sensor 1064 may use an Ethernet interface as I / O.
[0151] As a non-limiting example, the Autonomous Mobile Robot (AMR) 1000B may include three LiDAR sensors 1064. For example, the topmost LiDAR sensor 1064 may include a beam or 3D LiDAR sensor (e.g., a 360-degree horizontal and 90-degree vertical FOV or sensing field), and the front and rear LiDAR sensors may include planar or 2D LiDAR sensors (e.g., a 180-degree horizontal FOV or sensing field).
[0152] As a non-limiting embodiment, the AMR 1000B may further include eight cameras 1068, such as a front stereo camera (e.g., 120-degree FOV), a rear stereo camera (e.g., 120-degree FOV), a left stereo camera (e.g., 120-degree FOV), a right stereo camera (e.g., 120-degree FOV), a front fisheye camera (e.g., 202 degrees ± 3 degrees FOV), a rear fisheye camera (e.g., 202 degrees ± 3 degrees FOV), a left fisheye camera (e.g., 202 degrees ± 3 degrees FOV), and a right fisheye camera (e.g., 202 degrees ± 3 degrees FOV).
[0153] The AMR 1000B may also include a charging port, charging port contacts, status indicators, one or more (e.g., four) RGB LEDs, one or more IMU sensors 1066, a magnetometer, and a barometer. The AMR 1000B is capable of high-precision time synchronization between sensors using hardware timestamps and PTPs over Ethernet with sensor acquisition times of less than 10 microseconds. In an embodiment, the AMR 1000B provides simultaneous camera capture within 100 microseconds of a single hardware trigger on all cameras 1068, and can write sensor captures to disk at a rate of 4 GB / s for writing packets (e.g., writing to ROSbags of the Robot Operating System (ROS)). Therefore, the AMR 1000B is capable of running ROS (e.g., NVIDIA's Isaac ROS), can be remotely operated (as described herein), can map the environment, and can navigate the environment using the vision cameras 1068, LiDAR 1064, and / or other sensor types or modalities.
[0154] The humanoid robot 1000C may include (as a non-limiting example) a LiDAR sensor 1064. For example, the LiDAR sensor 1064 may include a beam or a 3D LiDAR sensor (e.g., a 360-degree horizontal and 90-degree vertical FOV or sensing field), or it may include a planar or 2D LiDAR sensor (e.g., a 180-degree horizontal FOV or sensing field).
[0155] As a non-limiting embodiment, the humanoid robot 1000C may also include four cameras 1068, such as a front stereo camera (e.g., 120-degree FOV), a rear stereo camera (e.g., 120-degree FOV), a front fisheye camera (e.g., 202-degree ± 3-degree FOV), and a rear fisheye camera (e.g., 202-degree ± 3-degree FOV).
[0156] As a non-limiting embodiment, the humanoid robot 1000C may also include four ultrasonic sensors 1062, such as a left arm ultrasonic sensor, a right arm ultrasonic sensor, a left leg ultrasonic sensor, and a right leg ultrasonic sensor.
[0157] The humanoid robot 1000C may also include any number of actuators, such as those allowing control and manipulation of joints. For example, the humanoid robot 1000C may include actuators that allow for various degrees of freedom (DoF) depending on the design. In a non-limiting embodiment, the humanoid robot 1000C may have a total of 40 degrees of freedom (DoF) (e.g., 6DoF x2 for arms, 6DoF x2 for hands, 6DoF x2 for legs, 2DoF for torso, and 2DoF for neck). Actuators can convert energy into physical motion, thereby allowing actions such as joint movement, locomotion, and grasping / manipulation. For example, motors and servos can be used to perform joint movements to control the rotation of joints in an arm or manipulator and allow for reaching, grasping, and manipulating objects. Locomotion can be achieved by moving through the environment using wheels, tracks, or other mobility devices (robot legs). Grasping and manipulation can be performed using end effectors or hands / fingers, which can be equipped with actuators to grasp objects, apply forces, and perform specific tasks. In some examples, the humanoid robot 1000C may include position and orientation sensors, such as encoders, gyroscopes, etc., to determine the robot 1000C's position in space, thereby enabling position determination and motion tracking. In embodiments, the humanoid robot 1000C may include force and pressure sensors to detect environmental interactions, enabling the robot 1000C to grasp objects with appropriate force and avoid obstacles along its path. Perception sensors (e.g., cameras, LiDAR, RADAR, ultrasound, sonar (SONAR), etc.) may be used in conjunction with tactile sensors to enable the robot 1000C to perceive objects, shapes, and textures, and to know when to begin and stop touching (in conjunction with force sensors that regulate the force used during touching). As a non-limiting example, the humanoid robot 1000C may have a height of approximately 1-2 meters (e.g., 1.7 meters or 5 feet 6 inches), a weight of 50-70 kilograms, be able to move at speeds of 8 km / h or higher, and be able to carry a payload of 20-100 kilograms, depending on the system design and requirements.
[0158] In embodiments, the humanoid robot 1000C may include a dialogue system—such as a dialogue system driven by a language model (e.g., LLM, VLM, MMLM, VLA, etc.)—to help understand the environment, reason, and communicate with humans, animals, devices, and / or other robots, and / or make planning, control, and navigation decisions. Therefore, in addition to performing various tasks, the humanoid robot 1000C can also use onboard sensors, microphones, and speakers to understand speech, audio, and visual cues, while also being able to communicate with the environment.
[0159] Referring to camera 1068 of machine 1000, the camera type of camera 1068 may include, but is not limited to, digital cameras applicable to components and / or systems of machine 1000. For vehicle 1000A implementation, camera 1068 may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 30 frames per second (fps), 60 fps, 120 fps, 240 fps, etc., depending on the embodiment. The camera may use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red colorless colorless colorless (RCCC) color filter array, a red colorless colorless blue (RCCB) color filter array, a red blue green colorless (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a colorless pixel camera (e.g., a camera with an array of RCCC, RCCB, and / or RBGC color filters) can be used to improve light sensitivity.
[0160] The field of view includes cameras (e.g., front-facing cameras) in the area in front of the machine 1000, which can be used for surround view to help identify the path and obstacles ahead, and, with the help of one or more 1036 and / or control SoCs, provide information crucial for generating an occupancy grid and / or determining preferred machine movement, trajectory, and / or path. The front-facing camera can be used to perform some of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance (these aspects are applied to the explicit detection or prediction of door openings for collision detection and avoidance), using data from such cameras. The front-facing camera can also be used in ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.
[0161] Various cameras can be used in front-mounted configurations, including, for example, monocular camera platforms that include complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 1068B, which can be used to perceive objects entering the field of view from the periphery (e.g., pedestrians, warehouse vehicles, other robots, pedestrian traffic, or bicycles). Furthermore, any number of long-range cameras 1068E (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. The long-range camera 1068E can also be used for object detection and classification, as well as basic object tracking.
[0162] Any number of stereo cameras 1068A can also be included in front-mounted and / or other (e.g., rear-mounted) configurations. In at least one embodiment, one or more stereo cameras 1068A may include an integrated control unit that includes a scalable processing unit that can provide programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the machine 1000 environment, including distance estimates of midpoints in the image (e.g., parallax or depth images). Alternative stereo cameras 1068A may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1068A may be used in addition to those described herein, or as an alternative to those described herein. For example, in some embodiments, a camera other than a stereo camera (e.g., two monocular cameras with at least partially overlapping fields of view) can be used to perform stereo depth estimation.
[0163] Cameras with a field of view including portions of the side environment of machine 1000 (e.g., side-view cameras) can be used, for example, for surround view, providing information for creating and updating occupancy grids, and generating side collision warnings and / or indicating to AMR1000B or humanoid robot 1000C, for example, the presence of objects, features, and / or people on the side. For example, a surround camera 1068D can be mounted on machine 1000. The surround camera 1068D can include a wide-angle camera 1068B, a fisheye camera, a 360-degree camera, etc. For example, four fisheye cameras can be mounted on the front, rear, and sides of machine 1000. In an alternative arrangement, machine 1000 can use three surround cameras 1068D (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a front-facing camera) as a fourth surround-view camera.
[0164] A camera 1068 (e.g., a rear-view camera) having a field of view including a portion of the environment behind the machine 1000 can be used to understand objects, features, people, and / or other information behind the machine 1000, such as for parking assistance, surround view, rear collision warning, planning, control, and navigation determination, and / or to create and update occupancy grids, BEV images representing the environment, height maps, etc. A wide variety of cameras 1068 can be used, including but not limited to those also suitable for use as front cameras (e.g., long-range and / or mid-range cameras 1068E, stereo cameras 1068A, infrared cameras 1068C, etc.), rear cameras, side cameras, downward cameras, upward cameras, and / or similar cameras 1068, as described herein.
[0165] Similarly, for LiDAR sensor 1064, RADAR sensor 1060, ultrasonic sensor 1062 and / or other sensor modes or types, the location and placement of the sensors and their corresponding fields of view or sensing fields can be determined based on the use case, implementation or design of the particular machine 1000.
[0166] For example, machine 1000 includes a RADAR sensor 1060, which can be used by machine 1000 for long-range object detection, even in dark and / or inclement weather conditions. In embodiments, the RADAR functional safety level can be ASIL B. RADAR sensor 1060 can be controlled and accessed for object tracking data using CAN and / or bus 1002 (e.g., for transmitting data generated by RADAR sensor 1060), and in some examples, raw data can be accessed using Ethernet. Various types of RADAR sensors can be used. For example, but not limited to, RADAR sensor 1060 can be suitable for front, rear, and side radar applications. In some examples, a pulse Doppler RADAR sensor is used.
[0167] The RADAR sensor 1060 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range lateral coverage, etc. In some examples, long-range radar can be used for adaptive cruise control (ACC) functions. A long-range RADAR system can provide a wide field of view, for example, within a 250m range, achieved through two or more independent scans. The RADAR sensor 1060 can help distinguish between static and moving objects; ADAS systems can use it for emergency braking assistance and forward collision warning, and robots can use it to detect dynamic objects in various environments (e.g., low-light or no-light environments). A long-range RADAR sensor can include a single static multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the machine 1000's surroundings at higher speeds while minimizing peripheral interference (e.g., traffic from adjacent lanes). The other two antennas expand the field of view, enabling rapid detection of objects entering or leaving the machine's direct path (e.g., lanes).
[0168] Mid-range RADAR systems can include ranges of up to 1060m (front) or 80m (rear), and field of view of up to 42 degrees (front) or 150 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of a side surface (e.g., rear bumper), allowing continuous monitoring of blind spots behind and beside a machine (e.g., vehicle, robot, etc.) using two beams. Therefore, short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0169] Machine 1000 may also include an ultrasonic sensor 1062. The ultrasonic sensor 1062 may be located at the front, rear, and / or sides of machine 1000 and may be used to assist near-field sensing, such as for parking assistance, collision avoidance (e.g., for robot parts), and / or to create and update occupancy grids, evidence mesh maps (EGMs), height maps, BEV images, and / or other representations of objects and features in the machine 1000 environment. Multiple ultrasonic sensors 1062 may be used, and different ultrasonic sensors 1062 may be used for different detection ranges (e.g., 2.5 m, 4 m). For example, the ultrasonic sensor 1062 may operate at an ASIL B functional safety level.
[0170] Machine 1000 may include a LiDAR sensor 1064. The LiDAR sensor 1064 can be used for object and feature detection, pedestrian and other robot detection, emergency braking, collision avoidance, simultaneous localization and mapping (SLAM), free space detection, and / or other functions. In embodiments, the LiDAR sensor 1064 may be of functional safety level ASIL B. In some examples, machine 1000 may include multiple LiDAR sensors 1064 (e.g., two, four, six, etc.), which can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0171] In some examples, the LiDAR sensor 1064 may be able to provide a 360-degree field of view of objects and their distances. For example, a commercially available LiDAR sensor 1064 may have a advertised range of approximately 1000m, an accuracy of 2cm-3cm, and support for 1000Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 1064 may be used. In such examples, the LiDAR sensor 1064 can be implemented as a small device that can be embedded in the front, rear, side, top, and / or corner of the machine 1000. In such examples, the LiDAR sensor 1064 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of up to 200m even for low-reflectivity objects. The front-mounted LiDAR sensor 1064 can be configured with a horizontal field of view between 45 and 135 degrees.
[0172] In some examples, LiDAR technology, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser flash as a transmission source, illuminating the environment around a vehicle up to approximately 200m away. The flash LiDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the distance from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment with each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the machine 1000. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras with no moving parts other than a fan (e.g., non-scanning LiDAR devices). Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of 3D distance point clouds and co-registration intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 1064 may be less susceptible to motion blur, vibration, and / or shock.
[0173] Figure 10BThis is an illustration of the location of sensors and components of an example autonomous or semi-autonomous vehicle 1000A (also referred to herein as "vehicle 1000", "self-vehicle 1000", "self-machine 1000", or "machine 1000") according to some embodiments of this disclosure. While vehicle 1000A is shown in the figure, this is not intended to be limiting, and similar components and / or sensors may be included on any other machine type without departing from the scope of this disclosure. For example, similar sensors and / or components may be used in vehicles, cars, trucks, buses, first-response vehicles, shuttle buses, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, ships, construction vehicles, underwater vehicles, robots (e.g., AMRs, humanoid robots, robotic arms, end effectors, forklifts, etc.), drones, aircraft, vehicles coupled to trailers (e.g., semi-trailers for hauling goods), and / or another type of vehicle or machine (e.g., driverless and / or capable of accommodating one or more passengers).
[0174] Figure 10CThis is a block diagram of an example system architecture for a machine 1000 (e.g., an autonomous or semi-autonomous vehicle 1000A, an autonomous mobile robot (AMR) 1000B, a humanoid robot 1000C, and / or other types of machines) according to some embodiments of this disclosure. It should be understood that such and other arrangements described herein are presented as examples only. Other arrangements, components, features, and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of the arrangements shown, and some elements may be omitted entirely. Furthermore, many of the arrangements, components, features, elements, etc., described herein are functional entities that can be implemented as discrete or distributed components or together with other components, and can be implemented in any suitable combination and location (e.g., on a local device, vehicle, or edge machine, in the field (e.g., a locally hosted server), in a remote location (e.g., in one or more computing or server devices in one or more data centers in the cloud) and / or other locations). The various functions performed by the entities described herein can be performed by hardware, firmware, and / or software. For example, various functions can be implemented using one or more processors (e.g., central processing unit (CPU), graphics processing unit (GPU), microprocessor, microcontroller, embedded processor, digital signal processor (DSP), image signal processor (ISP), physical processing unit (PPU), field-programmable gate array (FPGA), accelerators (e.g., deep learning accelerator (DLA), deep learning accelerator cluster (XNN), neural network accelerator (NNA) and / or neural processing unit (NPU), programmable vision accelerator (PVA), optical flow accelerator (OFA), etc.), application-specific integrated circuit (ASIC), data processing unit (DPU), quantum processor, etc.) executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be used with... Figures 10A to 10E Example machine 1000 Figure 11 Example calculation of ecosystem 1100, Figure 12 Example Generative Language Model System 1200 and / or Figure 13 The example computing device 1300 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the same task.
[0175] Figure 10CEach component, feature, and system of machine 1000 is shown connected via bus 1002 (or referred to as "machine communication network 1002" or simply "communication network 1002"). Bus 1002 may include a Controller Area Network (CAN) data interface (or referred to herein as "CAN bus"). CAN can be a network within machine 1000 used to help control various features and functions of machine 1000, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button position, and / or other vehicle status indicators. CAN bus may conform to ASIL B standard. In some embodiments, in addition to or as an alternative to the CAN bus, bus 1002 may include FlexRay, embedded buses (e.g., SPI, I2C), local interconnect links (LIN), NVIDIA's NVLink, USB (2.0, 3.0 and above), radio frequency (RF), Ethernet (e.g., 10BASE / 100BASE, 1000BASE, 10G, etc.), and / or other communication protocols or functions. Furthermore, while a single line is used to represent bus 1002, this is not limiting. For example, there can be any number of buses 1002, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1002 may be used to perform different functions and / or for redundancy. For example, a first bus 1002 may be used for a collision avoidance function, while a second bus 1002 may be used for actuation control. In any example, each bus 1002 may communicate with any component of machine 1000, and two or more buses 1002 may communicate with the same component. In some examples, each computer or computing engine within each SoC 1004, each controller 1036, and / or machine 1000 can access the same input data (e.g., input from sensors in the machine 1000) and can be connected to a common bus, such as the CAN bus.
[0176] Machine 1000 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, batteries, side mirrors, and / or other parts of the vehicle or machine. Machine 1000 may include a propulsion system 1050, such as an internal combustion engine, a hybrid power plant, an all-electric motor, a hydrogen fuel cell engine, and / or another type of propulsion system. Propulsion system 1050 may be connected to the drivetrain of machine 1000, which may include a transmission to enable propulsion of machine 1000. Propulsion system 1050 may be controlled in response to a signal received from throttle / accelerator 1052.
[0177] Steering system 1054 may include a steering wheel and / or other steering mechanisms (e.g., remote steering and / or local steering) for steering machine 1000 (e.g., along a desired path or route) while propulsion system 1050 is in operation (e.g., when the vehicle is traveling). Steering system 1054 may receive signals from steering actuator 1056. In some embodiments, a steering wheel or other steering mechanism may be omitted, for example, for machine 1000 capable of fully automated (e.g., level 5) functionality.
[0178] The brake sensor system 1046 can be used to operate the vehicle brakes in response to signals received from the brake actuator 1048 and / or the brake sensor.
[0179] Machine 1000 may include one or more controllers 1036, such as those described herein. Figure 10AThe controller 1036 is described. It can be used for a variety of functions and can be coupled to any of the various other components and systems of the machine 1000. For example, the controller 1036 can be used to control the machine 1000, artificial intelligence performed on the machine 1000, infotainment of the machine 1000, etc. For example, one controller 1036 can be used for some or all of the functions, or different controllers 1036 can be used for different functions, for example, to ensure availability and safety separation between various controllers for different tasks. For example, the controller 1036 can use a system-calculated plan (e.g., the path or trajectory of vehicle 1000A or AMR 1000B, or the movement, component trajectory, motion position or displacement, etc. of joints or components (e.g., manipulators, end effectors, limbs, hands, fingers, legs, feet, etc.) of humanoid robot 1000C to control the machine 1000 in the environment. In some cases, the controller 1036 may include a proportional-integral-derivative (PID) controller, a fuzzy logic controller, a neural controller (e.g., a controller embodied as one or more neural networks), a force control controller, a programmable logic controller (PLC), and / or other types of controllers. For example, in the humanoid robot 1000C, the controller 1036 can act as the brain, responsible for analyzing sensor data, making decisions, and sending commands to the actuators. The controller 1036 may include a low-level controller that handles basic motor control, ensuring accurate and precise movement of the individual joints and actuators. The controller 1036 may also include a high-level controller to coordinate multiple actuators and sensors, plan complex movements, and adapt to constantly changing environments.
[0180] In some embodiments, controller 1036 may include an artificial intelligence controller that can use AI algorithms (e.g., DNN, MLM, etc.) to learn, make decisions, and autonomously execute tasks of machine 1000. In some embodiments, controller 1036 may use a fixed open-loop control algorithm and may not adjust its actions based on the environment. In other embodiments, closed-loop control may be used, which incorporates a feedback mechanism to monitor the robot's performance and make necessary adjustments. In examples, controller 1036 may implement reactive control to directly respond to sensor inputs, enabling rapid reflexive actions and real-time changes. Furthermore, in some examples, deliberate control may be implemented, using internal models and planning algorithms to generate advanced actions, which may be suitable for complex tasks requiring reasoning, decision-making, and long-term planning.
[0181] The controller 1036 may include one or more system-on-chip (SoC) 1004 ( Figure 10C and Figure 10DThe controller 1036 may provide signals (e.g., signals representing commands or messages) to one or more components and / or systems of the machine 1000, including components such as CPUs, GPUs, and accelerators. Although the controller 1036 is listed separately from the SoC 1004, this is not intended to be limiting, and in some embodiments, one or more components of the SoC 1004 may perform the operations of the controller 1036. For example, the controller may send signals to operate machine brakes via one or more brake actuators 1048, to operate steering system 1054 via one or more steering actuators 1056, to operate propulsion system 1050 via one or more throttle / accelerators 1052, and so on. The controller 1036 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous or semi-autonomous navigation and movement and / or assist human operators using the machine 1000. Controller 1036 may include a first controller 1036 for autonomous control and navigation functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy in emergency situations, and / or other controllers. For example, hardware for safety monitoring and other safety functions (e.g., functional safety islands) may be discrete or partitioned (physically or by processing separation) relative to hardware for processing sensor data to make perception and vehicle control decisions. Similarly, hardware for controlling in-vehicle infotainment and / or in-cabin monitoring (e.g., controllers, SOCs, etc.) may be separate or independent from hardware for vehicle perception and control. In some examples, a single controller 1036 may handle two or more of the above functions, two or more controllers 1036 may handle a single function, and / or any combination thereof.
[0182] The controller 1036 may provide signals for controlling one or more components and / or systems of the machine 1000 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data can be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensor 1058 (e.g., Global Positioning System sensor), RADAR sensor 1060, ultrasonic sensor 1062, LiDAR sensor 1064, Inertial Measurement Unit (IMU) sensor 1066 (e.g., accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 1096, camera 1068 (e.g., stereo camera 1068A, wide-angle camera 1068B (e.g., fisheye camera), infrared camera 1068C, surround camera 1068D (e.g., 360-degree camera), long-range and / or medium-range camera 1068E and / or other types of cameras), speed sensor 1044 (e.g., for measuring the speed of machine 1000), vibration sensor 1042, steering sensor 1040, braking sensor (e.g., as part of braking sensor system 1046), actuators and / or other types of sensors.
[0183] One or more controllers 1036 may receive input (e.g., represented by input data) from the instrument panel 1032 of the machine 1000 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1034 (e.g., screen, head-up display, mirror display, facial display, robot display, etc.), sound alarms, loudspeakers, speakers, and / or via other components of the machine 1000. Outputs may include parameters such as machine speed, rate, time, and... Figure 10C The system may display map data corresponding to map 1022 (e.g., from a navigation map, standard definition (SD) map, high definition (“HD” map, etc.), location data (e.g., the location of machine 1000, such as its location on map 1022), direction, the location of other vehicles (e.g., occupancy map, elevation map, bird's-eye view (BEV) image, grid, etc.), information about objects perceived by the system and their states, system status information, and so on. For example, HMI display 1034 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.), and / or information about driving maneuvers that the vehicle has performed, is performing, or will perform (e.g., changing lanes now, exiting from exit 34B in two miles, etc.).
[0184] Machine 1000 may include one or more System-on-a-Chip (SoC) 1004 (in Figure 10D(described in more detail below). SoC 1004 may include CPU 1006, GPU 1008, processor 1010, cache 1012, accelerator 1014, data storage 1016, and / or other components and features. SoC 1004 can be used to process and provide data for various operations, such as navigation, planning, reasoning, inference, perception, control, and / or actuation operations of machine 1000 on various platforms and systems. For example, SoC 1004 can process real-time perception data (e.g., from cameras, LiDAR, RADAR, ultrasound, etc.) and map data corresponding to one or more maps 1022 (e.g., HD maps, SD maps, navigation maps, occupancy maps, etc.) to perform or assist in performing various operations of machine 1000. When using maps and / or AI, maps and / or AI (e.g., model parameter updates, fine-tuning, etc.) are transmitted via network interface 1024 from one or more servers (e.g., [unclear]). Figure 10E Refresh and / or update server 1078 (e.g., one or more servers in a cloud-based data center).
[0185] Despite Figures 10A to 10E The SoC 1004 is shown, however, additional or alternative components and / or architectures, such as multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-package (SiPs), field-programmable gate arrays (FPGAs), heterogeneous integration (HI), and single-board computers (SBCs), may be used without departing from the scope of this disclosure. For example, depending on the type of machine 1000, the purpose of machine 1000, the model of machine 1000, and the capabilities required by machine 1000, one or more SoCs 1004s and / or alternative architectures and / or components may be used to meet a particular implementation.
[0186] Machine 1000 may include CPU 1018 (e.g., a discrete CPU or dCPU) which may be coupled to SoC 1004 via a high-speed interconnect (e.g., PCIe). CPU 1018 may include, for example, an x86 processor. CPU 1018 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC 1004, and / or monitoring the status and health of controller 1036 and / or infotainment SoC 1030.
[0187] Machine 1000 may include GPU 1020 (e.g., a discrete GPU or dGPU) which may be coupled to SoC 1004 via a high-speed interconnect (e.g., NVIDIA's NVLink). GPU 1020 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs from sensors of machine 1000 (e.g., sensor data).
[0188] The machine 1000 may also include a network interface 1024, which may include one or more wireless antennas 1026 and / or modems (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1024 can be used to establish wireless connectivity with the cloud (e.g., with server 1078 and / or other network devices) via the Internet, and with other vehicles and / or computing devices (e.g., passenger client devices). For communication with other vehicles, direct links and / or indirect links (e.g., across networks and via the Internet) can be established between the two vehicles. A vehicle-to-vehicle communication link can be used to provide a direct link. The vehicle-to-vehicle communication link can provide the machine 1000 with information about vehicles near the machine 1000 (e.g., vehicles in front, to the side, and / or behind the machine 1000). This functionality can be part of the machine 1000's cooperative adaptive cruise control function.
[0189] Network interface 1024 may include a SoC that provides modulation and demodulation functions and enables controller 1036 to communicate over a wireless network. Network interface 1024 may include a radio frequency (RF) front-end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. Frequency conversion can be performed by well-known processes and / or using superheterodyne processes. In some examples, the RF front-end functionality may be provided by a separate chip. For example, network interface 1024 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), 5G, 6G, and / or other cellular and / or wireless communication standards. The wireless antenna 1026 can also enable communication between objects in the environment (such as vehicles, mobile devices, etc.) using local area networks (such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc.) and / or low power wide area networks (“LPWAN”) (such as LoRaWAN, SigFox, etc.).
[0190] Machine 1000 may also include data memory 1028, which may include off-chip (e.g., outside of SoC 1004) storage. Data memory 1028 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, Flash, hard disk and / or other components and / or devices capable of storing at least one bit of data.
[0191] Machine 1000 may also include a GNSS sensor 1058. The GNSS sensor 1058 (e.g., a GPS, an auxiliary GPS sensor, a differential GPS (DGPS) sensor, etc.) is used to assist in mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1058 can be used, such as, but not limited to, GPS with a USB connector having an Ethernet-to-serial (RS-232) bridge.
[0192] Machine 1000 may also include an IMU sensor 1066. In some examples, the IMU sensor 1066 may be located at the center of the rear axis of machine 1000. The IMU sensor 1066 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1066 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1066 may include an accelerometer, a gyroscope, and a magnetometer.
[0193] In some embodiments, the IMU sensor 1066 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1066 can enable the machine 1000 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 1066, without requiring input from a magnetic sensor. In some examples, the IMU sensor 1066 and the GNSS sensor 1058 can be combined in a single integrated unit.
[0194] The vehicle may include one or more microphones 1096 placed inside and / or around the machine 1000. The microphones 1096 may be used for emergency vehicle detection and identification, etc.
[0195] Machine 1000 may also include vibration sensor 1042. Vibration sensor 1042 can measure vibrations of machine parts, such as the arm or leg of humanoid robot 1000C, or the axle of vehicle 1000A or AMR 1000B. For example, changes in vibration may indicate changes in roads, walking, or traversable surfaces. In another example, when two or more vibration sensors 1042 are used, differences between vibrations can be used to determine friction or slippage on surfaces (e.g., when the vibration difference is between an electrically driven shaft and a freely rotating axle).
[0196] Machine 1000 may include ADAS system 1038, for example, when machine 1000 is vehicle 1000A. In some examples, ADAS system 1038 may include a dedicated SoC. ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision or collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), blind spot monitoring (BSM), rear cross traffic warning (RCTW), pedestrian detection, driver monitoring, collision warning system (CWS), traffic sign recognition, speed limit detection, automatic parking, lane centering (LC), high beam safety system and / or other features and functions.
[0197] Machine 1000 may also include an infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include one or more discrete components, such as multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-package (SiP), heterogeneous integration (HI), single-board computers (SBCs), etc. The infotainment SoC 1030 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., wireless, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assist, radio data system, vehicle-related information (e.g., fuel level, total driving distance, brake fluid level, fuel level, door opening / closing, air filter information, etc.) to the machine 1000. For example, the infotainment SoC 1030 may be a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, head-up display (HUD), HMI display 1034, telematics device, control panel (e.g., for controlling various components, features, and / or systems and / or interacting with various components, features, and / or systems), and / or other components. 1030 can also be used to provide information to vehicle users (e.g., visual and / or auditory), such as information from ADAS system 1038, autonomous driving information (e.g., planned vehicle maneuvers, trajectories), surrounding environment information (e.g., intersection information, vehicle information, road information, etc.) and / or other information.
[0198] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 can communicate with other devices, systems, and / or components of the machine 1000 via bus 1002 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1030 may be coupled to a monitoring MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1036 (e.g., the main computer and / or backup computer of the machine 1000). In such an example, the infotainment SoC 1030 may place the machine 1000 into a driver-safe stop mode, as described herein.
[0199] In some embodiments, the infotainment system may provide a digital or virtual assistant, which may be voice-only or may have visual components (e.g., in the form of a digital human or digital avatar). The assistant may provide basic functions such as sending text messages, adjusting vehicle settings, controlling music or video, navigation, etc., and / or provide more advanced functions, such as those supported by one or more language models (e.g., large language model (LLM), visual language model (VLM), multimodal language model (MMLM), etc.). For example, the driver and / or occupants may interact with the assistant in a manner similar to how a user interacts with a language model, such as asking general questions, specific questions, requesting restaurants, gas stations, and / or other recommendations and / or locations, inquiring about vehicle functions or troubleshooting (e.g., asking for tire pressure information, oil change information, battery swap information, etc.). Therefore, the machine 1000 (whether it is a vehicle 1000A, AMR 1000B, humanoid robot 1000C, or other types of machine) may include a locally stored language model and / or communicate with a remotely hosted language model (e.g., via one or more APIs) to provide users of the machine 1000 with more detailed and in-depth communication capabilities.
[0200] In some examples, the infotainment SoC 1030, SoC 1004, and / or another SoC or computing / processing system can perform in-cabin driver and / or occupant monitoring. For example, the computing system can perform facial recognition, and vehicle owner recognition can use data from cameras and / or other sensors to identify the presence of the authorized driver and / or owner of the machine 1000. An always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in security mode, to disable the vehicle when the owner leaves. Thus, SoC 1004 provides security against theft and / or hijacking.
[0201] In some embodiments, one or more neural networks running on another or a dedicated SoC (e.g., an in-vehicle infotainment or in-vehicle monitoring SoC) can be used to monitor in-cabin monitoring camera sensors. This other or dedicated SoC is configured to recognize in-cabin events and respond accordingly. The in-cabin system can activate cellular services and make phone calls via lip reading, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. The in-cabin system may also include one or more in-cabin AI agents or assistants that can interact with one or more LLMs, VLMs, MMLMs, etc., in the cloud using one or more APIs or plugins. For example, the in-cabin AI agents or assistants can provide directions, vehicle or machine feedback information, answer general questions, handle music / video and / or other requests, activate windows, doors, and / or other vehicle components, etc. Therefore, one or more dedicated SoCs and / or processor groups can be used to perform in-cabin infotainment and / or in-cabin monitoring (e.g., as an occupant monitoring system (OMS)) for Machine 1000.
[0202] Machine 1000 may also include instrument panel 1032 (e.g., digital instrument panel, electronic instrument panel, digital dashboard, etc.). Instrument panel 1032 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). Instrument panel 1032 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn signal indicator, gearshift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting control, safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1030 and instrument panel 1032. In other words, instrument panel 1032 may be included as part of infotainment SoC 1030, and vice versa.
[0203] Figure 10D A computing system according to at least some embodiments of this disclosure (about Figure 10C A block diagram of an example architecture (a subset of the systems described). Although illustrated as SoC 1004, this is not intended to be limiting, and the computing system may additionally or alternatively include multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-package (SiP), heterogeneous integration (HI), single-board computers (SBCs), and / or other components and / or architectures without departing from the scope of this disclosure.
[0204] The SoC 1004 can be an end-to-end platform with a flexible architecture spanning automation levels 2-5, or it can be specifically designed for a particular automation level (e.g., a first SoC 1004 for levels 2 to 2++, a second SoC 1004 for level 3, a third SoC 1004 for level 4, and so on), providing a comprehensive functional safety architecture that leverages and effectively utilizes computer vision, neural network inference, robot planning, control and navigation, ADAS technologies, and more, offering versatility and redundancy to provide a flexible and reliable platform for driving or robot control software stacks, as well as deep learning tools. The SoC 1004 can be faster, more reliable, and even more energy-efficient and space-saving than traditional systems. For example, when the accelerator 1014 is used in conjunction with the CPU 1006, GPU 1008 and data storage 1016, it can provide a fast and efficient platform for Level 2-5 autonomous vehicles, as well as for the safety planning, navigation and control of AMR 1000B, humanoid robot 1000C and / or other robot or machine types.
[0205] In some embodiments, for example, SoC 1004 includes a GPU 1008 having 2,000 or more cores (e.g., 2,048 cores), 60 or more tensor cores (e.g., 64 tensor cores) and a maximum GPU frequency exceeding 1 GHz (e.g., 1.3 GHz); a CPU 1006 having 10 or more cores (e.g., 12 cores), 64-bit, 3 MB L2 cache memory and 6 MB L3 cache memory and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz); and one or more deep learning accelerators (DLA), deep learning accelerator clusters (XNN), neural network accelerators (NNA), or neural processing units (NPU) 1009 (e.g., 2 DLA / XNN / NNA / NPU 1009); and a vision accelerator—e.g., a programmable vision accelerator (PVA) 1007, or a single SoC 1004) that may be capable of achieving AI performance of 275 trillion operations per second (TOPS). For example, NVIDIA's Jetson AGX Orin 64GB SoC meets these standards and achieves this performance.
[0206] Similarly, in the following embodiments, SoC 1004 includes a GPU 1008 having 1700 or more cores (e.g., 1792 cores), 50 or more tensor cores (e.g., 56 tensor cores) and a maximum GPU frequency exceeding 900MHz (e.g., 930MHz); a CPU 1006 having 8 or more cores (e.g., 8 cores), 64-bit architecture, 2MB L2 cache, and 4MB L3 cache and a maximum frequency of 2GHz or more (e.g., 2.2GHz); one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs), or neural processing units (NPUs) 1009 (e.g., two DLAs / XNNs / NNAs / NPUs 1009); and a vision accelerator—e.g., a programmable vision accelerator (PVA) 1007, or a single SoC 1004—potentially capable of 200 trillion operations per second (TOPS) of AI performance. For example, NVIDIA's Jetson AGX Orin 32 GB SoC meets these standards and achieves this performance.
[0207] In some embodiments, for example, SoC 1004 may include a GPU 1008 having 1,000 or more cores (e.g., 1,024 cores), 28 or more tensor cores (e.g., 32 tensor cores) and a maximum GPU frequency exceeding 900 MHz (e.g., 1,173 MHz); a CPU 1006 having 8 or more cores (e.g., 8 cores), 64-bit architecture, 2 MB L2 cache, and 4 MB L3 cache and a maximum frequency of 2 GHz or higher (e.g., 2 GHz); one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs), or neural processing units (NPUs) 1009 (e.g., one DLA / XNN / NNA / NPU 1009); and a vision accelerator—e.g., a programmable vision accelerator (PVA) 1007 and a single SoC 1004, etc.) that may be capable of AI performance of 157 trillion operations per second (TOPS). For example, NVIDIA's Jetson AGX Orin NX 16 GB SoC meets these standards and achieves this performance.
[0208] In various embodiments, for example, SoC 1004 may include a GPU 1008 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 32 tensor cores), and a maximum GPU frequency exceeding 900MHz (e.g., 1020MHz), and a CPU 1006 with 6 or more cores (e.g., 6 cores), 64-bit architecture, 1.5MB L2 cache, and 4MB L3 cache, and a maximum frequency of 1.5GHz or higher (e.g., 1.7GHz), then a single SoC 1004 may be able to achieve AI performance of 67 trillion operations per second (TOPS). For example, NVIDIA's Jetson Orin Nano 8 GB SoC meets these criteria and achieves this performance.
[0209] SoC 1004 may include one or more CPUs 1006. In some embodiments, CPU 1006 may include CPU clusters or CPU complexes (hereinafter alternatively referred to as "CCPLEX"). CPU 1006 may include multiple cores and / or (e.g., L2, L3) caches. For example, in some embodiments, CPU 1006 may include twelve cores arranged in a one-to-many processor configuration. In some embodiments, CPU 1006 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 3MB L2 cache). CPU 1006 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, thereby allowing any combination of clusters of CPU 1006 to be active at any given time.
[0210] SoC 1004 may include any type and number of GPUs 1008. For example, in some embodiments, an integrated GPU (referred to herein as an "iGPU") may be used. GPU 1008 may be programmable and capable of efficiently handling parallel workloads. In some embodiments, GPU 1008 may use an enhanced tensor instruction set. GPU 1008 may include one or more streaming microprocessors, each of which may include a cache (e.g., an L1 cache with a capacity of at least 96KB), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with a storage capacity of 512KB). In some embodiments, GPU 1008 may include at least eight streaming microprocessors. GPU 1008 may use a computing application programming interface (API). Furthermore, GPU 1008 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0211] The GPU 1008 is power-optimized for automotive, robotics, and / or other embedded applications to achieve optimal performance. For example, the GPU 1008 can be manufactured using a FinFET process. However, this is not a limitation; the GPU 1008 can also be manufactured using other semiconductor manufacturing or fabrication processes. Each streaming microprocessor can contain multiple mixed-precision processing cores, which are divided into multiple modules. For example (but not limited to), 64 PF32 cores and 32 PF64 cores can be divided into four processing modules. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix operations, an instruction cache (e.g., L0), a thread bundle scheduler, dispatch units, and / or a register file (e.g., 64KB). Furthermore, the streaming microprocessor can contain independent parallel integer and floating-point data paths to efficiently execute workloads that combine computation and addressing operations. Streaming microprocessors can include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can also include combined L1 data caches and shared memory units to improve performance while simplifying programming.
[0212] The GPU 1008 may include high-bandwidth memory (HBM) and / or (e.g., 16GB) HBM2 memory subsystems to provide peak memory bandwidth of approximately 900GB / s in some examples. In some examples, in addition to HBM memory, or as an alternative to HBM memory, synchronous graphics random access memory (SGRAM), such as graphics double data rate type 5 synchronous random access memory (GDDR5), may be used.
[0213] The GPU 1008 may include unified memory technology, including access counters, to more accurately migrate memory pages to the processor with the highest access frequency, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support may be used, enabling the GPU 1008 to directly access the page tables of the CPU 1006. In these examples, when a memory management unit (MMU) miss occurs in the GPU 1008, it can send an address translation request to the CPU 1006. The CPU 1006 responds to this request, looks up the virtual-to-physical address mapping for that address in its page tables, and sends the translation result back to the GPU 1008. Therefore, unified memory technology can provide a single unified virtual address space for the memory of the CPU 1006 and the GPU 1008, thereby simplifying the programming of the GPU 1008 and the porting of applications to the GPU 1008.
[0214] SoC 1004 may include any number of caches 1012, including the caches described herein. For example, cache 1012 may include L0 cache, L1 cache, L2 cache, L3 cache (e.g., accessible to both CPU 1006 and GPU 1008, for example, simultaneously connected to both CPU 1006 and GPU 1008), etc. Cache 1012 may include a write-back cache that can track the state of cache lines, for example, by using one or more cache coherence protocols (e.g., MEI, MESI, MSI, etc.). Depending on a specific embodiment, the cache (e.g., L3) may include 4MB or more, but smaller or larger cache sizes may also be used.
[0215] SoC 1004 may include one or more arithmetic logic units (ALUs) 1065, which can be used to perform any of the processing related to various tasks or operations of machine 1000, such as computer vision, machine learning or deep learning processing, world model management, etc. Furthermore, SoC 1004 may also include one or more floating-point units (FPUs) 1067 or other mathematical coprocessors or numerical coprocessors for performing mathematical operations within the system. For example, SoC 1004 may include one or more FPUs 1067, which are integrated as execution units within CPU 1006 and / or GPU 1008.
[0216] SoC 1004 may include one or more accelerators 1014 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 1004 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. Large on-chip memory 1015 (e.g., 4MB SRAM, 32GB and / or 64GB 256-bit LPDDR5 (read / write speed 204.8GB / s), 8GB and / or 16GB 128-bit LPDDR5 (read / write speed 102.4GB / s), and / or other memory types and capacities) can enable the hardware acceleration cluster to accelerate neural network processing, transformer processing, optical flow processing, vision processing, and / or other computations or processing. The hardware acceleration cluster can be used to assist GPU 1008 and offload some of the tasks of GPU 1008 (e.g., freeing up more computation cycles of GPU 1008 to perform other tasks). As an example, the accelerator 1014 can be used for specific workloads (e.g., perceptual, convolutional neural network (CNN), deep neural network (DNN), language model (LLM, VLM, MMLM, VLA, etc.), transformer model, diffusion model, encoder-only model, encoder-decoder model, etc.), which must be stable enough to be accelerated.
[0217] Accelerator 1014 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA) 1009 (alternately referred to herein as “Deep Learning Accelerator Cluster (XNN) 1009”, “Neural Network Accelerator (NNA) 1009”, or “Neural Processing Unit (NPU) 1009”). DLA 1009 may include one or more tensor processing units (TPUs) 1041, which may be configured to provide additional, for example, trillions of operations per second, for deep learning applications and inference. TPUs 1041 may be accelerators configured and optimized for image processing functions, such as for CNNs, RCNNs, DNNs, etc. DLA 1009 may be further optimized for a specific set of neural network types and floating-point operations as well as inference. The DLA is designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperforms CPUs. TPUs 1041 may perform several functions, including single-instance convolution functions, for example, supporting INT8, INT16, and FP16 data types as features and weights, and post-processor functions. Although TPU 1041 is described as being included in DLA 1009, this is not intended to be limiting; TPU 1041 may be included in additional or alternative accelerator 1014 and / or other components, and / or may be included as a discrete processing component.
[0218] The DLA 1009 can quickly and efficiently execute neural networks on processed or unprocessed data to achieve a variety of functions, including but not limited to: object and feature recognition and detection using data from one or more sensor modalities (e.g., vehicles, pedestrians, other robots, lane lines, road boundary lines, debris, potholes, boxes, warehouse items, etc.); distance estimation using data from one or more sensor modalities; emergency vehicle detection and identification using data from microphones and / or vision-based sensors; facial recognition; pick-up and place operations; maneuvering operations; occupant monitoring; vehicle owner identification; and / or other in-cabin operations using data from in-cabin cameras and / or other sensor types; and / or safety and / or safety-related events, to name a few.
[0219] The DLA 1009 can perform any function of the GPU 1008, and by using inference accelerators, for example, designers can anchor either the DLA 1009 or the GPU 1008 for any function. For instance, designers can centralize the processing of DNNs and floating-point operations on the DLA 1009, leaving other functions to the GPU 1008 and / or other accelerators 1014. The DLA 1009 can be used to run any type of network to enhance control and security, including, for example, neural networks that output a confidence metric for each object detection.
[0220] Accelerator 1014 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA) 1007, which may alternatively be referred to herein as a computer vision accelerator or generally as a vision accelerator. PVA 1007 may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), semi-autonomous driving, autonomous driving, robotics applications, safety and supervision applications, augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) applications. PVA 1007 provides a balance between performance and flexibility. For example, each PVA 1007 may include (e.g., but not limited to) any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) systems, Pixel Processing Engines (PPEs), Vector Processors or Vector Processing Units (VPUs), and / or other components. The PVA engine may include an Advanced Very Long Instruction Word (VLIW) or Single Instruction Multiple Data (SIMD) digital signal processor. PVA 1007 may be optimized for image processing and computer vision algorithm acceleration tasks. For example, the PVA 1007 offers excellent performance with extremely low power consumption and can be used asynchronously and concurrently as part of a heterogeneous computing pipeline with other accelerators in CPU 1006, GPU 1008 and / or systems (e.g., vehicles, robots, etc.).
[0221] The PVA 1007 may include one or more (e.g., two) Vector Processing Subsystems (VPSs), each of which may include one or more Vector Processing Unit (VPU) cores, one or more Decoupled Lookup Units (DLUTs), one or more shared or vector memory (VMEMs), and one or more instruction caches (I-caches). The VPU core may be the main processing unit and may include a vector SIMD VLIW DSP 1043 optimized for computer vision. The VPU core can fetch instructions via the I-cache and access data via the VMEM. The DLUT may include dedicated hardware components that enhance the efficiency of parallel lookup operations. For example, the DLUT allows parallel lookups using a single copy of the lookup table by performing these lookups in a decoupled pipeline independent of the main processor pipeline. By doing so, the DLUT can minimize or reduce memory usage and increase throughput while avoiding data-related memory library conflicts, ultimately leading to improved overall system performance. The VPU VMEM can provide local data storage for the VPU, allowing for the efficient implementation of various image processing and computer vision algorithms. The VPU VMEM can support access from external VPS hosts, such as Direct Memory Access (DMA) and the CPU 1006 (e.g., an ARM Cortex-R5 processor), thereby facilitating data exchange with the CPU 1006 and other system-level components. The VPU I-cache can provide instruction data to the VPU upon request, request missing instruction data from system memory, and / or maintain temporary instruction storage for the VPU. For each VPU task, the CPU 1006 can configure the DMA system, optionally prefetching the VPU program into the VPU I-cache, and / or initiating each VPU-DMA pair to process the task. The PVA 1007 may also include L2 SRAM memory shared between one or more (e.g., two) sets of VPS and DMA. In some embodiments, one or more (e.g., two) DMA devices are used to move data between external memory, PVA L2 memory, VMEM (e.g., one per VPS), CPU tightly coupled memory (TCM), DMA descriptor memory, and / or PVA-level configuration registers. In lightly loaded systems, two parallel DMA accesses to DRAM can achieve read / write bandwidth of up to 15 GB / s, while in heavily loaded systems, this bandwidth can reach up to 10 GB / s. Regarding compute capacity, INT8 gigabyte multiply-accumulate operations per second (GMAC) can be 2048 or greater, excluding DLUT. FP32 GMACs can be 32 per PVA instance.
[0222] The RISC core can interact with image sensors (e.g., the image sensor of any camera described herein), image signal processors, etc. Each RISC core may include any amount of memory. The RISC core can use any of a variety of protocols, depending on the implementation. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0223] The DMA system enables components of the PVA 1007 to access system memory independently of the CPU 1006. DMA can support any number of functions optimized for the PVA 1007, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0224] A vector processor, or VPU, can be a programmable processor designed to efficiently and flexibly execute computer vision algorithms and provide signal processing capabilities. In some examples, the PVA 1007 may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA 1007 and may include one or more vector processing units (VPUs), one or more pixel processing engines (PPEs) (which may include a 2D layout of interconnected (e.g., for north, south, east, and west communication) processing elements), one or more instruction caches, and / or one or more shared or vector memories (e.g., VMEM). The VPU core may include a digital signal processor, such as a single-instruction, multiple-data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.
[0225] In some embodiments, each vector processor may include an instruction cache and may be coupled to dedicated memory. Therefore, in some examples, each vector processor may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA 1007 may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA 1007 may execute the same computer vision algorithm, but for different regions of an image. In other examples, the vector processors included in a particular PVA 1007 may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on consecutive images or portions of an image. Among other things, the hardware acceleration cluster may include any number of PVA 1007s, and each PVA may include any number of vector processors. Furthermore, the PVA 1007 may include additional error correction code (ECC) memory to enhance overall system security.
[0226] Accelerator 1014 (e.g., a hardware accelerator cluster) has broad applications for autonomous and semi-autonomous machine control. PVA 1007 can be a programmable vision accelerator used in critical processing stages in perception, robot understanding and reasoning, ADAS, semi-autonomous and autonomous vehicles, etc. PVA 1007's capabilities are well-suited for algorithmic domains requiring predictable processing, featuring low power consumption and low latency. In other words, PVA 1007 performs well in semi-intensive or intensive rule computation, even on small datasets requiring predictable runtime, low latency, and low power consumption. Therefore, in the context of autonomous vehicles and robotic platforms, PVA 1007 is designed to run classic computer vision algorithms, as they are highly efficient in object detection and integer mathematical operations.
[0227] For example, according to one embodiment of this technology, the PVA 1007 is used to perform computer stereo vision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended to limit it. Many Level 3-5 autonomous driving applications require on-the-fly motion estimation / stereo matching (e.g., moving structures, pedestrian recognition, lane detection, etc.). The PVA 1007 can perform computer stereo vision functions on input from two monocular cameras.
[0228] In some examples, the PVA 1007 can be used to perform dense optical flow, providing processed RADAR data based on the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, the PVA 1007 is used for time-of-flight depth processing, for example, providing processed time-of-flight data by processing the raw time-of-flight data.
[0229] While the VPU, DMA, RISC Core, VMEM, and decoupled coprocessors (e.g., DLUT) are described as being included within the PVA1007, this does not imply limitation. In some embodiments, these components may be included in alternative or additional processing components and / or accelerators 1014, and / or may be included as discrete components of SoC 1004 and / or other computing system architectures.
[0230] In some examples, SoC 1004 may include a real-time ray tracing hardware accelerator (RTA) 951, which can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time or near-real-time visualization simulations, for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR, RADAR, LiDAR, camera and / or other sensor modalities in a simulation, for general wave propagation simulation, for comparison with LiDAR data for localization, for generating realistic training data for training neural networks, and / or other functions and uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations. For example, machine 1000 (or another machine or device) may perform simulations in a simulation environment, and one or more light transport simulation algorithms (e.g., ray tracing, path tracing, etc.) may be used to generate the simulation environment. Therefore, the ray tracing accelerator 1051 and / or a ray tracing-optimized GPU 1008 (e.g., NVIDIA's RTX GPU) may be used to accelerate these ray tracing algorithms.
[0231] Accelerator 1014 (e.g., in a hardware acceleration cluster) may include one or more optical flow accelerators (OFAs) 1011. For example, OFA 1011 can be used to calculate optical flow and stereo disparity between sensor data frames (e.g., images). Optical flow can be accelerated on OFA 1011 for purposes such as object detection and tracking, and / or for stereo depth estimation, where stereo disparity is calculated between stereo image frames (e.g., two or more frames captured using two or more image sensors with at least partially overlapping fields of view).
[0232] SoC 1004 may include one or more Camera Serial Interfaces (CSIs) 1023. For example, CSI 1023 may include a Mobile Industry Processor Interface (MIPI) Camera Serial Interface (CSI) for receiving video and input from cameras, a high-speed interface, and / or a video input block available for camera and associated pixel input functions. SoC 1004 may also include a software-controllable input / output controller and may be used to receive I / O signals not assigned to a specific role. For example, CSI 1023 may include MIPI CSI-2 connectors, such as a 16-channel MIPI CSI-2 connector, D-PHY 2.1 (up to 40Gbps), and C-PHY 2.0 (up to 164Gbps) to support 16 virtual channels and 6 or more cameras; an 8-channel MIPI CSI-2 connector, D-PHY 2.1 (up to 20Gbps) to support 8 virtual channels and 4 or more cameras; and / or 2x MIPI CSI-2, 22-pin camera connectors, depending on the embodiment and implementation.
[0233] Accelerator 1014 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network (CVNOC) 1063 and SRAM for providing high-bandwidth, low-latency SRAM to accelerator 1014. In some examples, on-chip memory may include at least 4 MB of SRAM, such as, but not limited to, eight field-configurable memory blocks accessible by PVA 1007, OFA 1011, DLA 1009, and / or other accelerators 1014. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory 1015 may be used. PVA 1007, OFA 1011, DLA 1009, and / or other accelerators 1014 may access memory via a backbone providing high-speed memory access to accelerator 1014. The backbone may include an on-chip computer vision network that interconnects accelerator 1014 to memory (e.g., using APB).
[0234] CVNOC 1063 may include an interface that determines whether accelerator 1014 provides ready and valid signals before transmitting any control signals / addresses / data. Such an interface can provide separate stages and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface may conform to ISO 26262 or IEC 61508 standards, although other standards and protocols may be used.
[0235] SoC 1004 may include a data repository 1016 and / or memory 1015. The data repository 1016 may be on-chip memory 1015 of SoC 1004, which may store neural networks and / or other algorithms to be executed on CPU 1006, GPU 1008, and / or one or more accelerators 1014. In some examples, the capacity of the data repository 1016 may be large enough to store multiple neural network instances for redundancy and security. The data repository 1016 may include, for example, L2 and / or L3 cache 1012. The memory 1015 may include SRAM, LPDDR5, and / or other memory types. For example, the memory 1015 may include 4MB SRAM, 32GB and / or 64GB 256-bit LPDDR5 (204.8GB / s), 8GB and / or 16GB 128-bit LPDDR5 (102.4GB / s), and / or other memory types and sizes. References to data repository 1016 may include references to memories associated with PVA 1007, OFA 1011, DLA 1009 and / or other accelerators 1014, as described herein.
[0236] Data repository 1016 may include various storage types, such as eMMC, NVMe, etc. For example, SoC 1004 may include storage in the form of an embedded multimedia card (eMMC) (e.g., 64GB eMMC 5.1) and / or an SD card slot, with external NVM Express (NVMe) capability, for example, via M.2 Key M. For example, data repository 1016 and / or other storage may be accessed via, for example, NVMe using PCI Express (PCIe), RDMA, TCP, and / or other protocols.
[0237] SoC 1004 may include one or more processors 1010 (e.g., embedded processors). Processor 1010 may include a boot and power management processor (BPMP) 1053, which may be a dedicated processor and subsystem for handling boot power and management functions, as well as related security implementations. BPMP 1053 may be part of the SoC 1004 boot sequence and may provide runtime power management services. BPMP 1053 may provide clock and voltage programming, assistance with low-power state transitions of the system, management of SoC 1004 thermal and temperature sensors, and / or management of SoC 1004 power states. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 1004 may use the ring oscillator to detect the temperature of CPU 1006, GPU 1008, accelerator 1014, and / or other components. If the temperature is determined to exceed the threshold, the BPMP 1053 may enter the temperature fault routine and put the SoC 1004 into a lower power state and / or put the machine 1000 into a driver safety stop mode (e.g., to safely stop the machine 1000).
[0238] The processor 1010 may also include a set of embedded processors that can be used as an audio processing engine (APE) 1055. The APE 1055 can be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces, as well as a wide and flexible audio I / O interface. In some examples, the APE 1055 is a dedicated processor core with a digital signal processor and dedicated RAM.
[0239] The processor 1010 may also include an always-on processor engine (AOPE) 1057, which can provide the necessary hardware features to support low-power sensor management and wake-up use cases. AOPE 1057 may include a processor core, tightly coupled RAM, support for peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0240] Processor 1010 may also include a security processor 1013 (or "security island 1013"), which may include a security cluster engine comprising a dedicated processor or processor subsystem for handling security management for automotive, robotic, and / or other applications. Security processor 1013 and / or the security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and as a single core, with comparison logic to detect any differences in their operation. In some embodiments, security processor 1013 may include discrete processors such that failure of other system components may not affect the performance and availability of security processor 1013.
[0241] The processor 1010 may also include a real-time or near-real-time sensor engine (SE) 1059, which may include a dedicated processor subsystem for handling real-time or near-real-time camera, LiDAR, RADAR and / or other sensor modal management.
[0242] The processor 1010 may also include one or more image signal processors (ISPs) 1027, which may include high dynamic range signal processors and / or hardware engines as part of one or more sensor processing pipelines.
[0243] The processor 1010 may include a video image synthesizer (VIC) 961, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to produce the final image of the player window. The VIC 1061 can perform lens distortion correction on the wide-angle camera 1068B, the surround camera 1068D, the cabin monitoring camera sensor, and / or other camera sensors with distorted fields of view.
[0244] VIC 1061 may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, noise reduction appropriately weights spatial information, thereby reducing the weight of information provided by adjacent frames. When an image or part of an image does not contain motion, temporal noise reduction performed by the video image synthesizer can use information from previous images to reduce noise in the current image.
[0245] The VIC 1061 can also be configured to perform stereoscopic correction on input stereoscopic camera frames. The video image compositor can also be used for user interface compositing when the operating system desktop is in use and the GPU 1008 does not need to continuously render new surfaces. Even when the GPU 1008 is powered on and actively performing 3D rendering, the video image compositor can be used to offload the GPU 1008 to improve performance and responsiveness.
[0246] SoC 1004 may also include various peripheral interfaces for input / output (I / O) 1025, for example, to enable communication with peripheral devices, audio codecs, power management and / or other devices. SoC 1004 can be used to process data from cameras (e.g., via a gigabit multimedia serial link and / or Ethernet connection), from sensors (e.g., LiDAR sensor 1064, RADAR sensor 1060, etc., connected via Ethernet), from bus 1002 (e.g., speed of machine 1000, steering wheel position, etc.), and from GNSS sensor 1058 (e.g., connected via Ethernet or CAN bus). SoC 1004 may also include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine and can be used to free CPU 1006 from routine data management tasks. In some embodiments, the SoC 1004 I / O 1025 may include headers (e.g., 40-pin headers or 40-pin extension headers) supporting Universal Asynchronous Receiver / Transmitter (UART), Serial Peripheral Interface (SPI), Inter-Integrated Sound (I2S), Inter-Integrated Sound (I2C), Controller Area Network (CAN), Pulse Width Modulation (PWM), Digital Microphone Interface (DMIC), Digital Speaker Station (DSPK), General Purpose I / O (GPIO), etc.; automation headers (e.g., 12-pin automation headers); audio panel headers (e.g., 10-pin audio panel headers); Joint Test Action Group (JTAG) headers (e.g., 10-pin JTAG headers); fan headers (e.g., 4-pin fan headers); RTC battery spare connectors (e.g., 2-pin battery spare connectors); microSD slot; DC power jack; power, force, resume, and reset buttons; one or more display connectors (e.g., DisplayPort (DP), such as DP 1.4A (+MST), eDP 1.41, HDMI 2.1, and / or 4K30 multi-model DP). 1.2 (+MST) connector) and / or other I / O 1025 elements, components or features.
[0247] SoC 1004 may include machine intranetting capabilities using, for example, Ethernet (e.g., automotive Ethernet), SERDES, Controller Area Network (CAN), FlexRay, Local Interconnect Network (LIN), Low Voltage Differential Signaling (LVDS), Media-Oriented System Transport (MOST), another network type, and / or combinations thereof. For example, SoC 1004 may include RJ45 connectors with up to 10GbE, 1GbE connectors, and / or other network connector types.
[0248] So104 may include one or more digital signal processors (DSPs) 943. For example, DSP 1043 may include a dedicated or specialized microprocessor chip optimized for digital signal processing, such as audio signal processing, telecommunications, digital image processing, RADAR, SONAR, LiDAR and / or other sensor processing, speech recognition and / or other applications.
[0249] SoC 1004 may include one or more video encoders 1019 and / or one or more video decoders 1021. For example, video encoder 1019 may include a hardware-based (e.g., as part of GPU 1008) video encoder (e.g., supporting H.264, H.265, etc., and conforming to HEVC standards, such as NVIDIA's NVENC) that can process image input (e.g., as YUV, RGB, etc.) to generate a video bitstream. Video decoder 1021 may include a video decoder engine that provides fully accelerated hardware video decoding capabilities (e.g., supporting decoding of various bitstream formats such as AV1, H.264, H.265, VP8, VP9, MPEG-1, MPEG-2, MPEG-4, VC-1, etc., and conforming to HEVC standards, such as NVIDIA's NVDEC). In some examples, video decoder 1021 may be hardware-based (e.g., as part of GPU 1008).
[0250] SoC 1004 may include one or more General Purpose Computing Acceleration Clusters (GCACs) 1029. For example, GCAC 1029 may include various processor types that can be used to accelerate computing, such as one or more Vector Microcode Processors (VMPs) 1033, one or more Multi-Threaded Processing Clusters (MPCs) 1031, one or more Programmable Macroarrays (PMAs) 1035, and / or one or more other processor types. For example, GCAC 1029 may include PMA 1035, two VMPs 1033, and two MPCs 1031.
[0251] SoC 1004 may include one or more vector microcode processors (VMPs) 1033. In an embodiment, VMP 1033 may include a wide vector (Very Long Instruction Word (VLIW) and Single Instruction Multiple Data (SIMD)) machine that performs a variety of operations, such as short integral-type operations common in computer vision and deep learning algorithms.
[0252] SoC 1004 may include one or more multi-threaded processing clusters (MPCs) 1031. MPCs 1031 may include processing clusters that are more general-purpose than GPUs and more efficient than CPUs in some embodiments. For example, MPC 1031 may include a multi-threaded processor that allows multiple threads to share resources and execute instructions concurrently.
[0253] SoC 1004 may include one or more programmable macro arrays (PMAs) 1035. PMA 1035 may include coarse-grained reconfigurable architecture (CGRA) dataflow machines, which have a unique architecture that can deliver powerful performance on intensive computer vision and deep learning algorithms that may not be achievable in classic digital signal processing (DSP) architectures.
[0254] SoC 1004 may include one or more Display Processing Units (DPUs) 945 for performing hardware-accelerated image processing. For example, DPU 1045 may retrieve pixel data from memory 1015 and send it to display peripherals via a standard interface. Therefore, DPU 1045 can handle display processing and rendering for displays within and / or on the machine.
[0255] SoC 1004 may include one or more application processing units (APUs) 1039. For example, APU 1039 may include a quad-core or dual-core processor with 48KB / 32KB L1 cache (with parity and ECC) and 1MB L2 cache with ECC. APU 1039 may support NEON instructions and single-precision and double-precision floating-point operations.
[0256] The SoC 1004 may include one or more Real-Time Processing Units (RTPUs) 1069. The RTPU 1069 may include a dual-core processor with 32KB / 32KB L1 cache and a 256KB TCM with ECC. The RTPU 1069 can support single-precision and double-precision floating-point operations.
[0257] SoC 1004 may include one or more built-in self-test (BIST) components 1037. For example, BIST component 1037 may include a memory BIST (MBIST) for testing the system's memory and / or a logic BIST (LBIST) for testing the system's logic. BIST component 1037 may include embedded logic for directly testing the system's logic and / or memory.
[0258] SoC 1004 may include one or more dynamically reconfigurable processors (DRPs) 1071. For example, DRP 1071 can be used to accelerate various computational operations. For instance, in one embodiment, DRP 1071 may be combined with a MAC unit to function as an AI accelerator. In another embodiment, DRP 1071 can execute an application while dynamically switching the circuit connection configuration of the arithmetic unit (e.g., ALU) on the chip each operating clock according to what needs to be processed. Because it uses only the necessary arithmetic circuitry, DRP 1071 can consume less power than a CPU and achieve higher speeds. Furthermore, compared to a CPU, which suffers from performance degradation due to frequent accesses to external memory for cache misses and other reasons, DRP 1071 can pre-build the necessary data paths in the hardware, thereby reducing performance degradation and operating speed variations (jitter) caused by memory accesses. DRP 1071 may include a dynamic loading function that switches circuit connection information each time the algorithm changes, enabling processing with limited hardware resources, even in robotics / automotive applications that require processing multiple algorithms.
[0259] In some embodiments, accelerator 1014 may include an OpenCV accelerator for accelerating OpenCV processing, OpenCV being an open-source industry-standard library for computer vision processing. In some embodiments, the combination of one or more DRP 1071s deployed as AI accelerators with an OpenCV accelerator can enhance AI computation and image processing algorithms to enable complex and computationally intensive operations such as visual simultaneous localization and mapping (SLAM).
[0260] Compared to traditional systems, the techniques described herein, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously (e.g., at least partially in parallel) and / or sequentially, and the results combined to achieve Level 2-5 autonomous driving capabilities and / or autonomous robot motion, control, planning, and / or navigation operations. Furthermore, since the SoC 1004 can include various computing engines (e.g., processor 1010, CPU 1006, GPU 1008, accelerator 1014, etc.), tasks can be distributed among the computing engines, and in some cases, common-cause failures are avoided due to the discrete footprint of the computing engines. Additionally, since the SoC 1004 can include a dedicated safety processor 1013 (or safety island 1013), critical safety or redundant operations can be performed without the common-cause failures of the main processing components or computing engines of the SoC 914. Due to these features, the underlying systems of the SoC 1004 and / or machine 1000 may be able to meet higher safety levels—such as the Automotive Safety Integrity Level (ASIL) D of the ISO 26262 standard.
[0261] Figure 10E According to some embodiments of this disclosure, cloud-based servers (e.g., servers such as those described herein in a data center) and Figure 10A The following is a system diagram illustrating communication between an example autonomous or semi-autonomous vehicle or machine 1000. System 1076 may include server 1078, network 1090, and machine 1000. Server 1078 may include multiple GPUs 1084(A)-1084(H) (collectively referred to herein as GPU 1084), switches 1082(A)-1082(H) (e.g., PCIe 4.0 / 5.0 switches, M.2 slots, Thunderbolt, USB4, NVIDIA's NVLink, NVIDIA's NVSwitch, GPUDirectRDMA, GPUDirect Storage, etc.), CPUs 1080(A)-1080(B) (collectively referred to herein as CPU 1080), accelerators, and / or other processor types. GPU 1084, CPU 1080, and PCIe switches may interconnect with high-speed interconnects, such as, but not limited to, NVIDIA-developed NVLink interface 1088 and / or PCIe connection 1086. In some examples, the GPU 1084 is connected via NVLink and / or NVSwitch SoC, and the GPU 1084 and PCIe switch 1082 are connected via PCIe interconnect. While the figure shows eight GPUs 1084, two CPUs 1080, and two PCIe switches, this is not limiting. According to embodiments, each server 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, each server 1078 may include eight, sixteen, thirty-two, and / or more GPUs 1084.
[0262] Server 1078 may receive sensor data from machine 1000 via network 1090, indicating information about new or previously unexplored locations, and / or sensor data indicating changes to previously seen / stored locations (e.g., unexpected or altered road conditions, such as recently started road construction). Server 1078 may transmit neural network 1092, updated neural network 0192, map information 1094, etc., including information about traffic and road conditions, to machine 1000 via network 1090. Updates to map information 1094 may include updates to HD maps 1022, SD maps, navigation maps, etc., such as information about construction sites, potholes, detours, floods, and / or other obstacles. In some examples, neural network 1092, updated neural network 1092, map information 1094, and / or other information may come from new training and / or experience, represented in data received from any number of machines 1000 in the environment, and / or based on training performed in a data center (e.g., using server 1078 and / or other servers).
[0263] Server 1078 can be used to train a machine learning model (e.g., a neural network) based on training data. Training data can be generated by machine 1000 and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is unlabeled and / or unprocessed (e.g., the neural network does not require supervised learning). Training can be performed according to any one or more machine learning techniques, including but not limited to the following categories: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by machine 1000 (e.g., transmitted to machine 1000 via network 1090), and / or the machine learning model can be used by server 1078 to remotely monitor and / or control machine 1000.
[0264] In some examples, server 1078 can receive data from machine 1000 and apply the data to a state-of-the-art real-time neural network for real-time intelligent inference. Server 1078 may include a deep learning supercomputer powered by GPU 1084 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1078 may include a deep learning infrastructure in a data center that uses only CPU power.
[0265] The deep learning infrastructure of server 1078 may be capable of rapid, real-time inference and can use this capability to assess and verify the health of the processor, software, and / or related hardware in machine 1000. For example, the deep learning infrastructure may receive periodic updates from machine 1000, such as image sequences and / or objects located by machine 1000 in the image sequence (e.g., through computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with the objects identified by machine 1000. If the results do not match and the infrastructure concludes that the AI in machine 1000 has malfunctioned, server 1078 may send a signal to machine 1000 instructing the fail-safe computer of machine 1000 to take over control, notify the occupants, and perform safety maneuvers or operations, such as slowing down, returning control to the driver, stopping, and / or pulling over / closing the vehicle.
[0266] For inference, server 1078 may include GPU 1084 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-driven server and inference acceleration enables real-time response. In other examples, such as in scenarios where performance is less critical, servers driven by CPUs, FPGAs, and other processors can be used for inference.
[0267] Computational ecosystem for generating, training, and deploying AI Figure 11 This is a system diagram illustrating a three-computer ecosystem 1100 according to at least some embodiments of the present disclosure, including a first computing system 1102 for generating or creating artificial intelligence (AI) (e.g., AI training and validation data), a second computing system 1104 for training the AI, and a third computing system 1106 for deploying AI at the edge (which may include or correspond to...). Figures 10A to 10E (SoC 1004). For example, to develop and deploy materialized or physical AI, a three-computer ecosystem 1100 can be used, comprising three accelerated computer systems to handle physical AI training, simulation, and runtime (e.g., edge deployment). These systems can use scalable, physical-based simulations of the Machine 1000 and its world to generate training data and train multimodal base models (and / or other model types). By doing so, simulations of the Machine 1000 can be performed at scale, allowing skills (e.g., robotic skills) to be improved, tested, and optimized in virtual worlds that simulate the laws of physics (e.g., using NVIDIA's OMNIVERSE), thereby helping to reduce the cost of real-world data acquisition and ensuring that the Machine 1000 can operate safely in a controlled environment.
[0268] Computing system 1104 (e.g., NVIDIA's DGX platform) can be used to train and fine-tune powerful foundational and generative AI models. Models such as general foundational models (e.g., NVIDIA's Project GR00T) can be used to enable robots and other machines 1000 to understand natural language and mimic actions by observing human movements. Computing system 1104 may include a platform that integrates software, infrastructure, and expertise into a modern, unified AI development and training solution. Computing system 1104 may include individual computing devices 1110 (e.g., NVIDIA's DGX B200, H200, etc.) and / or any number of computing devices 1110 (e.g., NVIDIA's DGX SuperPOD) within data center infrastructure 1112.
[0269] For example, a standalone computing device 1110 may include GPUs (e.g., 8 GPUs with a total GPU memory of 1,440 GB) and CPUs (e.g., 2 CPUs with a total of 112 cores, 2.1 GHz or 4 GHz (with enhancements)), providing up to 72 petaFLOPS of training capability and 144 petaFLOPS of inference capability. The computing device 1110 may include memory (e.g., 4 TB of memory) and storage (e.g., 2 x 1.9 TB NVMe M.2 OS storage and 8 x 3.84 TB NVMe U.2 internal storage). The computing device 1110 may include various networking and network management components, such as OSFP ports (e.g., 4 OSFP ports) for servicing single-port intelligent host channel adapters (e.g., 8 single-port ConnextX-7 Virtual Protocol Interconnect (VPI)), providing up to 400 GB / s of Infiniband / Ethernet. The computing device 1110 may also include, for example, a dual-port quad small pluggable (QSFFP) data processing unit (DPU) (e.g., two dual-port QSFP112 DPUs – such as NVIDIA’s BlueField-3 DPU) providing up to 400Gb / s of InfiniBand / Ethernet. The computing device 1110 may include an onboard network interface card (NIC) (e.g., a 10Gb / s onboard NIC with RJ45), a dual-port Ethernet NIC (e.g., a 100GB / s dual-port Ethernet NIC), and / or a host board management controller (MBC) (e.g., with RJ45). In some embodiments, the NIC for the computing device 1110 may include a SuperNIC (e.g., NVIDIA’s ConnectX-8 SuperNIC) to provide up to 800Gb / s of data throughput for in-network computing acceleration engines, thereby providing the performance and robust feature set required to support trillion-parameter-scale AI factories and scientific computing workloads. In other embodiments, computing device 1110 may include a smart host channel adapter (HCA) (e.g., NVIDIA's ConnectX-7) to provide ultra-low latency and 400Gb / s throughput for computing acceleration engines within the network.
[0270] Data center infrastructure 1112 may include: any number of computing devices 1110, and an operating system (OS) (e.g., a DGX OS extension of a Linux distribution) for maximizing system uptime, security, and reliability; network / storage acceleration libraries and management for accelerating end-to-end infrastructure performance; cluster management for scaling and managing a single node (e.g., a single computing device 1110) to thousands of nodes; job scheduling and orchestration to ensure the smooth execution of each developer's jobs; AI workflow management and machine learning operations (MLOps) for moving more models from prototypes to production; and enterprise software to accelerate developer success.
[0271] Computational system 1102 (e.g., NVIDIA's OVX server) can provide a development and simulation platform for testing and optimizing physics-based AI using APIs and frameworks for simulation (e.g., NVIDIA's DriveSIM, ISAAC Sim, ISAAC Gym, ISAAC Lab, etc.). Computational system 1102 allows developers to use simulation frameworks to simulate and validate robot models, and / or generate large amounts of physics-based synthetic data to guide model training. Computational system 1102 can support learning frameworks that power robot reinforcement learning and imitation learning to accelerate robot policy training and refinement. For example, computational system 1102 can be used to generate any number of simulations 1108, such as in NVIDIA's OMNIVERSE. The Computing System 1102 can be used to optimize and accelerate the entire software stack, from training, fine-tuning, and deploying generative AI to supporting industrial digitization within content collaboration platforms, including APIs, software development kits (SDKs), and services. These platforms allow the integration of OpenUSD, ray tracing rendering technologies (such as NVIDIA's RTX), and generative physics AI into existing software tools and simulation workflows, such as those for industrial and robotics use cases (such as NVIDIA's OMNIVERSE). Therefore, the Computing System 1102 can host or support native OpenUSD software platforms, enabling enterprises to connect 3D pipelines and develop advanced real-time 3D applications for industrial digitization. With powerful ray tracing-accelerated AI and graphics capabilities, the Computing System 1102 delivers robust performance for workloads such as extended reality (XR), multi-user design collaboration, and digital twins. This allows for the creation of physically accurate models with high-fidelity ray tracing and path-tracing material rendering, large-scale, AI-enabled simulation operations, and the generation of realistic 3D synthetic data for training. The computing system 1102 may include a single computing device 1114 (e.g., an NVIDIA OVX L40S server) and / or any number of computing devices 1114 (e.g., NVIDIA OVX systems) in the data center infrastructure 1116.
[0272] Computing device 1114 (which may include a server) may include CPUs (e.g., two CPUs, each with 32 cores) and GPUs (e.g., four or eight GPUs, each including 48GB GDDR6 with ECC memory, 864GB / s memory bandwidth, a PCIe Gen4x16:64GB / s bidirectional interconnect interface, 18,176 CUDA cores, 142 ray tracing (RT) cores, and 568 Tensor cores). Computing device 1114 may include various networking and network management components, such as Intelligent Host Channel Adapters (HCAs) (e.g., two or four single-port ConnextX-7s, each with 200Gb / s, providing up to 800Gb / s InfiniBand / Ethernet), one or more DPUs (e.g., dual-port QSFP112 DPUs, such as the NVIDIA BlueField-3 DPU), providing up to 400Gb / s InfiniBand / Ethernet. In some embodiments, the NIC for computing device 1114 may include a SuperNIC (e.g., NVIDIA's ConnectX-8 SuperNIC) to provide up to 800Gb / s of data throughput for in-network compute acceleration engines, thereby providing the performance and robust feature set required to support trillion-parameter-scale AI factories and scientific computing workloads. In other embodiments, computing device 1114 may include a Smart Host Channel Adapter (HCA) (e.g., NVIDIA's ConnectX-7) to provide ultra-low latency and 400Gb / s throughput for in-network compute acceleration engines. Computing device 1114 may include host memory (e.g., 384Gb DDR5 ECC for four GPUs, or 768Gb DDR5 ECC for eight GPUs) and may include dual in-line memory module (DIMM) slots, host boot drives (e.g., 1TB NVMe), and / or host storage (e.g., 24TB NVMe).
[0273] Similar to data center infrastructure 1112, data center infrastructure 1116 can allow any number of computing devices 1114 to be combined into a cluster configuration according to a reference architecture.
[0274] Computing system 1106 can be used to deploy trained AI models on a runtime computer, such as the SoC 1004 described herein. For example, these computing systems 1106 can be designed for compact onboard computing needs, including ensembles of models such as control policies, vision, and language models deployed on an energy-efficient onboard edge computing system 1106. (See also: [link to document]). Figures 10A to 10EA more detailed description of the components, features, and capabilities of the computing system 1106.
[0275] Example Generative Model In at least some embodiments, language models such as Large Language Models (LLM), Visual Language Models (VLM), Multimodal Language Models (MMLM), Visual Language Action (VLA) models, and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD formats such as OpenUSD), and / or the like, based on context provided in input prompts or queries. In embodiments, these language models may be considered "large" because they are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—e.g., millions or billions of parameters. This disclosure allows for the implementation of LLM / VLM / MMLM, etc., for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), extracting insights from data (e.g., text, images, videos, etc.), and generating new text / images / videos / , etc., in a user-specified style, tone, and / or format. In some embodiments, the LLM / VLM / MMLM, etc., of this disclosure may be specifically designed for text processing, while in other embodiments, multimodal LLMs may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio (sound, synthesized speech, etc.), 2D and / or 3D data (e.g., USD format), and / or video. For example, a visual language model (VLM) or more specifically a multimodal language model (MMLM) may be implemented to accept images, videos, sensor data, audio, text, 3D designs (e.g., CAD), and / or other input data types and / or generate or output images, videos, audio, text, 3D designs, and / or other output data types.
[0276] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented using different techniques to understand and generate outputs (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, converter architectures (e.g., architectures relying on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. may also include one or more diffusion blocks (e.g., noise reduction blocks). The LLM / VLM / MMLM / etc. of this disclosure may include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc., including encoder and decoder components (e.g., T5 (Text-to-Text Transformer)), can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting and any architecture type (including, but not limited to, those described herein) can be implemented depending on the specific implementation and the task performed using LLM / VLM / MMLM / etc.
[0277] In various embodiments, LLM / VLM / MMLM / etc. can be trained using unsupervised learning, where LLM / VLM / MMLM / etc. learns patterns from a large amount of unlabeled text / audio / video / image / design / USD / etc. data. Due to the extensive training, in these embodiments, the model may not require task-specific or domain-specific training. An LLM / VLM / MMLM / etc. extensively pre-trained on a large amount of unlabeled data can be referred to as a base model and can excel at various tasks, such as question answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as cue tuning, fine-tuning, retrieval augmentation generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers to tune or adjust cues or labels to bias the language model towards a specific task or domain), and / or using optimization models for specific tasks and / or other fine-tuning or customization techniques for specific domains.
[0278] In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using LLM / VLM / MMLM / etc., and / or to prevent the output or presentation of information generated by LLM / VLM / MMLM / etc. (e.g., displays, audio outputs, etc.). In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with the model's inputs and / or outputs. For example, these "protective" models can be trained to identify "safe" or otherwise okay or desired inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a particular application / implementation. Therefore, the LLM / VLM / MMLM / etc. disclosed herein are unlikely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, insecure, out of scope, and / or unwanted for a particular application / implementation.
[0279] In some embodiments, an LLM / VLM / etc. can be configured or able to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations where the model is not ideally suited, the model may have instructions for accessing one or more plugins (e.g., third-party plugins) to help process the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such an example, when at least part of the prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least part of the response requires mathematical computation, the model can access one or more mathematical plugins or APIs to help solve the problem, and then the response from the plugins and / or APIs can be used in the model's output. This process can be repeated (e.g., recursively) an arbitrary number of iterations, using any number of plugins and / or APIs, until a response to each query / question / request / process / operation / etc. can be generated in response to the input prompt. Therefore, models can rely not only on their own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (such as APIs, plugins, etc.).
[0280] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple hints provided to the same language model or instances of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, the same input query and hints (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents, for example, providing more than one hint to constrain, guide, or otherwise influence the style, content, or character of the provided output. In one or more exemplary non-limiting embodiments, the same language model can be required to provide output corresponding to different roles, perspectives, characters, or different knowledge bases, as defined by the provided hints.
[0281] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated proxies of at least one language model, and / or provided to two or more prompts for at least one language model can be further processed, such as aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or proxy) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be required to generate or otherwise obtain output about the input source material, wherein the output is associated with the input source material. Such association can include, for example, generating captions or text portions embedded (e.g., as metadata) within the input source text or image. In one or more embodiments, the output of the language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, the language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, wherein the text or image is annotated to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a curatorial dataset, for example, but not limited to this.
[0282] Figure 12 This is a block diagram of an example generative language model system 1200 suitable for implementing at least some embodiments of the present disclosure. Figure 12 In the example shown, the generative language model system 1200 includes a retrieval augmentation (RAG) component 1292, an input processor 1205, a tokenizer 1210, an embedding component 1320, a plugin / API 1295, and a generative language model (LM) 1230 (which may include LLM, VLM, MMLM, VLA models, etc.).
[0283] At a high level, the input processor 1205 can receive input 1201, which includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, generic scene descriptor (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 1230 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, input 1201 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 1201 may include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., tabular format, JSON, or XML). In the generative LM In some implementations of 1230 capable of handling multimodal input, input 1201 can combine text (or text that may be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking raw input text as an example, input processor 1205 can prepare the raw input text in various ways. For example, input processor 1205 can perform various types of text filtering to remove noise (e.g., special characters, punctuation marks, HTML tags, stop words, portions of images, portions of audio, etc.) from the relevant text content. In examples involving stop words (common words that often have little semantic meaning), input processor 1205 can remove stop words to reduce noise and enable generative LM. 1230 focuses on more meaningful content. Input processor 1205 can apply text normalization, for example, by converting all characters to lowercase, removing accents, and / or handling special cases (such as abbreviations or abbreviations) to ensure consistency (e.g., converting ¼ to ¼). Similarly, input processor 1205 and / or post-processor can perform inverse text normalization (ITN) to convert plain language back to canonical or other forms (e.g., converting ¼ to ¼). These are just a few examples; other types of input and / or output processing can be applied.
[0284] In some embodiments, RAG component 1292 (which may include one or more RAG models, and / or may be performed using generative LM 1230 itself) may be used to retrieve additional information to be used as part of input 1201 or a prompt. RAGs can be used to enhance input to LLM / VLM / MMLM / etc. with external knowledge to make the answer to a specific question or query or request more relevant, for example, where specific knowledge is required. RAG component 1292 may obtain this additional information from one or more external sources (e.g., basic information such as basic text / images / videos / audio / USD / CAD / etc.), which can then be fed along with the prompt to LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.
[0285] For example, in some embodiments, in addition to the data retrieved using RAG component 1292, input 1201 may also be generated using query or model inputs (e.g., questions, requests, etc.). In some embodiments, input processor 1205 may analyze input 1201 and communicate with RAG component 1292 (or in embodiments, RAG component 1292 may be part of input processor 1205) to identify relevant text and / or other data to provide to generative LM 1230 as additional context or information sources, typically from which responses, answers, or outputs 1290 are identified. For example, when the input indicates that a user is interested in the required tire pressure for a particular brand and model of vehicle, RAG component 1292 may use a RAG model, for example, to perform a vector search in the embedding space to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the owner's manual for that particular vehicle brand and model. Similarly, when a user revisits the chatbot related to a specific product sale or service, the RAG component 1292 can retrieve previously stored conversation history (or at least its summary) and provide the previous conversation history, along with the current inquiry / request, as part of the generative LM 1230 as input 1201.
[0286] RAG component 1292 can use various RAG techniques. For example, naïve RAG can be used, where documents are indexed, chunked, and applied to an embedding model to generate embeddings corresponding to chunks. User queries can also be applied to this embedding model and / or another embedding model of RAG component 1292, and the embeddings of chunks can be compared with the embeddings of the query to identify the most similar / relevant embeddings to the query. These most similar / relevant embeddings can be provided to generative LM 1230 to generate output.
[0287] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.
[0288] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.
[0289] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which can lead to a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide the model with structured entity information by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, and may also be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can aggregate the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.
[0290] In any embodiment, the RAG component 1292 may implement plugins, APIs, user interfaces, and / or other functions to perform RAG. For example, LLM / VLM / MMLM / etc. may use graph RAG plugins to run queries on knowledge graphs to extract relevant information to feed into the model, and may use standard or vector RAG plugins to run queries on vector databases. For example, the graph database may interact with the plugin's REST interface, thus decoupling the graph database from the vector database and / or embedded models.
[0291] The tokenizer 1210 can segment (e.g., processed) text data into smaller units (tags) for subsequent analysis and processing. Depending on the implementation, the tags can represent individual words, sub-words, characters, audio / video / images, etc. Word-based tokenization divides the text into individual words, treating each word as a separate tag. Sub-word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 1230 to understand morphological changes and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate tag, enabling the generative LM 1230 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, the tokenizer 1210 can transform (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.
[0292] Embedding component 1320 can use any known embedding technique to transform discrete tokens into semantically meaningful (e.g., dense, continuous vector) representations. For example, embedding component 1320 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.
[0293] In some implementations where input 1201 includes image data / video data, etc., input processor 1201 may resize the data to a standard size compatible with the format of the corresponding input channel and / or normalize pixel values to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 1320 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where input 1201 includes audio data, input processor 1201 may resample the audio file to a consistent sampling rate for uniform processing, and embedding component 1320 may use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where input 1201 includes video data, input processor 1201 may extract frames or apply resizing to extracted frames, and embedding component 1320 may extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where input 1201 includes multimodal data, the embedded component 1320 can use techniques such as early fusion (concatenation), late fusion (sequential processing), and attention-based fusion (e.g., self-attention, cross-attention) to fuse representations of different types of data (e.g., text, images, audio, USD, video, design, etc.).
[0294] Other components of the generative LM 1230 and / or generative LM system 1100 may use different types of neural network architectures depending on the implementation scheme. For example, a transducer-based architecture (such as the one used in models like GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feedforward network that processes the output of the self-attention layer, applying a nonlinear transformation to the input representation and extracting higher-level features. Some non-limiting example architectures include transducers (e.g., encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning, linear time series modeling using a selective state-space modeling (SSM) architecture (e.g., the Mamba LLM architecture), etc. Therefore, depending on the implementation scheme and architecture, the embedded component 1320 can apply the encoded representation of the input 1201 to the generative LM 1230, and the generative LM 1230 can process the encoded representation of the input 1201 to generate an output 1290, which may include response text and / or other types of data.
[0295] As described herein, in some embodiments, the generative LM 1230 may be configured to access or use (or be able to access or use) plugins / APIs 1295 (which may include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations where the generative LM 1230 is not ideally suited, the model may have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 1292) to access one or more plugins / APIs 1295 (e.g., third-party plugins) to help process the current input. In such an example, when at least part of the prompt is related to a restaurant or weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs), sending at least part of the prompt related to a particular plugin / API 1295 to the plugin / API 1295, which may process the information and return an answer to the generative LM 1230, which may use the response to generate output 1290. This process can be repeated (e.g., recursively) an arbitrary number of iterations and repeated using any number of plugins / APIs 1295 until an output 1290 that resolves each query / question / request / process / action / etc. from input 1201 is generated. Therefore, the model can rely not only on its own knowledge acquired from training on a large dataset and / or from data retrieved using the RAG component 1292, but also on the expertise or optimized properties of one or more external resources (e.g., plugins / APIs 1295).
[0296] In some embodiments, one or more converter engines (TEs) can be implemented. Converter engines can use micro-tensor scaling to optimize performance and accuracy, such as enabling 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FP4) AI processing. For example, a converter engine can use 16-bit or 8-bit floating-point precision and 8-bit or 4-bit floating-point data formats, combined with software algorithms, to improve AI performance and capabilities. By reducing mathematical operations to 8 or 4 bits, TEs can train larger networks faster without compromising accuracy. For example, TEs can include libraries for accelerating converter models on processing devices such as GPUs to provide better performance in training and inference with lower memory utilization. When TEs are combined with other technologies, such as high-speed interconnects between nodes (e.g., using NVLink switches) and tensor cores (which enable mixed-precision computation, such as micro-scaling precision support), server clusters can be more capable of training massive networks (e.g., billions of parameters) at high speeds. Therefore, it can support tensor core precision for FP64, TF32, BF16, FP16, FP8, INT8, FP6 and FP4, as well as CUDA core precision for FP64, FP32, FP16 and BF16.
[0297] The LLM / VLM / MMLM / VLA and other architectures described herein are intended to be examples only, and other suitable architectures may be implemented within the scope of this disclosure.
[0298] Example computing device Figure 13This is a block diagram of an example computing device 1300 suitable for implementing some embodiments of the present disclosure. The computing device 1300 may include an interconnect system 1302 directly or indirectly coupled to the following devices: a memory 1304, one or more central processing units (CPUs) 1306, one or more graphics processing units (GPUs) 1208, a communication interface 1210, input / output (I / O) ports 1312, input / output components 1314, a power supply 1316, one or more presentation components 1318 (e.g., one or more displays, one or more speakers, etc.), and one or more logic units 1320. In at least one embodiment, one or more computing devices 1300 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1308 may include one or more vGPUs, one or more CPUs 1306 may include one or more vCPUs, and / or one or more logic units 1320 may include one or more virtual logic units. Thus, one or more computing devices 1300 may include discrete components (e.g., a full GPU dedicated to computing device 1300), virtual components (e.g., a portion of the GPU dedicated to computing device 1300), or a combination thereof.
[0299] although Figure 13 The various blocks are shown as connected via interconnect system 1302 using lines, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 1318 (such as a display device) may be considered I / O component 1314 (e.g., if the display is a touchscreen). As another example, CPU 1306 and / or GPU 1308 may include memory (e.g., memory 1304 may represent a storage device other than the memory of GPU 1308, CPU 1306, and / or other components). Therefore, Figure 13 The computing devices described are for illustrative purposes only. No distinction is made between such categories as “workstation,” “server,” “laptop computer,” “desktop computer,” “tablet computer,” “client device,” “mobile device,” “handheld device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are considered within the scope of this category. Figure 13 Within the scope of computing devices.
[0300] Interconnect system 1302 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1302 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Fast Peripheral Component Interconnect (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 1306 may be directly connected to memory 1304. Further, CPU 1306 may be directly connected to GPU 1308. Where there is a direct or point-to-point connection between components, interconnect system 1302 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required to be included in computing device 1300.
[0301] The memory 1304 may include any computer-readable medium from a variety of computer-readable media. The computer-readable medium may be any available medium accessible by the computing device 1300. The computer-readable medium may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, the computer-readable medium may include computer storage media and communication media.
[0302] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1304 may store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Universal Disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 1300. As used herein, computer storage media does not include the signal itself.
[0303] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and includes any information transmission medium. The term "modulated data signal" can refer to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the above should also be included within the scope of computer-readable media.
[0304] CPU 1306 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1300 to perform one or more of the methods and / or processes described herein. Each CPU 1306 may contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. CPU 1306 may contain any type of processor and may contain different types of processors depending on the type of computing device 1300 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1300, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as math coprocessors), computing device 1300 may also include one or more CPUs 1306.
[0305] In addition to or in lieu of one or more CPUs 1306, one or more GPUs 1308 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1300 to perform one or more of the methods and / or processes described herein. One or more GPUs 1308 may be integrated GPUs (e.g., with one or more CPUs 1306) and / or one or more GPUs 1308 may be discrete GPUs. In embodiments, one or more GPUs 1308 may be coprocessors of one or more CPUs 1306. GPUs 1308 may be used by computing device 1300 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPUs 1308 may be used for general-purpose computing on a GPU (GPGPU). GPUs 1308 may include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. GPUs 1308 may produce pixel data of an output image in response to rendering commands (e.g., rendering commands received from CPUs 1306 via a host interface). GPU 1308 may include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 1304. GPU 1308 may include two or more GPUs operating in parallel (e.g., via links). The links may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 1308 may produce pixel data or GPGPU data for different portions of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for an analog image). Each GPU may include its own memory or may share memory with other GPUs.
[0306] In addition to or in lieu of CPU 1306 and / or GPU 1308, logic unit 1320 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1300 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPUs 1306, one or more GPUs 1308, and / or one or more logic units 1320 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 1320 may be a portion of one or more CPUs 1306 and / or GPUs 1308 and / or integrated into one or more CPUs 1306 and / or GPUs 1308, and / or one or more logic units 1320 may be discrete components or otherwise external to CPUs 1306 and / or GPUs 1308. In an embodiment, one or more of the logic units 1320 may be coprocessors of one or more of the CPU 1306 and / or one or more of the GPU 1308.
[0307] Examples of logic unit 1320 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree lateral unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), deep learning accelerator cluster (XNN), neural processing unit (NPU), neural network accelerator (NNA), programmable vision accelerator (PVA) (which may include one or more direct memory access (DMA) systems), and one or more vision or vector processing units (VPUs). This includes one or more pixel processing engines (PPEs) (e.g., a 2D array of processing elements, each of which communicates north, south, east, and west with one or more other processing elements in the array), one or more decoupled accelerators or units (e.g., decoupled lookup table (DLUT) accelerators or units), vision processing units (VPUs), optical flow accelerators (OFAs), field-programmable gate arrays (FPGAs), neuromorphic chips, quantum processing units (QPUs), associative processing units (APUs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnects (PCIs) or fast peripheral component interconnects (PCIe) elements, etc.
[0308] Communication interface 1310 may include one or more receivers, transmitters, and / or transceivers enabling computing device 1300 to communicate with other computing devices via electronic communication networks (including wired and / or wireless communications). Communication interface 1310 may include components and functions for enabling communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or wirelessband), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, logic unit 1320 and / or communication interface 1310 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 1302 to one or more GPUs 1308 (e.g., memory of one or more GPUs 1308).
[0309] I / O port 1312 enables computing device 1300 to be logically coupled to other devices including I / O component 1314, one or more presentation components 1318, and / or other components, some of which may be built into (e.g., integrated into) computing device 1300. Illustrative I / O component 1314 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, etc. I / O component 1314 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, pen recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 1300. Computing device 1300 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. Additionally, the computing device 1300 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables motion detection. In some examples, the computing device 1300 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0310] Power supply 1316 may include a hardwired power supply, a battery power supply, or a combination thereof. Power supply 1316 may provide power to computing device 1300 to enable the components of computing device 1300 to operate.
[0311] Presentation component 1318 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Presentation component 1318 may receive data from other components (e.g., GPU 1308, CPU 1306, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0312] Example network environment A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 13 It is implemented on one or more instances of one or more computing devices 1300—for example, each device may include similar components, features, and / or functions of one or more computing devices 1300. Furthermore, in the case of implementing back-end devices (e.g., servers, NAS, etc.), the back-end devices may be included as part of a data center, such as, but not limited to, those described herein.
[0313] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0314] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.
[0315] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or application may respectively include network-based service software or applications. In embodiments, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer may be, but is not limited to, a free and open-source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").
[0316] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0317] One or more client devices may include the information described in this article. Figure 13 At least some of the components, features, and functions of one or more example computing devices 1300 described. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, chat kiosk, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0318] Other variations also fall within the spirit and scope of this disclosure. Therefore, while the disclosed technology can be modified and alternatively constructed in various ways, certain exemplary embodiments have been shown in the accompanying drawings and described in detail above. However, it should be understood that this disclosure is not intended to limit the scope to the specific forms or multiple specific forms disclosed, but rather to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of this disclosure as defined in the appended claims.
[0319] In describing the disclosed embodiments (particularly in the following claims), the terms “a,” “an,” “the,” and similar pronouns should be interpreted to cover both singular and plural forms unless otherwise stated herein or the context clearly indicates otherwise, and should not be considered as limiting the terms. The terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including but not limited to”) unless otherwise stated. When the word “connection” is unmodified and refers to a physical connection, it should be interpreted as partially or wholly contained within, attached to, or linked together, even with intervening elements. The enumeration of numerical ranges herein is intended only as a convenient method to individually refer to each individual value falling within a range, unless otherwise stated herein, and each individual value is incorporated into the specification as if it had been individually enumerated herein. In at least one embodiment, unless otherwise stated or the context clearly indicates otherwise, the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set containing one or more members. Furthermore, unless otherwise stated or the context indicates otherwise, a “subset” of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and the corresponding set can be equal.
[0320] Conjunctive phrases such as "at least one of A, B, and C" or "at least one of A, B, and C" are generally understood, depending on the context, to indicate that an item, term, etc., can be A, B, or C, or any non-empty subset of the set A, B, and C, unless explicitly stated otherwise or contradicted by the context. For example, in an illustrative example of a set containing three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}. Therefore, such conjunctive phrases are generally not intended to imply that some embodiments require the simultaneous inclusion of at least one A, at least one B, and at least one C. Furthermore, unless explicitly stated otherwise or contradicted by the context, the term "multiple" indicates a plural state (e.g., "multiple items" means multiple items). In at least one embodiment, the number of multiple items is at least two, but the number can be more if explicitly stated or determined by the context. Furthermore, unless otherwise stated or the context clearly indicates otherwise, the phrase “based on” means “at least partially based on” rather than “completely based on”.
[0321] The operations of the processes described herein can be performed in any suitable order unless otherwise stated herein or otherwise explicitly contradicted by the context. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented in the form of code (e.g., executable instructions, one or more computer programs, or one or more application programs) that execute cooperatively on one or more processors, or by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program containing multiple instructions executable by one or more processors.
[0322] In at least one embodiment, the computer-readable storage medium is a non-volatile computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes a non-volatile data storage circuitry system (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, code (e.g., executable code or source code) is stored on a collection of one or more non-volatile computer-readable storage media on which executable instructions (or other memory for storing executable instructions) are stored, causing the computer system to perform the operations described herein when these instructions are executed by one or more processors of a computer system (i.e., as a result of execution). In at least one embodiment, the collection of non-volatile computer-readable storage media includes multiple non-volatile computer-readable storage media, and one or more individual non-volatile storage media do not contain all the code, while the multiple non-volatile computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed in such a manner that different instructions are executed by different processors—for example, instructions are stored on a non-transitory computer-readable storage medium, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes the remaining instructions. In at least one embodiment, different components of the computer system have independent processors, and different processors execute different subsets of instructions.
[0323] In at least one embodiment, the arithmetic logic unit is a set of combinational logic circuit systems that receives one or more inputs to produce a result. In at least one embodiment, the arithmetic logic unit is used by a processor to implement mathematical operations such as addition, subtraction, or multiplication. In at least one embodiment, the arithmetic logic unit is used to implement logical operations, such as logical AND / OR or XOR operations. In at least one embodiment, the arithmetic logic unit is stateless and consists of physical switching elements (e.g., semiconductor transistors arranged to form logic gates). In at least one embodiment, the arithmetic logic unit may operate internally as a stateful logic circuit with an associated clock. In at least one embodiment, the arithmetic logic unit may be constructed as an asynchronous logic circuit whose internal state is not maintained in an associated set of registers. In at least one embodiment, the processor uses the arithmetic logic unit to combine operands stored in one or more registers of the processor and produce an output, which may be stored by the processor in another register or memory location.
[0324] In at least one embodiment, after processing instructions retrieved by the processor, the processor provides one or more inputs or operands to the arithmetic logic unit (ALU), causing the ALU to produce a result at least partially based on instruction codes provided to the ALU. In at least one embodiment, the instruction codes provided by the processor to the ALU are at least partially based on instructions executed by the processor. In at least one embodiment, combinational logic in the ALU processes the inputs and produces an output, which is placed on a bus within the processor. In at least one embodiment, the processor selects a destination register, memory location, output device, or output storage location on the output bus to control the processor via a clock signal so that the result produced by the ALU is sent to the desired location.
[0325] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and the computer system is configured with corresponding hardware and / or software to enable the execution of these operations. Furthermore, the computer system implementing at least one embodiment of this disclosure may be a single device; in another embodiment, it may be a distributed computer system comprising multiple devices operating in different ways, performing the operations described herein, such that a single device does not perform all operations.
[0326] Any and all examples or exemplary language provided herein (e.g., "for example") are used only to better illustrate embodiments of this disclosure and, unless otherwise stated, do not constitute a limitation on the scope of this disclosure. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of this disclosure.
[0327] In the specification and claims, the terms “coupled” and “connected” and their derivatives may be used. It should be understood that these terms are not synonymous with each other. More specifically, in some examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also indicate that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0328] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “operation,” and “determining” refer to the operations and / or processes of a computer or computing system or similar electronic computing component that process and / or transform data represented as physical quantities (e.g., electronic quantities) in the registers and / or memory of the computing system into other data similarly represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0329] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory and transforms that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" are used interchangeably herein, as a system can contain one or more methods, and a method can be considered a system.
[0330] This document may refer to the acquisition, reception, or input of analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of acquiring, receiving, or inputting analog and digital data can be implemented in various ways, such as receiving data as a parameter of a function call or application programming interface (API) call. In at least one embodiment, the process of acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data through a serial or parallel interface. In at least one embodiment, the process of acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data from a providing entity to an receiving entity via a computer network. The provision, output, transmission, sending, or presentation of analog or digital data may also be mentioned. In at least one embodiment, the process of providing, outputting, transmission, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter of a function call, a parameter of an application programming interface (API), or a parameter of an inter-process communication mechanism.
[0331] Although exemplary implementations of the described technology are described herein, other architectures can be used to implement the described functionality, and all such architectures are within the scope of this disclosure. Furthermore, while specific assignments of responsibilities may have been defined above for ease of description, various functions and responsibilities may be assigned and divided in different ways depending on the specific circumstances.
[0332] Furthermore, although the subject matter has been described using language specific to structural features and / or method steps, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or steps described. Rather, the specific features and steps described are disclosed only as exemplary forms for implementing the claims.
[0333] The description in this article is supported by at least the following example clauses: 1. At least one processor, the at least one processor being configured to: predict the door opening status of a vehicle in part based on bounding boxes applied to a representation of the vehicle, the representation of the vehicle being used for a first machine learning ML model, the first machine learning ML model being trained using bounding boxes for different vehicles having one or more closed or open doors; and classify the door opening status into a specific type in part based on at least one portion of the representation used for a second machine learning ML model, the second machine learning ML model being trained using the categories of the different types of open doors of the different vehicles.
[0334] 2. At least one processor as described in Clause 1, configured to: provide the representation in the form of two-dimensional 2D information using a two-dimensional 2D sensor; and to indicate the opening status of the vehicle's doors in the form of three-dimensional 3D information using a two-dimensional ...
Claims
1. At least one processor, said at least one processor being configured to: The door opening status of a vehicle is predicted in part based on bounding boxes applied to a representation of the vehicle, the vehicle representation being used in a first machine learning (ML) model trained using bounding boxes for different vehicles having one or more closed or open doors; and The door opening status is classified into specific types, in part based on at least a portion of the representation used for the second machine learning ML model, which is trained using the categories of the different types of open doors for the different vehicles.
2. The at least one processor as claimed in claim 1, wherein it is configured to: The representation is provided in the form of two-dimensional 2D information using a two-dimensional 2D sensor; and Using a 2D-to-3D conversion subsystem, partly based on the output of the second machine learning (ML) model, the opening status of the vehicle's doors is shown in 3D information form.
3. The at least one processor as claimed in claim 2, wherein, The at least one processor is included in the 2D-to-3D conversion subsystem, which is used to generate or support a top-down or bird's-eye view of the vehicle's door opening position (BEV representation).
4. The at least one processor as claimed in claim 2, wherein, The two-dimensional (2D) sensor is an onboard camera of an autonomous or semi-autonomous vehicle.
5. The at least one processor as claimed in claim 2, wherein, The at least one processor is included in or supports the driving subsystem of an autonomous or semi-autonomous vehicle, and wherein the at least one processor is further configured to perform or suggest a response to the door opening status.
6. The at least one processor as described in claim 5 is further configured to: As part of the driving subsystem, actions such as detouring, slowing down, stopping, or maintaining distance from the vehicle are performed as part of the response to the door opening status.
7. The at least one processor as claimed in claim 1, wherein, The specific type is one of the following: left passenger door open, right passenger door open, left driver's door open, right driver's door open, left door open, right door open, rear door open, sunroof or panoramic sunroof open, or engine hood open.
8. The at least one processor as claimed in claim 1, wherein, The at least one portion of the representation is a portion of the representation having an extended dimension relative to the bounding box in the first machine learning ML model, or a predetermined portion of the representation defined from the edge of at least one dimension of the bounding box.
9. The at least one processor as claimed in claim 1, wherein, The at least one portion of the representation includes one of the following: Referencing the bounding box and taking the extended dimension relative to the bounding box; or A predetermined width from the edge of the bounding box.
10. The at least one processor as claimed in claim 1, further configured to: The first machine learning (ML) model is trained using first data generated from bounding boxes of one or more of the closed or open doors surrounding the different vehicles, wherein, The first machine learning (ML) model is used to generate an output representing the prediction of the opening status of the vehicle's doors; and A second machine learning (ML) model is trained using second data generated from different bounding boxes specifically surrounding different types of open doors of the different vehicles, wherein the second ML model is used to generate output representing the classification of the door opening status into a specific type for the vehicle from the different types.
11. The at least one processor as claimed in claim 1, further configured to: The representation is provided in the form of two-dimensional 2D information using a two-dimensional 2D sensor; Using a depth sensor and the representation in the form of 2D information, at least a portion of the representation in the form of 3D information is generated; and The 2D information is used as input to a second machine learning (ML) model trained using the 2D information and the 3D information to perform classification of the door opening status in 3D space.
12. One or more processors for training a machine learning (ML) model to determine a specific type of door opening condition of a vehicle using different categories associated with different types of door opening conditions of different vehicles.
13. The processors of claim 12 are further configured to train additional machine learning (ML) models for predicting the door opening status of the vehicle, in part based on different bounding boxes for different vehicles having one or more closed or open doors.
14. The one or more processors as claimed in claim 12, wherein, The specific type is one of the following: left passenger door open, right passenger door open, left driver's door open, right driver's door open, left door open, right door open, rear door open, sunroof or panoramic sunroof open, or engine hood open.
15. One or more processors for determining, in part, a specific type of door opening of a vehicle based on a machine learning ML model trained using different categories associated with different types of door openings of different vehicles.
16. The one or more processors as claimed in claim 15, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for the autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models (MMLMs); A system for performing operations using one or more visual-language-action (VLA) models; A system for using or deploying one or more inference microservices; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
17. A method, the method comprising: The door opening status of the vehicle is predicted in part based on bounding boxes applied to a representation of the vehicle, the representation of the vehicle being used for a first machine learning ML model, which is trained using bounding boxes for different vehicles having one or more doors that are closed or open. as well as The door opening status is classified into specific types, partly based on at least a portion of the representation used for the second machine learning ML model, which is trained using the categories of the different types of open doors for the different vehicles.
18. The method of claim 17, further comprising: The first machine learning ML model is trained using first data generated from bounding boxes of one or more of the closed or open doors of the different vehicles, wherein the first machine learning ML model is used to generate an output representing the prediction of the door opening status of the vehicles; and A second machine learning (ML) model is trained using second data generated from different bounding boxes specifically surrounding different types of open doors of the different vehicles, wherein the second ML model is used to generate output representing the classification of the door opening status into a specific type for the vehicle from the different types.
19. The method of claim 17, further comprising one or more of the following: A top-down or bird's-eye view of the door opening position of the vehicle, as shown in the BEV representation, is generated or supported. The response to the door opening is executed by the driving subsystem of an autonomous or semi-autonomous vehicle having at least one processor; The suggested response to the door being opened; Use one of the following as a specific type of door opening condition: left passenger door open, right passenger door open, left driver's door open, right driver's door open, left door open, right door open, rear door open, sunroof or panoramic sunroof open, or hood open; or The action may be to take an alternative route, drive slowly, stop, or keep a distance from the vehicle, as a response.
20. The method of claim 17, wherein, The representation is partially included in the two-dimensional 2D information from the two-dimensional 2D sensor, and the method further includes: At least a portion of the representation in the form of three-dimensional 3D information is generated using a depth sensor and the two-dimensional 2D information, wherein the two-dimensional 2D information is used as input to a second machine learning (ML) model trained using the two-dimensional 2D information and the three-dimensional 3D information to perform classification of the door opening status in three-dimensional 3D space; or Using a 2D to 3D conversion subsystem, partly based on the output of the second machine learning (ML) model, the opening status of the vehicle's doors is shown in the form of 3D information.