Optical flow based frame interpolation to synchronize between sensors

By using optical flow-based frame interpolation technology, the problem of frame data synchronization between sensors was solved, achieving efficient multi-sensor data synchronization and improving the accuracy of perception and decision-making processes as well as the utilization rate of hardware resources.

CN121397159APending Publication Date: 2026-01-23NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410986191.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively synchronize frame data between multiple sensors, leading to differences in frame rates, inconsistent exposure times, and readout time deviations, which affect the accuracy and reliability of the perception and decision-making process.

Method used

By employing optical flow-based frame interpolation technology, interpolated frames are generated to achieve synchronization by detecting motion and temporal differences between sensor data. Optical flow accelerators and programmable vision accelerators are used to improve hardware utilization and reduce blur and inconsistency.

Benefits of technology

It achieves efficient synchronization of multi-sensor data, improves the accuracy and reliability of perception and decision-making processes, reduces ambiguity and inconsistencies between frames, and enhances hardware resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397159A_ABST
    Figure CN121397159A_ABST
Patent Text Reader

Abstract

The invention discloses frame interpolation based on optical flow to synchronize between sensors. In various examples, systems and methods are disclosed that perform motion detection, such as optical flow determination, across image frames to synchronize asynchronous frames with respect to a target time of the asynchronous frames. For example, an image frame from a sensor may be processed by an optical flow accelerator to detect a displacement across the image frame, and the displacement may be used to interpolate a modified frame at a target time. This may be used to perform data collection and combination operations, such as splicing and / or reconstruction. Synchronization may be performed according to sensor data from sensors such as cameras, LIDAR sensors, and / or RADAR sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Camera frame synchronization in sensor systems, such as those used in autonomous or semi-autonomous vehicles or machines, is useful for ensuring the accuracy and reliability of perception, prediction, and decision-making processes. It enables the system to have a coherent representation of the vehicle's environment, which is beneficial for safe and efficient navigation. However, factors such as variations in exposure and readout times, as well as frame rate differences between individual sensors, can make effective synchronization challenging. Summary of the Invention

[0002] Embodiments of this disclosure relate to optical flow-based frame interpolation for synchronization between cameras. For example, systems and methods are disclosed that facilitate synchronization between data streams from cameras based on motion-based image processing (e.g., optical flow-based frame interpolation). This can include: generating interpolated frames from data from a first camera to synchronize with timestamps from a second camera, and generating interpolated frames to increase the frame rate from a given camera. For example, optical flow-based techniques can adjust and / or smooth motion between frames, which can reduce blur and / or inconsistencies. In some embodiments, frame interpolation can be performed using an optical flow accelerator (OFA), which can allow for greater hardware footprint and / or more efficient hardware utilization. This can be used for applications including, but not limited to, data collection, surround-view camera stitching, 3D reconstruction, latency (e.g., transmission delay) compensation, increased frame rate, peripheral blind spot enhancement, and perceptual fusion.

[0003] At least one aspect involves one or more processors comprising one or more circuits. One or more circuits can detect an asynchronous condition in sensor data from a sensor. In response to detecting the asynchronous condition, one or more circuits can determine motion associated with a first frame (e.g., an input frame) and a second frame of sensor data. One or more circuits can generate a third frame based at least on the motion and the first frame.

[0004] In some implementations, one or more circuits detect motion by detecting optical flow between a first frame and a second frame. The sensor may be a first sensor, and one or more circuits may detect asynchronous situations in response to detecting the time difference between a timestamp of the first frame and a timestamp of a fourth frame from at least one second sensor.

[0005] In some implementations, one or more circuits detect an asynchronous condition in response to detecting that a missing frame is not received from the sensor at the expected timestamp of the missing frame. One or more circuits may generate a third frame with a timestamp equal to that of the fourth frame from the second sensor.

[0006] In some implementations, one or more circuits generate a third frame with a timestamp following that of the first and second frames. One or more circuits may detect motion based at least on the displacement of pixels representing an object from a first position in the first frame to a second position in the second frame. The sensor may include a camera, a light detection and ranging (LIDAR) system, or a radio frequency detection and ranging (RADAR) system.

[0007] At least one aspect relates to a system. The system may include one or more processing units and one or more memory units storing instructions, which, when executed by the one or more processing units, cause the one or more processing units to perform an operation. The operation may include: detecting an asynchronous condition of sensor data from a sensor; in response to detecting the asynchronous condition, determining motion associated with a first frame and a second frame of sensor data; and generating a third frame based at least on the motion and the first frame.

[0008] In some implementations, one or more processing units are used to detect (e.g., for detecting; configured to detect) motion by detecting optical flow between a first frame and a second frame. The sensor may be a first sensor, and one or more processing units may detect asynchronous situations in response to detecting the time difference between the timestamp of the first frame and the timestamp of a fourth frame from at least one second sensor.

[0009] In some implementations, one or more processing units are used to detect an asynchronous condition in response to detecting that a missing frame has not been received from the sensor at the expected timestamp of the missing frame. The sensor may be a first sensor, and one or more processing units may generate a third frame to have a timestamp equal to that of a fourth frame from a second sensor.

[0010] In some implementations, one or more processing units generate a third frame with a timestamp following that of the first and second frames. The processing units may detect motion based at least on the displacement of pixels representing an object from a first position in the first frame to a second position in the second frame. The sensor may include a camera, a light detection and ranging (LIDAR) system, or a radio frequency detection and ranging (RADAR) system.

[0011] At least one aspect relates to a method. The method may include: detecting an asynchronous condition in sensor data from a sensor using one or more processors. The method may include: in response to detecting the asynchronous condition, determining, using one or more processors, motion associated with a first frame and a second frame of sensor data. The method may include: generating a third frame using one or more processors, at least based on the motion and the first frame.

[0012] In some implementations, the sensor is a first sensor, and detecting asynchronous conditions includes detecting at least a threshold time difference between the first frame and a fourth frame from a second sensor. The method may include generating a third frame, which involves interpolating the timestamp of the first frame to the timestamp of the fourth frame.

[0013] In some implementations, generating the third frame includes interpolating the first frame to the timestamp of the missing frame. The missing frame may be a frame that was not received from the sensor within a threshold of the expected time for receiving the missing frame.

[0014] The processors, systems, and / or methods described herein can be implemented by or included in at least one of the following: control systems for autonomous or semi-autonomous machines; perception systems for autonomous or semi-autonomous machines; systems comprising one or more virtual machines (VMs); systems implemented using robots; systems for performing deep learning operations; systems for performing simulation operations; systems for performing collaborative content creation of 3D assets; and systems for generating synthetic data.

[0015] Systems for performing digital twin operations; systems implemented using edge devices; systems including one or more visual language models (VLMs); systems including one or more large language models (LLMs); systems including one or more multimodal language models; systems for performing conversational AI operations; systems for performing optical transmission simulations; systems implemented at least partially in a data center; or systems implemented at least partially using cloud computing resources. Attached Figure Description

[0016] The following describes in detail, with reference to the accompanying drawings, the system and method for optical flow-based frame interpolation for synchronization between cameras, wherein:

[0017] Figure 1 This is a block diagram of an example synchronization system according to some embodiments of this disclosure;

[0018] Figure 2 This is a schematic diagram illustrating an example of frame interpolation to perform synchronization according to some embodiments of this disclosure;

[0019] Figure 3 This is a schematic diagram illustrating an example of inserting interpolated frames to replace missing frames according to some embodiments of this disclosure;

[0021] Figure 4 This is a flowchart of a frame interpolation method according to some embodiments of this disclosure;

[0022] Figure 5A These are illustrations of example autonomous vehicles based on some embodiments of this disclosure;

[0023] Figure 5B It is based on some implementation schemes of this disclosure. Figure 5A Examples of camera positions and fields of view for autonomous vehicles;

[0024] Figure 5C It is based on some implementation schemes of this disclosure. Figure 5A A block diagram of an example system architecture for an example autonomous vehicle;

[0025] Figure 5D It is a cloud-based server and according to some embodiments of this disclosure Figure 5A A system diagram illustrating communication between autonomous vehicles;

[0026] Figure 6 This is a block diagram of an example computing device applicable to implementing some embodiments of this disclosure; and

[0027] Figure 7 This is a block diagram of an example data center applicable to some implementation schemes of this disclosure. Detailed Implementation

[0028] Systems and methods related to optical flow-based frame interpolation for synchronization between cameras are disclosed. Although this disclosure may relate to an example autonomous vehicle 500 (which may be alternatively referred to herein as "vehicle 500," "self-vehicle 500," "machine 500," or "self-machine 500"), its examples pertain to... Figures 5A-5D The description herein is provided for purposes of reference only and is not intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, submarines, drones, and / or other vehicle types. Furthermore, while this disclosure may describe the synchronization between data streams from sensors of autonomous or semi-autonomous machines, this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and supervision, autonomous or semi-autonomous machine applications, and / or any other applications that may utilize sensor data synchronization and / or fusion. Technical Field

[0029] Synchronization can include at least two techniques: hardware synchronization and software synchronization. Hardware synchronization can include using hardware capabilities to ensure simultaneous data sampling on multiple sensors. Software synchronization can include time synchronization. Time synchronization (e.g., time alignment) can be achieved by providing a uniform reference time to all sensors, thereby ensuring consistency in data collection timing. In some implementations, synchronization can include spatial synchronization, such as including processes to ensure that data collected from different sensors accurately corresponds to the same location or object in the environment.

[0030] Hardware synchronization can rely on one or more of a unified time reference, a triggering mechanism, and calibration and / or testing. A unified time reference is used to ensure precise coordination between sensors, including but not limited to imaging sensors such as cameras, LiDAR, and radar. For example, a high-precision clock source can be used as the unified time reference. The triggering mechanism can be external (e.g., outside one or more sensors) to simultaneously initiate the acquisition process on all cameras. This can be achieved via a hardware trigger line, where a central control unit sends an FSYNC signal to all cameras, providing simultaneous image acquisition. Prior to system deployment, the synchronization mechanism can be rigorously calibrated and tested to minimize time discrepancies and achieve microsecond-level accuracy.

[0031] Software synchronization can rely on one or more algorithms for timestamp management and efficient use of synchronization dates. These algorithms can leverage the temporal and spatial alignment of data from multiple sensors to perform tasks such as sensor fusion, object detection, and decision-making.

[0032] Various such techniques may be limited in terms of the accuracy and / or reliability of achievable synchronization. Timestamp management and / or synchronization may involve synchronizing clock sources, as the internal clocks of each sensor may drift in inconsistent ways. Hardware synchronization can be used to ensure that trigger times are identical, for example, ensuring that each sensor is triggered on the same rising edge (given that a synchronized clock source may only ensure consistent intervals). Software synchronization may involve maintaining time gaps to be compatible with readout time differences between sensors that cannot be synchronized.

[0033] However, since the exposure time varies for each sensor, discrepancies may still exist between them. Inconsistencies in exposure time can be influenced by many factors, such as different batches, brands, and resolutions, all of which affect exposure time and make perfect timestamp synchronization difficult.

[0034] For example, sensors, including those in a camera, can have at least two triggering modes: trigger exposure time and trigger readout. Even if the sensors are triggered simultaneously (e.g., via a hardware FSYNC signal), differences in exposure time can result in different final exposure times, leading to asynchronous situations where one sensor has a longer exposure time than the others, thus capturing more information.

[0035] Some systems can use readout time control to facilitate synchronization, such as by maintaining the synchronization time of data readout from the sensor. However, readout time cannot always take into account different imaging times (e.g., those obtained from different exposure lengths) and may be useless if the sensor performs other operations (e.g., color conversion) after exposure and before readout.

[0036] Failure to properly synchronize frames in camera data can lead to significant problems in sensor data processing, including, for example, stitching of surround views. For instance, in an environment with multiple sensors, dynamic objects (such as other vehicles and pedestrians) can move rapidly. If the cameras are not synchronized, the captured images may show discrepancies in the positions of these dynamic objects due to timing differences. This can include object misalignment, variations in brightness and contrast (such as exposure differences), and / or visible seams. Furthermore, in multi-sensor fusion sensing systems, various sensors (such as LiDAR, cameras, and millimeter-wave radar) are combined to create a comprehensive understanding of the environment. Their frame rates may differ; for example, a LiDAR might have a frame rate of 10 frames per second (fps), while a camera might have a frame rate of 25 fps / 28 fps / 30 fps / 60 fps. When the frame rates of multiple sensors are not evenly distributed, their data may only be synchronized within a certain range.

[0037] The systems and methods according to this disclosure can address such considerations, allowing for more efficient synchronization between sensors, even if the sensors may have different frame rates, exposure times, exposure lengths, and / or post-processing operations. For example, the system can perform optical flow techniques on frames from asynchronous sensors to detect motion between frames, and can generate new frames at selected timestamps based on the detected motion to provide synchronization between sensors, such as between frames from different sensors. This can allow for higher quality output from processes that rely on data from multiple sensors (e.g., frames of images and / or similar images from multiple sensors). For example, image data from one or more cameras, LiDAR sensors, and / or radar sensors can be combined, for example, by stitching them together to form a surround view (e.g., in a vehicle's surround view system (SVS)), for example, but not limited to, with reduced blur, inconsistencies, seams, and / or exposure variations. Such operations can have useful applications, including but not limited to data collection, autonomous driving, and smart cockpit integration. In some implementations, at least some of the operations described herein (including optical flow operations) can be performed using vehicle components with relatively low utilization hardware, such as programmable visual accelerator (PVA) and / or optical flow accelerator (OFA) hardware, thereby allowing for higher overall utilization without increasing the burden on CPU and / or GPU resources. Furthermore, PVA and OVA hardware can have greater power and / or energy efficiency relative to CPU and / or GPU components, which is useful in applications including, but not limited to, vehicle and portable electronic device applications. Systems and methods according to this disclosure can be used in end-to-end (E2E) implementations, such as synchronizing frames before downstream operations (e.g., functions for 2D stitching or 3D reconstruction) use them. Systems and methods according to this disclosure can be implemented for applications such as virtual reality (VR) and augmented reality (AR), motion analysis, video effects production, and / or applications relying on precise time alignment from multiple video sources.

[0038] refer to Figure 1 , Figure 1This is an example system 100 for performing optical flow-based frame interpolation according to some embodiments of this disclosure. It should be understood that such and other arrangements described herein are illustrative only. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be implemented by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be used with… Figures 5A-5D Example of autonomous vehicles 500 Figure 6 Example computing device 600 and / or Figure 7 The example data center 700 uses similar components, features, and / or functions to perform this.

[0039] System 100 may include or be coupled to multiple sensors 104, such as a first sensor 104a and a second sensor 104b. Sensors 104 may include one or more imaging sensors. Sensors 104 may include, for example, but not limited to, any one or more cameras, video cameras, LiDAR sensors, RADAR sensors, and / or ultrasonic sensors. System 100 may receive and / or process data from sensors 104 as a data stream. System 100 may output a synchronous frame stream for operations including, but not limited to, 2D stitching and 3D reconstruction.

[0040] Sensor 104 can detect sensor data (e.g., image data) representing the environment surrounding sensor 104 and / or one or more objects in the environment. For example, sensor 104 can output sensor data in two-dimensional and / or three-dimensional data structures representing the characteristics of the environment and / or one or more objects, such as frames of sensor data (e.g., image frames).

[0041] Sensor 104 can output sensor data at a rate (e.g., frame rate). Sensor 104 can detect sensor data based on an exposure time that is at least one of a start time or an end time of data acquisition. Sensor 104 can have a readout time at which frames of sensor data are sampled and / or retrieved, and sensor 104 can perform processing on the sensor data before readout. The frame can have a timestamp indicating the time of sensor data detection.

[0042] As explained above, various such timing considerations can lead to asynchronous situations between the sensor data output by the individual sensors 104. For example, the first sensor 104a may output frames (of sensor data) at a first frame rate, while the second sensor 104b may output frames at a second frame rate different from the first frame rate. Sensors 104 may experience situations where no frames are output in a given instance, such as frames being dropped.

[0043] System 100 may include at least one frame interpolator 112. Frame interpolator 112 may process frames from sensor 104 to determine whether interpolation is available to resolve asynchronous situations. Frame interpolator 112 may generate interpolated frames (e.g., modifications to frames from sensor 104) in response to detected situations (e.g., asynchronous situations between sensor 104 and / or data from sensor 104). Frame interpolator 112 may be used for phased data synchronization; for example, data from the camera may be synchronized before synchronization with other sensors (e.g., LiDAR and / or RADAR sensors). Frame interpolator 112 may be used for forward interpolation of frames (e.g., modifying a given frame at a first timestamp to a target timestamp after the first timestamp) or backward interpolation of frames (e.g., modifying a given frame to a target timestamp before the first timestamp). Frame interpolator 112 can generate frames to increase the frame rate from one or more sensors 104, for example, by modifying the first frame forward to the timestamp between the first and second frames or modifying the second frame backward to the timestamp between the first and second frames based on motion detected between the first and second frames.

[0044] Frame interpolator 112 may include frame evaluator 116. Frame evaluator 116 may evaluate frames from sensor 104 to detect at least one condition associated with asynchronous situations, including, for example, detecting a time difference (e.g., offset) between frames and / or detecting missing frames. For example, frame evaluator 116 may identify a first timestamp of a first frame from sensor 104a and a second timestamp of a second frame from second sensor 104b. Frame evaluator 116 may compare the first timestamp with the second timestamp to detect a time difference between the first and second frames. Frame evaluator 116 may detect a condition based at least on the time difference, for example, in response to a time difference greater than a threshold difference. The threshold difference may be in the millisecond range or smaller. The threshold difference may be less than the frame rate, for example, a fraction of the frame rate. In some embodiments, the threshold difference is zero, for example, system 100 operates to eliminate offsets between frames. Regarding missing frames, frame evaluator 116 can determine that a frame (e.g., from second sensor 104b) was not received within a expected time period (e.g., according to the frame rate of sensor 104b; after receiving the first frame from first sensor 104a). Frame evaluator 116 can determine the time difference and can generate an identifier for sensor 104 associated with synchronization, for example, to indicate that frame generation should be performed on a second frame from second sensor 104; the second frame may be the target frame.

[0045] In response to the detection of a situation associated with asynchronous conditions, frame evaluator 116 may induce operations to mitigate asynchronous conditions and / or achieve synchronization (e.g., frames with the same timestamp and / or a target timestamp), such as inducing the generation of synchronized frames and / or frames at the target timestamp of missing frames. This can allow system 100 (e.g., output generator 140) to generate output data (e.g., fused data) based on information from multiple sensors 104 and with fewer errors relative to the real-world environment surrounding the sensors 104.

[0046] like Figure 1As depicted, system 100 may include at least one motion detector 120. Motion detector 120 may detect indications of motion between at least two frames from a given sensor 104 (e.g., from sensor 104b). For example, motion detector 120 may detect indications of motion between consecutive frames from sensor 104b. Indications of motion may include movement and / or displacement (e.g., distance, such as pixel distance), such as displacement from the position of one or more first pixels in a first frame from sensor 104b and displacement from the position of one or more second pixels in a second frame from sensor 104b, where the first pixels represent a feature in the first frame (e.g., a feature of an object or environment) and the second pixels represent that feature. For example, motion detector 120 may track the position of features in a frame to detect indications of motion. Motion detector 120 may detect indications of motion based on the displacement of pixels representing multiple features from the first frame to the second frame. In some implementations, motion detector 120 can detect motion without processing individual pixels or referencing individual pixels, for example, based at least on the displacement of object positions and / or blocks in a frame, which may or may not be directly mapped to pixels. Motion detector 120 can generate motion indications (e.g., flow vectors) to represent a synthesis of displacements of multiple features from a first frame to a second frame. Motion detector 120 can apply any one or more optical flow techniques, object or block matching or tracking techniques, machine learning model-based techniques, motion estimation and motion compensation (MEMC) techniques, or various combinations thereof to detect motion.

[0047] The motion detector 120 can output motion indications as vectors, for example, representing the motion using X and Y coordinates of displacement. The motion detector 120 can also output motion indications as a flow map, which includes one or more vectors representing the detected motion.

[0048] In some embodiments, the motion detector 120 includes at least one preprocessor 124. The preprocessor 124 can perform any of a variety of operations on a first and a second frame to which motion is to be detected, to facilitate the generation of motion indications, such as facilitating the generation of optical flow information and / or flow graphs. For example, the preprocessor 124 can generate an image data pyramid representing the characteristics of the frame at multiple levels (e.g., resolution levels). The preprocessor 124 and / or one or more of its components or operations can be implemented by at least one of a PVA, CPU, video image synthesizer (VIC), or CUDA (e.g., as described with reference to FIG. 5). The preprocessor 124 can perform operations such as, but not limited to, dedistortion and normalization.

[0049] Motion detector 120 may include at least one optical flow detector 128. Optical flow detector 128 may process data from the first and second frames, such as processing an image data pyramid, to determine indications of motion. For example, optical flow detector 128 may apply any one or more optical flow algorithms to detect motion of features represented by the first and second frames, such as generating a flow graph. For example, optical flow detector 128 may generate a vector representing motion from the first frame to the second frame, such as representing pixel displacement from the first frame to the second frame. The flow graph may assign displacements (e.g., represented by X and Y coordinates) to one or more pixels of a flow graph frame, which may have the same dimensions as the frame from the second sensor 104b. As further described herein, system 100 may determine a modified frame (e.g., an intermediate frame between the first and second frames) based at least on the flow graph and a time offset, such as processing the flow graph according to a time ratio indicated by the time offset.

[0050] The motion detector may include at least one post-processor 132. The post-processor 132 may be used to perform post-processing on the output of the optical flow detector 128. For example, the post-processor 132 may perform smoothing or noise reduction on the output of the optical flow detector 128. In some embodiments, the post-processor 132 includes a median filter, which may be used to modify one or more values ​​of the output of the optical flow detector 128 based on the median of associated values ​​from neighboring frames. In some embodiments, the optical flow detector 128 performs filtering operations. The post-processor 132 may perform operations such as mixing or inspecting the output from the optical flow detector 128.

[0051] In some implementations, one or more components of motion detector 120 may be implemented by hardware components (e.g., accelerators) that may be coupled to the CPU and / or GPU resources of system 100. For example, preprocessor 124 may be implemented as a programmable vision accelerator (PVA), while optical flow detector 128 may be implemented as an optical flow accelerator (OFA). Using such targeted hardware allows for more efficient use of hardware resources and / or lower power requirements. For example, optical flow operations performed by motion detector 120 may be implemented on the OFA (e.g., high occupancy and / or full OFA occupancy), which can save CPU and GPU resources; the utilization of PVA and OFA on the vehicle side may be low, so these components can achieve greater utilization for performing operations of system 100. Therefore, system 100 can achieve greater overall hardware utilization compared to systems that may rely on CPU resources for software-side synchronization.

[0052] Further reference Figure 1System 100 may include at least one frame generator 136. Frame generator 136 may generate modified frames (e.g., a third frame), such as to modify a target frame from a second sensor 104b, which is associated with asynchronous conditions detected by frame evaluator 116, such as frames that are out of sync with and / or missing frames from a first frame from a first sensor 104a.

[0053] Frame generator 136 can generate modified frames based at least on a time difference (e.g., determined by frame evaluator 116, and / or the timestamp for which the modified frame is to be generated) and a motion indication (e.g., determined by motion detector 120). For example, frame generator 136 can generate a modified frame as a new frame for the second sensor 104b, synchronized with the first frame from the first sensor 104a, by interpolating the target frame to a target time using the motion indication and the time difference. The target time can be the time of the first timestamp of the first frame from the first sensor 104a and / or the time when the missing frame should have been received.

[0054] For example, frame generator 136 can generate a modified frame by modifying the position of features in a target frame based at least on motion indications and time differences, such as interpolating the modified position of features based at least on motion indications and time differences. For example, if the motion indication includes a flow graph, frame generator 136 can determine the amount by which to move a feature from a position (e.g., a pixel) in the target frame to a position in the modified frame based on the flow graph and time difference. By generating modified frames based on motion indications, system 100 can achieve high video quality because it can enable a higher synchronization rate of frames from sensor 104. Furthermore, system 100 can perform interpolation as software synchronization to address situations where hardware synchronization may not achieve perfectly accurate synchronization.

[0055] In some implementations, system 100 uses modified frames generated by frame generator 136 to allow the output rate of one or more sensors 104 to decrease. For example, system 100 may determine that one or more sensors 104 are operating in a reduced frame rate mode, system 100 may select one or more timestamps relative to the timestamps of the outputs of one or more sensors 104 to achieve a target frame rate, and interpolate frames to the selected timestamps to increase the frame rate to the target frame rate (in various such implementations, system 100 may not perform time offset detection on frames from multiple sensors, e.g., where system 100 processes frames from a given sensor to determine motion indications from the processed frames and generates interpolated frames based on the motion indications and the selected timestamps, e.g., without relying on information from other sensors 104).

[0056] In some implementations, system 100 generates modified frames to account for output latency, such as processing latency and / or transmission latency. For example, in response to detecting at least a threshold latency (e.g., latency) associated with receiving or outputting data at one or more points (e.g., along a data processing pipeline associated with system 100), system 100 may cause modified frames to be generated at a target time for the expected data. This may include, for example, generating output image frames for display.

[0057] Further reference Figure 1 System 100 may include at least one output generator 140 or be coupled to at least one output generator 140. Output generator 140 may generate outputs based at least on synchronized frames from sensor 104. For example, output generator 140 may generate one or more outputs based at least on a first frame from first sensor 104a and a modified frame generated by frame generator 136 (e.g., a modification of a target frame from second sensor 104b). Output generator 140 may generate outputs for operations including, but not limited to, displaying data, computer vision tasks, perception tasks, or vehicle control tasks. In some embodiments, output generator 140 generates an output combining the first frame and the modified frame. For example, output generator 140 may generate a surround view, such as generating a 2D stitch of the first frame and the modified frame. Output generator 140 may generate a 3D reconstruction of the environment represented by data from sensor 104.

[0058] Figure 2 Examples of a process 200 for synchronizing offset frames according to some embodiments of this disclosure are depicted. System 100 may perform one or more operations of process 200.

[0059] like Figure 2 As depicted, along timeline 204, the first sensor 104a can output frames 1-1, 1-2, and 1-3, while the second sensor 104b can output frames 2-1, 2-2, and 2-3. For example, sensors 104a and 104b can output frames at a frame rate of 30 fps, such as... Figure 2 The timestamps depicted are 0ms for frame 1-1, 33ms for frame 1-2, and 66ms for frame 1-3. Frames 2-1, 2-2, and 2-3 are offset from frames 1-1, 1-2, and 1-3 (e.g., as shown in the image). Figure 2 The depicted offset is approximately 10-15 ms. In various implementations, process 200 (and references) Figure 3 The described process 300) can be based on greater than or less than as per reference. Figure 2 The frame rate (and timestamp) described is used to perform the operation, including different frame rates between two or more sensors 104.

[0060] System 100 can receive frames from sensors 104a and 104b and detect one or more offsets between frames 1-1, 2-1, 1-2, 2-2, and / or frames 1-3, 2-3. For example, system 100 can determine one or more time differences between frames 1-1, 2-1, 1-2, 2-2, and / or frames 1-3, 2-3, which can indicate an asynchronous condition. In response to the detection of an offset (e.g., frame offset and / or asynchrony from the second sensor 104b), system 100 can process consecutive frames from the second sensor 104b to detect indications of motion with respect to consecutive frames, such as indications of motion between frames 2-1, 2-2, 2-2, 2-3, between frame 2-1 and the previous frame, and / or between frame 2-3 and the frame following frame 2-3. For example, system 100 can provide any of such consecutive frame pairs as input to optical flow detector 128 so that optical flow detector 128 generates an indication of motion, such as generating a flow graph representing displacement between consecutive frames.

[0061] System 100 can generate one or more modified frames 2-1' (e.g., such that frame 2-1 is the first frame, frame 2-2 is the second frame, modified frame 2-1' is the third frame, and frame 1-1 used for offset detection is the fourth frame), 2-2', 2-3' based at least on indications of offset and motion. For example, frame 2-1 can be interpolated to frame 2-1' at timestamp 0 ms, frame 2-2 to frame 2-2' at timestamp 33 ms, and / or frame 2-3 to frame 2-3' at timestamp 66 ms. Therefore, system 100 can align frames from the second sensor 104b with frames from the first sensor 104a to achieve precise synchronization.

[0062] Figure 3 Examples of a process 300 for synchronizing missing frames according to some embodiments of this disclosure are depicted. System 100 may perform one or more operations of process 300. For example, in a scenario with a large number of sensors 104 (e.g., more than three sensors, more than five sensors, more than ten sensors), the probability that at least one sensor 104 will drop a frame at a given timestamp may be non-negligible. By implementing process 300, system 100 can prevent situations where, in addition to missing frames, frames from all sensors 104 might need to be dropped.

[0063] For example, such as Figure 3As shown, the first sensor 104a outputs frames 1-1, 1-2, and 1-3 at timestamps 0ms, 33ms, and 66ms, while the second sensor 104b outputs frames 2-1 and 2-3 synchronized with the corresponding frames 1-1 and 1-3 at timestamps 0ms and 66ms. However, the expected frame from the second sensor 104b at timestamp 33ms is missing.

[0064] In response to the detection of a missing frame, system 100 can use frame 2-1 (or frame 2-3, for example, to use frame 2-3 as the original frame, for example, relative to) Figure 3 The example shown, in another direction, generates a modified frame 2-2' for the missing frame 2-2 by detecting motion indications (e.g., performing optical flow), for example by processing frame 2-1 and the frame preceding frame 2-1. System 100 can generate a modified frame 2-2' for the missing frame 2-2 at timestamp 33ms, based at least on the motion indications and timestamp 33ms (and / or an offset of 33ms from, for example, timestamp 0ms of frame 2-1 to timestamp 33ms). For example, in a first mode (e.g., prediction mode), system 100 can, for example, use frame 2-1 as the original frame to be modified to determine the modified frame 2-2', and generate (e.g., using frame generator 136) the modified frame 2-2' based on the flow graph of frame 2-1 and the previous frame (e.g., determined by optical flow generator 128). In the second mode (e.g., non-predictive mode), system 100 may, for example, use frame 2-1 as the original frame to be modified to determine the modified frame 2-2', and generate the modified frame 2-2' based on the flow graph of frame 2-3 and frame 2-1.

[0065] The following describes an example of a process for increasing the frame rate of frames output by sensor 104a according to some embodiments of the present disclosure. System 100 can generate frame 1-1' (e.g., at a selected timestamp between timestamps of frames 1-1 and 1-2) based at least on frames 1-1 and 1-2, and can generate frame 1-2' (e.g., at a selected timestamp between timestamps of frames 1-2 and 1-3) based at least on frames 1-2 and 1-3. For example, system 100 can provide frames 1-1 and 1-2 as input to optical flow detector 128 to generate an indication of motion between frames 1-1 and 1-2, and can determine frame 1-1' based on the selected timestamp and the indication of motion (and can determine frame 1-2' based on the selected timestamp and the indication of motion).

[0066] Now for reference Figure 4Each block of the method 400 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a standalone application, service, or managed service (independently or in combination with another managed service), or a plug-in to another product, to name a few. Furthermore, method 400 is illustrated by way of example regarding... Figure 1 System 100 is described herein. However, these methods may be performed additionally or alternatively by any system or combination of systems, including but not limited to the system described herein. Method 400 may be performed at least in part using hardware accelerators, such as in combination with a CPU and / or GPU. Method 400 may include references Figure 2 and / or Figure 3 Describe one or more operations.

[0067] Figure 4 This is a flowchart illustrating a method 400 for synchronizing sensor data according to some embodiments of the present disclosure. Method 400 may be performed in response to receiving and / or sampling sensor data from one or more sensors. Method 400 may be performed in response to a request to adjust (e.g., increase) the frame rate of the data.

[0068] At block B402, method 400 includes: retrieving image frames from one or more sensors. For example, image frames from multiple of one or more cameras, LiDAR sensors, and / or RADAR sensors may be received at one or more timestamps. In some embodiments, a first frame from a first sensor may be asynchronous with a second image frame from a second sensor, e.g., time-offset from the first image frame. In some embodiments, one or more frames from one or more sensors may be missing, e.g., not output and / or received within a threshold time range relative to an expected time, such as an expected time associated with the frame rate of one or more frames. Image frames may be retrieved periodically, e.g., based on readouts from the sensors and / or the sensor's frame rate. Image frames may be retrieved in response to a request for an image frame.

[0069] At box B404, method 400 includes detecting motion between image frames from a given sensor. For example, in response to detecting an asynchronous situation relative to image frames from a given sensor, multiple image frames from the given sensor (e.g., two consecutive image frames) can be processed to detect motion, such as detecting displacement across image frames. This can include, for example, detecting the displacement of pixels representing one or more features in the image frames. For example, motion can represent the amount of displacement of a feature over time between timestamps of the image frames. In some embodiments, image frames may be applied to optical flow algorithms and / or hardware performing optical flow (e.g., an optical flow accelerator) to detect motion.

[0070] At box B406, method 400 includes: generating the target frame based on the detected motion and the target time of the target frame. The target time can be a timestamp of a frame from another sensor, for example, synchronizing a frame from a given sensor with a frame from another sensor. The target time can also be a timestamp indicating that a missing frame is expected to have been received.

[0071] For example, detected motion can be used to interpolate one or more image frames from a given sensor to a target time, such as from the corresponding timestamps of one or more image frames to the target time. For instance, in cases where motion detection involves a flow graph that assigns displacement vectors to one or more pixels in an image-like data structure, one or more frames can be modified by interpolating the positions of features associated with pixels in one or more frames to the modified positions in the target frame based on displacements belonging to the displacement vectors. In some implementations, the target frame is used for additional processing and / or output generation, such as stitching, reconstruction, and / or display operations.

[0072] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, submarines, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a wide range of purposes, including, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0073] The disclosed implementations include various systems such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems including one or more large language models (LLMs), systems including one or more visual language models (VLMs), systems for hosting real-time streaming applications, systems for presenting one or more of virtual reality content, augmented reality content, or mixed reality content, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0074] Example autonomous vehicles

[0075] Figure 5AThis is an illustration of an example autonomous vehicle 500 according to some embodiments of the present disclosure. The autonomous vehicle 500 (or, alternatively, referred to herein as “vehicle 500”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, submarines, robotic vehicles, drones, aircraft, trailer-coupled vehicles (semi-tractor-trailer trucks for hauling goods) and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in its "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 500 may be able to implement one or more functions that meet Level 3 through Level 5 of autonomous driving standards. For example, depending on the embodiment, vehicle 500 may be able to implement driver assistance (Level 1), semi-automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein can include any and / or all types of autonomy of the vehicle 500 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, auxiliary autonomy, semi-autonomy, primary autonomy, or other specified autonomy.

[0076] Vehicle 500 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 550 may be connected to the drivetrain of vehicle 500, which may include a transmission, to enable propulsion of vehicle 500. Propulsion system 550 may be controlled in response to receiving a signal from throttle / accelerator 552.

[0077] A steering system 554, which may include a steering wheel, can be used to steer the vehicle 500 (e.g., along a desired path or route) when the propulsion system 550 is operating (e.g., when the vehicle is in motion). The steering system 554 may receive signals from the steering actuator 556. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0078] The brake sensor system 546 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 548 and / or the brake sensor.

[0079] It may include one or more System-on-a-Chip (SoC) 504 ( Figure 5C One or more controllers 536, including and / or one or more GPUs, may provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 500. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 548, to operate steering system 554 via one or more steering actuators 556, and to operate propulsion system 550 via one or more throttles / accelerators 552. One or more controllers 536 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 500. One or more controllers 536 may include a first controller 536 for autonomous driving functions, a second controller 536 for functional safety functions, a third controller 536 for artificial intelligence functions (e.g., computer vision), a fourth controller 536 for infotainment functions, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 536 can handle two or more of the functions described above, and two or more controllers 536 can handle a single function, and / or any combination thereof.

[0080] One or more controllers 536 may provide signals for controlling one or more components and / or systems of vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, a Global Navigation Satellite System sensor (“GNSS”) 558 (e.g., a Global Positioning System sensor), a RADAR sensor 560, an ultrasonic sensor 562, a LIDAR sensor 564, an Inertial Measurement Unit (IMU) sensor 566 (e.g., an accelerometer, gyroscope, magnetic compass, magnetometer, etc.), a microphone 596, a stereo camera 568, a wide-angle camera 570 (e.g., a fisheye camera), an infrared camera 572, a surround camera 574 (e.g., a 360-degree camera), a long-range and / or medium-range camera 598, a speed sensor 544 (e.g., for measuring the rate of vehicle 500), a vibration sensor 542, a steering sensor 540, a braking sensor (e.g., as part of a braking sensor system 546), and / or other sensor types. System 100 can be used to process sensor data from any one or more such sensors, for example, to synchronize and / or improve the accuracy of the sensor data with respect to a target time.

[0081] One or more of the controllers 536 may receive inputs (e.g., represented by input data) from the instrument cluster 532 of the vehicle 500 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 534, an auditory signaling device, a speaker, and / or via other components of the vehicle 500. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 5C Information such as a high-definition (“HD”) map 522, location data (e.g., the location of vehicle 500 on the map), direction, the location of other vehicles (e.g., occupying a grid), and information about objects and their states perceived by controller 536, etc. For example, HMI display 534 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0082] Vehicle 500 further includes a network interface 524, which can communicate via one or more networks using one or more wireless antennas 526 and / or a modem. For example, network interface 524 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 526 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (LPWANs such as LoRaWAN, SigFox, etc.).

[0083] Figure 5B For use in accordance with some embodiments of this disclosure Figure 5A This is an example of the camera position and field of view of an autonomous vehicle 500. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 500.

[0084] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 500. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.

[0085] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0086] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) parts to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0087] A camera with a field of view that includes the environment in front of the vehicle 500 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 536 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (“ACC”), and / or other functions such as traffic sign recognition.

[0088] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 570, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 5B The diagram shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 570 can be present on vehicle 500. Furthermore, any number of remote cameras 598 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, particularly for objects for which neural networks have not yet been trained. Remote cameras 598 can also be used for object detection and classification, as well as basic object tracking. In some embodiments, vehicle 500 includes ten or more (e.g., eleven, fourteen) cameras.

[0089] One or more stereo cameras 568 may also be included in the front-mounted configuration. In at least one embodiment, one or more stereo cameras 568 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 568 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 568 may be used in addition to those described herein or alternatively.

[0090] Cameras with a field of view including the side portion of the environment of vehicle 500 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 574 (e.g., ... Figure 5B The four surround cameras 574 shown can be mounted on the vehicle 500. The surround cameras 574 can include a wide-angle camera 570, a fisheye camera, a 360-degree camera, and / or the like. Four examples are provided; the four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 574 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0091] Cameras with a field of view that includes the environment behind vehicle 500 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range cameras 598, stereo cameras 568, infrared cameras 572, etc.).

[0092] Figure 5C For use in accordance with some embodiments of this disclosure Figure 5AThe example autonomous vehicle 500 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0093] Figure 5C Each component, feature, and system in vehicle 500 is illustrated as being connected via bus 502. Bus 502 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 500 used to assist in the control of various features and functions of vehicle 500, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0094] Although bus 502 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 502 is represented by a single line, this is not intended to be limiting. For example, any number of buses 502 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 502 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 502 may be used for a collision avoidance function, and a second bus 502 may be used for drive control. In any example, each bus 502 may communicate with any component of vehicle 500, and two or more buses 502 may communicate with the same component. In some examples, each SoC 504, each controller 536, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 500) and may be connected to a common bus such as a CAN bus.

[0095] Vehicle 500 may include one or more controllers 536, such as those described herein. Figure 5A The controllers described herein. Controller 536 can be used for a wide variety of functions. Controller 536 can be coupled to any other different components and systems of vehicle 500 and can be used for the control of vehicle 500, artificial intelligence of vehicle 500, infotainment and / or the like for vehicle 500.

[0096] Vehicle 500 may include one or more System-on-Chip (SoC) 504s. SoC 504 may include a CPU 506, GPU 508, processor 510, cache 512, accelerator 514, data storage 516, and / or other components and features not shown. SoC 504 can be used to control vehicle 500 across a wide variety of platforms and systems. For example, one or more SoCs 504s may be combined with an HD map 522 in a system (e.g., the system of vehicle 500), the HD map being transmitted via a network interface 524 from one or more servers (e.g., [server name missing]). Figure 5D One or more servers (578) receive map refresh and / or updates.

[0097] CPU 506 may include CPU clusters or CPU complexes (alternatively referred to herein as "CCPLEX"). CPU 506 may include multiple cores and / or L2 cache. For example, in some embodiments, CPU 506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 506 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). CPU 506 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 506 can be active at any given time.

[0098] CPU 506 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. CPU 506 can further implement enhanced algorithms for managing power states, wherein allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.

[0099] GPU 508 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 508 may be programmable and efficient for parallel workloads. In some examples, GPU 508 may use an enhanced tensor instruction set. GPU 508 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 508 may include at least eight streaming microprocessors. GPU 508 may use a computation application programming interface (API). Furthermore, GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0100] In automotive and embedded applications, the GPU 508 can be power-optimized for optimal performance. For example, the GPU 508 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 508 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations for efficient execution of workloads. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.

[0101] The GPU 508 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.

[0102] The GPU 508 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 508 to directly access the CPU 506 page tables. In such examples, when the GPU 508 Memory Management Unit (MMU) experiences a miss, an address translation request can be transferred to the CPU 506. In response, the CPU 506 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 508. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 506 and the GPU 508, simplifying GPU 508 programming and porting applications to the GPU 508.

[0103] In addition, the GPU 508 may include access counters that track how frequently the GPU 508 accesses the memory of other processors. These access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0104] SoC 504 may include any number of caches 512, including those described herein. For example, cache 512 may include an L3 cache available to both CPU 506 and GPU 508 (e.g., it is connected to both CPU 506 and GPU 508). Cache 512 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.

[0105] SoC 504 may include one or more arithmetic logic units (ALUs) that can be used to perform processing of any of a variety of tasks or operations related to vehicle 500, such as processing a DNN. Additionally, SoC 504 may include a floating-point unit (FPU) or other mathematical coprocessor or digital coprocessor type for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated as execution units within CPU 506 and / or GPU 508.

[0106] SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 508 and offload some tasks from GPU 508 (e.g., freeing up more cycles of GPU 508 to perform other tasks). As an example, accelerator 514 can be used for targeted workloads (e.g., perceptrons, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0107] Accelerator 514 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0108] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0109] The DLA can perform any function of the GPU 508, and by using inference accelerators, for example, a designer can make either the DLA or the GPU 508 target any function. For example, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 508 and / or other accelerators 514.

[0110] Accelerator 514 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0111] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.

[0112] DMA enables PVA components to access system memory independently of the CPU 506. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0113] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.

[0114] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0115] Accelerator 514 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 514. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both the PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA may access memory via a backbone that provides high-speed memory access to the PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects the PVA and DLA to memory.

[0116] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.

[0117] In some examples, SoC 504 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.

[0118] Accelerators 514 (e.g., hardware accelerator clusters) have broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.

[0119] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras. PVA can be used for preprocessing of optical flow operations (e.g., generating pyramids that will be processed by an optical flow accelerator (OFA)).

[0120] In some examples, PVA can be used to perform intensive optical flow, providing processed RADAR data from the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.

[0121] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as its input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 566 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 564 or RADAR sensor 560), etc.

[0122] SoC 504 may include one or more data storage units 516 (e.g., memory). Data storage units 516 may be on-chip memory of SoC 504, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, data storage units 516 may be large enough to store multiple instances of the neural network. Data storage units 512 may include L2 or L3 cache 512. References to data storage units 516 may include references to memory associated with PVA, DLA, and / or other accelerators 514 as described herein.

[0123] SoC 504 may include one or more processors 510 (e.g., embedded processors). Processor 510 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 504 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 504 thermal and temperature sensor management, and / or SoC 504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 504 may use the ring oscillator to detect the temperature of CPU 506, GPU 508, and / or accelerator 514. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 504 into a lower power state and / or place vehicle 500 into a driver-safe parking mode (e.g., safely stop vehicle 500).

[0124] The processor 510 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.

[0125] The processor 510 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0126] Processor 510 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.

[0127] The processor 510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0128] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0129] Processor 510 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 570, the surround camera 574, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.

[0130] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.

[0131] The video image compositer can also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 508 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 508 is powered on and activated, performing 3D rendering, the video image compositer can be used to offload the GPU 508 to improve performance and responsiveness.

[0132] SoC 504 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. SoC 504 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.

[0133] SoC 504 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 504 can be used to process data from cameras and sensors (e.g., LIDAR sensor 564, RADAR sensor 560, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 502 (e.g., vehicle 500 speed, steering wheel position, etc.), and data from GNSS sensor 558 (connected via Ethernet or CAN bus). SoC 504 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 506 from routine data management tasks.

[0134] SoC 504 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. SoC 504 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with CPU 506, GPU 508, and data storage 516, accelerator 514 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.

[0135] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages ​​such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0136] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 520) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.

[0137] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions," along with a light, can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network, which informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 508.

[0138] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 500. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 504 provides security against theft and / or carjacking.

[0139] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 596 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 504 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 558. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 562, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.

[0140] The vehicle may include a CPU 518 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 may include, for example, an x86 processor. The CPU 518 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 504, and / or monitoring the status and health of the controller 536 and / or the infotainment SoC 530.

[0141] Vehicle 500 may include a GPU 520 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 520 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 500 (e.g., sensor data).

[0142] Vehicle 500 may further include a network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 524 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 578 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 500 with information about vehicles approaching vehicle 500 (e.g., vehicles in front, to the side, and / or behind vehicle 500). This functionality can be part of vehicle 500's cooperative adaptive cruise control function.

[0143] Network interface 524 may include a SoC that provides modulation and demodulation functions and enables controller 536 to communicate via a wireless network. Network interface 524 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0144] Vehicle 500 may further include data storage 528, which may include off-chip (e.g., off-chip SoC 504) storage devices. Data storage 528 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0145] Vehicle 500 may further include GNSS sensor 558. GNSS sensor 558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for assisted mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 558 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0146] Vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 can be used by vehicle 500 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 560 can use CAN and / or bus 502 (e.g., to transmit data generated by the RADAR sensor 560) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 560 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0147] The RADAR sensor 560 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 560 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 500's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 500's lane.

[0148] As an example, a mid-range RADAR system can include a range of up to 560m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.

[0149] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0150] Vehicle 500 may further include ultrasonic sensors 562. Ultrasonic sensors 562, which may be positioned at the front, rear, and / or sides of vehicle 500, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 562 can be used, and different ultrasonic sensors 562 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 562 can operate at functional safety level ASIL B.

[0151] Vehicle 500 may include a LIDAR sensor 564. The LIDAR sensor 564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 564 may be of functional safety level ASIL B. In some examples, vehicle 500 may include multiple LIDAR sensors 564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0152] In some examples, the LiDAR sensor 564 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 564 may have an advertising range of, for example, approximately 500m, with an accuracy of 2cm-3cm, and support for 500Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 564 may be used. In such examples, the LiDAR sensor 564 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 500. In such examples, the LiDAR sensor 564 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 564 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0153] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-browsing LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LiDAR, and because a flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 564 is less susceptible to motion blur, vibration, and / or shock.

[0154] The vehicle may further include an IMU sensor 566. In some examples, the IMU sensor 566 may be located at the center of the rear axle of the vehicle 500. The IMU sensor 566 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 566 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 may include an accelerometer, a gyroscope, and a magnetometer.

[0155] In some embodiments, the IMU sensor 566 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 566 can enable the vehicle 500 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 566 without input from a magnetic sensor. In some examples, the IMU sensor 566 and the GNSS sensor 558 can be combined into a single integrated unit.

[0156] The vehicle may include a microphone 596 placed in and / or around the vehicle 500. Among other things, the microphone 596 may be used for emergency vehicle detection and identification.

[0157] The vehicle may further include any number of camera types, including stereo camera 568, wide-angle camera 570, infrared camera 572, surround camera 574, long-range and / or mid-range camera 598, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 500. The camera types used depend on the embodiment and the requirements of the vehicle 500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 500. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 5A and Figure 5B It was described in more detail.

[0158] Vehicle 500 may further include vibration sensor 542. Vibration sensor 542 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 542 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).

[0159] Vehicle 500 may include ADAS system 538. In some examples, ADAS system 538 may include SoC. ADAS system 538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0160] The ACC system can use a RADAR sensor 560, a LIDAR sensor 564, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 500 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0161] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or network connection (e.g., via the Internet) through network interface 524 and / or wireless antenna 526. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 500 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 500, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.

[0162] The Forward Collision Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.

[0163] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.

[0164] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0165] The LKA system is a variation of the LDW system. If vehicle 500 begins to leave the lane, the LKA system provides corrective steering input or braking to vehicle 500.

[0166] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system can utilize a rear-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0167] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0168] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as they alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 500, in the event of conflicting results, the vehicle 500 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 may be a backup and / or auxiliary computer used to provide perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and varied software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from the ADAS system 538 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0169] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0170] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of SoC 504 and / or be included as components of SoC 504.

[0171] In other examples, ADAS system 538 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0172] In some examples, the output of ADAS system 538 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if ADAS system 538 issues a forward collision warning because an object is immediately in front, the perception block can use this information when recognizing the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0173] Vehicle 500 may further include an infotainment SoC 530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 530 may include a combination of hardware and software that can be used to provide vehicle 500 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 530 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 534, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 530 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from the ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0174] The infotainment SoC 530 may include GPU functionality. The infotainment SoC 530 can communicate with other devices, systems, and / or components of the vehicle 500 via bus 502 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 530 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 536 (e.g., the primary and / or backup computer of the vehicle 500). In such an example, the infotainment SoC 530 may place the vehicle 500 into a driver-safe parking mode as described herein.

[0175] Vehicle 500 may further include an instrument cluster 532 (e.g., a digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 532 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.

[0176] Figure 5D For cloud-based servers and according to some embodiments of this disclosure Figure 5A This is a system diagram illustrating communication between example autonomous vehicles 500. System 576 may include server 578, network 590, and vehicles including vehicle 500. Server 578 may include multiple GPUs 584(A)-584(H) (collectively referred to herein as GPU 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switch 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPU 580). GPUs 584, CPUs 580, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 586, such as, but not limited to, NVLink interfaces 588 developed by NVIDIA. In some examples, GPUs 584 are connected via NVLink and / or NVSwitch SoCs, and GPUs 584 and PCIe switches 582 are connected via PCIe interconnects. Although eight GPUs 584, two CPUs 580, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, each of the servers 578 may include eight, sixteen, thirty-two, and / or more GPUs 584.

[0177] Server 578 can receive image data from vehicles via network 590, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 578 can also transmit neural network 592, updated neural network 592, and / or map information 594, including information about traffic and road conditions, to vehicles via network 590. Updates to map information 594 may include updates to HD map 522, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 592, updated neural network 592, and / or map information 594 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 578 and / or other servers).

[0178] Server 578 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 590), and / or the machine learning model can be used by server 578 to remotely monitor the vehicle.

[0179] In some examples, server 578 can receive data from a vehicle and apply that data to a state-of-the-art real-time neural network for real-time intelligent inference. Server 578 may include a deep learning supercomputer powered by GPU 584 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 578 may include a deep learning infrastructure in a data center that uses only CPU power.

[0180] The deep learning infrastructure of server 578 may be capable of rapid, real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 500. For example, the deep learning infrastructure may receive periodic updates from vehicle 500, such as image sequences and / or objects located in those image sequences by vehicle 500 (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 500. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 500 has malfunctioned, then server 578 may transmit a signal to vehicle 500 instructing its fail-safe computer to take control, notify passengers, and complete a safe stopping operation.

[0181] For inference, server 578 may include GPU 584 and one or more programmable inference accelerators (such as NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.

[0182] Example computing device

[0183] Figure 6 The block diagram is provided for an example computing device 600 suitable for implementing some embodiments of the present disclosure. The computing device 600 may include an interconnect system 602 directly or indirectly coupled to the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., a display), and one or more logic units 620. In at least one embodiment, the computing device 600 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 608 may include one or more vGPUs, one or more CPUs 606 may include one or more vCPUs, and / or one or more logic units 620 may include one or more virtual logic units. Therefore, computing device 600 may include discrete components (e.g., a complete GPU dedicated to computing device 600), virtual components (e.g., a portion of the GPU dedicated to computing device 600), or combinations thereof. Computing device 600 may be used to implement at least some components of system 100.

[0184] although Figure 6 The various boxes are shown connected via an interconnect system 602 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 618, such as a display device, may be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device other than the memory of GPU 608, CPU 606, and / or other components). In other words, Figure 6 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 6 Within the scope of computing devices.

[0185] Interconnect system 602 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 602 may include one or more link or bus types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 606 may be directly connected to memory 604. Furthermore, CPU 606 may be directly connected to GPU 608. In cases where there is a direct or point-to-point connection between components, interconnect system 602 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required in computing device 600.

[0186] Memory 604 may include any of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by computing device 600. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.

[0187] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any way or by any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 600. As used herein, computer storage media does not include the signal itself.

[0188] Computer storage media may contain computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transport medium. The term "modulated data signal" can refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0189] CPU 606 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. Each of CPU 606 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 606 may include any type of processor and may include different types of processors depending on the type of computing device 600 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 600, the processor may be an advanced RISC mechanism (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 600 may also include one or more CPUs 606.

[0190] In addition to or replacing CPU 606, GPU 608 may also be configured to execute at least some computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. One or more GPUs 608 may be integrated GPUs (e.g., having one or more CPUs 606) and / or one or more GPUs 608 may be discrete GPUs. In embodiments, one or more GPUs 608 may be coprocessors of one or more CPUs 606. Computing device 600 may use GPU 608 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 608 may be used for general-purpose computing on a GPU (GPGPU). GPU 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPU 608 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received via a host interface from CPU 606). GPU 608 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 604. GPU 608 may include two or more GPUs operating in parallel (e.g., via links). The links may connect the GPUs directly (e.g., using NVLINK) or via a switch (e.g., using NVSwitch). When combined, each GPU 608 may generate different portions of pixel data or GPGPU data for different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.

[0191] In addition to or replacing CPU 606 and / or GPU 608, logic unit 620 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 600 to perform one or more methods and / or processes described herein. In embodiments, CPU 606, GPU 608, and / or logic unit 620 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 620 may be part of and / or integrated into one or more CPUs 606 and / or one or more GPUs 608, and / or one or more logic units 620 may be discrete components of CPU 606 and / or GPU 608 or otherwise external thereto. In embodiments, one or more logic units 620 may be processors of one or more CPUs 606 and / or one or more GPUs 608.

[0192] Examples of logic unit 620 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating-point unit (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) elements, etc.

[0193] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 610 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) for directly transmitting data received via the network and / or interconnect system 602 to one or more GPUs 608 (e.g., the memory of GPU 608).

[0194] I / O port 612 enables computing device 600 to be logically coupled to other devices, including I / O component 614, presentation component 618, and / or other components, some of which may be built into (e.g., integrated into) computing device 600. Illustrative I / O component 614 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, browsers, printers, wireless devices, and so on. I / O component 614 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 600 (described in more detail below). Computing device 600 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, computing device 600 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 600 to render immersive augmented reality or virtual reality.

[0195] Power supply 616 may include a hard-wired power supply, a battery power supply, or a combination thereof. Power supply 616 may supply power to computing device 600 so that the components of computing device 600 can operate.

[0196] The presentation component 618 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 618 may receive data from other components (such as GPU 608, CPU 606, DPU, etc.) and output that data (e.g., as images, videos, sounds, etc.).

[0197] Example Data Center

[0198] Figure 7 An example data center 700 is shown, which can be used in at least one embodiment of this disclosure. The data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0199] like Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“nodes CR”) 716(1)-716(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR 716(1)-716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and cooling modules, etc. In some embodiments, one or more node CRs of nodes CR 716(1)-716(N) may correspond to servers having one or more of the aforementioned computing resources. Furthermore, in some embodiments, node CRs

[0200] 716(1)-716(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the nodes CR716(1)-716(N) may correspond to virtual machines (VMs).

[0201] In at least one embodiment, the grouped computing resources 714 may include individual groups (not shown) of nodes CR716 housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. Individual groups of nodes CR716 within the grouped computing resources 714 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several nodes CR716, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0202] Resource coordinator 712 may be configured or otherwise controlled to control one or more nodes CR716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may include a Software Design Infrastructure (“SDI”) management entity for data center 700. Resource coordinator 712 may include hardware, software, or some combination thereof.

[0203] In at least one embodiment, such as Figure 7As shown, framework layer 720 may include a job scheduler 733, a configuration manager 734, a resource manager 736, and a distributed file system 738. Framework layer 720 may include a framework for software 732 supporting software layer 730 and / or one or more applications 742 of application layer 740. Software 732 or application 742 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 738 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 733 may include a Spark driver for facilitating the scheduling of workloads supported by various layers of data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as software layer 730 and framework layer 720 including Spark and a distributed file system 738 for supporting large-scale data processing. The resource manager 736 is able to manage cluster or grouped computing resources mapped to or allocated to support the distributed file system 738 and the job scheduler 733. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 at data center infrastructure layer 710. The resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.

[0204] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes CR716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet webpage search software, email virus browsing software, database software, and streaming video content software. The software 732 may be used to implement one or more operations of the system 100.

[0205] In at least one embodiment, the application layer 740 may include one or more applications 742 that can be used by at least a portion of the nodes CR716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0206] In at least one embodiment, any of the configuration manager 734, resource manager 736, and resource coordinator 712 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can alleviate the risk of data center operators of data center 700 making potentially poor configuration decisions and can prevent underutilization and / or skewed portions of the data center.

[0207] Data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 700. In at least one embodiment, by using weight parameters calculated through one or more training techniques, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks, such as, but not limited to, those described herein, using the resources described above with respect to data center 700.

[0208] In at least one embodiment, the data center 700 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0209] Example network environment

[0210] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 6 The implementation is carried out on one or more instances of computing device 600—for example, each device may include similar components, features, and / or functions of computing device 600. Furthermore, in the case of implementing back-end devices (e.g., servers, NAS, etc.), the back-end devices may be included as part of data center 700, examples of which are described herein. Figure 7 To describe in more detail.

[0211] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks, or networks within multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0212] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.

[0213] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or applications may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software network application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").

[0214] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0215] Client devices may include those described in this article. Figure 6 The example computing device 600 described includes at least some components, features, and functions. By way of example and not limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming equipment or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0216] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.

[0217] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" could include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0218] This document describes in detail the subject matter of this disclosure to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

Claims

1. One or more processors, comprising: One or more circuits are used for: Detect asynchronous sensor data from the sensor; In response to the detection of the asynchronous situation, motion associated with the first frame and the second frame of the sensor data is determined; The third frame is generated based at least on the motion and the first frame; as well as Perform one or more machine-related operations based on at least the generated third frame.

2. The processor of claim 1 or more, wherein the one or more circuits are configured to detect the motion by detecting optical flow between the first frame and the second frame.

3. The processor of claim 1 or more, wherein the sensor is a first sensor, and the one or more circuits are configured to detect the asynchronous condition in response to detecting a time difference between a timestamp of the first frame and a timestamp of a fourth frame from at least one second sensor.

4. The processor of claim 1 or more, wherein the one or more circuits are configured to detect the asynchronous condition in response to detecting that the missing frame is not received from the sensor at the expected timestamp of the missing frame.

5. The processor of claim 1 or more, wherein the sensor is a first sensor, and the one or more circuits are configured to generate the third frame to have a timestamp equal to that of the fourth frame from the second sensor.

6. One or more processors according to claim 1, wherein the one or more circuits are configured to generate the third frame to have a timestamp following the timestamps of the first frame and the second frame.

7. One or more processors according to claim 1, wherein the one or more circuits are configured to detect the motion based at least on the displacement of at least one pixel representing an object or feature from a first position in the first frame to a second position in the second frame.

8. One or more processors according to claim 1, wherein the sensor comprises a camera, a light detection and ranging LIDAR system or a radio frequency detection and ranging RADAAR system.

9. One or more processors according to claim 1, wherein the one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system containing one or more virtual machines (VMs); Systems implemented using robots; A system used to perform deep learning operations; A system used to perform simulation operations; A system for collaborative content creation of 3D assets; A system for generating synthetic data; Systems used to perform digital twin operations; Systems implemented using edge devices; A system that includes one or more Visual Language Models (VLMs); Systems that include one or more large language model LLMs; A system that includes one or more multimodal language models; Systems used to perform conversational AI operations; A system for performing optical transmission simulation; A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

10. A system comprising: One or more processing units are configured to trigger the execution of an operation, the operation comprising: Detect asynchronous sensor data from the sensor; In response to detecting the asynchronous condition, determine the motion associated with the first frame and the second frame of the sensor data; and The third frame is generated based at least on the motion and the first frame.

11. The system of claim 10, wherein the one or more processing units are configured to detect the motion by detecting optical flow between the first frame and the second frame.

12. The system of claim 10, wherein the sensor is a first sensor, and the one or more processing units are configured to detect the asynchronous situation in response to detecting a time difference between the timestamp of the first frame and the timestamp of a fourth frame from at least one second sensor.

13. The system of claim 10, wherein detecting the asynchronous condition is in response to detecting that the missing frame was not received from the sensor at the expected timestamp of the missing frame.

14. The system of claim 10, wherein the sensor is a first sensor, and the operation further comprises: The third frame is generated to have a timestamp equal to that of the fourth frame from the second sensor.

15. The system of claim 10, wherein the operation further comprises: The third frame is generated to have a timestamp following the timestamps of the first and second frames.

16. The system of claim 10, wherein the detection of the motion is based at least on the displacement of at least one pixel representing an object or feature from a first position in the first frame to a second position in the second frame.

17. The system of claim 10, wherein the sensor comprises a camera, a light detection and ranging LIDAR system or a radio frequency detection and ranging RADAAR system.

18. The system of claim 10, wherein the system comprises at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system containing one or more virtual machines (VMs); Systems implemented using robots; A system used to perform deep learning operations; A system used to perform simulation operations; A system for collaborative content creation of 3D assets; A system for generating synthetic data; Systems used to perform digital twin operations; Systems implemented using edge devices; A system that includes one or more Visual Language Models (VLMs); Systems that include one or more large language model LLMs; A system that includes one or more multimodal language models; Systems used to perform conversational AI operations; A system for performing optical transmission simulation; A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

19. A method comprising: Asynchronous detection of sensor data from sensors using one or more processors; In response to the detection of the asynchronous situation, the one or more processors are used to determine the motion associated with the first frame and the second frame of the sensor data; as well as The third frame is generated using one or more processors, based at least on the motion and the first frame.

20. The method of claim 19, wherein the sensor is a first sensor, and detecting the asynchronous condition comprises: Detecting at least a threshold time difference between the first frame and the fourth frame from the second sensor, and generating the third frame includes: interpolating the timestamp of the first frame to the fourth frame.

21. The method of claim 19, wherein generating the third frame comprises: The first frame is interpolated to the timestamp of the missing frame that was not received from the sensor.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2