Architecture and instruction set for multi-dimensional data processing
By using a multi-dimensional data processing architecture and instruction set, multiple processors are used to process image data in parallel, generate multiple features and perform measurements. This solves the speed mismatch problem in traditional processing systems, achieves fast and accurate image feature recognition, and improves image processing speed and system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional processing systems suffer from speed mismatch when handling complex datasets, leading to a decline in overall system performance, especially in image and graphics processing where they cannot quickly and accurately process large amounts of image data.
It adopts a multi-dimensional data processing architecture and instruction set, processes multiple instances of frame data in parallel through multiple processors, generates multiple features and measures them in multiple dimensions, and uses a pixel processing engine and vector processing unit to perform feature recognition operations, reducing redundant data loading operations and achieving parallel execution across multiple processors.
It significantly improves image processing speed, enhances responsiveness in low-latency environments, exceeds the capabilities of conventional processors, and achieves fast and accurate image feature recognition.
Smart Images

Figure CN121635962A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 2024111228640, filed on August 15, 2024, entitled "Architecture and Instruction Set for Multidimensional Data Processing". Technical Field
[0002] This implementation generally relates to microprocessor devices, including but not limited to architectures and instruction sets for multidimensional data processing, architectures for reusing frame data across multidimensional data processors, and parallelization architectures for multidimensional data processors. Background Technology
[0003] There is a growing expectation that computing processors can handle increasingly complex datasets at faster speeds. However, traditional processing systems may vary significantly in processing speed or data transfer rate, leading to a mismatch between processing components. This can degrade overall system performance and reduce or eliminate the ability to directly perform various types of computational processes, such as image processing or graphics processing. Summary of the Invention
[0004] This technical solution can improve processing speed in low-latency applications while maintaining the integrity of image feature recognition at higher speeds. For example, in image processing environments related to autonomous or semi-autonomous navigation (e.g., driving, robot control, etc.), it is necessary to process large amounts of image data quickly and accurately to maintain a reliable and up-to-date model of the physical environment. For instance, embodiments of this disclosure can provide high-speed and accurate image feature recognition of input frame data that exceeds the capabilities of CPU processing or typical GPU processing. Therefore, a technical solution for an architecture and instruction set for multidimensional data processing is provided.
[0005] At least one aspect relates to a system. The system may include multiple processors. The system can determine multiple instances of frame data, each instance being individually modified according to at least one of multiple dimensions, the multiple instances of frame data being provided to a corresponding processor among the multiple processors, each processor being configured to execute input arranged in the multiple dimensions. The system can generate multiple features based at least partially on the multiple instances of frame data, each feature corresponding to a frame data instance from the multiple instances of frame data. The system can generate a measure of a first frame data in multiple dimensions, indicating visual attributes, based at least partially on the multiple features.
[0006] At least one aspect relates to a system-on-a-chip (SoC). The SoC may include at least one graphics processing unit (GPU) and multiple processors. The system can determine multiple instances of frame data, each instance modified according to at least one of multiple dimensions. The system can generate multiple features, each corresponding to a specific instance of the frame data, based at least in part on the multiple instances of the frame data. The system can generate a metric for a first frame of data in multiple dimensions, indicating visual attributes, based at least in part on the multiple features.
[0007] At least one aspect relates to a method executed by multiple processors. The method may include: determining multiple instances of frame data, each instance being individually modified according to at least one of multiple dimensions; the multiple instances of frame data being provided to a corresponding processor among a plurality of processors, each processor being configured to execute input arranged in the multiple dimensions. The method may include: generating multiple features, at least partially based on the multiple instances of frame data, each feature corresponding to a frame data instance from the multiple instances of frame data. The method may include generating a measure of a first frame data in the multiple dimensions, the measure indicating visual attributes, at least partially based on the multiple features.
[0008] At least one aspect relates to a system. The system may include a plurality of processors configured to execute inputs in multiple dimensions. The system may include a memory device at least partially integrated with the plurality of processors, the plurality of processors generating a first frame output via a first feature recognition operation in the multiple dimensions based at least partially on input including first frame data arranged in the multiple dimensions, storing the first frame output in the memory device, the first frame output corresponding to an edge of the first frame data along one of the multiple dimensions, generating a second output of at least one second feature recognition operation in the multiple dimensions based at least partially on input including second frame data, and generating second frame data indicating features of an image based at least partially on the second output and a portion of the first output.
[0009] At least one aspect relates to a System-on-a-Chip (SoC). The SoC may include at least one GPU and multiple processors for generating a first frame output via a first feature recognition operation in multiple dimensions based at least in part on an input including first frame data arranged in multiple dimensions, storing the first frame output in a memory device, the first frame output corresponding to an edge of the first frame data along one of the multiple dimensions, generating a second output of at least one second feature recognition operation in multiple dimensions based on an input including second frame data, and generating second frame data indicating features of an image based at least in part on the second output and a portion of the first output.
[0010] At least one aspect relates to a method executed by multiple processors. The method may include: generating a first frame output via a first feature recognition operation in multiple dimensions based on an input including first frame data arranged in multiple dimensions, to perform input in multiple dimensions. The method may include: storing the first frame output in a memory device, the first frame output corresponding to an edge of the first frame data along one of the dimensions of the multiple dimensions. The method may include: generating a second output of at least one second feature recognition operation in multiple dimensions based on an input including second frame data. The method may include: generating second frame data indicating image features based on the second output and a portion of the first output.
[0011] At least one aspect relates to a system. The system may include multiple processors. The system can determine that a first frame of data corresponds to a first time, and a second frame of data corresponds to a second time following the first time. The system can provide the first frame of data to a first processor among the multiple processors configured to execute input arranged in a two-dimensional configuration. The system can: in parallel with providing the first frame of data to the first processor, provide the second frame of data to a second processor among the multiple processors configured to execute input arranged in a one-dimensional configuration. The system can: in parallel with executing the second processor based on the second frame of data, execute the first processor based on the first frame of data.
[0012] At least one aspect relates to a System-on-a-Chip (SoC). The SoC may include at least one GPU and multiple processors. The SoC may determine that a first frame of data corresponds to a first time interval, and a second frame of data corresponds to a second time interval following the first time interval. The SoC may provide the first frame of data to a first processor among the multiple processors configured to execute inputs arranged in two dimensions. The SoC may: provide the second frame of data to a second processor among the multiple processors configured to execute inputs arranged in one dimension, in parallel with providing the first frame of data to the first processor. The SoC may: execute the first processor based on the first frame of data, in parallel with executing the second processor based on the second frame of data.
[0013] At least one aspect relates to a method executed by a plurality of processors. The method may include: determining that a first frame of data corresponds to a first time, and that a second frame of data corresponds to a second time following the first time. The method may include: providing the first frame of data to a first processor among the plurality of processors configured to execute input arranged in two dimensions. The method may include: in parallel with providing the first frame of data to the first processor, providing the second frame of data to a second processor among the plurality of processors configured to execute input arranged in one dimension. The method may include: in parallel with executing the second processor based on the second frame of data, executing the first processor based on the first frame of data. Attached Figure Description
[0014] These and other aspects and features of this embodiment are illustrated by way of example in the figures discussed herein. This embodiment may be directed to, but is not limited to, the examples shown in the figures discussed herein. Therefore, this disclosure is not limited to any figure or portion thereof shown or referenced herein, or any aspect described herein with respect to any figure shown or referenced herein.
[0015] Figure 1 An example computing environment according to some embodiments of the present disclosure is depicted, in which one or more devices use a system-on-a-chip (SoC) to process data;
[0016] Figure 2 An example loading memory architecture according to some embodiments of this disclosure is described;
[0017] Figure 3 An example method for detecting a two-dimensional maximum value according to some embodiments of the present disclosure is described;
[0018] Figure 4A Example two-dimensional block states according to some embodiments of this disclosure are depicted;
[0019] Figure 4B An example first transformation state of a two-dimensional block according to some embodiments of the present disclosure is depicted;
[0020] Figure 4C An example second transformation state of a two-dimensional block according to some embodiments of the present disclosure is depicted;
[0021] Figure 4D An example state of a horizontal image block according to some embodiments of the present disclosure is depicted;
[0022] Figure 4E An example state of multiple vertical blocks according to some embodiments of the present disclosure is depicted;
[0023] Figure 5 An example processor parallelization architecture for a single frame is described according to some embodiments of the present disclosure;
[0024] Figure 6 An example processor parallelization architecture for multiple frames is described according to some embodiments of the present disclosure;
[0025] Figure 7A Example methods for architectures and instruction sets for multidimensional data processing according to some embodiments of this disclosure are described;
[0026] Figure 7B Example methods for architectures and instruction sets for multidimensional data processing according to some embodiments of this disclosure are described;
[0027] Figure 7CExample methods for architectures and instruction sets for multidimensional data processing according to some embodiments of this disclosure are described;
[0028] Figure 7D Example methods for architectures and instruction sets for multidimensional data processing according to some embodiments of this disclosure are described;
[0029] Figure 8 Example methods for a parallelized architecture for multidimensional data processing according to some embodiments of the present disclosure are described;
[0030] Figure 9A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;
[0031] Figure 9B According to some embodiments of this disclosure Figure 9A Examples of camera positions and field of view for autonomous vehicles;
[0032] Figure 9C According to some embodiments of this disclosure Figure 9A A block diagram of an example system architecture for an example autonomous vehicle;
[0033] Figure 9D Cloud-based servers and according to some embodiments of this disclosure Figure 9A A system diagram illustrating communication between autonomous vehicles;
[0034] Figure 10 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0035] Figure 11 This is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0036] This document describes various aspects of the present technical solution with reference to the accompanying drawings, which are illustrative examples of the present technical solution. The drawings and examples below are not intended to limit the scope of the present technical solution to this embodiment or a single embodiment, and other embodiments can be implemented according to this embodiment, for example, by interchange of some or all of the described or illustrated elements. When certain elements of this embodiment can be implemented partially or entirely using known components, only the portions of these known components necessary for understanding this embodiment are described, and detailed descriptions of other portions of these known components are omitted to avoid obscuring the embodiment. Unless expressly stated herein, the terminology in the specification and claims should not be given uncommon or special meanings. Furthermore, this technical solution and this embodiment encompass existing and future known equivalents of the known components mentioned herein by description, illustration, or example.
[0037] This disclosure aims at least to improve the computational efficiency of determining image features in frame data by reducing the number of data loading operations and moving the loaded data in multiple dimensions (e.g., rows and columns of image data) to reduce processing latency and computational processing caused by redundant data loading operations. Therefore, this technical solution can significantly improve the speed of image processing, thereby enhancing responsiveness in low-latency environments beyond the capabilities of manual processing. Specifically, this technical solution can obtain multidimensional frame data (e.g., 2x2 image data blocks) and can provide portions of this frame data to multiple pixel processing engines (PPEs). Each PPE can process the frame data in multiple dimensions to provide technical improvements to increase the speed of image processing beyond the capabilities of general-purpose or conventional processors. For example, a pixel processing engine (PPE) can process input data according to a two-dimensional (e.g., NxN) data structure including multiple rows and columns. The system according to this disclosure can perform various feature recognition operations (e.g., Harris corner and non-maximum suppression) through one or more PPEs, which can provide accelerated NxN data execution on frame data. Therefore, this technical solution can achieve technical improvements, at least accelerating the hardware-level image feature recognition across multiple PPE processors as described in this article, but is not limited thereto.
[0038] This disclosure aims at least to reduce the computation of image feature recognition by reusing portions of image data in multiple feature recognition operations. In one aspect, the system may include one or more processors and memory devices (e.g., L2 caches), wherein portions of frame data may be temporarily stored in the memory devices and provided via the memory devices to one or more processors. For example, a processor may correspond to a PPE (Presentation Processing Element), which can process input data according to a two-dimensional (e.g., NxN) data structure comprising multiple rows and multiple columns. The PPE can perform various feature recognition operations across rows and columns, including, for example, performing Harris corner operations on an NxN block. The PPE may store a portion of the processed block (e.g., 1xN rows) into the memory devices, where this portion corresponds to the halo of the first frame. The PPE (or another PPE of the system) can perform a second feature recognition operation (e.g., non-maximum suppression) on a second frame of data (e.g., for different portions of the image data) with different frame data sizes (e.g., (N-1)xN), and can combine the output of the second feature recognition operation with a portion of the processed block to create the output of the second feature recognition operation of the full block size (NxN) via operations smaller than the full block size (e.g., (N-1)xN). Therefore, this technical solution can achieve technical improvements by reusing the data generated by the various feature recognition operations described herein, at least accelerating hardware-level image feature recognition, but is not limited thereto.
[0039] This disclosure aims at least to provide parallel execution across multiple processor types to maximize the utilization of each of the multiple processors, including parallelization of image feature processing and image feature preprocessing and loading operations. A processor may include multiple processing components (e.g., processor cores or allocated portions thereof), each of which can perform various feature recognition operations according to various data structures. For example, a vector processing unit (VPU) can process input data according to a one-dimensional (e.g., 1xN) data structure including a single row. For example, a PPE can process input data according to a two-dimensional (e.g., NxN) data structure including multiple rows and multiple columns. A system according to this disclosure can parallelize the execution of various feature recognition operations (e.g., Harris corner and non-maximum suppression) on a PPE that can be accelerated by NxN data execution, and various image processing operations (e.g., masking, packing, and halo preparation) on a VPU performing 1xN data execution. Therefore, this technical solution can achieve technical improvements through parallel processing across PPE and VPU processors as described herein, at least accelerating hardware-level image feature recognition, but is not limited thereto.
[0040] In one aspect, feature detection in image or video data may include computing features based on one or more image processing operations (e.g., Harris corner detection, SIFT, SURF). For example, a programmable one-dimensional (1D) single-instruction multiple-data (SIMD) processor (e.g., a VPU). For instance, a VPU may include multiple SIMD channels that can support vector memory operations and vector mathematics operations suitable for computationally intensive applications. For example, vision processing in autonomous or semi-autonomous driving or robot control requires higher resolution and more cameras, leading to a significant increase in computational workload. Therefore, feature detection processing running solely on a VPU device may consume up to 80% of the execution time. Therefore, this technical solution at least involves a PPE device to coordinate the execution of feature processing operations with one or more PPE devices to accelerate execution. Although this document discusses Harris corner image recognition as an example, this disclosure is not limited to Harris corner operations.
[0041] refer to Figure 1 Environment 100 may include processor 102, memory 104, instruction switch 106, memory 108 (sometimes referred to as dynamic random access memory or DRAM), and functional blocks 110a and 110b (unless otherwise specified, individually referred to as functional block 110, collectively referred to as functional block 110). In some embodiments, processor 102, memory 104, instruction switch 106, memory 108, and functional block 110 may be interconnected via wired and / or wireless connections (e.g., establishing connections for communication, etc.). In some embodiments, components of environment 100 may be included in a system-on-a-chip (SoC). For example, components of environment 100 may be included in one or more SoCs that combine some or all of the components of environment 100 to form an integrated circuit. Environment 100 may serve as... Figure 9A-11 Any functional block is included or can be used to implement Figure 9A-11 Any function block.
[0042] Processor 102 may include one or more processors, such as one or more central processing units (CPUs), graphics processing units (GPUs), microprocessors, microcontrollers, etc. Processor 102 may be interconnected with an instruction cache (not explicitly shown) that stores instructions for execution by processor 102. In some embodiments, processor 102 may be configured to output to and from... Figure 1Data associated with the configuration and / or control of one or more devices. For example, processor 102 may be configured to output data associated with the configuration of direct memory access (DMA) hardware sequencer 114a and / or DMA hardware sequencer 114b to control DMA transfers to and from vector memory (VMEM) 112a and / or VMEM 112b in function blocks 110a and 110b, respectively.
[0043] Memory 104 (sometimes referred to as an L2 buffer or L2 cache) may include a storage device interconnected with DMA hardware sequencers 114a and / or DMA hardware sequencer 114b of functional block 110. In some embodiments, memory 104 may be configured to receive and store data from DMA hardware sequencers 114a and / or DMA hardware sequencer 114b of functional block 110, as described herein. In some embodiments, memory 104 may have one or more (e.g., two) banks that allow simultaneous read or write requests. For example, memory 104 may have a first bank associated with DMA hardware sequencer 114a and a second bank associated with DMA hardware sequencer 114b.
[0044] Instruction switch 106 may include one or more processors configured to scan memory 108, receive data from memory 108, cause data stored in memory 108 and / or the local memory of instruction switch 106 to be loaded into VMEM 112, etc. For example, instruction switch 106 may be coupled to memory 108 and / or include internal memory storing instructions relating to operating one or more devices of the corresponding functional block 110. In one example, instruction switch 106 may be configured to fetch and provide data associated with an instruction to perform one or more DMA transfers as described herein. In another example, instruction switch 106 may be configured to fetch and provide data associated with an instruction to perform one or more operations specific to one or more devices of functional block 110. In an illustrative example, instruction switch 106 may be configured to fetch and provide data associated with an instruction to perform one or more filtering operations, and instruction switch 106 may transfer data to cache 120 of the corresponding functional unit 110. In this illustrative example, the corresponding cache 120 can be configured to transfer (e.g., load) instruction-associated data to VPU 116 or PPE 118 to enable the corresponding device to perform one or more filtering operations.
[0045] Memory 108 may include storage devices interconnected with DMA hardware sequencers 114a and / or DMA hardware sequencer 114b of functional block 110. In some embodiments, memory 108 may receive and store data generated by a robot (e.g., Figures 9A-9D The memory 108 may be configured to receive data based at least in part on direct interconnection with one or more sensors or indirect interconnection with one or more sensors (e.g., via communication via a CAN bus and / or the like) during robot operation. In these examples, the sensor data may include image data associated with one or more images generated by one or more cameras, LiDAR data associated with one or more LiDAR data associated with one or more point clouds generated by one or more LiDAR sensors, radar data associated with one or more radar images generated by one or more radar sensors, etc. In some embodiments, the memory 108 may be configured to provide (e.g., transfer) the sensor data stored therein to one or more components of functional block 110. For example, during processing of one or more images generated by one or more cameras of the robot, DMA hardware sequencers 114a and / or DMA hardware sequencer 114b may retrieve image data from memory 108 and store the image data in VMEM 112a and / or VMEM 112b, respectively. In some embodiments, memory 108 may receive and store data from DMA hardware sequencer 114a and / or DMA hardware sequencer 114b of function block 110. For example, DMA hardware sequencer 114a and / or DMA hardware sequencer 114b may provide image data updated at least in part based on image data processing to memory 108, and memory 108 may store the updated image data in memory 108.
[0046] Functional block 110 may include VMEM 112a, 112b, DMA hardware sequencers 114a, 114b, vector processing units (VPU) 116a, 116b, pixel processing engines (PPE) 118a, 118b, caches 120a, 120b, 120c, 120d, and decoupled lookup tables (DLUT) 122a, 122b. For clarity, unless otherwise specified, each device will be referred to as VMEM 112, DMA hardware sequencer 114, VPU 116, PPE 118, cache 120, and DLUT 122, and collectively referred to as VMEM 112, DMA hardware sequencer 114, VPU 116, PPE 118, cache 120, and DLUT 122. Although some interconnections are shown in the figure, it should be understood that the connections shown in the figure are for simplicity, and one or more devices of functional block 110 may be interconnected with one or more other devices of functional block 110, unless otherwise explicitly stated.
[0047] VMEM 112 may include a storage device interconnected with the corresponding DMA hardware sequencer 114, VPU 116, PPE 118, and cache 120 of processor 102 and function block 110. In some embodiments, VMEM 112 may receive and store sensor data obtained from memory 108. For example, VMEM 112 may receive and store sensor data obtained from memory 108 by DMA hardware sequencer 114. Alternatively, VMEM 112 may receive and store sensor data obtained from memory 108 via instruction switch 106. In some embodiments, VMEM 112 may be interconnected with PPE 118 via a decoupled load / store unit (DLSU) 124. As described herein, DLSU 124 may be configured to buffer data transferred between VMEM 112 and PPE 118 to reduce latency associated with communication between VMEM 112 and PPE 118.
[0048] The DMA hardware sequencer 114 may include one or more processors that control the execution of one or more instructions. For example, the DMA hardware sequencer 114 may receive instructions from processor 102, a corresponding VPU 116 or PPE 118, and / or storage devices (e.g., devices associated with the DMA hardware sequencer 114, such as internal or external memory, not explicitly shown), and the DMA hardware sequencer 114 may coordinate with the corresponding VPU 116 and / or PPE 118 to perform one or more operations during instruction execution. In an illustrative example, the DMA hardware sequencer 114 may receive instructions that cause it to retrieve data (e.g., sensor data and / or similar data) from memory 108 and store the data in a corresponding VMEM 112. In some embodiments, the DMA hardware sequencer 114 may perform one or more operations based at least in part on the data retrieved from memory 108. For example, the DMA hardware sequencer 114 can fill frames (e.g., image frames), manipulate addresses, manage overlapping data, manage different traversal orders, consider different frame sizes, and so on. In some embodiments, the DMA hardware sequencer 114 can receive signals (e.g., from VPU 116 or PPE 118) indicating that one or more operations have been performed on data stored in VMEM 112, update one or more descriptors at least in part based on updates to the data, and perform operations on the data again.
[0049] VPU 116 may include one or more processors that execute one or more instructions. For example, VPU 116 may receive instructions from processor 102, and the corresponding VPU 116 may coordinate with DMA hardware sequencer 114 and / or PPE 118 to perform one or more operations during instruction execution. In an illustrative example, VPU 116 may receive instructions from processor 102 that cause VPU 116 to trigger the corresponding DMA hardware sequencer 114 to retrieve sensor data from memory 108 and store the sensor data in the corresponding VMEM 112. In the example, VPU 116 may process the data stored in the corresponding VMEM 112 and write the data back to VMEM 112. In these examples, the data written by VPU 116 to the corresponding VMEM 112 may include updated sensor data and / or data generated at least in part based on the analysis performed by VPU 116 on the sensor data, including the location of objects or features within a frame, classification indicating the type of object or agent, etc. In some embodiments, VPU 116 may provide (e.g., send, transmit, transfer, etc.) signals to the corresponding DMA hardware sequencer 114 to cause the DMA hardware sequencer 114 to update one or more descriptors (described herein). For example, VPU 116 may send signals to the corresponding DMA hardware sequencer 114 to cause the DMA hardware sequencer 114 to update one or more descriptors at least in part based on data written by VPU 116 to the corresponding VMEM 112.
[0050] PPE 118 may include one or more processors that execute one or more instructions. For example, PPE 118 may receive instructions from processor 102, and each PPE 118 may coordinate with DMA hardware sequencer 114 and / or VPU 116 to perform one or more operations during instruction execution. In an illustrative example, PPE 118 may receive instructions from processor 102 that cause PPE 118 to trigger the corresponding DMA hardware sequencer 114 to retrieve (e.g., receive, acquire, capture, etc.) sensor data from memory 108 and store the sensor data in the corresponding VMEM 112. In the example, PPE 118 may process the data stored in the corresponding VMEM 112 and write the data back to the VMEM 112. In these examples, the data written by PPE 118 to the corresponding VMEM 112 may include updated sensor data and / or data generated at least in part based on analysis performed by PPE 118 on the sensor data, including intra-frame object or feature locations, classifications indicating the type of object or agent, etc. In some embodiments, PPE 118 may signal to the corresponding DMA hardware sequencer 114 to update one or more descriptors (described herein). For example, PPE 118 may signal to the corresponding DMA hardware sequencer 114 to update one or more descriptors at least in part based on data written by PPE 118 to the corresponding VMEM 112.
[0051] Cache 120 may include a storage device interconnected with VMEM 112 and / or instruction switch 106. As described above, cache 120 may receive instruction-associated data from instruction switch 106 and load instructions into one or more devices of function block 110 to cause one or more devices to operate according to the instructions. DLUT 122 may include a processor and / or memory configured to store one or more lookup tables. In some embodiments, DLUT 122 may be configured to implement communication between processor 102 and one or more components of function block 110. For example, DLUT 122 may be configured to communicate with processor 102 and / or Figure 1 The processor 102 communicates with one or more memory devices (e.g., memory 108 and / or memory 104). The DLUT 122 can then manage the communication between the processor 102 and... Figure 1The DLSU 124 can include storage and retrieval processes between one or more memory devices. It may include storage devices interconnected with the VMEM 112 and PPE 118 of a given functional block 110. For example, the DLSU 124 may receive and store sensor data obtained by the VMEM 112 from the memory 108. Alternatively, the DLSU 124 may receive and store data provided as output by the PPE 118.
[0052] Figure 2 An example loading memory architecture according to this disclosure is described. For example... Figure 2 As shown in the example, the load memory architecture 200 may include at least a system processor 210, a processor direct I / O channel 212, a processor memory I / O channel 214, a register I / O channel 222, local memory 230, a decoupled load memory unit (DLSU) 240, and a data memory 250.
[0053] System processor 210 can execute one or more instructions related to load-memory architecture 200. System processor 210 may include electronic processors, integrated circuits, etc., including one or more of digital logic, analog logic, digital sensors, analog sensors, communication buses, volatile memory, and non-volatile memory. System processor 210 may include, but is not limited to, at least one microcontroller core, microprocessor core, central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PSU), etc. System processor 210 or load-memory architecture 200 typically includes one or more communication bus controllers to enable communication between system processor 210 and other components of load-memory architecture 200. System processor 210 may include processor register 220.
[0054] Processor register 220 may store one or more instructions that can be executed by system processor 210 according to one or more processing elements of system processor 210. For example, each processor register 220 may store a word of a predetermined length that system processor 210 can execute. For example, the word may have a predetermined length corresponding to the bit length capacity of system processor 210 or its components. For example, the predetermined length may be 8 bits, 32 bits, or 64 bits, but is not limited thereto. For example, system processor 210 may execute a given word from a register of processor register 220 in a given cycle. For example, system processor 210 and processor register 220 may be integrated into a common die, chip, or package, but are not limited thereto.
[0055] Vector processing using SIMD is a common technique in programmable processors used to accelerate applications with data-level parallelism. SIMD instructions operate on a register file, requiring that the number of vector elements processed by the data path matches the number of vector elements provided to or written to the register file. The size of the register file should provide high bandwidth for the data path while having low read / write latency.
[0056] The width of the register file is adjusted to match the SIMD data path so that the compiler can efficiently schedule operations, thereby maximizing the utilization of SIMD data path resources. Due to low access latency requirements, vector register files and SIMD data paths are implemented to be physically close, so it is common practice to resize both together. On the other hand, data memory is built independently and may be shared across multiple processing engines.
[0057] Processor Direct I / O channel 212 provides communication between system processor 210 and local memory 230. For example, processor direct I / O channel 212 can communicatively couple system processor 210 and local memory 230. Processor direct I / O channel 212 can transfer one or more instructions, signals, conditions, states, etc. between system processor 210 and one or more local memories 230. Processor direct I / O channel 212 may include one or more digital, analog, or other communication channels, lines, traces, etc. For example, processor direct I / O channel 212 may include a first communication bus having at least one serial or parallel communication line among multiple communication lines with a communication interface integrated with system processor 210. Processor Memory I / O channel 214 provides communication between system processor 210 and DLSU 240. For example, processor memory I / O channel 214 can communicatively couple system processor 210 and DLSU 240. Processor memory I / O channel 214 can transfer one or more instructions, signals, conditions, statuses, etc., between one or more of the system processor 210 and DLSU 240. Processor memory I / O channel 214 may include one or more digital, analog, or similar communication channels, lines, traces, etc. For example, processor memory I / O channel 214 may include a second communication bus having at least one serial or parallel communication line among multiple communication lines with a communication interface integrated with the system processor 210.
[0058] Register I / O channel 222 can provide communication between processor register 220 and DLSU 240. For example, register I / O channel 222 can communicatively couple processor register 220 and DLSU 240. Register I / O channel 222 can transfer one or more instructions, signals, conditions, states, etc., between one or more of processor register 220 and DLSU 240. Register I / O channel 222 may include one or more digital, analog, or similar communication channels, lines, traces, etc. For example, register I / O channel 222 may include a third communication bus having at least one serial or parallel communication line among multiple communication lines with a communication interface integrated with processor register 220. Local memory 230 may store one or more instructions for operating system processor 210 components and operating components operatively coupled to system processor 210. For example, one or more instructions may include one or more of firmware, software, hardware, operating system, embedded operating system, etc. For example, local memory 230 may correspond to, but is not limited to, a solid-state memory device, a flip-flop array, a register array, or any combination thereof.
[0059] DLSU 240 provides data between system processor 210 and data memory 250 to mitigate or prevent system processor 210 from waiting for data transfer with data memory 250. System processor 210 can configure DLSU 240 to adapt to the latency of data memory 250. For example, DLSU 240 can be configured to prefetch a predetermined number of instructions from data memory 250 at a predetermined frequency and provide the prefetched instructions to system processor 210 at a rate corresponding to system processor 210 to prevent system processor 210 from stalling. Therefore, DLSU 240 can decouple fetch operations from data memory 250 from read / write operations of system processor 210 to provide technical improvements that enable faster processors to reliably read and write slower memory at the silicon level. The decoupled load-memory unit (DLSU) 240 may include load stream processor 242 and memory stream processor 244. The decoupled load-memory unit (DLSU) 240 may include one or more logic or electronic devices, including but not limited to integrated circuits, logic gates, flip-flops, gate arrays, programmable gate arrays, etc.
[0060] In one aspect, the DLSU 240 implements a load stream processor 242 and a storage stream processor 244, thereby allowing multiple load / store operations to be issued from the processor pipeline and processed by the DLSU 240 in the background. The size adjustment of the buffers (e.g., queues) of the load stream processor 242 and the storage stream processor 244 can be based on the number of loads / stores in flight to support continuous active operation of the system processor 210, where the DLSU uses the data memory 250 for fetching in the background. For example, if a load queue or storage queue is full and the processor issues another vector load or store operation, the processor may pause until the queue can accept data. For example, the DLSU may include one or more queue structures to ensure that loads and stores are processed in the order they are issued by the processor, to provide technical improvements to ensure memory ordering, but is not limited to this.
[0061] Load stream processor 242 can load one or more instructions from data memory 250 within at least one given cycle. For example, load stream processor 242 can load multiple instructions within a cycle of data memory 250, where the number of instructions corresponds to the number of instructions that system processor 210 can execute. Load stream processor 242 may include one or more logic or electronic devices, including but not limited to integrated circuits, logic gates, flip-flops, gate arrays, programmable gate arrays, etc. Load stream processor 242 may be fabricated on a silicon device (e.g., a die or wafer) of DLSU 240, or otherwise integrated with DLSU 240. Therefore, load stream processor 242 can provide a technological improvement in low-latency on-die prefetching of data from low-speed devices (e.g., data memory 250) to high-speed devices (e.g., system processor 210).
[0062] The storage stream processor 244 can store one or more instructions into the data memory 250 in at least one given cycle. For example, the storage stream processor 244 can store multiple instructions in one cycle of the data memory 250, where the number of instructions corresponds to the number of instructions that the system processor 210 can execute. The storage stream processor 244 may include one or more logic or electronic devices, including but not limited to integrated circuits, logic gates, flip-flops, gate arrays, programmable gate arrays, etc. The storage stream processor 244 may be fabricated on a silicon device (e.g., a die or wafer) of the DLSU 240, or otherwise integrated with the DLSU 240. Therefore, the storage stream processor 244 can provide a technological improvement in low-latency on-die pre-buffered data from high-speed devices (e.g., the system processor 210) to low-speed devices (e.g., the data memory 250).
[0063] Data memory 250 can store data associated with system processor 210. Data memory 250 may include one or more hardware memory devices for storing binary data, digital data, etc. Data memory 250 may include one or more electrical components, electronic components, programmable electronic components, reprogrammable electronic components, integrated circuits, semiconductor devices, flip-flops, arithmetic units, etc. Data memory 250 may include at least one of non-volatile memory devices, solid-state memory devices, flash memory devices, or NAND memory devices. Data memory 250 may include one or more addressable memory regions disposed on one or more physical memory arrays. The physical memory arrays may include NAND gate arrays disposed on at least one of, for example, a specific semiconductor device, an integrated circuit device, and a printed circuit board device. For example, data memory 250 may be fabricated on a silicon device (e.g., a die or wafer) of DLSU 240, or otherwise integrated with DLSU 240. The data memory can have a memory hardware architecture (e.g., NAND memory integrated into the die) with a larger memory capacity than local memory 230 to accommodate storing large amounts of data at batch processing speeds lower than the system processor 210. For example, input and output data structures that are much larger than the capacity of local memory 230 or processor register 220 can be stored in data memory 250. Therefore, the size of data memory 250 can be adjusted to allow for a larger capacity, but at the cost of bandwidth and access latency. Therefore, DLSU 240 can provide a technical solution for moving data between processor register 220 and data memory 250, thereby providing a technical improvement to maintain data throughput at a level sufficient to prevent system processor 210 from stalling.
[0064] In one aspect, DLSU 240 can provide data between a second processor and a stream processor with a second delay, while providing data between a first processor and a memory device with a first delay. For example, the first processor corresponds to DLSU 240, the memory device corresponds to data memory 250, and the second processor corresponds to system processor 210. For example, the first delay corresponds to the processing speed of data memory 250, and the second delay corresponds to the processing speed of system processor 210. For example, the processing speed of data memory 250 is lower than the processing speed of system processor 210.
[0065] In one aspect, the system is configured to provide multiple instructions for data within one cycle. For example, the number of instructions corresponds to the number of instructions that the system processor 210 can execute during a time interval between read or write operations to the data memory 250. For example, if the system processor 210 can execute 10 instructions within the time required to fetch data from the data memory 250, the DLSU can be configured to prefetch 10 instructions in a single read request to the data memory 250. In another aspect, the number of instructions is based on at least one of the size of the memory device or the length of the stream processor's buffer. For example, the size of the memory device may correspond to the arrangement of data in the data memory 250 according to one or more rows and columns of the data memory 250. The DLSU 240 can be configured to fetch and load data, or buffer and store data, with the data memory 250 according to the size of the data memory 250, to provide technological improvements in fast data loading and storage at the processor hardware level, thereby eliminating processor pauses due to memory latency.
[0066] System processor 210 can execute one or more instructions related to memory architecture 200. System processor 210 may include electronic processors, integrated circuits, etc., including one or more of digital logic, analog logic, digital sensors, analog sensors, communication buses, volatile memory, non-volatile memory, etc. System processor 210 may include, but is not limited to, at least one microcontroller unit (MCU), microprocessor unit (MPU), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), embedded controller (EC), etc. System processor 210 may include memory operable to store or store one or more instructions for components of operating system processor 210 and components operably coupled to system processor 210. For example, one or more instructions may include one or more of firmware, software, hardware, operating system, and embedded operating system. System processor 210 or memory architecture 200 typically includes one or more communication bus controllers to enable communication between system processor 210 and other elements of memory architecture 200.
[0067] Figure 3An example method for detecting two-dimensional maximum values according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can perform method 300. In one aspect, method 300 is directed to an example process of offloading Harris corner and NMS functions from the VPU for execution on a PPE. In one aspect, the PPE may include one or more processing elements (PEs). For example, each PE corresponds to a discrete semiconductor processor associated with a position in the X and Y axes. Generally, each PE may correspond to a 2D array of PPEs and may receive corresponding data in two dimensions to enable two-dimensional concurrent processing of the PPE. For example, each PPE may include its own set of corresponding PEs arranged in a 2D configuration or associated with a 2D configuration described herein. For example, as camera resolution increases, the size of video and image frame data also increases accordingly. In response, the computational workload of feature detection algorithms increases significantly, resulting in a significant increase in the execution time (e.g., geometric rate) of feature processing operations (e.g., Harris corner and NMS). As described herein at least with respect to method 300, one or more PPE devices can enhance the VPU. Here, by employing techniques that include 2D parallelism in image feature recognition operations, 2D processing via the SIMD architecture of PPE devices can provide significant improvements in reducing computation time in latency-sensitive deployments. For example, PPE is suitable for computationally intensive applications that operate on 2D data (such as Harris corners and NMS).
[0068] In one aspect, Harris corner detection and NMS can operate at the pixel level. For example, Harris corner detection operations can be performed on image pixels to identify the location and characteristics of visual attributes (e.g., corners) in an image. For example, NMS can be performed on Harris corner detection output corners to identify the location and characteristics of local maxima as final corners. Computationally, a CPU can process one pixel at a time, a VPU can process a 1D pixel vector at a time, and a PPE can process a 2D pixel block at a time. Method 300 addresses an example instruction set for transferring feature processing data between VPU and PPE devices to offload Harris corner detection and NMS operations from the VPU to the PPE. For example, a PE of a PPE can receive a 2D block according to a VPU transfer instruction set (VXfer). VXfer can also be referred to as the PPE instruction set. The VXfer instruction set controls data movement between adjacent PEs. For example, a PPE or VPU moves data between adjacent PEs in a northward direction via a northward VPU transfer instruction (VXferNorthW). For example, a PPE or VPU moves data between adjacent PEs in a southward direction via a southward VPU transfer instruction (VXferSouthW). Similarly, a PPE or VPU moves data between adjacent PEs in an eastward direction via an eastward VPU transfer instruction (VXferEastW). And vice versa. By using this feature on two data blocks, data can be reused to avoid redundant loading. For example, performing an NMS operation on a 5x5 pixel block via one or more VPUs may involve loading 5 rows and 5 columns of pixel data for each NMS operation, while performing the same NMS operation on the same 5x5 pixel block via one or more PEs may involve one data block and a total of 8 VXfer instructions. Therefore, the number of loading operations can be significantly reduced, and the execution speed of feature processing and recognition can be greatly improved. As described herein, at least a VPU, PPE, or a PE of a PPE can execute any of the instructions in method 300, as illustrated in the examples.
[0069] At 310, method 300 can load input. For example, the VXfer instruction set may include a load instruction (VLoad) that takes input data (IN) and an identifier or address of the target location at the PPE device as input, including, for example, an input buffer associated with one or more target PE devices. Therefore, method 300 can execute the instruction according to Equation 1:
[0070] VLoad(IN,Input_Buffer) (Equation 1)
[0071] At 320, method 300 can perform one or more two-dimensional shifts based on the loaded input to provide one or more blocks to one or more adjacent PEs. For example, the VXfer instruction set may include the VXferEastW instruction, which takes input data (IN), one or more shift-in pixel values (SP) of the block to be shifted in as input, and provides a 2D block as output, which includes shift inputs combined with shift-in pixel values (In_E1) in the eastward direction. For example, VXferEastW may include a shift-in pixel value (S0) corresponding to zero. The shift-in value S0 may indicate that data from adjacent blocks (e.g., from the eastward direction) is not provided or is unavailable. For example, the S0 value may correspond to an "empty" or "placeholder" value to implement the XOR operation described herein. For example, VXferEastW may include a shift-in pixel value corresponding to In_E1 to indicate that data from adjacent blocks in the eastward direction is provided. VXferEastW using In_E1 can be executed after VXferEastW using S0 as input to further shift data into the adjacent block in the eastward direction. For example, the VXfer instruction set may include the VXferWestW instruction, which takes IN and SP as input and provides a 2D block as output, the 2D block including shift input combined with the shifted-in pixel value (In_W1) in the westward direction. For example, VXferWestW may include S0 to indicate that data from the adjacent block in the westward direction is not provided or is unavailable. For example, VXferWestW may include the shifted-in pixel value corresponding to In_W1 to indicate that data from the adjacent block in the westward direction is provided. VXferWestW using In_E1 can be executed after VXferWestW using S0 as input to further shift data into the adjacent block in the westward direction. Therefore, method 300 can execute one or more instructions according to Equations 2-5:
[0072] VXferEastW(IN, S0)>>In_E1 (Formula 2)
[0073] VXferWestW(IN, S0)>>In_W1 (Formula 3)
[0074] VXferEastW(In_E1, S0)>>In_E2 (Equation 4)
[0075] VXferWestW(In_W1, S0)>>In_W2 (Equation 5)
[0076] At 330, method 300 can determine the block maximum value of one or more 2D blocks (e.g., by transferring one block to each PE). For example, the VXfer instruction set can include a local maximum instruction (VMaxW) that takes at least one of IN, In_E1, In_W1, In_E2, or In_W2 as input and provides an identifier of one or more pixels corresponding to the local maximum value (HMax) of the block as output. VMaxW can also take HMax from a previous VMaxW instruction operation as input to update the HMax value based on multiple blocks, thereby providing a 2D local maximum value based on block data processed rapidly via multiple PEs. Therefore, method 300 can execute one or more instructions according to Equations 6-9:
[0077] VMaxW(IN, In_E1)>>HMax (Equation 6)
[0078] VMaxW(HMax, In_W1)>>HMax (Equation 7)
[0079] VMaxW(HMax, In_E2)>>HMax (Equation 8)
[0080] VMaxW(HMax, In_W2)>>HMax (Equation 9)
[0081] At 340, method 300 can move the maximum value of the block based on HMax. For example, the VXfer instruction set can include the VXferNorthW instruction, which takes HMax and SP as input and provides the local maximum value (HMax_N1) of the block to be moved north as output. For example, VXferNorthW can include S0 to indicate that data from neighboring blocks to the north is not provided or unavailable. For example, VXferNorthW using HMax_N1 as input can be executed after VXferNorthW using S0 as input to further move data into neighboring blocks to the north. For example, the VXfer instruction set can include the VXferSouthW instruction, which takes HMax and SP as input and provides the local maximum value (HMax_S1) of the block to be moved south as output. For example, VXferSouthW can include S0 to indicate that data from neighboring blocks to the south is not provided or unavailable. For example, a VXferSouthW with HMax_S1 as input can be executed after a VXferSouthW with S0 as input to further move data into the adjacent block in the southward direction. Therefore, method 300 can execute one or more instructions according to Equation 10-13:
[0082] VXferNorthW(HMax, S0)>> HMax_N1 (Equation 10)
[0083] VXferSouthW(HMax, S0)>> HMax_S1 (Equation 11)
[0084] VXferNorthW(HMax_N1, S0)>> HMax_N2 (Equation 12)
[0085] VXferSouthW(Hmax_S1, S0)>> HMax_S2 (Equation 13)
[0086] At 350, method 300 can determine at least one common maximum value based on one or more local maximum values. For example, VMaxW can take at least one of HMax, HMax_N1, HMax_S1, HMax_N2, and HMax_S2 as input and provide an identifier of one or more pixels corresponding to the common maximum value (Max) of the block as output. VMaxW can also take Max from the previous VMaxW instruction operation as input to update the Max value based on multiple blocks, thereby providing a 2D common maximum value based on block data processed rapidly via multiple PEs. Therefore, method 300 can execute one or more instructions according to Equations 14-17:
[0087] VMaxW(HMax, HMax_N1)>> Max (Equation 14)
[0088] VMaxW(Max, HMax_S1)>> Max (Equation 15)
[0089] VMaxW(Max, HMax_N2)>> Max (Equation 16)
[0090] VMaxW(Max, HMax_S2)>> Max (Equation 17)
[0091] At 360, method 300 can determine whether the common maximum value Max is the maximum value of the current input IN. For example, the VXfer instruction set can include a comparator instruction (VCmpEQW) that takes at least one of Max and IN as input and provides a binary indication as output whether Max is the common maximum value (Is_Max) of IN. For example, if Max reaches or exceeds a threshold indicating the current or expected common maximum value of IN, VCmpEQW can provide Is_Max with a true value as output. For example, if Max does not reach or exceed the threshold indicating the current or expected common maximum value of IN, VCmpEQW can provide Is_Max with a false value as output. For example, the VXfer instruction set can include a multiplexed output instruction (VMuxW) that takes at least one of Is_Max, IN, and S0 as input and provides an output (OUT) whose value depends on Max according to the binary indication of whether Is_Max is the common maximum value of IN. At 362, method 300 can provide an empty output. For example, VMuxW can pass S0 to the output (OUT) based on the multiplexing state given an input with a value of False for Is_Max. At 364, method 300 can provide an output equal to the received input. For example, VMuxW can pass IN to OUT based on the multiplexing state given an input with a value of True for Is_Max. At 370, method 300 can store the output. For example, the VXfer instruction set can include a store instruction (VStore) that takes as input an identifier or address of the target location at the OUT and PE devices, such as including an output buffer (Output_Buffer) associated with one or more target PE devices. Therefore, method 300 can execute one or more instructions according to Equations 15-17:
[0092] VCmpEQW(Max, IN)>> Is_Max (Equation 15)
[0093] VMuxW(Is_Max, IN, S0)>> OUT (Equation 16)
[0094] VStore(OUT,Output_Buffer) (Equation 17)
[0095] Figure 4AAn example two-dimensional block state according to this disclosure is depicted. In one aspect, this technical solution can provide technical improvements by adjusting one or more configurations of the tile size of the image data block and reusing data at the block and tile levels to reduce the use of computational resources (e.g., energy, processor resources, heat dissipation capabilities). For example, the input data of a feature detection algorithm can be divided into multiple tiles, each tile including halos from the entire image. This technical solution can reduce computational waste by performing various block operations that combine the reuse of edge data. For example, when Harris corners and NMS have a 5x5 dependency, four additional rows and columns are required in each of the four directions. For example, when the NMS output tile size is defined as [width] x [height], the Harris corner output tile size can be defined as [width + 4] x [height + 4], and the Harris corner input size can be correspondingly [width + 8] x [height + 8]. Therefore, this technical solution can minimize or eliminate the halo loss of 8 pixels in the vertical and horizontal directions. For example, the Harris corner output can be provided as input to the NMS operation. Therefore, this technical solution, at least according to Figure 4, can provide technical improvements to reduce block-level computational waste. In the 2D array architecture corresponding to the PPE, a sub-block can correspond to a set of pixels mapped to the corresponding PE of the PPE. For example, for an 8-column x 10-row PPE array, a sub-block can be 8 x 10 words (e.g., 32-bit pixels) or 16 x 10 half-words (e.g., 16-bit pixels). For example, processing a sub-block in the vertical direction can produce two rows of output, where four rows are lost each during the Harris corner and NMS operations. To minimize pixel loss in the vertical direction, two sub-blocks in the vertical direction can be mapped to the corresponding PE of the PPE and processed simultaneously, performing parallel operations in the horizontal direction. Figure 4A As shown in the example, the two-dimensional block state 400A can include at least sub-blocks 410A, 420A, 430A, and 440. Figure 4A This can correspond to the reuse of one or more pixels in the vertical direction.
[0096] Sub-block 410A may correspond to a 2D block comprising at least a portion of image data corresponding to a frame. Sub-block 420A may at least partially correspond to sub-block 410A in one or more aspects of structure and operation, and may be located at a position adjacent to sub-block 410A in the vertical direction (e.g., having a defined logical position). Sub-block 430A may at least partially correspond to sub-block 410A in one or more aspects of structure and operation, and may be located at a position adjacent to sub-block 410A in the horizontal direction (e.g., having a defined logical position). Sub-block 440 may at least partially correspond to sub-block 410A in one or more aspects of structure and operation, may be located at a position adjacent to sub-block 430A in the vertical direction (e.g., having a defined logical position), and may be located at a position adjacent to sub-block 420A in the horizontal direction (e.g., having a defined logical position).
[0097] Figure 4B An example first transformation state of a two-dimensional block according to this disclosure is depicted. For example... Figure 4B As shown in the example, the first transformation state of the two-dimensional block 400B may include at least sub-block 410B, sub-block 420B and sub-block 430A. Figure 4B This can correspond to the reuse of one or more pixels in the vertical direction, and can correspond to... Figure 4A The state following state 400A is illustrated in the example. Sub-block 410B may correspond to a 2D block that includes at least a portion of the image data corresponding to the frame following the feature processing operation. For example, sub-block 410B may correspond to data transformed according to Harris corner operation or NMS operation.
[0098] Sub-block 420B may correspond to a 2D block comprising at least a portion of image data corresponding to a frame after feature processing operations on sub-block 410B. Sub-block 420B may include a first frame output 422 and a second frame data 424. The first frame output 422 may correspond to a portion of data transformed according to the data of sub-block 410B. For example, the first frame output 422 may correspond to a halo or edge of a Harris corner operation performed on the data of sub-block 410B. For example, a halo of a Harris corner operation performed on the data of sub-block 410B may be detected as the lower edge of a pixel frame of sub-block 410B. For example, based on the determination that sub-block 420B is located below and adjacent to sub-block 410B in the vertical direction, the first frame output 422 may be applied to the top of sub-block 420B. The second frame data 424 may correspond to a portion of image data of an image adjacent to a portion of the data provided to sub-block 410A. Therefore, subblock 420B can reuse the data generated by the Harris corner operation for subblock 410B, and can reuse this data as input to subblock 420B, thereby reducing computational waste by avoiding the repeated generation of the first frame output 422 for operations on multiple subblocks. For example, the first frame output 422 and the second frame data 424 can be provided as input to an NMS operation, where the first frame output 422 is generated based on the Harris corner operation.
[0099] Figure 4C An example second transformation state of a two-dimensional block according to this disclosure is depicted. For example... Figure 4C As shown in the example, the second transformation state of the two-dimensional block 400C may include at least sub-block 410C and sub-block 430C. Figure 4C It can correspond to the reuse of one or more pixels in the horizontal direction, and can correspond to, for example... Figure 4A The state following state 400A is illustrated in the example. Sub-block 410C may correspond to a 2D block including at least a portion of the image data corresponding to the frame after the feature processing operation. For example, sub-block 410C may correspond to data transformed according to the Harris corner operation or the NMS operation.
[0100] Sub-block 430C may correspond to a 2D block including at least a portion of image data corresponding to a frame after feature processing operations on sub-block 410C. Sub-block 430C may include a first frame output 432 and a second frame data 434. The first frame output 432 may correspond to a portion of data transformed according to the data of sub-block 410C. For example, the first frame output 432 may correspond to a halo or edge of a Harris corner operation performed on the data of sub-block 410C. For example, a halo of a Harris corner operation performed on the data of sub-block 410C may be detected as the right edge of a pixel frame of sub-block 410C. For example, based on the determination that sub-block 430C is located to the right of sub-block 410C in the horizontal direction and is adjacent to sub-block 410C, the first frame output 432 may be applied to the left portion of sub-block 430C. The second frame data 434 may correspond to a portion of image data of an image adjacent to a portion of the data provided to sub-block 410A. Therefore, sub-block 430C can reuse the data generated by the Harris corner operation for sub-block 410C, and can reuse this data as input to sub-block 430C, thereby reducing computational waste by avoiding repeatedly generating the first frame output 432 for operations on multiple sub-blocks. For example, the first frame output 432 and the second frame data 434 can be provided as input to the NMS operation, where the first frame output 432 is generated based on the Harris corner operation.
[0101] Figure 4D An example state of a horizontal image block according to this disclosure is depicted. For example... Figure 4D As illustrated by the example, the state of the horizontal image block 400D may include a frame portion 450 and an extended image pixel region 460. The frame portion 450 may correspond to the edge of an image feature. For example, the frame portion 450 may correspond to the halo or more smaller image pixel regions of the extended image pixel region 460. The frame portion 450 may include a right frame portion 452 and a left frame portion 454.
[0102] For example, the input data for a feature processing operation can be segmented into tiles from the entire image. Since the computation of a feature processing operation (e.g., Harris corner or NMS) includes not only image pixels within the tile but also halo pixels surrounding the tile, using larger tiles can reduce wasted computation by decreasing the computation of some halo pixels in adjacent tiles. For example, the system can adjust one or more parameters of the input tile size (e.g., WxH tile) based on the capacity of the VMEM to reduce wasted computation. An extended image pixel region 460 may correspond to the block processing architecture according to this disclosure to provide technical improvements by reducing wasted computation by reducing the number of redundant operations. The extended image pixel region 460 may include a first image pixel region 462, a second image pixel region 464, and a recycled frame portion 466.
[0103] The first image pixel region 462 may correspond to a block smaller than the extended image pixel region 460 in at least one dimension. For example, the first image pixel region 462 may have a square dimension. The second image pixel region 464 may correspond to a block smaller than the extended image pixel region 460 in at least one dimension. For example, the second image pixel region 464 may have a square dimension matching the square dimension of the image pixel region 462. The recovered frame portion 466 may correspond to a portion of the extended image pixel region 460, which can be recovered from halo processing to reduce redundant computation. For example, the recovered frame portion 466 may be processed by one or more PPE devices described herein according to a rectangular dimension larger than the square dimension of the first image pixel region 462 or the second image pixel region 464. The right frame portion 452 may correspond to a portion of the halo of the first image pixel region 462. The left frame portion 454 may correspond to a portion of the halo of the second image pixel region 464. For example, the left frame portion 454 may be horizontally adjacent to the right frame portion 452 (e.g., adjacent to the right edge of the right frame portion 452). Therefore, according to Figure 4D The operation of this technical solution can provide technical improvements, at least reducing redundant calculations of halo pixels in dark areas with large tile sizes.
[0104] Figure 4E Example states of several vertical blocks according to this disclosure are depicted. For example... Figure 4E As illustrated by example, the state of multiple vertical blocks 400E may include one or more vertically adjacent image pixel blocks 470, 472, and 474, and one or more NMS data blocks 480, 482, and 484. Image pixel blocks 470, 472, and 474 may correspond to blocks or sub-blocks as described herein. For example, image pixel blocks 470, 472, and 474 may each correspond to portions of an image to which corresponding Harris corner operations can be performed. Image pixel blocks 470, 472, and 474 may be incorporated into image pixel blocks 470, 472, and 474, as... Figure 4E The example illustrates this. For instance, one or more frame blocks 476 and 478 can be configured according to... Figure 4BVertical operations are incorporated into image pixel blocks 470, 472, and 474. NMS data blocks 480, 482, and 484 may correspond to the blocks or sub-blocks described herein. For example, NMS data blocks 480, 482, and 484 may each correspond to portions of the image to which corresponding NMS operations can be performed, including reused pixel data from Harris corner operations performed on image pixel blocks 470, 472, and 474, as well as frame blocks 476 and 478. Therefore, this technique can reduce redundant computation of halo pixels in dark areas with large tile sizes. For example, the bottom portion of the Harris corner output in one tile row can be reused in the top portion of the NMS input in the next tile row. Furthermore, the system can store the vertical context of one or more high Harris corner outputs in a tile row into an L2 cache and restore the vertical context as the NMS input in the next tile row, providing a technical improvement in reducing redundant computation.
[0105] Figure 5 An example processor parallelization architecture for a single frame, according to this disclosure, is described. For example... Figure 5 As illustrated by example, a processor parallelization architecture 500 for a single frame may include at least a PPE 502, a VPU 504, a first time point 512, a second time point 514, a third time point 522, and a fourth time point 542. The PPE 502 may correspond at least partially to the PPE 118A or 118B described herein in one or more aspects of its architecture and operation. The VPU 504 may correspond at least partially to the VPU 116A or 116B in one or more aspects of its architecture and operation. The PPE 502 may perform Harris corner operation 510 and non-maximum suppression (NMS) operation 540.
[0106] At a first time point 512, PPE 502 may begin executing Harris corner operation 510. Harris corner operation 510 may be executed on first frame data corresponding to the first frame in a given image frame sequence, wherein the given image frame sequence corresponds to video data comprising multiple consecutive image frames. At a second time point 514, PPE 502 may complete the execution of Harris corner operation 510. PPE 502 may convey an indication of completion of Harris corner operation 510 and may provide one or more instructions or data to VPU 504 in response to the completion of Harris corner operation 510. In response, VPU 504 may begin executing injection operation 520 at least in part based on the data generated by PPE 502 using Harris corner operation 510.
[0107] At the third time point 522, VPU 504 can complete the execution of injection operation 520 and can provide one or more instructions or data to PPE 502 in response to the completion of injection operation 520. In response, PPE 502 can begin executing NMS operation 540 at least partially based on the data generated by VPU 504 using injection operation 520. VPU 504 can also begin executing position calculation operation 530 at least partially based on the data generated by VPU 504 using injection operation 520. At the fourth time point 542, PPE 502 can complete the execution of NMS operation 540, and VPU 504 can complete the execution of position calculation operation 530. Therefore, PPE 502 and VPU 504 can parallelize the NMS operation 540 and the position calculation operation 530 of the first frame of data to provide technical improvements, increase the speed of image feature processing, and maintain the accuracy of image feature processing.
[0108] Figure 6 An example processor parallelization architecture for multiple frames according to this disclosure is described. Figure 6 As illustrated in the example, the processor parallelization architecture 600 for multiple frames may include at least the following: masking and packing operation 602, halo preparation operation 604, Harris corner point operation for subsequent frames 610, fifth time point 612, sixth time point 614, injection operation for subsequent frames 620, seventh time point 622, position calculation operation for subsequent frames 630, and non-maximum suppression operation for subsequent frames. (NMS) Operation 640 and the eighth time point 642.
[0109] At the fourth time point, VPU 504 can determine that the position calculation operation 530 of the first frame data has been completed, and can begin executing the masking and packing operation 602 of the first frame data. VPU 504 may provide one or more instructions or data to PPE 502 in response to the completion of the position calculation operation 530. In response, PPE 502 may begin performing Harris corner point operation 610 on the second frame data corresponding to the second frame in a given image frame sequence, where the second frame data corresponds to an image data frame in the video data following the first frame data.
[0110] At the fifth time point 612, VPU 504 can complete the masking and packing operation 602 on the first frame of data, and can begin the halo preparation operation 604 at least in part based on the data generated by VPU 504 using the masking and packing operation 602. In parallel, PPE 502 can continue to perform Harris corner point operation 610 on the second frame of data to provide technical improvements to accelerate feature recognition of video data across frames.
[0111] At the sixth time point 614, the VPU can complete the execution of the halo preparation operation 604, and the PPE 502 can complete the execution of the Harris corner point operation 610 on the second frame data. Therefore, the PPE 502 and VPU 504 can provide a technical solution for parallelizing specific image feature recognition operations to accelerate feature recognition of cross-frame video data. The PPE 502 can transmit an indication of the completion of the Harris corner point operation 610 and can provide one or more instructions or data to the VPU 504 in response to the completion of the Harris corner point operation 610. In response, the VPU 504 can begin executing the injection operation 620, at least in part, based on the data generated by the PPE 502 using the Harris corner point operation 610.
[0112] At the seventh time point 622, VPU 504 can complete the execution of injection operation 620 and can provide one or more instructions or data to PPE 502 in response to the completion of injection operation 620. In response, PPE 502 can begin executing NMS operation 640 at least partially based on the data generated by VPU 504 using injection operation 620. VPU 504 can also begin executing position calculation operation 630 at least partially based on the data generated by VPU 504 using injection operation 620. At the eighth time point 642, PPE 502 can complete the execution of NMS operation 640, and VPU 504 can complete the execution of position calculation operation 630. Therefore, PPE 502 and VPU 504 can parallelize NMS operation 640 for executing the second frame data and position calculation operation 630 for executing the second frame data to provide technical improvements, increase the speed of image feature processing, and maintain the accuracy of image feature processing.
[0113] Figure 7A An example method for multidimensional data processing architecture and instruction set according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can execute method 700a.
[0114] At 710, method 700a can modify multiple instances of frame data based on at least one of multiple dimensions. For example, the multiple instances can include multiple different data objects, where each instance is a different sub-block. For example, the multiple dimensions can include a horizontal (e.g., width) dimension and a vertical (e.g., height) direction. Method 700a can modify a first instance based on a second instance of the multiple instances of frame data, according to at least one of the multiple dimensions. For example, the system can determine a first instance of the multiple instances of frame data, which is structured in multiple dimensions. For example, the first instance can correspond to a first frame of frame data at a first time, relating to multiple frames of video data. The system can modify the first instance based at least in part on a second instance of the multiple instances of frame data, according to at least one of the multiple dimensions. At 712, method 700a can modify multiple instances. In one aspect, at least one dimension corresponds to at least one of a shift in a horizontal dimension of one or more bits or a shift in a vertical direction of one or more bits.
[0115] At 720, method 700a may provide multiple instances of frame data to corresponding processors among multiple processors. In one aspect, the method may include determining a first instance among the multiple instances of frame data, the first instance of frame data being structured in multiple dimensions. In one aspect, the method may include providing the first instance to a first processor among multiple processors, the first processor being configured to execute input arranged in multiple dimensions. In one aspect, the system may provide the first instance to the first processor among multiple processors, the first processor being configured to execute input arranged in multiple dimensions. In one aspect, the system may provide a second instance to a second processor among multiple processors, the second processor being configured to execute input arranged in multiple dimensions. At 722, method 700a may provide multiple instances to multiple processors, each processor being configured to execute input arranged in multiple dimensions.
[0116] At 730, method 700a can determine multiple instances of frame data, each instance being modified individually according to at least one of multiple dimensions. The multiple instances of frame data are respectively provided to a corresponding processor among multiple processors, each processor being configured to execute the input arranged in the multiple dimensions.
[0117] At 740, method 700a can generate multiple features based at least partially on multiple instances of frame data, each feature corresponding to a frame data instance from the multiple instances of frame data. In one aspect, method 700a can include determining a first feature among the multiple features, the first feature being structured in at least one of multiple dimensions. The method can include modifying the first feature based at least partially on the first instance of the multiple instances of frame data, according to at least one of the multiple dimensions. In one aspect, the system can determine a first feature among multiple features, the first feature being structured in at least one of multiple dimensions. The system can modify the first feature based at least partially on the first instance of the multiple instances of frame data, according to at least one of the multiple dimensions. In one aspect, the method can include providing the first feature to a first processor among multiple processors. The method can include providing a second feature to a second processor among multiple processors. In one aspect, the system can provide the first feature to the first processor among multiple processors. In one aspect, the method can include providing a second instance to the second processor among multiple processors, the second processor being configured to execute input arranged in multiple dimensions. The system can provide the second feature to the second processor among multiple processors. For example, features may correspond to, but are not limited to, attributes of an image identified according to the feature processing discussed herein (e.g., Harris corners, NMS). At 742, method 700a may generate multiple features based at least in part on multiple instances of frame data.
[0118] Figure 7B An example method for architecture and instruction set for multidimensional data processing according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can execute method 700b.
[0119] At 750, method 700b can generate a metric for the first frame of data in multiple dimensions. At 752, method 700b can generate a metric indicating visual attributes. For example, visual attributes can correspond to aspects of image data corresponding to physical features. For example, visual attributes can correspond to boundaries between shapes, objects, regions of different colors, regions of different brightness, or any combination thereof, but are not limited thereto. At 754, method 700b can generate a metric based on multiple features. At 756, method 700b can generate a metric. In one aspect, method 700b can include providing a metric as output in response to a metric satisfying a condition indicating the presence of features in the frame data. In one aspect, the system can provide a metric as output in response to a metric satisfying a condition indicating the presence of features in the frame data. For example, the metric can correspond to IN as described herein, but is not limited thereto. In one aspect, the method can include providing an output different from the metric in response to a metric not satisfying a condition indicating the presence of features in the frame data. In one aspect, the system can provide an output different from the metric in response to a metric not satisfying a condition indicating the presence of features in the frame data. For example, an output different from the metric can correspond to S0 or an empty output as described herein, but is not limited to this. In one aspect, the condition corresponds to a local maximum associated with a visual attribute, as described herein, and the feature indicates the visual attribute.
[0120] In one aspect, the system may include or correspond to a System-on-a-Chip (SoC). The SoC may include at least one GPU and multiple processors. The SoC (e.g., multiple processors) may determine a first instance of a plurality of instances of frame data, the first instance of frame data being structured in multiple dimensions. The SoC may modify the first instance based on a second instance of the plurality of instances of frame data, according to at least one of the multiple dimensions. The SoC may provide the first instance to a first processor among one or more processors, the first processor being configured to perform input arranged in multiple dimensions. The SoC may provide the second instance to a second processor among one or more processors, the second processor being configured to perform input arranged in multiple dimensions. In one aspect, the at least one dimension corresponds to at least one of a shift in a horizontal dimension of one or more bits or a shift in a vertical direction of one or more bits. The SoC may determine a first feature among a plurality of features, the first feature being structured in at least one of the multiple dimensions. The SoC may modify the first feature based on the first instance of the plurality of instances of frame data, according to at least one of the multiple dimensions. The SoC may provide the first feature to a first processor among one or more processors. The SoC may provide the second feature to a second processor among one or more processors. In one aspect, at least one dimension corresponds to at least one of a horizontal shift of one or more bits or a vertical shift of one or more bits. The SoC may include providing a metric as output in response to a metric satisfying a condition indicating the presence of a feature in the indicated frame data. The SoC may provide an output different from the metric in response to a metric not satisfying a condition indicating the presence of a feature in the indicated frame data.
[0121] Figure 7C An example method for architecture and instruction set for multidimensional data processing according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can execute method 700c. Method 700c can be executed by multiple processors.
[0122] In one aspect, method 700c may include configuring multiple processors to receive tile data in a format having a first tile dimension among multiple dimensions, the first tile dimension being greater than a second tile dimension among multiple dimensions, the data including first frame data and second frame data.
[0123] In one aspect, the system may be configured with multiple processors to receive tile data in a format having a first block dimension among multiple dimensions, the first block dimension being greater than a second block dimension among multiple dimensions, the data including first frame data and second frame data.
[0124] In one aspect, the method may include determining a first block dimension as a portion of a first tile dimension. The method may include determining a second block dimension as a portion of a second tile dimension. In one aspect, the system may determine a first block dimension as a portion of a first tile dimension. The system may determine a second block dimension as a portion of a second tile dimension. In one aspect, the method may include a first tile dimension in an direction corresponding to the first block dimension, and a second tile dimension in an direction corresponding to the second block dimension. In one aspect, the system may include a first tile dimension in an direction corresponding to the first block dimension, and a second tile dimension in an direction corresponding to the second block dimension.
[0125] At 760, method 700c can generate a first frame output. At 762, method 700c can generate a first frame output based on input including first frame data arranged in multiple dimensions. At 764, method 700c can generate a first frame output configured to perform input in multiple dimensions. At 766, method 700c can generate a first frame output via a first feature recognition operation in multiple dimensions. At 770, method 700c can store the first frame output in a memory device. At 772, method 700c can store the first frame output along the edge of the first frame data in one of the multiple dimensions. At 774, method 700c can store the first frame output.
[0126] Figure 7D An example method for multidimensional data processing architecture and instruction set according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can execute method 700d. Method 700d can be executed by multiple processors. In one aspect, the method may include determining the tile size of tile data, wherein the tile size corresponds to a first frame of data and a second frame of data, and the tile size has a first tile dimension greater than the first tile dimension and a second tile dimension greater than the second tile dimension. In one aspect, the system can determine the tile size of tile data, wherein the tile size corresponds to a first frame of data and a second frame of data, and the tile size has a first tile dimension greater than the first tile dimension and a second tile dimension greater than the second tile dimension.
[0127] At 780, method 700d can generate a second output of at least one second feature recognition operation. At 782, method 700d can generate a second output based on input including second frame data. At 784, method 700d can generate a second output of the second feature recognition operation in multiple dimensions. At 786, method 700d can generate a second output.
[0128] At 790, method 700d can generate second frame data indicating features of the image. In one aspect, the method may include: combining a portion of a second output and a first output along a second dimension different from the first dimension among a plurality of dimensions to form second frame data. In one aspect, the system may combine a portion of a second output and a first output along a second dimension different from the first dimension among a plurality of dimensions to form second frame data. In one aspect, the method may include: partitioning tile data into first frame data according to a first block dimension and a second block dimension. The method may include: partitioning tile data into second frame data according to a first block dimension and a second block dimension. In one aspect, the system may partition tile data into first frame data according to a first block dimension and a second block dimension. The system may partition tile data into second frame data according to a first block dimension and a second block dimension. At 792, method 700d can generate second frame data based on a second output. At 794, method 700d can generate second frame data based on a portion of the first output. At 796, method 700d can generate second frame data.
[0129] In one aspect, the first frame data corresponds to a first portion of image data, and the second frame data corresponds to a second portion of image data, which may include the edges of the first frame data. In one aspect, the first feature recognition operation performs Harris corner operations in multiple dimensions. In one aspect, the second feature recognition operation performs non-maximum suppression operations in multiple dimensions.
[0130] In one aspect, the SoC may include at least one GPU and multiple processors. The SoC may include configuring the multiple processors to receive tile data in a format having a first block dimension among multiple dimensions, the first block dimension being larger than a second block dimension among the multiple dimensions, the data including first frame data and second frame data. The SoC may include determining tile sizes of the tile data, wherein the tile sizes correspond to the first frame data and the second frame data, and the tile sizes have a first tile dimension larger than the first block dimension and a second tile dimension larger than the second block dimension. The SoC may include determining the first block dimension as a portion of the first tile dimension. The SoC may include determining the second block dimension as a portion of the second tile dimension. In one aspect, the first tile dimension is in a direction corresponding to the first block dimension, and the second tile dimension is in a direction corresponding to the second dimension. The SoC may include: combining portions of a second output and a first output along a second dimension different from the first dimension among the multiple dimensions to form second frame data. The SoC may include: partitioning the tile data into first frame data according to the first block dimension and the second dimension. The SoC may include: partitioning the tile data into second frame data according to the first block dimension and the second dimension. In one aspect, the first frame data corresponds to a first portion of image data, and the second frame data corresponds to a second portion of image data, which may include the edges of the first frame data. In one aspect, the first feature recognition operation performs Harris corner operations in multiple dimensions. In one aspect, the second feature recognition operation performs non-maximum suppression operations in multiple dimensions.
[0131] Figure 8 An example method for a parallelized architecture for multidimensional data processing according to this disclosure is described. At least environment 100, memory architecture 200, or any component thereof can execute method 800. Method 800 can be executed by multiple processors. At 810, method 800 can determine that a first frame of data corresponds to a first time. At 820, method 800 can determine that a second frame of data corresponds to a second time after the first time. For example, the first time may correspond to a first timestamp, and the second time may correspond to a second timestamp. For example, based on the determination that the value or absolute value of the second timestamp (e.g., a UNIX timestamp) is greater than the value or absolute value of the first timestamp (e.g., a different UNIX timestamp), the second time may be after the first time, where the UNIX timestamp is measured as the number of seconds incremented from a given epoch.
[0132] At 830, method 800 may provide first frame data to a first processor among a plurality of processors configured to execute input arranged in two dimensions. In one aspect, the method may execute one or more first instructions based on the first frame data. In one aspect, the system may execute one or more first instructions based on the first frame data. In one aspect, the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a halo preparation operation. In one aspect, the method may include: after providing the first frame data to the first processor, providing a third frame data to the first processor, the third frame data corresponding to the output of the first processor based on the first frame data. The method may include: after providing the first frame data to the first processor, having the first processor execute one or more instructions based on the first frame data. In one aspect, the system may provide a third frame data to the first processor after providing the first frame data to the first processor, the third frame data corresponding to the output of the first processor based on the first frame data. After providing the first frame data to the first processor, the system may have the first processor execute one or more instructions based on the first frame data. In one aspect, the one or more instructions correspond to a non-maximum suppression operation.
[0133] At 840, method 800 may provide the second frame data to a second processor among a plurality of processors configured to execute input arranged in one dimension. In one aspect, after executing one or more first instructions, the method may execute one or more second instructions based on the first frame data. In another aspect, after executing one or more first instructions, the system may execute one or more second instructions based on the first frame data. At 842, method 800 may provide the second frame data to the second processor in parallel with providing the first frame data to the first processor. In one aspect, the one or more second instructions correspond to a halo preparation operation.
[0134] At 850, method 800 can execute the first processor in parallel with the execution of the second processor. At 852, method 800 can execute the first processor based on the first frame data. At 854, method 800 can execute the second processor based on the second frame data. In one aspect, the method may include providing a fourth frame of data to the second processor, the fourth frame of data corresponding to the output of the second processor based on the second frame data, in parallel with the execution of one or more instructions based on the first frame data. In one aspect, the system can provide a fourth frame of data to the second processor, the fourth frame of data corresponding to the output of the second processor based on the second frame data, in parallel with the execution of one or more instructions based on the first frame data. For example, the fourth frame of data may be different from the first frame data, the second frame data, and the third frame data, and may correspond to a frame or a portion of video data that is different from the video data corresponding to the first frame data, the second frame data, or the third frame data.
[0135] In one aspect, the method may include: providing the first frame of data to a second processor after providing the first frame of data to a first processor. The method may also include: after providing the first frame of data to the first processor, having the second processor execute one or more instructions based on the first frame of data. In one aspect, the system may provide the first frame of data to a second processor after providing the first frame of data to the first processor. The system may also have the second processor execute one or more instructions based on the first frame of data after providing the first frame of data to the first processor. In one aspect, the one or more instructions correspond to an injection operation.
[0136] In one aspect, the system may include or correspond to a SoC. In one aspect, the SoC may include at least one GPU and multiple processors. The SoC may include executing one or more first instructions based on first frame data. In one aspect, the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a halo preparation operation. The SoC may, after executing one or more first instructions, have a first processor among the multiple processors execute one or more second instructions based on the first frame data. In one aspect, the one or more second instructions correspond to a halo preparation operation. The SoC may, after providing the first frame data to a first processor, provide the first frame data to a second processor among the multiple processors. The SoC may, after providing the first frame data to a first processor, have a second processor execute one or more instructions based on the first frame data. In one aspect, the one or more instructions correspond to an injection operation. The SoC may, after providing the first frame data to a first processor, provide a third frame data to the first processor, the third frame data corresponding to the output of the first processor based on the first frame data. The SoC may, after providing the first frame data to a first processor, have the first processor execute one or more instructions based on the first frame data. In one aspect, the one or more instructions correspond to a non-maximum suppression operation. The SoC can provide a fourth frame of data to a second processor in parallel with executing one or more instructions based on the first frame of data. The fourth frame of data corresponds to the output of the second processor based on the second frame of data.
[0137] In one aspect, the system may include multiple processors in a control system for autonomous or semi-autonomous machines. The system may include a perception system for autonomous or semi-autonomous machines. The system may include a system implemented using robotics. The system may include an aviation system. The system may include a medical system. The system may include a marine system. The system may include an intelligent area monitoring system. The system may include a system for performing deep learning operations. The system may include a system for performing simulation operations. The system may include a system for generating or presenting virtual reality (VR), augmented reality (AR), or mixed reality (MR) content. The system may include a system for performing digital twin operations. The system may include a system implemented using edge devices. The system may include a system combining one or more virtual machines (VMs). The system may include a system for generating synthetic data. The system may be implemented at least partially in a data center. The system may be a system for performing conversational artificial intelligence (AI) operations. The system may include a system for performing generative AI operations. The system may include a system for implementing language models. The system may include a system for implementing a visual language model (VLM). The system may include a system for implementing a large language model (LLM). The system may include a system for implementing a multimodal language model. The system may include a system for hosting one or more real-time streaming applications. The system may include a system for performing optical transport simulations. The system may include a system for performing collaborative content creation of 3D assets. In one aspect, the system may be implemented using cloud computing resources at least in part.
[0138] Example autonomous vehicles
[0139] Figure 9AThis is an illustration of an example autonomous vehicle 900 according to some embodiments of the present disclosure. The autonomous vehicle 900 (also referred to herein as “vehicle 900”) may include, but is not limited to, passenger vehicles such as automobiles, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, underwater vehicles, robotic vehicles, drones, aircraft, trailer-attached vehicles (e.g., semi-trailer trucks for towing cargo) and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are generally described according to the level of automation defined by the National Highway Traffic Safety Administration (NHTSA) of the U.S. Department of Transportation and the Society of Automotive Engineers (SAE) in their standard “Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles” (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 900 may be capable of having one or more functions according to Level 3-5 of the autonomous driving level. Vehicle 900 may be capable of functioning according to one or more Level 1-5 of the autonomous driving level. For example, vehicle 900 may be able to provide driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment. As used herein, the term "autonomy" may include any and / or all types of autonomy for vehicle 900 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, providing assisted autonomy, semi-autonomy, primary autonomy, or other names.
[0140] Vehicle 900 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 900 may include a propulsion system 950, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 950 may be connected to the drivetrain of vehicle 900, which may include a transmission, to enable propulsion of vehicle 900. Propulsion system 950 may be controlled in response to receiving a signal from throttle / accelerator 952.
[0141] A steering system 954, which may include a steering wheel, can be used to steer the vehicle 900 (e.g., along a desired path or route) when the propulsion system 950 is operating (e.g., when the vehicle is in motion). The steering system 954 may receive signals from the steering actuator 956. For fully automatic (level 5) functions, the steering wheel may be optional.
[0142] The brake sensor system 946 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 948 and / or the brake sensor.
[0143] It can include one or more System-on-a-Chip (SoC) 904 ( Figure 9C One or more controllers 936, including one or more GPUs, can provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 900. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 948, to operate steering system 954 via one or more steering actuators 956, and to operate propulsion system 950 via one or more throttles / accelerators 952. One or more controllers 936 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 900. One or more controllers 936 may include a first controller 936 for autonomous driving functions, a second controller 936 for functional safety functions, a third controller 936 for artificial intelligence functions (e.g., computer vision), a fourth controller 936 for infotainment functions, a fifth controller 936 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 936 can handle two or more of the functions described above, two or more controllers 936 can handle a single function, and / or any combination thereof.
[0144] One or more controllers 936 may provide signals for controlling one or more components and / or systems of vehicle 900 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensors 958 (e.g., Global Positioning System sensors), RADAR sensors 960, ultrasonic sensors 962, LIDAR sensors 964, inertial measurement unit (IMU) sensors 966 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 996, stereo cameras 968, wide-angle cameras 970 (e.g., fisheye cameras), infrared cameras 972, surround cameras 974 (e.g., 360-degree cameras), long-range and / or medium-range cameras 998, speed sensors 944 (e.g., for measuring the rate of vehicle 900), vibration sensors 942, steering sensors 940, braking sensors (e.g., as part of braking sensor system 946), and / or other sensor types.
[0145] One or more controllers 936 may receive inputs (e.g., represented by input data) from the instrument cluster 932 of the vehicle 900 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 934, an auditory signaling device, a speaker, and / or via other components of the vehicle 900. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 9C Information such as high-definition (“HD”) maps 922, location data (e.g., the location of vehicle 900 on the map), orientation, and the location of other vehicles (e.g., occupying grids), as well as information about objects and their states perceived by controller 936. For example, HMI display 934 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).
[0146] The vehicle 900 further includes a network interface 924, which can communicate via one or more networks using one or more wireless antennas 926 and / or a modem. For example, the network interface 924 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 926 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (“LPWAN”) such as LoRaWAN, SigFox, etc.
[0147] Figure 9B For use in accordance with some embodiments of this disclosure Figure 9A This is an example of the camera position and field of view of an example autonomous vehicle 900. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 900.
[0148] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 900. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-transparent (RCCC) color filter array, a red-transparent-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a high-resolution camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.
[0149] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0150] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) components to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.
[0151] A camera with a field of view that includes the environment in front of the vehicle 900 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 936 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (“ACC”), and / or other functions such as traffic sign recognition.
[0152] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 970, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 9B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 970 can exist on vehicle 900. Furthermore, any number of remote cameras 998 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 998 can also be used for object detection and classification, as well as basic object tracking.
[0153] Any number of stereo cameras 968 can also be included in the front-mounted configuration. In at least one embodiment, one or more stereo cameras 968 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 968 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 968 may be used in addition to those described herein or alternatively.
[0154] Cameras with a field of view including the side portion of the vehicle 900 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 974 (e.g., ... Figure 9B The four surround cameras 974 shown can be mounted on the vehicle 900. The surround cameras 974 can include wide-angle cameras 970, fisheye cameras, 360-degree cameras, and / or the like. Four examples are provided; the four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 974 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0155] A camera with a field of view that includes the environment behind the vehicle 900 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 998, stereo camera 968, infrared camera 972, etc.).
[0156] Figure 9C For use in accordance with some embodiments of this disclosure Figure 9A The example autonomous vehicle 900 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0157] Figure 9C Each component, feature, and system in vehicle 900 is illustrated as being connected via bus 902. Bus 902 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 900 used to assist in the control of various features and functions of vehicle 900, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0158] Although bus 902 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 902 is represented by a single line, this is not intended to be limiting. For example, any number of buses 902 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 902 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 902 may be used for collision avoidance functions, and a second bus 902 may be used for drive control. In any example, each bus 902 may communicate with any component of vehicle 900, and two or more buses 902 may communicate with the same component. In some examples, each SoC 904, each controller 936, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 900) and may be connected to a common bus such as a CAN bus.
[0159] Vehicle 900 may include one or more controllers 936, such as those described herein. Figure 9A The controllers described. Controller 936 can be used for a wide variety of functions. Controller 936 can be coupled to any other different components and systems of vehicle 900 and can be used for the control of vehicle 900, artificial intelligence of vehicle 900, infotainment and / or the like for vehicle 900.
[0160] Vehicle 900 may include one or more System-on-Chip (SoC) 904s. SoC 904 may include a CPU 906, GPU 908, processor 910, cache 912, accelerator 914, data storage 916, and / or other components and features not shown. SoC 904 can be used to control vehicle 900 across a wide variety of platforms and systems. For example, one or more SoCs 904s may be combined with an HD map 922 in a system (e.g., the system of vehicle 900), the HD map being transmitted via a network interface 924 from one or more servers (e.g., [server name missing]). Figure 9D One or more servers (978) receive map refresh and / or updates.
[0161] The CPU 906 may include a CPU cluster or a CPU complex (or, alternatively, referred to herein as "CCPLEX"). The CPU 906 may include multiple cores and / or L2 cache. For example, in some embodiments, the CPU 906 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 906 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). The CPU 906 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of CPU 906 clusters can be active at any given time.
[0162] The CPU 906 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. The CPU 906 can further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.
[0163] The GPU 908 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). The GPU 908 may be programmable and efficient for parallel workloads. In some examples, the GPU 908 may use an enhanced tensor instruction set. The GPU 908 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 908 may include at least eight streaming microprocessors. The GPU 908 may use a computation application programming interface (API). Furthermore, the GPU 908 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0164] In automotive and embedded applications, the GPU 908 can be power-optimized for optimal performance. For example, the GPU 908 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 908 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations for efficient execution of workloads. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.
[0165] The GPU 908 may include, in some examples, a high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem providing peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), may be used.
[0166] The GPU 908 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 908 to directly access the CPU 906 page tables. In such examples, when the GPU 908 Memory Management Unit (MMU) experiences a miss, an address translation request can be transferred to the CPU 906. In response, the CPU 906 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 908. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 906 and GPU 908, simplifying GPU 908 programming and porting applications to the GPU 908.
[0167] In addition, the GPU 908 may include access counters that track how frequently the GPU 908 accesses the memory of other processors. These access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.
[0168] SoC 904 may include any number of caches 912, including those described herein. For example, cache 912 may include an L3 cache available to both CPU 906 and GPU 908 (e.g., it is connected to both CPU 906 and GPU 908). Cache 912 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.
[0169] SoC 904 may include an arithmetic logic unit (ALU), which can be used to perform processing of any of a variety of tasks or operations related to vehicle 900—such as processing a DNN. Additionally, SoC 904 may include a floating-point unit (FPU)—or other mathematical coprocessor or digital coprocessor type—for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated as execution units within CPU 906 and / or GPU 908.
[0170] SoC 904 may include one or more accelerators 914 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 904 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 908 and offload some tasks from GPU 908 (e.g., freeing up more cycles of GPU 908 to perform other tasks). As an example, accelerator 914 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0171] Accelerator 914 (e.g., hardware acceleration clusters) may include a Deep Learning Accelerator (DLA). A DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. TPUs may be accelerators configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. DLAs may be further optimized for a specific set of neural network types and floating-point operations as well as inference. DLAs are designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperform CPUs. TPUs can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0172] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0173] The DLA can perform any function of the GPU 908, and by using inference accelerators, for example, designers can target either the DLA or the GPU 908 for any function. For instance, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 908 and / or other accelerators 914.
[0174] Accelerator 914 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. A PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. A PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0175] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0176] DMA enables PVA components to access system memory independently of the CPU 906. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0177] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., a VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.
[0178] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error-correcting code (ECC) memory to enhance overall system security.
[0179] Accelerator 914 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 914. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. PVA and DLA can access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects PVA and DLA to memory.
[0180] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.
[0181] In some examples, the SoC 904 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.
[0182] Accelerator 914 (e.g., hardware accelerator clusters) has broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.
[0183] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.
[0184] In some examples, PVA can be used to perform intensive optical flow, providing processed RADAR data from the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.
[0185] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 966 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 964 or RADAR sensor 960), etc.
[0186] The SoC 904 may include one or more data storage units 916 (e.g., memory). The data storage unit 916 may be on-chip memory of the SoC 904, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, the data storage unit 916 may be large enough to store multiple instances of the neural network. The data storage unit 912 may include an L2 or L3 cache 912. References to the data storage unit 916 may include references to memory associated with the PVA, DLA, and / or other accelerators 914 as described herein.
[0187] The SoC 904 may include one or more processors 910 (e.g., embedded processors). Processor 910 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 904 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 904 thermal and temperature sensor management, and / or SoC 904 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC 904 may use the ring oscillator to detect the temperature of the CPU 906, GPU 908, and / or accelerator 914. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place the SoC 904 into a lower power state and / or place the vehicle 900 into a driver-safe parking mode (e.g., safely stopping the vehicle 900).
[0188] The processor 910 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0189] The processor 910 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0190] The processor 910 may further include a secure cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0191] The processor 910 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0192] The processor 910 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0193] Processor 910 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 970, the surround camera 974, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.
[0194] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.
[0195] The video image compositer can also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 908 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 908 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 908 to improve performance and responsiveness.
[0196] The SoC 904 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. The SoC 904 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.
[0197] SoC 904 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 904 can be used to process data from cameras and sensors (e.g., LIDAR sensor 964, RADAR sensor 960, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 902 (e.g., vehicle 900 speed, steering wheel position, etc.), and data from GNSS sensor 958 (connected via Ethernet or CAN bus). SoC 904 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 906 from routine data management tasks.
[0198] The SoC 904 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 904 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 906, GPU 908, and data storage 916, the accelerator 914 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0199] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0200] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 920) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on a CPU complex.
[0201] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 908.
[0202] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 900. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 904 provides security against theft and / or carjacking.
[0203] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 996 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 904 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 958. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 962, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.
[0204] The vehicle may include a CPU 918 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 904 via a high-speed interconnect (e.g., PCIe). The CPU 918 may include, for example, an x86 processor. The CPU 918 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 904, and / or monitoring the status and health of the controller 936 and / or the infotainment SoC 930.
[0205] Vehicle 900 may include a GPU 920 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 904 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 920 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 900 (e.g., sensor data).
[0206] Vehicle 900 may further include a network interface 924, which may include one or more wireless antennas 926 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 924 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 978 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 900 with information about vehicles approaching vehicle 900 (e.g., vehicles in front, to the side, and / or behind vehicle 900). This functionality can be part of vehicle 900's cooperative adaptive cruise control function.
[0207] Network interface 924 may include a SoC that provides modulation and demodulation functions and enables controller 936 to communicate over a wireless network. Network interface 924 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0208] Vehicle 900 may further include data storage 928, which may include off-chip (e.g., outside of SoC 904) storage devices. Data storage 928 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0209] Vehicle 900 may further include a GNSS sensor 958. The GNSS sensor 958 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 958 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0210] Vehicle 900 may further include a RADAR sensor 960. The RADAR sensor 960 can be used by vehicle 900 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level can be ASIL B. The RADAR sensor 960 can use CAN and / or bus 902 (e.g., to transmit data generated by the RADAR sensor 960) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 960 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0211] The RADAR sensor 960 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 960 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 900's surroundings at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 900's lane.
[0212] As an example, a mid-range RADAR system can include a range of up to 960m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 950 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.
[0213] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0214] Vehicle 900 may further include ultrasonic sensors 962. Ultrasonic sensors 962, which may be positioned at the front, rear, and / or sides of vehicle 900, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 962 can be used, and different ultrasonic sensors 962 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 962 can operate at functional safety level ASIL B.
[0215] Vehicle 900 may include a LIDAR sensor 964. The LIDAR sensor 964 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 964 may be of functional safety level ASIL B. In some examples, vehicle 900 may include multiple LIDAR sensors 964 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0216] In some examples, the LiDAR sensor 964 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 964 may have an advertising range of, for example, approximately 900m, with an accuracy of 2cm-3cm, and support 900Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 964 may be used. In such examples, the LiDAR sensor 964 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 900. In such examples, the LiDAR sensor 964 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 964 may be configured for a horizontal field of view between 45 and 135 degrees.
[0217] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LiDAR, and because a flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 964 is less susceptible to motion blur, vibration, and / or shock.
[0218] The vehicle may further include an IMU sensor 966. In some examples, the IMU sensor 966 may be located at the center of the rear axle of the vehicle 900. The IMU sensor 966 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 966 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 966 may include an accelerometer, a gyroscope, and a magnetometer.
[0219] In some embodiments, the IMU sensor 966 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 966 can enable the vehicle 900 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 966 without input from a magnetic sensor. In some examples, the IMU sensor 966 and the GNSS sensor 958 can be combined into a single integrated unit.
[0220] The vehicle may include a microphone 996 placed in and / or around the vehicle 900. Among other things, the microphone 996 may be used for emergency vehicle detection and identification.
[0221] The vehicle may further include any number of camera types, including stereo camera 968, wide-angle camera 970, infrared camera 972, surround camera 974, long-range and / or mid-range camera 998, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 900. The camera types used depend on the embodiment and the requirements of the vehicle 900, and any combination of camera types can be used to provide the necessary coverage around the vehicle 900. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 9A and Figure 9B It was described in more detail.
[0222] Vehicle 900 may further include vibration sensor 942. Vibration sensor 942 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 942 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).
[0223] Vehicle 900 may include ADAS system 938. In some examples, ADAS system 938 may include SoC. ADAS system 938 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.
[0224] The ACC system can use a RADAR sensor 960, a LIDAR sensor 964, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 900 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 900 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.
[0225] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or network connection (e.g., via the Internet) through network interface 924 and / or wireless antenna 926. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 900 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 900, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0226] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.
[0227] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.
[0228] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0229] The LKA system is a variation of the LDW system. If vehicle 900 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 900.
[0230] The BSW system detects and warns the driver of vehicles in the blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system can utilize a rear-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0231] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0232] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as they alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 900, in the event of conflicting results, the vehicle 900 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 936 or the second controller 936). For example, in some embodiments, the ADAS system 938 may be a backup and / or auxiliary computer used to provide perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and varied software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from the ADAS system 938 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0233] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.
[0234] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of and / or be included as components of the SoC 904.
[0235] In other examples, ADAS system 938 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0236] In some examples, the output of the ADAS system 938 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if the ADAS system 938 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0237] Vehicle 900 may further include an infotainment SoC 930 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 930 may include a combination of hardware and software that can be used to provide vehicle 900 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 930 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 934, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 930 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from the ADAS system 938, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0238] The infotainment SoC 930 may include GPU functionality. The infotainment SoC 930 can communicate with other devices, systems, and / or components of the vehicle 900 via bus 902 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 930 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 936 (e.g., the primary and / or backup computer of the vehicle 900). In such an example, the infotainment SoC 930 may place the vehicle 900 into a driver-safe parking mode as described herein.
[0239] Vehicle 900 may further include instrument cluster 932 (e.g., digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). Instrument cluster 932 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). Instrument cluster 932 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 930 and instrument cluster 932. In other words, instrument cluster 932 may be included as part of infotainment SoC 930, or vice versa.
[0240] Figure 9D For cloud-based servers and according to some embodiments of this disclosure Figure 9A The diagram illustrates a system for communication between example autonomous vehicles 900. System 976 may include server 978, network 990, and vehicles including vehicle 900. Server 978 may include multiple GPUs 984(A)-984(H) (collectively referred to herein as GPU 984), PCIe switches 982(A)-982(H) (collectively referred to herein as PCIe switch 982), and / or CPUs 980(A)-980(B) (collectively referred to herein as CPU 980). GPUs 984, CPUs 980, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 986, such as, but not limited to, NVLink interface 988 developed by NVIDIA. In some examples, GPUs 984 are connected via NVLink and / or NVSwitch SoCs, and GPUs 984 and PCIe switches 982 are connected via PCIe interconnects. Although eight GPUs 984, two CPUs 980, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 978 may include any number of GPUs 984, CPUs 980, and / or PCIe switches. For example, each of the servers 978 may include eight, sixteen, thirty-two, and / or more GPUs 984.
[0241] Server 978 can receive image data from vehicles via network 990, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 978 can also transmit neural network 992, updated neural network 992, and / or map information 994, including information about traffic and road conditions, to vehicles via network 990. Updates to map information 994 may include updates to HD map 922, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 992, updated neural network 992, and / or map information 994 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 978 and / or other servers).
[0242] Server 978 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to: categories such as supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 990), and / or the machine learning model can be used by Server 978 to remotely monitor the vehicle.
[0243] In some examples, server 978 can receive data from vehicles and apply that data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 978 may include a deep learning supercomputer powered by GPU 984 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 978 may include a deep learning infrastructure in a data center that uses only CPU power.
[0244] The deep learning infrastructure of server 978 is capable of rapid, real-time inference and can be used to assess and verify the health of the processor, software, and / or associated hardware in vehicle 900. For example, the deep learning infrastructure can receive periodic updates from vehicle 900, such as image sequences and / or objects located within those image sequences by vehicle 900 (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them to those identified by vehicle 900. If the results do not match and the infrastructure concludes that the AI in vehicle 900 has malfunctioned, server 978 can transmit a signal to vehicle 900 instructing its fail-safe computer to take control, notify passengers, and complete a safe stopping operation.
[0245] For inference, the server 978 can include a GPU 984 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0246] Example computing device
[0247] Figure 10 This is a block diagram suitable for implementing some embodiments of the present disclosure of an example computing device 1000. The computing device 1000 may include an interconnect system 1002 directly or indirectly coupled to the following devices: a memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communication interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., displays), and one or more logic units 1020. In at least one embodiment, one or more computing devices 1000 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1008 may include one or more vGPUs, one or more CPUs 1006 may include one or more vCPUs, and / or one or more logic units 1020 may include one or more virtual logic units. Thus, one or more computing devices 1000 may include discrete components (e.g., a full GPU dedicated to computing device 1000), virtual components (e.g., a portion of the GPU dedicated to computing device 1000), or combinations thereof.
[0248] although Figure 10 The various blocks are shown as connected via an interconnect system 1002 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 1018, such as a display device, may be considered an I / O component 1014 (e.g., if the display is a touchscreen). As another example, the CPU 1006 and / or GPU 1008 may include memory (e.g., memory 1004 may represent a storage device other than the memory of the GPU 1008, CPU 1006, and / or other components). In other words, Figure 10 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 10 Within the scope of computing devices.
[0249] Interconnect system 1002 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1002 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 1006 may be directly connected to memory 1004. Furthermore, CPU 1006 may be directly connected to GPU 1008. In cases where there is a direct or point-to-point connection between components, interconnect system 1002 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required in computing device 1000.
[0250] The memory 1004 may include any of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by the computing device 1000. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.
[0251] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1004 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 1000. As used herein, computer storage media does not include the signal itself.
[0252] Computer storage media may include computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium. The term "modulated data signal" may refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0253] CPU 1006 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. Each of CPU 1006 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 1006 may include any type of processor and may include different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1000, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 1000 may also include one or more CPUs 1006.
[0254] In addition to or as a replacement for CPU 1006, one or more GPUs 1008 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. One or more GPUs 1008 may be integrated GPUs (e.g., having one or more CPUs 1006) and / or one or more GPUs 1008 may be discrete GPUs. In embodiments, one or more GPUs 1008 may be coprocessors of one or more CPUs 1006. Computing device 1000 may use GPUs 1008 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, one or more GPUs 1008 may be used for general-purpose computing on a GPU (GPGPU). One or more GPUs 1008 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPUs 1008 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received from CPU 1006 via a host interface). GPU 1008 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 1004. One or more GPUs 1008 may include two or more GPUs operating in parallel (e.g., via a link). The link may be directly connected to the GPUs (e.g., using NVLINK) or connected via a switch (e.g., using NVSwitch). When combined, each GPU 1008 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.
[0255] In addition to or as an alternative to CPU 1006 and / or GPU 1008, logic unit 1020 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. In embodiments, CPU 1006, GPU 1008, and / or logic unit 1020 may execute any combination of methods, processes, and / or portions thereof, discretely or jointly. One or more logic units 1020 may be part of and / or integrated into one or more of CPU 1006 and / or GPU 1008, and / or one or more logic units 1020 may be discrete components or otherwise separate from CPU 1006 and / or GPU 1008. In embodiments, one or more logic units 1020 may be coprocessors of one or more CPUs 1006 and / or one or more GPUs 1008.
[0256] Examples of logic unit 1020 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.
[0257] The communication interface 1010 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1000 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 1010 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 1020 and / or the communication interface 1010 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 1002 to (e.g., memory) one or more GPUs 1008.
[0258] I / O port 1012 enables computing device 1000 to be logically coupled to other devices, including I / O component 1014, presentation component 1018, and / or other components, some of which may be built into (e.g., integrated into) computing device 1000. Illustrative I / O component 1014 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, and so on. I / O component 1014 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 1000 (described in more detail below). Computing device 1000 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, the computing device 1000 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by the computing device 1000 to render immersive augmented reality or virtual reality.
[0259] The power supply 1016 may include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 1016 may supply power to the computing device 1000 so that the components of the computing device 1000 can operate.
[0260] The presentation component 1018 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 1018 may receive data from other components (such as GPU 1008, CPU 1006, DPU, etc.) and output that data (such as as images, videos, sounds, etc.).
[0261] Example Data Center
[0262] Figure 11 An example data center 1100 that may be used in at least one embodiment of this disclosure is shown. The data center 1100 may include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.
[0263] like Figure 11As shown, the data center infrastructure layer 1110 may include a resource coordinator 1112, grouped computing resources 1114, and node computing resources (“nodes CR”) 1116(1)-1116(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR 1116(1)-1116(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CR 1116(1)-1116(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR1116(1)-11161(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of nodes CR1116(1)-1116(N) may correspond to virtual machines (VMs).
[0264] In at least one embodiment, the grouped computing resources 1114 may include individual groups of nodes CR1116 housed within one or more racks (not shown), or multiple racks housed within a data center at different geographical locations (also not shown). Individual groups of nodes CR1116 within the grouped computing resources 1114 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several nodes CR1116, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0265] Resource coordinator 1122 may be configured or otherwise control one or more nodes CR1116(1)-1116(N) and / or grouped computing resources 1114. In at least one embodiment, resource coordinator 1122 may include a Software Design Infrastructure (“SDI”) management entity for data center 1100. Resource coordinator 1122 may include hardware, software, or some combination thereof.
[0266] In at least one embodiment, such as Figure 11As shown, framework layer 1120 may include a job scheduler 1133, a configuration manager 1134, a resource manager 1136, and / or a distributed file system 1138. Framework layer 1120 may include a framework for software 1132 supporting software layer 1130 and / or one or more applications 1142 of application layer 1140. Software 1132 or application 1142 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 1120 may be, but is not limited to, free and open-source software web application frameworks (such as Apache Spark) that can utilize distributed file system 1138 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark") is a type of resource. In at least one embodiment, the job scheduler 1133 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 1100. The configuration manager 1134 may be able to configure different layers, such as software layer 1130 and framework layer 1120 (which includes Spark and distributed file system 1138 for supporting large-scale data processing). The resource manager 1136 may be able to manage compute resources mapped to or allocated to clusters of distributed file system 1138 and job scheduler 1133 or to support clusters of distributed file system 1138 and job scheduler 1133. In at least one embodiment, the clustered or grouped compute resources may include grouped compute resources 1114 in data center infrastructure layer 1110. The resource manager 1136 may coordinate with resource coordinator 1112 to manage these mapped or allocated compute resources.
[0267] In at least one embodiment, the software 1132 included in software layer 1130 may include software used in at least a portion of the nodes CRs 1116(1)-1116(N), the grouped computing resources 1114, and / or the distributed file system 1138 of framework layer 1120. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0268] In at least one embodiment, the application 1142 included in the application layer 1140 may include one or more types of applications used at least in part by nodes CR1116(1)-1116(N), grouped computing resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.
[0269] In at least one embodiment, any of the configuration manager 1134, resource manager 1136, and resource coordinator 1112 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can free data center operators of data center 1100 from making potentially poor configuration decisions and may prevent underutilization and / or poor performance of the data center.
[0270] According to one or more embodiments described herein, data center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 1100 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1100 by using weight parameters computed through one or more training techniques, such as, but not limited to, those described herein.
[0271] In at least one embodiment, the data center 1100 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0272] Example network environment
[0273] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 10 This can be implemented on one or more instances of computing device 1000—for example, each device may include similar components, features, and / or functions of computing device 1000. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 1100, examples of which are described in this document. Figure 11 To describe in more detail.
[0274] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks, or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0275] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.
[0276] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or applications may respectively include network-based service software or applications. In embodiments, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software network application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").
[0277] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations, such as a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0278] Client devices may include those described in this article. Figure 10 The example computing device 1000 described includes at least some components, features, and functions. By way of example and not limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these devices described, or any other suitable device.
[0279] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0280] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" could include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0281] This document describes in detail the subject matter of this disclosure to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
Claims
1. A system comprising: a plurality of processors to: determine that first frame data corresponds to a first time and that second frame data corresponds to a second time after the first time; provide the first frame data to a first processor of the plurality of processors configured to execute on inputs arranged in two dimensions; provide the second frame data to a second processor of the plurality of processors configured to execute on inputs arranged in one dimension in parallel with providing the first frame data to the first processor; and execute the first processor in parallel with executing the second processor from the second frame data according to the first frame data.
2. The system of claim 1, wherein the plurality of processors are to: execute one or more first instructions based on the first frame data.
3. The system of claim 2, wherein the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a bloom preparation operation.
4. The system of claim 2, wherein the plurality of processors are to: execute one or more second instructions based on the first frame data after executing the one or more first instructions.
5. The system of claim 4, wherein the one or more second instructions correspond to a bloom preparation operation.
6. The system of claim 1, wherein the plurality of processors are to: provide the first frame data to the second processor after providing the first frame data to the first processor; and execute one or more instructions by the second processor based on the first frame data after providing the first frame data to the first processor.
7. The system of claim 6, wherein the one or more instructions correspond to an injection operation.
8. The system of claim 1, wherein the plurality of processors are to: provide third frame data to the first processor after providing the first frame data to the first processor, the third frame data corresponding to an output of the first processor based on the first frame data; and execute one or more instructions by the first processor based on the first frame data after providing the first frame data to the first processor.
9. The system of claim 8, wherein the one or more instructions correspond to a non-maximum suppression operation.
10. The system of claim 8, wherein the plurality of processors are to: provide fourth frame data to the second processor in parallel with executing the one or more instructions based on the first frame data, the fourth frame data corresponding to an output of the second processor based on the second frame data.
11. The system of claim 1, wherein the plurality of processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using a robot; an aerial system; a medical system; a rowing system; an intelligent area monitoring system; a system for performing a deep learning operation; Systems for performing simulation operations; Systems for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content; Systems for performing digital twin operations; Systems implemented using edge devices; Systems including one or more virtual machines (VMs); Systems for generating synthetic data; Systems implemented at least partially in a data center; Systems for performing conversational artificial intelligence (AI) operations; Systems for performing generative AI operations; Systems implementing language models; Systems implementing visual language models (VLMs); Systems implementing large language models (LLMs); Systems implementing multi-modal language models; Systems for hosting one or more real-time streaming applications; Systems for performing optical transport simulation; Systems for performing collaborative content creation of 3D assets; or Systems implemented at least partially using cloud computing resources.
12. A system on a chip (SoC) comprising: at least one graphics processing unit (GPU); and a plurality of processors to: determine that first frame data corresponds to a first time and that second frame data corresponds to a second time after the first time; provide the first frame data to a first processor of the plurality of processors configured to perform input arranged in two dimensions; provide the second frame data to a second processor of the plurality of processors configured to perform input arranged in one dimension in parallel with providing the first frame data to the first processor; and execute the first processor in parallel with executing the second processor from the second frame data.
13. The SoC of claim 12, wherein the plurality of processors are to: execute one or more first instructions based on the first frame data.
14. The SoC of claim 13, wherein the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a bloom preparation operation.
15. The SoC of claim 13, wherein the plurality of processors are to: execute one or more second instructions by the first processor based on the first frame data after executing the one or more first instructions.
16. The SoC of claim 15, wherein the one or more second instructions correspond to a bloom preparation operation, an injection operation, or a non-maximum suppression operation.
17. The SoC of claim 12, wherein the plurality of processors are to: provide the first frame data to the second processor after providing the first frame data to the first processor; and execute one or more instructions by the second processor based on the first frame data after providing the first frame data to the first processor.
18. The SoC of claim 12, comprising the at least one GPU to: provide third frame data to the first processor after providing the first frame data to the first processor, the third frame data corresponding to an output of the first processor based on the first frame data; and execute the first processor in parallel with executing the second processor from the third frame data. After providing the first frame data to the first processor, one or more instructions are executed by the first processor based on the first frame data.
19. The SoC of claim 12, wherein the plurality of processors are to: provide the fourth frame data to the second processor in parallel with executing the one or more instructions based on the first frame data, the fourth frame data corresponding to an output of the second processor based on the second frame data.
20. A method performed by a plurality of processors, the method comprising: determining that first frame data corresponds to a first time and that second frame data corresponds to a second time after the first time; providing the first frame data to a first processor of the plurality of processors configured to execute inputs arranged in two dimensions; providing the second frame data to a second processor of the plurality of processors configured to execute inputs arranged in one dimension in parallel with providing the first frame data to the first processor; and executing the first processor in parallel with executing the second processor according to the second frame data according to the first frame data.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2