Core-based integrated central processing unit with accelerator for video processing

By receiving the input video data stream at the demultiplexer of the central processing unit, and using the CPU and accelerator-based decoding unit for dynamic decoding configuration evaluation, combined with the unified memory access tunnel system, PCIe connection limitation and system delay problems are solved, and the performance and resource utilization of the server system are optimized.

CN120303658APending Publication Date: 2025-07-11META PLATFORMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079845.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2023-12-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In existing server systems, PCIe connections limit the sharing of address space between the central processing unit and the accelerator, and external I/O connections increase system delay, making it difficult to balance the ratio of CPU to accelerator, affecting efficiency.

Method used

By receiving the input video data stream at the CPU-based demultiplexer of the central processing unit, and using the CPU-based decoding unit to perform dynamic decoding configuration evaluation, combined with the unified memory access tunnel system, the accelerator allows the accelerator to directly access the unified memory, avoiding the use of the accelerator local memory, and implementing dynamic decoding and preprocessing.

Benefits of technology

The efficiency of the central processing unit and accelerator is improved, the dependence on accelerator memory is reduced, and the system performance and resource utilization are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303658A_ABST
    Figure CN120303658A_ABST
Patent Text Reader

Abstract

In some embodiments, a method includes receiving, at a central processing unit (CPU)-based demultiplexer of a CPU, an input video data stream; performing, at the CPU, an accelerator decoding configuration evaluation on an accelerator decoding configuration of the accelerator; and dynamically decoding the CPU-based demultiplexer output from the CPU-based demultiplexer using the CPU-based decoding unit and the accelerator-based decoding unit based on the accelerator decoding configuration evaluation. In some embodiments of the method, the accelerator decode configuration evaluation includes performing an accelerator-based decode unit hardware configuration evaluation on the accelerator-based decode unit, and performing a CPU-based decode unit software configuration evaluation on CPU-based decode unit software for the CPU-based decode unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure describes methods, systems, and computer-readable media storing code for performing the methods. Background Art

[0002] The background description provided herein is for the purpose of generally presenting the background of the disclosure. To the extent that the work of one or more of the presently-named inventors is described in this background art section, and in the specification, to the extent it may not have been publicly available at the time of filing the application, such work is neither expressly nor impliedly admitted to be prior art to the present disclosure.

[0003] Many modern server systems utilize a system-on-chip (SoC) that includes an accelerator connected to a central processing unit (CPU) using the Peripheral Component Interconnect Express (PCIe). Although PCIe is generally acceptable for offloading workloads, it has several limitations. For example, PCIe may not allow sharing of the address space between the CPU and the accelerator. Additionally, the accelerator typically requires high bandwidth and low latency access to a common memory. External input / output (I / O) connections such as PCIe or Compute ExpressLink (CXL) may increase the system latency because the aggregate bandwidth is limited due to the limited number of PCIe / CXL ports and channels. The limited number of ports and channels reduces the number of accelerators that can be attached to the CPU and makes it difficult to balance the CPU-to-accelerator ratio to match the application requirements, which has a negative impact on the efficiency of the CPU and the accelerator. Summary of the Invention

[0004] The summary of the invention provided herein is for the purpose of introducing some concepts in a simplified form that will be further described in the detailed description below. The summary of the invention is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0005] This disclosure describes methods, systems, and computer-readable media storing code for performing the methods. In one aspect, a method includes: receiving an input video data stream at a CPU-based demultiplexer of a central processing unit (CPU); performing an accelerator decoding configuration evaluation of an accelerator decoding configuration of an accelerator at the CPU; and based on the accelerator decoding configuration evaluation, dynamically decoding a CPU-based demultiplexer output from the CPU-based demultiplexer using a CPU-based decoding unit and an accelerator-based decoding unit. In some embodiments of the method, the accelerator decoding configuration evaluation includes: performing an accelerator-based decoding unit hardware configuration evaluation of the accelerator-based decoding unit, and performing a CPU-based decoding unit software configuration evaluation of the CPU-based decoding unit software for the CPU-based decoding unit.

[0006] In some embodiments, the method further includes: performing an accelerator-based preprocessing unit configuration evaluation of an accelerator-based preprocessing unit coupled to the CPU-based decoding unit and the accelerator-based decoding unit.

[0007] In some embodiments, the accelerator-based preprocessing unit configuration evaluation includes an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0008] In some embodiments, the method further includes: dynamically preprocessing decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using the accelerator-based preprocessing unit configuration evaluation.

[0009] In some embodiments, the method further includes: performing an encoding unit hardware configuration evaluation of an accelerator-based encoding unit.

[0010] In some embodiments, the method further includes: dynamically encoding a preprocessing output from the accelerator-based preprocessing unit using the accelerator-based encoding unit.

[0011] In another aspect according to the present disclosure, there is provided a system-on-chip including: a CPU-based demultiplexer; a CPU-based decoding unit and an accelerator-based decoding unit, the CPU-based decoding unit and the accelerator-based decoding unit being coupled to the CPU-based demultiplexer via an inter-die interconnect; and an accelerator-based preprocessing unit, the accelerator-based preprocessing unit being coupled to the CPU-based decoding unit and the accelerator-based decoding unit via the inter-die interconnect, wherein the system-on-chip dynamically decodes a CPU-based demultiplexer output from the CPU-based demultiplexer using the CPU-based decoding unit and the accelerator-based decoding unit based on an accelerator decoding configuration evaluation for video processing.

[0012] In some embodiments, the accelerator decoding configuration evaluation includes an accelerator-based decoding unit hardware configuration evaluation and a CPU-based decoding unit software configuration evaluation.

[0013] In some embodiments, the accelerator-based decoding unit hardware configuration evaluation includes determining whether the CPU-based demultiplexer output is mapped to the accelerator-based decoding unit hardware configuration of the accelerator-based decoding unit.

[0014] In some embodiments, the CPU-based decoding unit software configuration evaluation includes determining whether the CPU-based demultiplexer output is mapped to the CPU-based decoding unit software configuration of the CPU-based decoding unit.

[0015] In some embodiments, the system-on-chip is configured to dynamically preprocess the decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using an accelerator-based preprocessing unit.

[0016] In some embodiments, the system-on-chip dynamically preprocesses the decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using an accelerator-based preprocessing unit configuration evaluation.

[0017] In some embodiments, the accelerator-based preprocessing unit configuration evaluation includes an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0018] In some embodiments, the system-on-chip further includes: an accelerator-based encoding unit, the accelerator-based encoding unit being coupled to the accelerator-based preprocessing unit, wherein the system-on-chip is configured to dynamically encode the preprocessing output from the accelerator-based preprocessing unit using the accelerator-based encoding unit.

[0019] In some embodiments, the system-on-chip dynamically encodes the preprocessor output from the accelerator-based preprocessing unit using an accelerator-based encoding unit hardware configuration evaluation.

[0020] According to another aspect of the present disclosure, there is provided a server system, the server system including: a host central processing unit (CPU); an accelerator, the accelerator being coupled to the host CPU via an inter-die interconnect; a unified memory, the unified memory being coupled to the host CPU and the accelerator, wherein the server system is configured to dynamically share the decoding operation, the preprocessing operation, and the encoding operation using the host CPU and the accelerator based on an accelerator configuration evaluation performed by the host CPU.

[0021] In some embodiments, to share decoding operations, the host CPU performs an accelerator-based decoding unit hardware configuration evaluation and a CPU-based decoding unit software configuration evaluation.

[0022] In some embodiments, to share preprocessing operations, the host CPU performs an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0023] In some embodiments, to share encoding operations, the host CPU performs an accelerator-based encoding unit hardware configuration evaluation and a CPU-based encoding unit software configuration evaluation.

[0024] It will be appreciated that any feature described herein as being suitable for incorporation into one or more aspects or embodiments of the present disclosure is intended to be generalizable to any and all aspects and embodiments of the present disclosure. Those skilled in the art can understand other aspects of the present disclosure based on the description, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are merely exemplary and illustrative and are not limiting of the claims. Further features and advantages of the embodiments, as well as the structure and operation of various embodiments, are described in detail below with reference to the drawings. It should be noted that the methods and systems are not limited to the specific embodiments described herein. These embodiments are presented herein for illustrative purposes only. Additional embodiments will be apparent to one or more persons skilled in the relevant art based on the teachings contained herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1A is a block diagram showing a server system according to some embodiments.

[0026] Figure 1B is showing a server system according to some embodiments Figure 1A in the block diagram.

[0027] Figure 1C is showing a server system according to some embodiments Figure 1B in the block diagram of the process flow.

[0028] Figure 1D shows a unified memory access method utilized in Figures 1A to 1C according to some embodiments.

[0029] Figure 2 is a block diagram showing a system-on-chip according to some embodiments.

[0030] Figure 3 shows in more detail a system-on-chip according to some embodiments Figure 2 in the block diagram.

[0031] Figure 4 is a block diagram showing a shared pre - processing flow according to some embodiments.

[0032] Figure 5 is a block diagram showing a shared encoding flow according to some embodiments.

[0033] Figure 6 is a block diagram showing a system - on - chip according to some embodiments.

[0034] Figure 7 shows according to some embodiments Figure 6 a block diagram of the system - on - chip in

[0035] Figure 8 is a flowchart showing a side - band - based accelerator firmware authentication method according to some embodiments.

[0036] Figure 9 is an example process flow of using Figure 7 the side - band - based accelerator firmware authentication method in Figure 8 the system - on - chip in

[0037] Figure 10 is a flowchart showing a side - band - based accelerator firmware authentication method according to some embodiments. Detailed Description

[0038] Figure 1A is a block diagram showing server system 100 according to some embodiments. In some embodiments, server system 100 is configured to use the unified memory access tunnel system in SoC 120 to access unified memory 160 through the inter - die interface 167, so that accelerator 130 in SoC120 does not have to use accelerator memory during the execution of processing operations by accelerator 130. In some embodiments, the unified memory access tunnel system is hardware and / or executable code in accelerator 130 configured to generate a unified memory access tunnel packet structure, which is used by the unified memory access tunnel system to tunnel the advanced interconnect protocol into the inter - die interface protocol of inter - die interface 167. In some embodiments, by using the unified memory access tunnel packet structure, accelerator 130 can generate a unified memory access tunnel packet that allows accelerator 130 to transfer data through inter - die interface 167, so that during the execution of operations at accelerator 130, accelerator 130 uses unified memory 160 instead of accelerator memory to process operations.

[0039] In some embodiments, server system 100 includes a system-on-chip (SoC) 120, unified memory 160, and an external input / output (I / O) interconnect 139. In some embodiments, SoC 120 is coupled to unified memory 160 via external input / output (I / O) interconnect 139. In some embodiments, external I / O interconnect 139 can be, for example, a Compute Express Link (CXL) interconnect configured to couple unified memory 160 to SoC 120, a Peripheral Component Interconnect Express (PCIe) interconnect, or other type of external I / O interconnect.

[0040] In some embodiments, SoC 120 is a system-on-chip configured to use an on-chip fabric as the communication protocol within SoC 120, such as an Advanced eXtensible Interface (AXI), a Network on Chip (NoC), or an Advanced Computing Environment (ACE). In some embodiments, SoC 120 includes a host central processing unit (CPU) 171, an accelerator 130, and an inter-die interface 167. In some embodiments, host CPU 171 is coupled to accelerator 130 via inter-die interface 167. In some embodiments, inter-die interface 167 is a physical interface or connection between host CPU 171 and accelerator 130, and this physical interface or connection includes an inter-die interconnect (e.g., Figure 1B inter-die interconnects 161 to 164 among those shown). In some embodiments, each of inter-die interconnects 161 to 164 includes a sideband bus and an inter-die bus (mainband bus). In some embodiments, inter-die interface 167 can be configured to transfer data on inter-die interface 167 using the Chiplet Data Exchange (CDX) protocol. CDX is a high-speed low-latency protocol designed for chip-to-chip communication and optimized for chiplet interconnects. In some embodiments, CDX is configured to support high data rates, allowing for fast data transfer between multiple chiplets in SoC 120. In some embodiments, inter-die interface 167 can be, for example, a universal chiplet interconnect express (UCIe) interface or another type of inter-die interface for the embodiments described herein.

[0041] In some embodiments, host CPU 171 is a processor that is configured to perform the operations described herein (referenced in Figures 1A to 10 and described in more detail) in addition to performing standard CPU processing operations within SoC 120. In some embodiments, host CPU 171 can be a server-class CPU that is coupled to accelerator 130 integrated on SoC 120. In some embodiments, SoC 120 can include multiple host CPUs, for example, depending on the type of server system 100.

[0042] In some embodiments, accelerator 130 is a dedicated processing unit in SoC 120 that is configured to access unified memory 160 using a unified memory access tunneling system in addition to performing tasks specific to accelerator 130. In some embodiments, accelerator 130 is configured to access unified memory 160 using a shared address space mapped to unified memory 160. In some embodiments, the shared address space is a range of shared memory addresses associated with unified memory 160 that host CPU 171 and accelerator 130 can access to perform processing operations. In some embodiments, the unified memory access tunneling system of SoC 120 is configured to allow direct data transfer from accelerator 130 to unified memory 160 and vice versa through die-to-die interface 167 in addition to allowing accelerator 130 and host CPU 171 to utilize the shared address space in unified memory 160 (further described herein in Figures 1B to 1D ). In some embodiments, as Figure 1B illustrated by way of example in, accelerator 130 includes accelerator 131, accelerator 132, accelerator 133, and accelerator 134. In some embodiments, accelerator 130 can be, for example, a video accelerator, a graphics processing unit (GPU), a digital signal processor, or other type of accelerator that is configured to perform the operations described herein using a unified memory access tunneling system. In some embodiments, each accelerator in accelerator 130 can be configured to include a unified memory access tunneling system to perform the required unified memory access tunneling operations described herein.

[0043] In some embodiments, unified memory 160 is a memory shared between host CPU 171 and accelerator 130, and the memory is configured to be directly accessed by accelerator 130 using the unified memory access tunnel operations and shared address space described herein. In some embodiments, unified memory 160 can be, for example, a low power (LP) memory and / or other types of memory associated with a shared address space for use by accelerator 130 and host CPU 171. In some embodiments, as described above, the shared address space is a range of shared memory addresses that host CPU 171 and accelerator 130 can use to access executable code required for performing processing operations. In some embodiments, unified memory 160 can be a random access memory (RAM), such as a dynamic random access memory (DRAM), a static random access memory (SRAM), and a non-volatile random access memory (NVRAM), etc. In some embodiments, unified memory 160 can include a system memory and an external memory associated with host CPU 171 located on SoC 120, and the system memory and the external memory are combined into unified memory 160 (e.g., the host CPU memory and the external memory are combined into unified memory 160). In some embodiments, unified memory 160 can be considered unified because, for example, due to the transmission of the advanced interconnect protocol tunnel into the interdie interface protocol of interdie interface 167 as described herein, accelerator 130 may only need the shared address space to directly access the memory for processing.

[0044] In some embodiments, to access unified memory 160 using the shared address space, the unified memory access tunnel system is configured to transmit the advanced interconnect protocol tunnel into the interdie interface protocol of interdie interface 167. In some embodiments, SoC 120 uses the unified memory access tunnel system to allow accelerator 130 and host CPU 171 to access the shared address space of unified memory 160, and to deactivate the use of accelerator memory during accelerator processing. Therefore, accelerator 130 does not need to have an accelerator memory for accelerator 130 to perform processing operations. The methods and systems are further described herein with reference to Figures 1B to 1D The various methods and systems are further described.

[0045] Figure 1B Further illustrated in accordance with some embodiments of Figure 1AThe server system 100 therein. In some embodiments, the SoC 120 of the server system 100 includes an accelerator 131 equipped with a unified memory access tunneling system 116, which is configured to perform unified memory access tunneling operations that allow the accelerator 131 to directly access the unified memory 160 (e.g., LP memory 121 and / or memory 126) through the die-to-die interconnect 161, such that the accelerator 130 in the SoC 120 does not have to utilize accelerator memory during the execution of processing operations by the accelerator 131. In some embodiments, the unified memory access tunneling system 116 is hardware and / or executable code in the accelerator 131 that is configured to generate a unified memory access tunneling packet structure, which is used by the unified memory access tunneling system 116 to tunnel a high-level interconnect protocol into the die-to-die interface protocol of the die-to-die interconnect 161. In some embodiments, as previously mentioned, the high-level interconnect protocol can be, for example, the AXI or CXL / PCIe protocol. In some embodiments, the die-to-die interface protocol can be, for example, the UCIe protocol configured to directly connect the die to the host CPU 171 or other die-to-die interface protocols. In some embodiments, the operations performed by the unified memory access tunneling system 116 can be performed by the accelerator die-to-die interface controller 113 in the accelerator 131.

[0046] In some embodiments, as Figure 1B shown, the server system 100 includes an SoC 120, a memory 126, a low-power (LP) memory 121, LP memories 122, 123, and 124. In some embodiments, the SoC 120 includes accelerators 131, 132, 133, 134, and a host CPU 171. In some embodiments, the LP memory 121 includes a dedicated memory 129 and a shared memory 127. In some embodiments, the host CPU 171 includes a CPU core 159, a storage controller 118, a memory management unit (MMU) 111, a host CPU interface controller 112, Low Power Double Data Rate 5 (LP5) interconnects 151, 152, 153, 154, 155, 156, 157, 158, die-to-die interconnects 161, 162, 163, 164, an external input / output (I / O) interconnect 168, and an external I / O interconnect 169.

[0047] In some embodiments, the MMU 111 of the host CPU 171 is a memory management unit configured to receive virtual addresses provided to the host CPU 171 from the accelerator 130 or other devices or components external to the host CPU 171 in the server system 100 (e.g., via unified memory access tunneling packets).

[0048] In some embodiments, the memory controller 118 of the host CPU 171 is a memory controller configured to control access to the unified memory 160. In some embodiments, the memory controller 118 can be implemented in hardware, firmware, software, or any combination thereof. In some embodiments, the memory controller 118 is configured to read data from and write data to the unified memory 160.

[0049] In some embodiments, the host CPU interface controller 112 is a component in the host CPU 171 that, in addition to performing unified memory access de-tunneling operations using the unified memory access de-tunneling system 117 described herein, is configured to control communication between the host CPU 171 and other devices or subsystems within the SoC 120.

[0050] In some embodiments, the LP5 interconnects 151 to 158 of the host CPU 171 are LPDDR5 interconnects that connect storage devices (e.g., LP memories 121 to 124) to the SoC 120 using the Low Power Double Data Rate 5 (LPDDR5) technology standard. In some embodiments, the LPDDR5 standard defines the interface protocol (e.g., AXI protocol) utilized by the LP5 interconnects and the electrical signal characteristics for LPDDR5 storage devices, including signal voltage levels, timing, and bus width.

[0051] In some embodiments, the inter-die interconnects 161 to 164 of the host CPU 171 are interconnects configured to enable communication and data transfer between two or more separate integrated circuits (dies) within a single package or the SoC 120. In some embodiments, the inter-die interconnects 161 to 164 are UCIe interconnects or some other type of inter-die interconnect configured to operate according to an inter-die interconnect standard.

[0052] In some embodiments, the external I / O interconnect 168 of the host CPU 171 is, for example, a CXL interconnect or some other type of external I / O interconnect. In some embodiments, the external I / O interconnect 169 is, for example, a PCIe interconnect utilized in the SoC or some other type of external I / O interconnect.

[0053] In some embodiments, accelerator 131 is coupled to host CPU 171 via inter-die interconnect 161. In some embodiments, accelerator 132 is coupled to host CPU 171 via inter-die interconnect 162. In some embodiments, accelerator 133 is coupled to host CPU 171 via inter-die interconnect 163. In some embodiments, accelerator 134 is coupled to host CPU 171 via inter-die interconnect 164.

[0054] In some embodiments, host CPU 171 is coupled to memory 126 via external I / O interconnect 168. In some embodiments, host CPU 171 may also be coupled to memory 126 or other devices via external I / O interconnect 169. In some embodiments, host CPU 171 is coupled to LP memory 121 via LP5 interconnect 151. Similarly, in some embodiments, host CPU 171 is coupled to LP memories 122 to 124 via LP5 interconnects 152 to 154, respectively. In some embodiments, LP memory 121 is configured to communicate with host CPU 171 via LP5 interconnect 151. Similarly, in some embodiments, LP memories 122 to 124 are configured to communicate with host CPU 171 via LP5 interconnects 152 to 154, respectively.

[0055] In some embodiments, accelerator 131 is configured to communicate with memory 126 via inter-die interconnect 161 and external I / O interconnect 168. In some embodiments, accelerator 131 is configured to communicate with memory 126 via inter-die interconnect 161 and external I / O interconnect 169. In some embodiments, accelerator 131 is configured to communicate with host CPU 171 via inter-die interconnect 161. Similarly, accelerators 132 to 134 are configured to communicate with memory 126 via inter-die interconnects 162 to 164 and external I / O interconnect 168, respectively. In some embodiments, accelerators 132 to 134 are configured to communicate with host CPU 171 via inter-die interconnects 162 to 164 and external I / O interconnect 169, respectively. In some embodiments, accelerators 132 to 134 are configured to communicate with host CPU 171 via inter-die interconnects 162 to 164, respectively.

[0056] In some embodiments, accelerator 131 includes a memory controller 114 and an accelerator inter-die interface controller 113. In some embodiments, memory controller 114 is configured to manage memory access and communication between accelerator 131 and a memory (e.g., unified memory 160), such that accelerator 131 can effectively utilize the memory resources available to SoC 120. In some embodiments, accelerator inter-die interface controller 113 is a component in accelerator 131 that, in addition to performing the unified memory access tunneling operations described herein, is configured to control communication between accelerator 131 and other devices or subsystems within SoC 120. In some embodiments, accelerator inter-die interface controller 113 includes a unified memory access tunneling system 116. In some embodiments, as described above, unified memory access tunneling system 116 is configured to tunnel a high-level interconnect protocol (e.g., AXI protocol, PCIe protocol, or CXL protocol) into an inter-die interface protocol (e.g., UCIe protocol) of an inter-die interconnect (e.g., inter-die interconnect 161).

[0057] In some embodiments, the host CPU interface controller 112 of host CPU 171 includes a unified memory access detunneling system 117. In some embodiments, unified memory access detunneling system 117 is hardware and / or executable code (further described herein) configured to detunnel unified memory access tunnel data packets tunneled by unified memory access tunneling system 116.

[0058] In some embodiments, unified memory access tunneling system 116 and unified memory access detunneling system 117 are jointly configured to allow accelerator 131 to directly access unified memory 160, thereby bypassing the need to use accelerator memory for memory access during accelerator processing.

[0059] In some embodiments, in operation, accelerator 131 (via memory controller 114) initiates a memory access request to host CPU 171 to request access to unified memory 160. In some embodiments, as described above, memory controller 114 is a component within accelerator 131 that manages data flow to and from unified memory 160 and is responsible for generating memory access requests to host CPU 171. In some embodiments, unified memory 160 may include memory 126 (e.g., memory associated with host CPU 171) and / or LP memory 121. In some embodiments, the memory access request generated by memory controller 114 includes a virtual memory address associated with the type of data and memory access (e.g., read or write) requested.

[0060] In some embodiments, since the accelerator 131 directly and immediately accesses the unified memory 160 to perform processing operations, before sending a memory access request to the host CPU 171, the memory controller 114 notifies the unified memory access tunnel system 116 to perform a unified memory access tunnel operation for the memory access request on the inter-die interface protocol of the inter-die interconnect 161.

[0061] In some embodiments, the unified memory access tunnel system 116 receives the notification and the memory access request and starts the following process: performing the tunnel operations required to immediately perform the memory access request on the inter-die interconnect 161. In some embodiments, in order to tunnel the high-level interconnect protocol over the inter-die protocol, the unified memory access tunnel system 116 generates a unified memory access tunnel packet structure that maps to the inter-die interface protocol of the inter-die interconnect 161. In some embodiments, the unified memory access tunnel system 116 generates the unified memory access tunnel packet structure by modifying the inter-die interface protocol packet structure to include additional tunnel fields configured to allow the accelerator 131 to access the unified memory 160 using the high-level interconnect protocol without utilizing the memory associated with the accelerator 131. In some embodiments, after generating the unified memory access tunnel packet structure, the unified memory access tunnel system 116 starts the following process: tunneling the high-level interconnect protocol over the inter-die interface protocol using the modified inter-die interface protocol packet structure to generate a modified inter-die interface packet (e.g., a unified memory access tunnel packet).

[0062] In some embodiments, the unified memory access tunnel system 116 tunnels the high-level interconnect protocol associated with the memory access request over the inter-die interface protocol of the inter-die interconnect 161 by encapsulating the high-level interconnect protocol into a modified inter-die interface packet structure associated with the inter-die interconnect 161. In some embodiments, the unified memory access tunnel system 116 encapsulates the high-level interconnect protocol into the modified inter-die interface packet structure by including the high-level interconnect protocol information in the additional tunnel fields that have been added to the inter-die interface packet structure. In some embodiments, by encapsulating the high-level interconnect protocol (with additional high-level interconnect protocol information) into the modified inter-die interface packet structure associated with the inter-die interconnect 161, the accelerator 131 can forward the unified memory access tunnel packet (e.g., the modified inter-die interface packet) to the unified memory 160 through the inter-die interconnect 161. In some embodiments, the high-level interconnect protocol information can be extracted from the unified memory access tunnel packet and processed according to the high-level interconnect protocol.

[0063] In some embodiments, after tunneling a memory access request into an inter-die interface protocol data packet structure (inter-die interface protocol corresponding to the inter-die interconnect 161), the unified memory access tunneling system 116 provides a modified inter-die interface protocol data packet (e.g., a unified memory access tunneling data packet) to the unified memory access de-tunneling system 117 of the host CPU 171 via the inter-die interconnect 161.

[0064] In some embodiments, the unified memory access de-tunneling system 117 of the host CPU 171 receives a unified memory access tunneling data packet from the unified memory access tunneling system 116. In some embodiments, the unified memory access de-tunneling system 117 is configured to de-tunnel high-level interconnect protocol information from the unified memory access tunneling data packet provided by the accelerator 131 and extract the high-level interconnect protocol information from the unified memory access tunneling data packet provided by the accelerator 131. For example, in some embodiments, the unified memory access de-tunneling system 117 is configured to decode the tunnel fields of the unified memory access tunneling data packet to perform the operations indicated by these fields. For example, the unified memory access de-tunneling system 117 receives a unified memory access tunneling data packet from the unified memory access tunneling system 116 and decodes the write indication tunnel field by evaluating the bits located in the field (e.g., the memory opcode) to determine whether the accelerator 131 is requesting a write operation.

[0065] In some embodiments, when the write indication tunnel field indicates that a write operation is to be performed on the unified memory 160, the associated virtual address is provided to the MMU 111 of the host CPU 171, and the MMU converts the associated virtual address to a physical address of the unified memory 160. In some embodiments, the MMU 111 provides the physical address to the memory controller 114 of the host CPU 171. In some embodiments, the memory controller 118 of the host CPU 171 receives the physical address and determines whether the requested memory associated with the physical address is available in the unified memory 160. In some embodiments, when the memory controller 118 of the host CPU 171 determines that the requested memory is available in the unified memory 160, the host CPU 171 allows the accelerator 131 to write to the requested memory location.

[0066] In some embodiments, the unified memory access de-tunneling system 117 receives a unified memory access tunnel data packet from the unified memory access tunneling system 116 and decodes the read indication tunnel field by evaluating the read indication tunnel field to determine whether the bits located in the field (e.g., memory opcode) indicate that the accelerator 131 is requesting a read operation. In some embodiments, when the read indication tunnel field indicates that a read operation is to be performed by the accelerator 131, the associated virtual address for the read operation is provided to the MMU 111 of the host CPU 171, and the MMU converts the associated virtual address to a physical address of the unified memory 160. In some embodiments, the MMU 111 provides the physical address to the memory controller 118 of the host CPU 171.

[0067] In some embodiments, the memory controller 118 receives the physical address from the MMU 111 and determines whether the requested memory associated with the physical address is available for reading from the unified memory 160. In some embodiments, when the memory controller 114 of the host CPU 171 determines that the requested memory is available in the unified memory 160, the host CPU 171 allows the accelerator 131 to read from the requested memory location. In some embodiments, the memory controller of the host CPU 171 sends the requested data to the memory controller 114 of the accelerator 131 via the inter-die interconnect 161. In some embodiments, the memory controller 114 of the accelerator 131 provides the data to the component in the accelerator 131 that needs the requested data. In some embodiments, with the operations described herein, the accelerator 131 improves existing computer systems because the accelerator 131 can save accelerator power and energy in the accelerator 131 by primarily focusing on accelerator processing operations rather than operations typically associated with locally fetching data from accelerator memory.

[0068] Figure 1C The process flow of the server system 100 is shown in more detail. In some embodiments, Figure 1CThe process flow shown in the figure illustrates that the accelerator 131 utilizes the operations described herein to access the unified memory 160 (e.g., shared memory 127 and / or memory 126) for direct accelerator processing. In some embodiments, referring to process 149, the accelerator 131 is a video die that generates a memory access request to the shared memory 127. In some embodiments, the unified memory access tunnel system 116 tunnels the AXI protocol onto the UCIe protocol of the die-to-die interconnect 161 to generate a unified memory access tunnel data packet. In some embodiments, using the unified memory access de-tunnel system 117, the host CPU 171 de-tunnels the unified memory access tunnel data packet to perform the memory request action decoded by the unified memory access de-tunnel system 117, in which case the unified memory access de-tunnel system accesses data from the shared memory 127.

[0069] Similarly, referring to process 148, the accelerator 131 generates a memory access request to the memory 126. In some embodiments, the unified memory access tunnel system 116 tunnels the CXL protocol onto the UCIe protocol of the die-to-die interconnect 161 to generate a unified memory access tunnel data packet. In some embodiments, using the unified memory access de-tunnel system 117, the host CPU 171 de-tunnels the unified memory access tunnel data packet to perform the memory request action decoded by the unified memory access de-tunnel system 117. In some embodiments, by utilizing the operations described herein, the accelerator 131 is able to immediately and directly perform processing operations using the data accessed at the memory 126.

[0070] In some embodiments, processes 165 and 166 illustrate that the CPU core 159 accesses the memory 126 and the shared memory 127 respectively using the shared address space of the unified memory 160. As shown, the CPU core 159 and the accelerator 131 are able to access the same unified memory 160 memory using the shared address space described herein.

[0071] Figure 1DIllustrates a unified memory access method 179 according to some embodiments. The methods, process steps, or stages shown in the figures may be implemented as separate routines or processes, or as part of a larger routine or process. It should be noted that each process step or stage depicted may be implemented as other embodiments such as an apparatus, method, or system of a processor that executes a set of instructions. In some embodiments, at operation 185, the accelerator 131 generates a unified memory access tunnel packet structure that maps to the inter-die interface protocol of the inter-die interconnect 161. In some embodiments, at operation 187, the accelerator 131 tunnels the high-level interconnect protocol over the inter-die interface protocol using the unified memory access tunnel packet structure to generate a unified memory access tunnel packet. In some embodiments, at operation 189, the accelerator 131 accesses the unified memory 160 using the unified memory access tunnel packet, thereby allowing the accelerator 131 to perform processing operations immediately without an accelerator memory for accelerator processing operations.

[0072] Figure 2 Further illustrates according to some embodiments Figure 1A Block diagram of the SoC 120 in. In some embodiments, the SoC 120 includes a host CPU 171 coupled to the accelerator 131 via an inter-die interconnect 161. In some embodiments, the host CPU 171 includes a host CPU-based coprocessor unit 271, and the accelerator 131 includes an accelerator-based coprocessor unit 272. In some embodiments, the accelerator-based coprocessor unit 272 is the hardware and / or executable code in the accelerator 131 configured to perform the operations of the specified accelerator described herein. In some embodiments, the operations of the specified accelerator are operations to be performed by the accelerator-based coprocessor unit 272 specified by the host CPU 171 and / or the accelerator 131. For example, in some embodiments, the accelerator 131 may be a video codec converter, and the accelerator-based coprocessor unit 272 is configured to perform video codec conversion operations, such as the decoding operations, preprocessing operations, encoding operations, and postprocessing operations described and illustrated by way of example in Figure 3 The decoding operation, preprocessing operation, encoding operation, and postprocessing operation shown in.

[0073] In some embodiments, the host-CPU-based coprocessor unit 271 is hardware and / or executable code within the host CPU 171 that is configured to perform the operations of the designated host CPU described herein. In some embodiments, the operations of the designated host CPU are operations to be performed by the host-CPU-based coprocessor unit 271 as designated by the host CPU 171 and / or the accelerator 131. For example, in some embodiments, the host-CPU-based coprocessor unit 271 may be the following hardware and / or executable code within the host CPU 171: the hardware and / or executable code is configured to perform the demultiplexing operation, decoding operation, preprocessing operation, encoding operation, and multiplexing operation described herein and shown by way of example in Figure 3 In some embodiments, as further described in the text reference Figures 3 to 5 , the host-CPU-based coprocessor unit 271 and the accelerator-based coprocessor unit 272 are jointly configured to perform the processing operations routinely performed by the accelerator 131.

[0074] Figure 3 is a block diagram showing the Figure 2 SoC 120 in accordance with some embodiments. As previously described, in some embodiments, the host CPU 171 and the accelerator 131 of the SoC 120 are configured to utilize the host-CPU-based coprocessor unit 271 and the accelerator-based coprocessor unit 272 to perform the processing operations routinely performed by the accelerator 131 for the server system 100. In some embodiments, utilizing the host CPU 171 and the accelerator 131 to perform coprocessing operations (e.g., decoding operations, preprocessing operations, and encoding operations) allows the SoC 120 to maximize the efficiency of the server system 100 by performing operations that are more appropriately configured for the hardware and / or software located within the host CPU 171 or the accelerator 131.

[0075] In some embodiments, the host-CPU-based coprocessor unit 271 includes a demultiplexer 311, a decoder 312, a preprocessing unit 313, an encoder 314, and a multiplexer 315. In some embodiments, the accelerator-based coprocessor unit 272 includes a decoder 331, a preprocessing unit 332, an encoder 333, and a postprocessing unit 334. In some embodiments, the demultiplexer 311, the decoder 312, the preprocessing unit 313, the encoder 314, and the multiplexer 315 of the host-CPU-based coprocessor unit 271 are jointly configured to perform operations together with the decoder 331, the preprocessing unit 332, the encoder 333, and the postprocessing unit 334 of the accelerator-based coprocessor unit 272 to perform the operations described herein.

[0076] In some embodiments, the demultiplexer 311 is based on the following hardware and / or executable code in the coprocessor unit 271 of the host CPU: the hardware and / or executable code is configured to receive the input data stream 340 and divide the input data stream 340 into a plurality of output data streams defined by the host CPU 171 and / or the accelerator 131. In some embodiments, the input data stream 340 may be, for example, a digital video data stream that has been multiplexed by a video source. In some embodiments, the demultiplexer 311 is configured to divide the input data stream 340 into: (1) a data stream 341 directed to the host CPU decoder that is configured to be decoded by the decoder 312 of the host CPU 171; and (2) a data stream 342 directed to the accelerator decoder that is configured to be decoded by the decoder 331 of the accelerator 131.

[0077] In some embodiments, the data stream 341 directed to the host CPU decoder is a data stream configured for decoding operations to be performed by the decoder 312 (which may be, for example, a software-based decoder configured to perform software-based decoding operations). In some embodiments, the data stream 341 directed to the host CPU decoder may be a data stream that requires software-based decoding operations that can only be performed by the decoder 312. For example, since the decoder 331 of the accelerator 131 may be a hardware decoder configured to decode a specific type of hardware-specific data stream, when the input data stream (or a portion thereof) is not such a type of input data stream that can be decoded by the decoder 331 (e.g., a hardware-based decoder), the input data stream may be provided by the demultiplexer 311 to the decoder 312 (e.g., a software-based decoder) as the data stream 341 directed to the host CPU decoder. In some embodiments, portions of the input data stream 340 may be specified by the host CPU 171 and / or the accelerator 131 as the data stream 341 directed to the host CPU decoder or the data stream 342 directed to the accelerator decoder. In some embodiments, the host CPU 171 may use a selection signal provided to the demultiplexer 311 to indicate to the demultiplexer 311 the portion of the input data stream 340 that is designated for decoding by the decoder 312 of the host CPU 171 (e.g., the data stream 341 directed to the host CPU decoder) or the portion of the input data stream 340 that is designated for decoding by the decoder 331 of the accelerator 131. In some embodiments, after performing the demultiplexing operation at the demultiplexer 311, the demultiplexer 311 provides the data stream 341 directed to the host CPU decoder to the decoder 312 of the host CPU 171 and provides the data stream 342 directed to the accelerator decoder (e.g., the data stream 342 directed to the accelerator decoder) to the decoder 331 of the accelerator 131.

[0078] In some embodiments, referring to decoder 312 of host CPU 171, decoder 312 receives data stream 341 directed to the host CPU decoder from demultiplexer 311 and begins the process of decoding data stream 341 directed to the host CPU decoder. In some embodiments, decoder 312 is such a software decoder, hardware decoder, or a combination thereof: the software decoder, hardware decoder, or a combination thereof is configured to perform decoding operations specific to data stream 341 directed to the host CPU decoder provided by demultiplexer 311 (e.g., software-specific data streams that cannot be decoded by decoder 331 due to the hardware configuration of decoder 331). For example, due to the fixed hardware configuration of decoder 331 and the reconfigurable software configuration of decoder 312, decoder 312 can be configured to perform operations specific to data stream 341 directed to the host CPU decoder. In some embodiments, decoder 312 is a decoder configured to perform decoding operations specific to the processing attributes of host CPU 171 and / or decoding operations specific to the non-processing attributes of accelerator 131. In some embodiments, decoder 312 is configured to perform video decoding operations specific to the video data stream provided by demultiplexer 311 from host CPU 171. In some embodiments, after the decoding operation is performed at decoder 312, decoder 312 provides the decoded output data stream 344 to preprocessing unit 332 to preprocess the decoded output data stream 344.

[0079] In some embodiments, referring to decoder 331 of accelerator 131, decoder 331 receives data stream 342 directed to the accelerator decoder from demultiplexer 311 and begins the process of decoding data stream 342 directed to the accelerator decoder. In some embodiments, decoder 331 is a hardware decoder, software decoder, or a combination thereof configured to perform decoding operations specific to accelerator 131. In some embodiments, decoder 331 is a decoder configured to perform decoding operations specific to data stream 342 directed to the accelerator decoder provided by demultiplexer 311 from host CPU 171. For example, in some embodiments, due to the fixed hardware configuration of decoder 331, decoder 331 can be configured to decode only data streams mapped to the fixed hardware configuration of decoder 331. In some embodiments, decoder 331 is a video decoder configured to perform video decoding operations specific to the video data stream (e.g., data stream 342 directed to the accelerator decoder) provided by demultiplexer 311 from host CPU 171. In some embodiments, after the decoding operation is performed at decoder 331, decoder 331 provides the decoded output data stream 343 to preprocessing unit 332 to preprocess the decoded output data stream 343.

[0080] In some embodiments, the preprocessing unit 332 receives the decoded output data stream 343 from the decoder 331 and the decoded output data stream 344 from the decoder 312, and begins the process of performing a shared preprocessing operation with the preprocessing unit 313 of the host CPU 171. In some embodiments, the preprocessing unit 332 is hardware and / or executable code located in the accelerator 131, and the hardware and / or executable code is configured to perform the following operations: (1) evaluate the received decoded data stream to determine whether the received input data stream is configured as an accelerator-specific preprocessing data stream or a host-CPU-specific preprocessing data stream; (2) perform accelerator-specific preprocessing operations; and (3) share host-CPU-specific preprocessing operations with the preprocessing unit 313 of the host-CPU-based coprocessor unit 271. In some embodiments, the preprocessing unit 313 is hardware and / or executable code located in the host CPU 171, and the hardware and / or executable code is configured to perform host-CPU-specific processing operations on the received host-CPU-specific preprocessing data stream 346 from the accelerator 131. In some embodiments, the accelerator-specific preprocessing data stream is a data stream configured to be preprocessed by the preprocessing unit 332 of the accelerator 131. In some embodiments, the host-CPU-specific preprocessing data stream is a data stream configured to be preprocessed by the preprocessing unit 313 of the host CPU 171.

[0081] In some embodiments, the preprocessing unit 332 receives the decoded output data stream 343 and the decoded output data stream 344, and determines whether the received decoded data stream (or a portion thereof) is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the preprocessing unit 332 determines whether the received decoded data stream is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream by evaluating the preprocessing operation configuration associated with the received decoded data stream. In some embodiments, the preprocessing operation configuration serves as an indication of whether the received decoded data stream is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the preprocessing operation configuration can be evaluated by identifying a data stream identification (ID) in the received decoded data stream. In some embodiments, the data stream ID is a unique identifier used to identify and manage the data stream, and in this case, the data stream ID is associated with being an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the data stream ID can be assigned by the operating system of the SoC 120 and / or the accelerator 131 when the data stream is created, and is used by the accelerator 131 to identify and manage the data stream. In some embodiments, the accelerator 131 utilizes the data stream ID to schedule the preprocessing of the decoded data stream and determine whether to switch between multiple decoded data streams for preprocessing by the host CPU 171 or the accelerator 131 (and allocate resources such as memory and processing time for each data stream). In some embodiments, the data stream ID is either mapped to accelerator-specific operations configured to be executed by the preprocessing unit 332 of the accelerator 131 or mapped to host CPU-specific operations configured to be executed by the preprocessing unit 313.

[0082] In some embodiments, when the decoded data stream is identified by the preprocessing unit 332 as an accelerator-specific preprocessing data stream, the accelerator-specific preprocessing data stream remains at the preprocessing unit 332 for accelerator-specific preprocessing. In some embodiments, the preprocessing unit 332 uses accelerator-specific preprocessing operations to preprocess the accelerator-specific processing data stream to generate an accelerator-specific preprocessed output data stream 345.

[0083] In some embodiments, when the data stream is recognized as a host CPU specific preprocessing data stream by the preprocessing unit 332 after decoding, the preprocessing unit 332 provides this data stream as the host CPU specific preprocessing data stream 346 to the preprocessing unit 313 for host CPU specific preprocessing. In some embodiments, the preprocessing unit 313 preprocesses the host CPU specific preprocessing data stream 346 using host CPU specific preprocessing operations to generate a preprocessed output data stream 347. In some embodiments, the preprocessing unit 313 provides the preprocessed output data stream 347 to the preprocessing unit 332. In some embodiments, the preprocessing unit 332 receives the preprocessed output data stream 347 and provides the preprocessed output data stream 347 and the accelerator specific preprocessed output data stream 345 together as the preprocessed output data stream 348 to the encoder 333.

[0084] In some embodiments, the encoder 333 receives the preprocessed output data stream 348 from the preprocessing unit 332 and starts the process of performing a shared encoding operation with the encoder 314 of the host CPU 171. In some embodiments, the encoder 333 is a hardware and / or software encoder configured to perform the following operations: (1) evaluate the preprocessed output data stream 348 to identify an accelerator specific encoding data stream and a host CPU specific encoding data stream; (2) perform encoding operations specific to the encoder 333 (e.g., accelerator specific encoding operations); and (3) share the host CPU specific encoding operations with the encoder 314 of the host CPU - based coprocessing unit 271. In some embodiments, the encoder 314 is a software encoder and / or hardware encoder in the host CPU 171 that is configured to perform host specific encoding operations on the host CPU specific encoding data stream 349 provided by the encoder 333. For example, in some embodiments, the encoder 333 is a video encoder configured to perform accelerator specific video encoding operations with a fixed hardware configuration specific to the encoder 333. In some embodiments, the encoder 314 is a software video encoder configured to perform host CPU specific video encoding operations that: (1) cannot be performed by the encoder 333, for example, due to the fixed configuration of the encoder 333; or (2) can be performed more efficiently by the encoder 314 using the unique processing capabilities of the host CPU 171.

[0085] In some embodiments, the encoder 333 receives the pre - processed output data stream 348 from the pre - processing unit 332 and evaluates the pre - processed output data stream 348 to identify the accelerator - specific encoded data stream and the host CPU - specific encoded data stream. In some embodiments, the encoder 333 identifies the accelerator - specific encoded data stream or the host CPU - specific encoded data stream in the data stream by searching for specific tags in the pre - processed output data stream 348 that indicate whether a portion of the pre - processed output data stream 348 is an accelerator - specific encoded data stream or a host CPU - specific encoded data stream. In some embodiments, for example, the encoder 333 searches for the accelerator - specific encoded data stream tag and the host CPU - specific encoded data stream tag in the pre - processed output data stream 348.

[0086] In some embodiments, when the encoder 333 identifies the pre - processed output data stream 348 or a portion thereof as an accelerator - specific encoded data stream, the encoder 333 encodes the accelerator - specific encoded data stream at the encoder 333 to generate the accelerator - specific encoded output data stream 336. In some embodiments, when the encoder 333 identifies a portion of the pre - processed output data stream 348 as the host CPU - specific encoded data stream 349, the encoder 333 provides the host CPU - specific encoded data stream 349 to the encoder 314. In some embodiments, the encoder 314 receives the host CPU - specific encoded data stream 349 and encodes the host CPU - specific encoded data stream 349 using the host CPU - specific encoding operation provided by the encoder 314 of the host CPU 171. In some embodiments, after performing the host CPU - specific encoding operation, the encoder 314 provides the encoded output as the host CPU - specific encoded output 361 to the encoder 333. In some embodiments, the encoder 333 receives the host CPU - specific encoded output 361 from the encoder 314 and provides the host CPU - specific encoded output 361 and the accelerator - specific encoded output data stream 336 together as the encoded - pre - processed output data stream 365 to the post - processing unit 334.

[0087] In some embodiments, the post - processing unit 334 receives the encoded - pre - processed output data stream 365 and performs post - processing operations on the encoded - pre - processed output data stream 365. In some embodiments, the post - processing unit 334 is hardware and / or executable code configured to perform post - processing operations on the encoded - pre - processed output data stream 365 of the encoder 333, and the post - processing operations are, for example, data compression, error correction, or other post - processing operations. In some embodiments, the post - processing unit 334 provides the post - processed data stream 368 to the multiplexer 315 for further processing or storage by the SoC 120.

[0088] Figure 4is a block diagram showing a shared preprocessing flow of the SoC 120 according to some embodiments. In some embodiments, as Figure 4 As shown in the preprocessing flow in

[0089] the preprocessing unit 332 includes an accelerator-specific preprocessing unit 411 and an accelerator-specific preprocessing unit 413. In some embodiments, the preprocessing unit 313 includes a host CPU-specific preprocessing unit 412. In some embodiments, the accelerator-specific preprocessing unit 411 is hardware and / or executable code within the preprocessing unit 332 configured to perform accelerator-specific preprocessing operations on an accelerator-specific preprocessing data stream identified by the preprocessing unit 332. In some embodiments, the host CPU-specific preprocessing unit 412 is hardware and / or executable code within the preprocessing unit 313 configured to perform host CPU-specific preprocessing operations on a host CPU-specific preprocessing data stream identified by the preprocessing unit 332. In some embodiments, the accelerator-specific preprocessing unit 413 is hardware and / or executable code within the preprocessing unit 332 that, in addition to being configured to perform accelerator-specific preprocessing operations on an accelerator-specific preprocessing data stream identified by the preprocessing unit 332, is further configured to perform additional preprocessing operations on the output of the host CPU-specific preprocessing unit 412 (e.g., the preprocessed output data stream 347) and / or combine the output of the host CPU-specific preprocessing unit 412 with an accelerator-specific preprocessed output data stream (e.g., the accelerator-specific preprocessed output data stream 345).

[0090] Figure 5 is a block diagram showing a shared encoding flow of the SoC 120 according to some embodiments. In some embodiments, as Figure 5As shown in the shared encoding process, encoder 333 includes accelerator-specific encoding unit 511 and accelerator-specific encoding unit 513. In some embodiments, encoder 314 includes host CPU-specific encoding unit 512. In some embodiments, accelerator-specific encoding unit 511 is hardware and / or executable code within encoder 333 that is configured to perform accelerator-specific encoding operations on accelerator-specific encoding data streams identified by encoder 333. In some embodiments, host CPU-specific encoding unit 512 is hardware and / or executable code within encoder 314 that is configured to perform host CPU-specific encoding operations on host CPU-specific encoding data streams identified by encoder 333. In some embodiments, accelerator-specific encoding unit 513 is hardware and / or executable code within encoder 333 that, in addition to being configured to perform accelerator-specific encoding operations on accelerator-specific encoding data streams identified by encoder 333, is also configured to perform additional encoding operations on the output of host CPU-specific encoding unit 512 (e.g., host CPU-specific encoded output 361) and / or to combine the output of host CPU-specific encoding unit 512 with the accelerator-specific encoded output data stream (e.g., accelerator-specific encoded output data stream 336).

[0091] As part of the shared encoding process, encoder 333 identifies accelerator-specific encoding data streams and encodes the accelerator-specific encoding data streams at accelerator-specific encoding unit 511. In some embodiments, encoder 333 identifies host CPU-specific encoding data streams and provides the host CPU-specific encoding data streams for host CPU-specific encoding at host CPU-specific encoding unit 512. In some embodiments, encoder 333 receives the encoded output data stream from host CPU-specific encoding unit 512 and combines the encoded output data stream with the accelerator-specific encoded output data stream.

[0092] In some embodiments, existing computer systems are improved using the operations described herein because the SoC 120 is capable of dynamically switching between the hardware decoder in the accelerator 131 and the software decoder in the host CPU 171 to avoid hardware codec issues (e.g., new codec configurations, error concealment). In some embodiments, multiple preprocessing operations are dynamically partitioned between the preprocessor of the host CPU 171 and the preprocessor of the accelerator 131 to achieve more flexible algorithms and power usage. In some embodiments, critical decisions of the encoder (e.g., mode decision or frame parameters of the convex-hull approach) can be executed on the host CPU 171 rather than on the accelerator 131 to obtain better video quality and bitrate. In some embodiments, the fine-grained interaction between the host CPU 171 and the accelerator 130 (e.g., accelerator hardware) enables the framework described herein to be used for improved co-processing.

[0093] Figure 6 which shows an SoC 120 according to some embodiments Figure 1A in the block diagram. In some embodiments, the SoC 120 includes a host CPU 171, a memory 104, accelerators 131 to 134, and a flash device 140. In some embodiments, the flash device 140 is a storage device configured to store accelerator firmware associated with the accelerators 131 to 134 and host CPU firmware associated with the host CPU firmware. In some embodiments, the accelerator firmware can be stored in the form of an accelerator firmware image file, and the host CPU firmware can be stored in the form of a host CPU firmware image file. In some embodiments, the flash device 140 can be, for example, an embedded Multi-Media Controller (EMMC) device or a Universal Flash Storage (UFS) device.

[0094] In some embodiments, as described above, the accelerators 131 to 135 are coupled to the host CPU 171 using the inter-die interconnects 161 to 164. In some embodiments, each of the inter-die interconnects 161 to 164 includes a sideband bus and an inter-die bus (mainband bus) (e.g., Figure 7The inter-die bus 770 and the sideband bus 790 shown in []. In some embodiments, the sideband bus is a set of dedicated communication lines that are configured to transmit control and management information between the accelerators 131 to 134 and the host CPU 171, and are further configured to transmit accelerator firmware components (e.g., critical accelerator firmware components) between the accelerators 131 to 134 and the host CPU 171 based on the evaluation of the accelerator firmware components performed by the root of trust (ROT) 620 of the host CPU 171. In some embodiments, the inter-die bus (or mainband bus) is a set of communication lines optimized for high-speed data transmission that are configured to transmit data and payloads between the accelerators 131 to 134 and the host CPU 171, and are further configured to transmit accelerator firmware components that are considered non-critical firmware components, for example, between the host CPU 171 and the accelerators 131 to 134.

[0095] In some embodiments, the host CPU 171 includes a root of trust (ROT) 620. In some embodiments, the ROT 620 is a secure hardware module and / or executable code or trusted execution environment (TEE) within the host CPU 171 that is configured to instantaneously authenticate accelerator firmware based on the accelerator firmware evaluation of the accelerator firmware in addition to performing traditional root of trust operations in a trusted computing environment. In some embodiments, the accelerator firmware evaluation performed by the ROT 620 includes, for example, determining whether each part of the accelerator firmware associated with the accelerators coupled to the host CPU 171 is a critical accelerator firmware component or a non-critical accelerator firmware component of the accelerator firmware. In some embodiments, as further described herein Figures 7 to 10 Based on the results of the accelerator firmware evaluation, the host CPU 171 provides the associated accelerators with the certified critical accelerator firmware components for processing via the sideband bus that connects the accelerators to the host CPU 171, and provides the associated accelerators with the certified non-critical accelerator firmware components for processing via the inter-die bus that connects the accelerators to the host CPU 171.

[0096] Figure 7 shows according to some embodiments Figure 6Block diagram of the SoC 120 in. In some embodiments, the host CPU 171 is coupled to the flash device 140, the memory 104, and the accelerators 131 to 134. In some embodiments, the host CPU 171 is coupled to the accelerators 131 to 134 via the inter-die bus 770 and the sideband bus 790. In some embodiments, the accelerators 131 to 134 include accelerator embedded u-controllers and accelerator memories. For example, in some embodiments, the accelerator 131 may include an accelerator embedded u-controller 740 and an accelerator memory 741. In some embodiments, the accelerator embedded u-controller 740 is an embedded controller in the accelerator 131, and the embedded controller is configured to coordinate the data flow between the accelerator 131 and other components of the SoC 120. In some embodiments, the accelerator embedded u-controller 740 is configured to utilize the authentication control unit 781, and the authentication control unit may be configured to control the operations associated with the authentication operations performed by the host CPU 171 within the accelerator 131.

[0097] In some embodiments, the ROT 620 includes a security agent 721, and the security agent is configured to perform the sideband-based accelerator firmware authentication method described herein by using the accelerator firmware identification unit 756, the accelerator firmware authentication unit 752, the accelerator firmware parsing unit 753, and / or the accelerator firmware component size determination unit 754. In some embodiments, the accelerator firmware parsing unit 753 is hardware and / or executable code that is configured to parse or partition the accelerator firmware into multiple accelerator firmware components by examining the code structure of the accelerator firmware to identify the unique functional components or modules of the accelerator firmware and splitting the accelerator firmware into each uniquely identified functional component or module (e.g., accelerator firmware component).

[0098] In some embodiments, the accelerator firmware identification unit 756 is hardware and / or executable code configured to identify accelerator firmware components parsed by the accelerator firmware parsing unit 753 as critical accelerator firmware components (e.g., accelerator firmware components critical for performing processing operations of the accelerator 131) and non-critical accelerator firmware components (e.g., accelerator firmware components not critical for performing processing operations of the accelerator 131). In some embodiments, the accelerator firmware identification unit 756 identifies critical non-accelerator components and non-critical accelerator firmware components of the accelerator firmware based on accelerator-specific information stored in the accelerator firmware image file. In some embodiments, the accelerator-specific information may be included in the form of headers, sections, symbols, or other metadata that define the structure and organization of the accelerator firmware. For example, in some embodiments, in the Executable and Linkable Format (ELF) or the Common Object File Format (COFF), the accelerator firmware image file may include sections and symbols that define the various components of the firmware and their roles in the overall system. In some embodiments, the accelerator firmware identification unit 756 may utilize headers and metadata associated with these sections and symbols to identify critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware. In some embodiments, the accelerator firmware identification unit 756 may utilize accelerator-specific information to determine which components are critical and which are not based on the specific requirements of the accelerator 131 and the SoC 120. For example, the accelerator firmware identification unit 756 may identify the bootloader, drivers, and low-level software as critical accelerator firmware components because these accelerator firmware components may be required for the proper operation of the accelerator 131 and the overall system (e.g., the SoC 120). In some embodiments, the host CPU 171 may identify applications, libraries, and high-level software as non-critical accelerator firmware components because these accelerator firmware components provide additional functionality but are not strictly required for the operation of the accelerator.

[0099] In some embodiments, the accelerator firmware component size determination unit 754 is hardware and / or executable code configured to determine the accelerator firmware component size of the accelerator firmware and the size of the accelerator firmware. In some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 is configured to determine the accelerator firmware component size of the accelerator firmware and the size of the accelerator firmware by evaluating size information (such as a table of headers or content) provided by the accelerator firmware itself to determine the size and location of each component. In some embodiments, the size information may be included in the firmware image of the accelerator firmware and may be used by the ROT 620 to partition the firmware into discrete accelerator firmware components.

[0100] In some embodiments, the accelerator firmware authentication unit 752 is hardware and / or executable code configured to perform accelerator firmware authentication operations for accelerators 131 to 134. In some embodiments, the accelerator firmware authentication unit 752 includes an immediate accelerator firmware component authentication unit 761 and a delayed accelerator firmware component authentication unit 762. In some embodiments, the immediate accelerator firmware component authentication unit 761 is hardware and / or executable code configured to receive an accelerator firmware component (e.g., a critical accelerator firmware component) and immediately authenticate the accelerator firmware component. In some embodiments, immediate authentication of performance-critical firmware means that critical accelerator firmware components are authenticated immediately or instantaneously without latency at the host CPU 171 such that the authenticated accelerator firmware can be provided directly to the associated accelerator via a sideband bus (e.g., sideband bus 791 for accelerator 131) coupled to the associated accelerator.

[0101] In some embodiments, the delayed accelerator firmware component authentication unit 762 is hardware and / or executable code configured to authenticate accelerator components (e.g., non-critical accelerator firmware components) at a delayed time indicated or commanded by the host CPU 171. For example, in some embodiments, delayed authentication refers to authentication performed by the delayed accelerator firmware component authentication unit 762, which is delayed by the ROT 620 such that non-critical accelerator firmware components are authenticated only after the critical accelerator firmware components have been authenticated by the immediate accelerator firmware component authentication unit 761.

[0102] In some embodiments, as described above, the security agent 721 is configured to utilize the accelerator firmware identification unit 756, the accelerator firmware authentication unit 752, the accelerator firmware parsing unit 753, and / or the accelerator firmware component size determination unit 754 to perform the sideband-based accelerator firmware authentication method described herein. In some embodiments, the operation of the SoC 120 is described below with reference to Figures 8 to 10 the operation of the SoC 120 is described.

[0103] Figure 8FIG. 0 is a flowchart showing a sideband-based accelerator firmware authentication method 800 according to some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 800 is configured to authenticate the accelerator firmware immediately (e.g., instantaneously) or non-immediately (e.g., non-instantaneously) based on an accelerator firmware evaluation performed on the accelerator firmware by the security agent 721 of the ROT 620. In some embodiments, as part of the accelerator firmware evaluation, the security agent 721 divides the accelerator firmware into a plurality of accelerator firmware components, and these accelerator firmware components are considered critical accelerator firmware components or non-critical accelerator firmware components. In some embodiments, the accelerator firmware components considered critical are provided to the associated accelerator via the sideband bus of the inter-die interconnect for processing. In some embodiments, the accelerator firmware considered non-critical is provided to the associated accelerator via the inter-die bus of the inter-die interconnect for processing. The methods, process steps, or phases shown in the figures may be implemented as independent routines or processes, or as part of a larger routine or process. It should be noted that each of the depicted process steps or phases may be implemented as other embodiments such as an apparatus, method, or system including a processor that executes a set of instructions.

[0104] In some embodiments, at operation 810, the security agent 721 of the ROT 620 reads the accelerator firmware from the flash device 140 (e.g., non-volatile memory). In some embodiments, as described above, the accelerator firmware read from the flash device 140 may be associated with a specific accelerator (e.g., accelerator 131, etc.) and may be stored in the flash device 140 in the form of an accelerator firmware image file. In some embodiments, reading the accelerator firmware from the flash device 140 occurs during the system bootup of the SoC 120. In some embodiments, once the accelerator firmware is read from the flash device 140, the accelerator firmware is provided to the accelerator firmware parsing unit 753.

[0105] In some embodiments, at operation 815, the accelerator firmware parsing unit 753 of the security agent 721 receives the accelerator firmware from the flash device 140 and parses the accelerator firmware into a plurality of accelerator firmware components. In some embodiments, the accelerator firmware parsing unit 753 parses the accelerator firmware into a plurality of accelerator firmware components by examining the code structure of the accelerator firmware to identify the unique functional components or modules of the accelerator firmware, and splitting the accelerator firmware into each uniquely identified functional component or module. In some embodiments, the accelerator firmware parsing unit 753 identifies each accelerator firmware component by scanning the unique digital signature of the accelerator firmware representing each functional component or module. In some embodiments, after parsing the accelerator firmware into a plurality of accelerator firmware components, the accelerator firmware parsing unit 753 provides these accelerator firmware components to the accelerator firmware identification unit 756, and operation 815 proceeds to operation 820.

[0106] In some embodiments, at operation 820, the accelerator firmware identification unit 756 receives the plurality of accelerator firmware components from the accelerator firmware parsing unit 753 and evaluates the plurality of accelerator firmware components of the accelerator firmware to identify the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware. In some embodiments, the accelerator firmware identification unit 756 identifies the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware by analyzing the accelerator firmware metadata and other accelerator firmware code associated with each accelerator firmware component. For example, in some embodiments, the accelerator firmware identification unit 756 identifies the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware by analyzing the dependencies (e.g., mutual dependencies) of the accelerator firmware components, analyzing the metadata associated with each accelerator firmware component, and analyzing the previous versions of the accelerator firmware and each accelerator firmware component. For example, in some embodiments, the accelerator firmware identification unit 756 checks the dependencies between the accelerator firmware components in the accelerator firmware by determining which accelerator firmware components are necessary (e.g., critical) for the normal operation of other accelerator firmware components and which components are not necessary (e.g., non-critical) for the normal operation of other components. In some embodiments, the accelerator firmware identification unit 756 identifies the dependencies of different accelerator firmware components by scanning the accelerator firmware code for any inter-component communication mechanisms and examining the inter-component communication mechanisms to determine the type of inter-component communication dependencies (e.g., function calls or shared data structures). In another example, in some embodiments, the accelerator firmware identification unit 756 uses the metadata to identify the critical accelerator firmware components and non-critical accelerator firmware components by scanning the metadata to find version numbers or annotations associated with each accelerator firmware component that indicate the importance (e.g., critical or non-critical) of the accelerator firmware component.

[0107] In some embodiments, the accelerator firmware identification unit 756 identifies critical and non-critical accelerator firmware components by comparing the current version of the accelerator firmware components with the previous version of the accelerator firmware components to identify any changes in the accelerator firmware components. In some embodiments, no change from the previous version of the accelerator firmware components to the current version of the accelerator firmware components may indicate that the accelerator firmware components are not critical accelerator firmware components, while a change from the previous version of the accelerator firmware components to the current version of the accelerator firmware components may indicate that the accelerator firmware components are critical accelerator firmware components. In some embodiments, after the accelerator firmware identification unit 756 identifies a component as a non-critical or critical accelerator firmware component, the accelerator firmware identification unit 756 provides the non-critical accelerator firmware components to the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752, and provides the critical accelerator firmware components to the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752. In some embodiments, before providing the non-critical accelerator firmware components to the delayed accelerator firmware component authentication unit 762, the non-critical accelerator firmware components may be stored in a secure area (e.g., a secure boot ROM) of the memory of the ROT 620.

[0108] In some embodiments, at operation 825, the in-time accelerator firmware component authentication unit 761 receives the critical accelerator firmware component and authenticates the critical accelerator firmware component in real time. In some embodiments, as described above, in-time authentication of performance-critical firmware means authenticating the critical accelerator firmware component immediately without latency at the host CPU 171, such that the authenticated accelerator firmware can be directly provided to the associated accelerator via the sideband bus (e.g., sideband bus 791 for accelerator 131) coupled to the associated accelerator. In some embodiments, the accelerator firmware authentication unit 752 is configured to format the packet structure of the accelerator firmware component data packet sent to the associated accelerator via the sideband bus, such that the packet structure indicates to the associated accelerator that an accelerator firmware component is being transmitted via the sideband bus. For example, in some embodiments, the accelerator firmware authentication unit 752 is configured to format the packet structure of the accelerator firmware data packet sent to accelerator 131 via sideband bus 791, such that the packet structure indicates to accelerator 131 that a critical accelerator firmware component is being transmitted in the data packet. In some embodiments, bit positions in the packet structure of the accelerator firmware component data packet can indicate to accelerator 131 that a critical accelerator firmware component is being transmitted in the data packet. In some embodiments, bit positions in the packet structure of the accelerator firmware data packet can indicate to the accelerator that a non-critical accelerator firmware component is being transmitted in via the inter-die bus (e.g., inter-die bus 771 associated with accelerator 131). In some embodiments, after in-time authenticating the accelerator firmware component, operation 825 proceeds to operation 830.

[0109] In some embodiments, at operation 830, the host CPU 171 provides the authenticated critical accelerator firmware component to accelerator 131 via sideband bus 791. In some embodiments, after providing the authenticated critical accelerator firmware component to accelerator 131 via sideband bus 791, operation 830 proceeds to operation 835.

[0110] In some embodiments, at operation 835, accelerator 131 receives the accelerator firmware component via sideband bus 791 and executes the critical accelerator firmware component. In some embodiments, once the accelerator firmware component is received, the authentication control unit 781 of accelerator 131 is configured to scan the data packet for a bit indicator indicating that the received data packet is an accelerator firmware component. In some embodiments, accelerator 131 is configured to scan the data packet for a bit indicator indicating that an associated non-critical accelerator firmware component is being transmitted via inter-die bus 771 for execution by accelerator 131. In some embodiments, after accelerator 131 executes the critical accelerator firmware component, operation 835 proceeds to operation 840.

[0111] In some embodiments, referring back to operation 820, when the accelerator firmware identification unit 756 regards the accelerator firmware component as a non-critical accelerator firmware component, at operation 850, the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 receives the non-critical accelerator firmware component and authenticates the non-critical accelerator firmware component. In some embodiments, the delayed accelerator firmware component authentication unit 762 uses delayed authentication to authenticate the non-critical accelerator firmware component. In some embodiments, delayed authentication refers to the authentication performed by the delayed accelerator firmware component authentication unit 762, which is delayed by the ROT 620 such that the non-critical accelerator firmware component is authenticated after the critical accelerator firmware component has been authenticated by the immediate accelerator firmware component authentication unit 761. In some embodiments, after authenticating the critical accelerator firmware component at the delayed accelerator firmware component authentication unit 762, operation 850 proceeds to operation 855.

[0112] In some embodiments, at operation 855, after authenticating the critical accelerator firmware component at the delayed accelerator firmware component authentication unit 762, the ROT 620 provides the authenticated non-critical accelerator firmware component to the memory 104 for storage. In some embodiments, after storage in the memory 104, operation 855 proceeds to operation 860.

[0113] In some embodiments, at operation 860, the authentication control unit 781 of the accelerator embedded u-controller 740 installs the non-performance-critical firmware from the memory 104 into the accelerator memory 741 of the accelerator 131 via the inter-die bus 771. In some embodiments, operation 860 proceeds to operation 840, at which the authenticated non-performance-critical firmware provided via the inter-die bus 771 is executed by the accelerator 131.

[0114] Figure 9 is a block diagram showing an example process flow utilized in the Figure 7 SoC 120 in accordance with some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 800 is utilized in the example process flow in the Figure 9 and is configured to immediately (e.g., instantaneously) authenticate a first accelerator firmware component (e.g., the bootloader accelerator firmware component) based on an accelerator firmware assessment of the accelerator firmware, and non-immediately (e.g., non-instantaneously) delay the authentication of a second accelerator firmware component (e.g., the body of the accelerator firmware).

[0115] In some embodiments, at step S1, the security agent 721 of ROT 620 reads the accelerator firmware from the flash device 140 at startup. After reading the accelerator firmware from the flash device 140, the accelerator firmware parsing unit 753 and the accelerator firmware identification unit 756 parse the accelerator firmware into a plurality of accelerator firmware components (e.g., a bootloader accelerator firmware component and a main body accelerator firmware component), and identify each accelerator firmware component of the accelerator firmware (e.g., a critical accelerator firmware component and a non-critical accelerator firmware component). In some embodiments, after the accelerator firmware parsing unit 753 and the accelerator firmware identification unit 756 have parsed the accelerator firmware into a bootloader accelerator firmware component and a main body accelerator firmware component, and have identified each accelerator firmware component as a critical accelerator firmware component and a non-critical accelerator firmware component, the accelerator firmware authentication unit 752 immediately authenticates the bootloader accelerator firmware component ("bootloader") at the immediate accelerator firmware component authentication unit 761. In some embodiments, after delaying the authentication of the main body accelerator firmware component of the accelerator firmware (e.g., delaying until the bootloader accelerator firmware component is authenticated), the delayed accelerator firmware component authentication unit 762 authenticates the main body accelerator firmware component. In some embodiments, an unauthenticated accelerator firmware component (e.g., the main body of the accelerator firmware component) can be temporarily stored in a secure area of the memory of ROT 620, which secure area is, for example, a secure boot ROM or a secure enclave within a trusted execution environment (TEE). In some embodiments, the unauthenticated accelerator firmware component is stored in the secure boot memory until the accelerator firmware authentication unit 752 authenticates the unauthenticated accelerator firmware component.

[0116] In some embodiments, at step S2A, immediately following the authentication of the bootloader accelerator firmware component, the security agent 721 provides or pushes the bootloader accelerator firmware component to the accelerator memory 741 of the accelerator 131 via the sideband bus 791. In some embodiments, the security agent 721 directly provides the bootloader accelerator firmware component to the accelerator memory 741 via the sideband bus 791 without the help of a storage controller located in the accelerator 131. In some embodiments, the security agent 721 provides the bootloader to the storage controller in the accelerator 131 via the sideband bus 791 before the bootloader is written to the accelerator memory 741.

[0117] In some embodiments, at step S2B, the security agent 721 writes the main accelerator firmware component of the accelerator firmware into the memory 104 for transmission to the accelerator 131 at step S4. In some embodiments, at step S3, after the interdie bus 771 becomes operable, the host CPU 171 performs device initialization and releases the reset of the accelerator embedded u-controller 740. In some embodiments, releasing the reset of the accelerator embedded u-controller 740 means that the host CPU 171 sends a reset signal to the accelerator embedded u-controller, and this reset signal enables the accelerator 131 to start executing instructions and control the internal operations of the accelerator 131. In some embodiments, releasing the reset of the accelerator embedded u-controller 740 enables the accelerator 131 to start operating and performing the expected functions of the accelerator 131, and is typically the initial step in the startup process of the entire SoC 120.

[0118] In some embodiments, at step S4, after releasing the reset of the accelerator embedded u-controller, the accelerator embedded u-controller 740 executes the bootloader accelerator firmware component, downloads the authenticated main accelerator firmware component from the memory 104 via the interdie bus 771, and executes the authenticated main accelerator firmware component at the accelerator 131.

[0119] Figure 10 is a flowchart showing a sideband-based accelerator firmware authentication method 1000 according to some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 1000 is a method implemented by the SoC 120, which is configured to use the size of the accelerator firmware component to determine whether the accelerator firmware component is to be instantaneously authenticated and provided to the accelerator 131 via the sideband bus 791. In some embodiments, for example, according to the design of the accelerator or the SoC 120, the accelerator firmware component can be, for example, a bootloader accelerator firmware component, a main accelerator firmware component, or other non-bootloader accelerator firmware components. The methods, process steps, or phases shown in the drawings can be implemented as independent routines or processes, or as part of a larger routine or process. It should be noted that each depicted process step or phase can be implemented as other embodiments such as a device, method, or system including a processor that executes a set of instructions.

[0120] In some embodiments, at operation 1010, the security agent 721 of ROT 620 reads the accelerator firmware from the flash device 140 at startup. In some embodiments, at operation 1015, after reading the accelerator firmware from the flash device 140, the accelerator firmware parsing unit 753 of the security agent 721 receives the accelerator firmware and parses the accelerator firmware into a plurality of accelerator firmware components. In some embodiments, at operation 1020, after parsing the accelerator firmware into a plurality of accelerator firmware components, the accelerator firmware component size determination unit 754 of the security agent 721 determines the size of each accelerator firmware component of the accelerator firmware and the size of the accelerator firmware. In some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 determines the size of the accelerator firmware components and the size of the accelerator firmware by evaluating the size information provided by the firmware itself to determine the size and location of each component. In some embodiments, the size information may be included in the firmware image of the accelerator firmware and may be used by ROT 620 to divide the firmware into discrete multiple accelerator firmware components. In some embodiments, ROT 620 may also utilize a heuristic or algorithm to determine the size of the accelerator firmware components. For example, in some embodiments, ROT 620 may estimate the size of the accelerator firmware components based on the storage amount required to execute the expected functions of the accelerator firmware components. In some embodiments, as described above, the accelerator firmware components may be, for example, a bootloader accelerator firmware component or a non-bootloader accelerator firmware component associated with the accelerator in the SoC 120.

[0121] In some embodiments, at operation 1020, the accelerator firmware component size determination unit 754 of the security agent 721 determines whether the size of each individual accelerator firmware component is less than the immediate authentication size threshold 755. For example, in some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 determines whether the size of the bootloader accelerator firmware component is less than the immediate authentication size threshold 755. In some embodiments, the immediate authentication size threshold 755 is a threshold value that the accelerator firmware component size determination unit 754 uses to determine whether the accelerator firmware component is to be immediately authenticated by the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752, or whether the accelerator firmware component is to be authenticated by the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 at a delayed time. In some embodiments, the immediate authentication size threshold 755 may be a byte size value of 10 gigabytes, 20 gigabytes, or some other byte size value that can be used as an immediate authentication threshold to determine whether an accelerator firmware component is to be immediately authenticated by the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752.

[0122] In some embodiments, at operation 1050, when the accelerator firmware component sizing unit 754 determines that the size of the accelerator firmware component is less than the immediate authentication size threshold 755, the security agent 721 uses the immediate accelerator firmware component authentication unit 761 to immediately authenticate the accelerator firmware component. In some embodiments, for example, when the accelerator firmware identification unit 756 determines that the size of the bootloader accelerator firmware component is below the immediate authentication size threshold, the security agent 721 uses the immediate accelerator firmware component authentication unit 761 to immediately authenticate the bootloader accelerator firmware component at the security agent 721 of the host CPU 171 without storing the bootloader accelerator firmware component in the memory 104. In some embodiments, since the size of the bootloader accelerator firmware component can be relatively small compared to the overall size of the accelerator firmware, the bootloader accelerator firmware component is the component in the accelerator firmware that is immediately authenticated by the host CPU 171.

[0123] In some embodiments, at operation 1055, after the immediate accelerator firmware component authentication unit 761 immediately authenticates the accelerator firmware component, the accelerator firmware component is pushed into the accelerator memory 741 via the sideband bus 791. For example, in some embodiments, after the immediate accelerator firmware component authentication unit 761 immediately authenticates the bootloader accelerator firmware component, the host CPU 171 pushes the bootloader into the accelerator memory 741 of the accelerator 131 via the sideband bus 791.

[0124] In some embodiments, at operation 1070, the accelerator 131 receives the accelerator firmware component via the sideband bus 791 and executes the critical accelerator firmware component. In some embodiments, since the speed of the data flow in the sideband bus 791 is generally lower than the speed of the data flow in the inter-die bus 771, the use of the sideband bus 791 as described herein improves the performance of the SoC 120 by using the sideband bus 791 for actions that were not previously used for the sideband bus 791, making the SoC 120 more efficient than other SoCs or computer systems.

[0125] In some embodiments, returning to reference operation 1020, when the accelerator firmware component sizing unit 754 determines that the size of the accelerator firmware component is not less than the immediate authentication size threshold 755, at operation 1025, the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 receives the accelerator firmware component (whose size is not less than the immediate authentication size threshold 755) and authenticates the accelerator firmware component. In some embodiments, after authenticating the accelerator firmware component, operation 1025 proceeds to operation 1030.

[0126] In some embodiments, at operation 1030, the host CPU 171 stores the accelerator firmware component in the memory 104. In some embodiments, at operation 1035, the accelerator firmware component is downloaded from the memory 104 to the accelerator 131 via the inter-die bus 771 and executed by the accelerator 131.

[0127] In some embodiments, when the accelerator firmware component size determination unit 754 determines that all accelerator firmware component sizes are greater than the immediate authentication size threshold 755, the entire authenticated accelerator firmware is provided to the accelerator 131 via the inter-die bus 771 for execution by the accelerator 131.

[0128] In some embodiments, using the embodiments described herein, the efficiency of the SoC 120 is partially improved because the sideband bus (e.g., sideband bus 790), which is not typically used for accelerator firmware component transfer, is used to transfer critical accelerator firmware components while the host CPU 171 is still processing non-critical accelerator firmware components. This allows the accelerator to first process the critical accelerator firmware components until the non-critical accelerator firmware components are provided to the accelerator via the inter-die bus (e.g., inter-die bus 770). Thus, in some embodiments, the systems and methods described herein improve and provide advantages over other methods, such as reducing the time and resources required for firmware authentication during startup or firmware updates, and utilizing resources (e.g., sideband bus 790) that are not fully utilized to transfer accelerator firmware.

[0129] In some embodiments, a method includes: receiving an input video data stream at a CPU-based demultiplexer of a central processing unit (CPU); performing an accelerator decoding configuration evaluation of an accelerator at the CPU; and dynamically decoding a CPU-based demultiplexer output from the CPU-based demultiplexer using a CPU-based decoding unit and an accelerator-based decoding unit based on the accelerator decoding configuration evaluation.

[0130] In some embodiments of the method, the accelerator decoding configuration evaluation includes: performing an accelerator-based decoding unit hardware configuration evaluation of the accelerator-based decoding unit, and performing a CPU-based decoding unit software configuration evaluation of the CPU-based decoding unit software for the CPU-based decoding unit.

[0131] In some embodiments, the method further includes: performing an accelerator-based preprocessing unit configuration evaluation of an accelerator-based preprocessing unit coupled to the CPU-based decoding unit and the accelerator-based decoding unit.

[0132] In some embodiments of the method, the accelerator-based preprocessing unit configuration evaluation includes an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0133] In some embodiments, the method further includes: dynamically preprocessing decoder outputs from a CPU-based decoding unit and an accelerator-based decoding unit using the accelerator-based preprocessing unit configuration evaluation.

[0134] In some embodiments, the method further includes: performing an encoding unit hardware configuration evaluation on an accelerator-based encoding unit.

[0135] In some embodiments, the method further includes: dynamically encoding the preprocessing output from the accelerator-based preprocessing unit using the accelerator-based encoding unit.

[0136] In some embodiments, a system-on-chip includes: a central processing unit (CPU)-based demultiplexer; a CPU-based decoding unit and an accelerator-based decoding unit, the CPU-based decoding unit and the accelerator-based decoding unit being coupled to the CPU-based demultiplexer via an inter-die interconnect; and an accelerator-based preprocessing unit, the accelerator-based preprocessing unit being coupled to the CPU-based decoding unit and the accelerator-based decoding unit via an inter-die interconnect, wherein the system-on-chip dynamically decodes a CPU-based demultiplexer output from the CPU-based demultiplexer using the CPU-based decoding unit and the accelerator-based decoding unit based on an accelerator decoding configuration evaluation for video processing.

[0137] In some embodiments of the system-on-chip, the accelerator decoding configuration evaluation includes an accelerator-based decoding unit hardware configuration evaluation and a CPU-based decoding unit software configuration evaluation.

[0138] In some embodiments of the system-on-chip, the accelerator-based decoding unit hardware configuration evaluation includes determining whether the CPU-based demultiplexer output is mapped to the accelerator-based decoding unit hardware configuration of the accelerator-based decoding unit.

[0139] In some embodiments of the system-on-chip, the CPU-based decoding unit software configuration evaluation includes determining whether the CPU-based demultiplexer output is mapped to the CPU-based decoding unit software configuration of the CPU-based decoding unit.

[0140] In some embodiments of the system-on-chip, the system-on-chip is configured to dynamically preprocess decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using the accelerator-based preprocessing unit.

[0141] In some embodiments of the system-on-chip, the system-on-chip utilizes an accelerator-based preprocessing unit configuration evaluation to dynamically preprocess decoder outputs from a CPU-based decoding unit and an accelerator-based decoding unit.

[0142] In some embodiments of the system-on-chip, the accelerator-based preprocessing unit configuration evaluation includes an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0143] In some embodiments, the system-on-chip further includes: an accelerator-based encoding unit, the accelerator-based encoding unit being coupled to the accelerator-based preprocessing unit, wherein the system-on-chip is configured to dynamically encode preprocessing outputs from the accelerator-based preprocessing unit using the accelerator-based encoding unit.

[0144] In some embodiments of the system-on-chip, the system-on-chip utilizes an accelerator-based encoding unit hardware configuration evaluation to dynamically encode preprocessor outputs from the accelerator-based preprocessing unit.

[0145] In some embodiments, a server system includes: a host central processing unit (CPU); an accelerator, the accelerator being coupled to the host CPU via an inter-die interconnect; a unified memory, the unified memory being coupled to the host CPU and the accelerator, wherein the server system is configured to dynamically share decoding operations, preprocessing operations, and encoding operations using the host CPU and the accelerator based on an accelerator configuration evaluation performed by the host CPU.

[0146] In some embodiments of the server system, to share decoding operations, the host CPU performs an accelerator-based decoding unit hardware configuration evaluation and a CPU-based decoding unit software configuration evaluation.

[0147] In some embodiments of the server system, to share preprocessing operations, the host CPU performs an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

[0148] In some embodiments of the server system, to share encoding operations, the host CPU performs an accelerator-based encoding unit hardware configuration evaluation and a CPU-based encoding unit software configuration evaluation.

Claims

1. A method, comprising: Receiving an input video data stream at a CPU-based demultiplexer of a central processing unit (CPU); Performing an accelerator decoding configuration evaluation on an accelerator decoding configuration of an accelerator at the CPU; And Based on the accelerator decoding configuration evaluation, dynamically decoding a CPU-based demultiplexer output from the CPU-based demultiplexer using a CPU-based decoding unit and an accelerator-based decoding unit.

2. The method according to claim 1, wherein: The accelerator decoding configuration evaluation includes: performing an accelerator-based decoding unit hardware configuration evaluation on the accelerator-based decoding unit, and performing a CPU-based decoding unit software configuration evaluation on CPU-based decoding unit software for the CPU-based decoding unit.

3. The method according to claim 1 or 2, further comprising: Performing an accelerator-based preprocessing unit configuration evaluation on an accelerator-based preprocessing unit coupled to the CPU-based decoding unit and the accelerator-based decoding unit.

4. The method according to claim 3, wherein: The accelerator-based preprocessing unit configuration evaluation includes an accelerator-based preprocessing unit hardware configuration evaluation and a CPU-based preprocessing unit software configuration evaluation.

5. The method according to claim 3 or 4, further comprising: Dynamically preprocessing decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using the accelerator-based preprocessing unit configuration evaluation.

6. The method according to any one of the preceding claims, further comprising: Performing an encoding unit hardware configuration evaluation on an accelerator-based encoding unit.

7. The method according to claim 6, further comprising: Dynamically encoding a preprocessing output from the accelerator-based preprocessing unit using the accelerator-based encoding unit.

8. A system on a chip, comprising: A CPU-based demultiplexer; A CPU-based decoding unit and an accelerator-based decoding unit, the CPU-based decoding unit and the accelerator-based decoding unit being coupled to the CPU-based demultiplexer via an inter-die interconnect; And An accelerator-based preprocessing unit, the accelerator-based preprocessing unit being coupled to the CPU-based decoding unit and the accelerator-based decoding unit via the inter-die interconnect, wherein the system on a chip, based on an accelerator decoding configuration evaluation, dynamically decodes a CPU-based demultiplexer output from the CPU-based demultiplexer using the CPU-based decoding unit and the accelerator-based decoding unit for video processing.

9. The system on a chip according to claim 8, wherein: The accelerator decoding configuration evaluation includes an accelerator-based decoding unit hardware configuration evaluation and a CPU-based decoding unit software configuration evaluation.

10. The system on a chip according to claim 9, wherein: The accelerator-based decoding unit hardware configuration evaluation includes determining whether the output of the CPU-based demultiplexer is mapped to the accelerator-based decoding unit hardware configuration of the accelerator-based decoding unit.

11. The system-on-chip according to claim 9 or 10, wherein: The CPU-based decoding unit software configuration evaluation includes determining whether the output of the CPU-based demultiplexer is mapped to the CPU-based decoding unit software configuration of the CPU-based decoding unit.

12. The system-on-chip according to any one of claims 8 to 11, wherein: i. The system-on-chip is configured to dynamically preprocess the decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using the accelerator-based preprocessing unit; and / or, ii. The system-on-chip dynamically preprocesses the decoder outputs from the CPU-based decoding unit and the accelerator-based decoding unit using accelerator-based preprocessing unit configuration evaluation; and preferably, wherein the accelerator-based preprocessing unit configuration evaluation includes accelerator-based preprocessing unit hardware configuration evaluation and CPU-based preprocessing unit software configuration evaluation.

13. The system-on-chip according to any one of claims 8 to 12, further comprising: An accelerator-based encoding unit, the accelerator-based encoding unit being coupled to the accelerator-based preprocessing unit, wherein the system-on-chip is configured to dynamically encode the preprocessing output from the accelerator-based preprocessing unit using the accelerator-based encoding unit; and / or preferably wherein the system-on-chip dynamically encodes the preprocessor output from the accelerator-based preprocessing unit using accelerator-based encoding unit hardware configuration evaluation.

14. A server system, comprising: A host central processing unit (CPU); An accelerator, the accelerator being coupled to the host CPU via an inter-die interconnect; A unified memory, the unified memory being coupled to the host CPU and the accelerator, wherein the server system is configured to dynamically share decoding operations, preprocessing operations, and encoding operations using the host CPU and the accelerator based on accelerator configuration evaluation performed by the host CPU.

15. The server system according to claim 14, wherein, One or more of the following: i. To share decoding operations, the host CPU performs accelerator-based decoding unit hardware configuration evaluation and CPU-based decoding unit software configuration evaluation; ii. To share preprocessing operations, the host CPU performs accelerator-based preprocessing unit hardware configuration evaluation and CPU-based preprocessing unit software configuration evaluation; iii. To share encoding operations, the host CPU performs accelerator-based encoding unit hardware configuration evaluation and CPU-based encoding unit software configuration evaluation.