Core-based integrated central processing unit with accelerator

By designing a unified memory access tunnel system in the server system, allowing the accelerator to access the unified memory using the shared address space, the problem of PCIe connection limitation and external I/O connection increased latency is solved, and the efficiency of the CPU and accelerator is improved.

CN120226009APending Publication Date: 2025-06-27META PLATFORMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079843.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2023-12-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In existing server systems, PCIe connections limit the use of shared address space between the CPU and the accelerator, and external I/O connections increase system delay, making it difficult to balance the ratio of CPU and accelerator to match application requirements, affecting efficiency.

Method used

Design a system on chip to access the tunnel system through a unified memory, allowing the accelerator to access the unified memory using the shared address space, avoiding the use of the accelerator memory for processing operations, and achieving efficient interoperability between the CPU and the accelerator.

Benefits of technology

By unified memory access tunneling system, the limitation of shared address space usage between the CPU and the accelerator is solved, the system delay is reduced, the efficiency of the CPU and the accelerator is improved, and the application requirements can be better matched.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226009A_ABST
    Figure CN120226009A_ABST
Patent Text Reader

Abstract

In some embodiments, a system on chip includes: a central processing unit (CPU); an accelerator coupled to the CPU via a first inter-die interconnect; and a unified memory coupled to the CPU via a second inter-die interconnect. In some embodiments, to prevent the accelerator from using accelerator memory for processing operations, the accelerator utilizes a unified memory access tunneling system located in the accelerator to tunnel an advanced interconnect protocol associated with a second inter-die interconnect to an inter-die interconnect protocol associated with a first inter-die interconnect, the unified memory access tunneling system is configured to allow access to the unified memory using the shared address space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a system-on-chip (SOC) method and a server system for unified memory access. Background Art

[0002] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. To the extent that work by one or more of the inventors, currently named, is described in this background art section, and aspects of the description that may not otherwise qualify as prior art at the time of filing the application, are neither expressly nor implicitly admitted to be prior art with respect to the present disclosure.

[0003] Many modern server systems utilize a system-on-chip (SoC) that includes an accelerator connected to a central processing unit (CPU) using the Peripheral Component Interconnect Express (PCIe). Although PCIe is generally acceptable for offloading workloads, it has several limitations. For example, PCIe may not allow sharing of the address space between the CPU and the accelerator. Additionally, the accelerator typically requires high bandwidth and low latency to access a common memory. External input / output (I / O) connections such as PCIe or Compute Express Link (CXL) may increase the system latency because the aggregate bandwidth is limited due to the limited number of PCIe / CXL ports and channels. The limited number of ports and channels reduces the number of accelerators that can be attached to the CPU and makes it difficult to balance the CPU-to-accelerator ratio to match the application requirements, which has a negative impact on the efficiency of the CPU and the accelerator. Summary of the Invention

[0004] According to a first aspect of the present disclosure, there is provided a system-on-chip including: a central processing unit (CPU); an accelerator coupled to the CPU via a first inter-die interconnect; and a unified memory coupled to the CPU via a second inter-die interconnect, wherein, to prevent the accelerator from using an accelerator memory for processing operations, the accelerator utilizes a unified memory access tunnel system located in the accelerator to tunnel a high-level interconnect protocol associated with the second inter-die interconnect to an inter-die interconnect protocol associated with the first inter-die interconnect, and the unified memory access tunnel system is configured to allow access to the unified memory using a shared address space.

[0005] In some embodiments, a unified memory access tunneling system tunnels an advanced interconnect protocol to an inter-die interconnect protocol using a unified memory access tunneling data packet.

[0006] In some embodiments, the unified memory access tunneling data packet includes a plurality of tunneling fields configured to allow an accelerator and a CPU to access a shared address space mapped to a unified memory.

[0007] In some embodiments, the unified memory access tunneling system provides a unified memory access tunneling data packet to the CPU to access the unified memory.

[0008] In some embodiments, a unified memory access detunneling system in the CPU receives the unified memory access tunneling data packet.

[0009] In some embodiments, the unified memory access detunneling system detunnels the unified memory access tunneling data packet to access the shared address space of the unified memory.

[0010] In some embodiments, the advanced interconnect protocol is an Advanced eXtensible Interface (AXI) protocol or a Compute Express Link (CXL) protocol.

[0011] In some embodiments, the inter-die interconnect protocol is a universal chiplet interconnect express (UCIe) protocol.

[0012] In some embodiments, based on the tunneling of the advanced interconnect protocol to the inter-die interconnect protocol, the accelerator does not need to use an accelerator memory.

[0013] According to a second aspect of the present disclosure, a method is provided, the method including: generating, at an accelerator, a unified memory access tunneling data packet protocol structure mapped to an inter-die interconnect protocol; generating, at the accelerator, a unified memory access tunneling data packet using the unified memory access tunneling data packet structure, the unified memory access tunneling data packet being used to tunnel an advanced interconnect protocol to an inter-die interconnect protocol associated with the inter-die interconnect; and directly accessing a unified memory using the unified memory access tunneling data packet.

[0014] In some embodiments, the unified memory access tunneling data packet includes a plurality of tunneling fields configured to allow an accelerator and a central processing unit (CPU) to access a shared address space mapped to a unified memory.

[0015] In some embodiments, the method further includes: providing a unified memory access tunneling packet to a unified memory access de-tunneling system.

[0016] In some embodiments, the unified memory access de-tunneling system in the CPU receives the unified memory access tunneling packet.

[0017] In some embodiments, the method further includes: de-tunneling the unified memory access tunneling packet by using the unified memory access de-tunneling system to access a shared address space of the unified memory.

[0018] In some embodiments, the advanced interconnect protocol is an Advanced eXtensible Interface (AXI) protocol or a Compute Express Link (CXL) protocol.

[0019] In some embodiments, the die-to-die interconnect is a Universal Chiplet Interconnect Express (UCIe) interconnect.

[0020] In some embodiments, based on tunneling from the advanced interconnect protocol to the die-to-die interconnect protocol, the accelerator does not need to use local memory.

[0021] According to a third aspect of the present disclosure, there is provided a server system, the server system includes: an accelerator interface controller; a die-to-die interconnect coupled to the accelerator interface controller; and a central processing unit (CPU) interface controller coupled to the die-to-die interconnect, wherein the accelerator interface controller is configured to tunnel an advanced interconnect protocol to a die-to-die interconnect protocol associated with the die-to-die interconnect to access a unified memory.

[0022] In some embodiments, the die-to-die interconnect is a Universal Chiplet Interconnect Express (UCIe) interconnect, and the advanced interconnect protocol is an Advanced eXtensible Interface (AXI) protocol or a Compute Express Link (CXL) protocol.

[0023] In some embodiments, the CPU interface controller is located in a host CPU, and the accelerator interface controller is located in a video accelerator.

[0024] It will be appreciated that any features described herein as being suitable for incorporation into one or more aspects or embodiments of the present disclosure are intended to be generalizable to any and all aspects and embodiments of the present disclosure. Those skilled in the art can understand other aspects of the present disclosure based on the specification, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are merely exemplary and explanatory, and are not limiting of the claims.

[0025] The following further describes the features and advantages of the embodiments, as well as the structures and operations of various embodiments, with reference to the accompanying drawings. It should be noted that the methods and systems are not limited to the specific embodiments described herein. These embodiments are presented herein for illustrative purposes only. Based on the teachings contained herein, additional embodiments will be apparent to one or more persons skilled in the relevant art. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1A is a block diagram showing a server system according to some embodiments.

[0027] Figure 1B is showing a server system according to some embodiments Figure 1A in the block diagram.

[0028] Figure 1C is showing a server system according to some embodiments Figure 1B in the block diagram of the process flow.

[0029] Figure 1D shows a unified memory access method utilized in Figures 1A to 1C according to some embodiments.

[0030] Figure 2 is a block diagram showing a system-on-chip according to some embodiments.

[0031] Figure 3 is a more detailed block diagram showing a system-on-chip in Figure 2 according to some embodiments.

[0032] Figure 4 is a block diagram showing a shared preprocessing flow according to some embodiments.

[0033] Figure 5 is a block diagram showing a shared encoding flow according to some embodiments.

[0034] Figure 6 is a block diagram showing a system-on-chip according to some embodiments.

[0035] Figure 7 shows a system-on-chip in Figure 6 according to some embodiments.

[0036] Figure 8 is a flowchart showing an edge-band based accelerator firmware authentication method according to some embodiments.

[0037] Figure 9 is in a system-on-chip according to some embodiments Figure 7 utilized in the system-on-chip Figure 8An example process flow of an edgeband-based accelerator firmware authentication method.

[0038] Figure 10 FIG. 4 is a flowchart illustrating an edgeband-based accelerator firmware authentication method according to some embodiments. DETAILED DESCRIPTION

[0039] Figure 1A FIG. 5 is a block diagram illustrating a server system 100 according to some embodiments. In some embodiments, the server system 100 is configured to utilize a unified memory access tunneling system in a system-on-chip (SoC) 120 to access a unified memory 160 through an inter-die interface 167, such that an accelerator 130 in the SoC 120 does not have to utilize an accelerator memory during processing operations performed by the accelerator 130. In some embodiments, the unified memory access tunneling system is hardware and / or executable code in the accelerator 130 configured to generate a unified memory access tunneling packet structure that is used by the unified memory access tunneling system to tunnel a high-level interconnect protocol into an inter-die interface protocol of the inter-die interface 167. In some embodiments, by utilizing the unified memory access tunneling packet structure, the accelerator 130 is able to generate a unified memory access tunneling packet that allows the accelerator 130 to transfer data through the inter-die interface 167, such that during operation execution at the accelerator 130, the accelerator 130 utilizes the unified memory 160 rather than an accelerator memory to process operations.

[0040] In some embodiments, the server system 100 includes a system-on-chip (SoC) 120, a unified memory 160, and an external input / output (I / O) interconnect 139. In some embodiments, the SoC 120 is coupled to the unified memory 160 via the external input / output (I / O) interconnect 139. In some embodiments, the external I / O interconnect 139 can be, for example, a Compute Express Link (CXL) interconnect, a Peripheral Component Interconnect Express (PCIe) interconnect, or other type of external I / O interconnect configured to couple the unified memory 160 to the SoC 120.

[0041] In some embodiments, SoC 120 is a system-on-chip configured to use an on-chip fabric as a communication protocol within SoC 120, such as an Advanced eXtensible Interface (AXI), a Network on Chip (NoC), or an Advanced Computing Environment (ACE). In some embodiments, SoC 120 includes a host central processing unit (CPU) 171, an accelerator 130, and an inter-die interface 167. In some embodiments, the host CPU 171 is coupled to the accelerator 130 via the inter-die interface 167. In some embodiments, the inter-die interface 167 is a physical interface or connection between the host CPU 171 and the accelerator 130, and this physical interface or connection includes an inter-die interconnect (e.g., Figure 1B the inter-die interconnects 161 to 164 in Figure 1B ). In some embodiments, each of the inter-die interconnects 161 to 164 includes a sideband bus and an inter-die bus (mainband bus). In some embodiments, the inter-die interface 167 may be configured to transfer data on the inter-die interface 167 using a Chiplet Data Exchange (CDX) protocol. CDX is a high-speed, low-latency protocol designed for chip-to-chip communication and optimized for chiplet interconnects. In some embodiments, CDX is configured to support high data rates, allowing for fast data transfer between multiple chiplets in SoC 120. In some embodiments, the inter-die interface 167 may be, for example, a Universal Chiplet Interconnect Express (UCIe) interface or another type of inter-die interface for the embodiments described herein.

[0042] In some embodiments, the host CPU 171 is a processor that is configured to perform the operations described herein (described in more detail with reference to Figures 1A to 10 ), in addition to performing standard CPU processing operations within SoC 120. In some embodiments, the host CPU 171 may be a server-class CPU coupled to an accelerator 130 integrated on SoC 120. In some embodiments, depending on the type of server system 100, for example, SoC 120 may include multiple host CPUs.

[0043] In some embodiments, accelerator 130 is a dedicated processing unit in SoC 120. In addition to performing tasks specific to accelerator 130, the dedicated processing unit is configured to access unified memory 160 using a unified memory access tunneling system. In some embodiments, accelerator 130 is configured to access unified memory 160 using a shared address space mapped to unified memory 160. In some embodiments, the shared address space is a range of shared memory addresses associated with unified memory 160 that host CPU 171 and accelerator 130 can access to perform processing operations. In some embodiments, the unified memory access tunneling system of SoC 120 is configured not only to allow accelerator 130 and host CPU 171 to utilize the shared address space in unified memory 160, but also to allow direct data transfer from accelerator 130 to unified memory 160 through inter-die interface 167, and vice versa (further described herein with reference to Figures 1B to 1D ). In some embodiments, as Figure 1B illustrated by way of example, accelerator 130 includes accelerator 131, accelerator 132, accelerator 133, and accelerator 134. In some embodiments, accelerator 130 can be, for example, a video accelerator, a graphic processing unit (GPU), a digital signal processor, or other type of accelerator configured to perform the operations described herein using a unified memory access tunneling system. In some embodiments, each accelerator in accelerator 130 can be configured to include a unified memory access tunneling system to perform the required unified memory access tunneling operations described herein.

[0044] In some embodiments, unified memory 160 is a memory shared between host CPU 171 and accelerator 130, and the memory is configured to be directly accessed by accelerator 130 using the unified memory access tunnel operations and shared address space described herein. In some embodiments, unified memory 160 can be, for example, a low power (LP) memory and / or other types of memory associated with a shared address space used by accelerator 130 and host CPU 171. In some embodiments, as described above, the shared address space is a range of shared memory addresses that host CPU 171 and accelerator 130 can use to access executable code required for performing processing operations. In some embodiments, unified memory 160 can be a random access memory (RAM), such as a dynamic random access memory (DRAM), a static random access memory (SRAM), and a non-volatile random access memory (NVRAM), etc. In some embodiments, unified memory 160 can include a system memory and an external memory associated with host CPU 171 located on SoC 120, and the system memory and the external memory are combined into unified memory 160 (e.g., the host CPU memory and the external memory are combined into unified memory 160). In some embodiments, unified memory 160 can be considered unified because, for example, due to the inter-die interface protocol that tunnels the high-level interconnect protocol to inter-die interface 167, accelerator 130 may only need the shared address space to directly access the memory for processing.

[0045] In some embodiments, to access unified memory 160 using the shared address space, the unified memory access tunnel system is configured to tunnel the high-level interconnect protocol into the inter-die interface protocol of inter-die interface 167. In some embodiments, SoC 120 uses the unified memory access tunnel system to allow accelerator 130 and host CPU 171 to access the shared address space of unified memory 160, and to deactivate the use of accelerator memory during accelerator processing. Thus, accelerator 130 does not need to have an accelerator memory for accelerator 130 to perform processing operations. The various methods and systems are further described herein with reference to Figures 1B to 1D which are further described.

[0046] Figure 1B which is further illustrated in accordance with some embodiments of Figure 1AThe server system 100 therein. In some embodiments, the SoC 120 of the server system 100 includes an accelerator 131 equipped with a unified memory access tunneling system 116, which is configured to perform unified memory access tunneling operations that allow the accelerator 131 to directly access the unified memory 160 (e.g., LP memory 121 and / or memory 126) through the die - to - die interconnect 161, so that the accelerator 130 in the SoC 120 does not have to utilize the accelerator memory during the execution of processing operations by the accelerator 131. In some embodiments, the unified memory access tunneling system 116 is hardware and / or executable code in the accelerator 131 configured to generate a unified memory access tunneling packet structure, which is used by the unified memory access tunneling system 116 to tunnel an advanced interconnect protocol into the die - to - die interface protocol of the die - to - die interconnect 161. In some embodiments, as described above, the advanced interconnect protocol can be, for example, the AXI or CXL / PCIe protocol. In some embodiments, the die - to - die interface protocol can be, for example, the UCIe protocol configured to directly connect the die to the host CPU 171 or other die - to - die interface protocols. In some embodiments, the operations performed by the unified memory access tunneling system 116 can be performed by the accelerator die - to - die interface controller 113 in the accelerator 131.

[0047] In some embodiments, as Figure 1B shown, the server system 100 includes an SoC 120, a memory 126, a low - power (LP) memory 121, LP memories 122, 123, and 124. In some embodiments, the SoC 120 includes accelerators 131, 132, 133, 134, and a host CPU 171. In some embodiments, the LP memory 121 includes a dedicated memory 129 and a shared memory 127. In some embodiments, the host CPU 171 includes a CPU core 159, a storage controller 118, a memory management unit (MMU) 111, a host CPU interface controller 112, Low Power Double Data Rate 5 (LP5) interconnects 151, 152, 153, 154, 155, 156, 157, 158, die - to - die interconnects 161, 162, 163, 164, external input / output (I / O) interconnects 168, and external I / O interconnects 169.

[0048] In some embodiments, the MMU 111 of the host CPU 171 is a memory management unit configured to receive virtual addresses provided to the host CPU 171 from the accelerator 130 or other devices or components external to the host CPU 171 in the server system 100 (e.g., via unified memory access tunneling packets).

[0049] In some embodiments, the memory controller 118 of the host CPU 171 is a memory controller configured to control access to the unified memory 160. In some embodiments, the memory controller 118 can be implemented in hardware, firmware, software, or any combination thereof. In some embodiments, the memory controller 118 is configured to read data from and write data to the unified memory 160.

[0050] In some embodiments, the host CPU interface controller 112 is a component in the host CPU 171 that, in addition to performing unified memory access de-tunneling operations using the unified memory access de-tunneling system 117 described herein, is also configured to control communication between the host CPU 171 and other devices or subsystems within the SoC 120.

[0051] In some embodiments, the LP5 interconnects 151 - 158 of the host CPU 171 are LPDDR5 interconnects that connect storage devices (e.g., LP memories 121 - 124) to the SoC 120 using the Low Power Double Data Rate 5 (LPDDR5) technology standard. In some embodiments, the LPDDR5 standard defines the interface protocol (e.g., AXI protocol) utilized by the LP5 interconnects and the electrical signal characteristics for LPDDR5 storage devices, including signal voltage levels, timing, and bus width.

[0052] In some embodiments, the inter-die interconnects 161 - 164 of the host CPU 171 are interconnects configured to enable communication and data transfer between two or more individual integrated circuits (dies) within a single package or the SoC 120. In some embodiments, the inter-die interconnects 161 - 164 are UCIe interconnects or some other type of inter-die interconnect configured to operate according to an inter-die interconnect standard.

[0053] In some embodiments, the external I / O interconnect 168 of the host CPU 171 is, for example, a CXL interconnect or some other type of external I / O interconnect. In some embodiments, the external I / O interconnect 169 is, for example, a PCIe interconnect utilized in the SoC or some other type of external I / O interconnect.

[0054] In some embodiments, accelerator 131 is coupled to host CPU 171 via inter-die interconnect 161. In some embodiments, accelerator 132 is coupled to host CPU 171 via inter-die interconnect 162. In some embodiments, accelerator 133 is coupled to host CPU 171 via inter-die interconnect 163. In some embodiments, accelerator 134 is coupled to host CPU 171 via inter-die interconnect 164.

[0055] In some embodiments, host CPU 171 is coupled to memory 126 via external I / O interconnect 168. In some embodiments, host CPU 171 may also be coupled to memory 126 or other devices via external I / O interconnect 169. In some embodiments, host CPU 171 is coupled to LP memory 121 via LP5 interconnect 151. Similarly, in some embodiments, host CPU 171 is coupled to LP memories 122 to 124 via LP5 interconnects 152 to 154, respectively. In some embodiments, LP memory 121 is configured to communicate with host CPU 171 via LP5 interconnect 151. Similarly, in some embodiments, LP memories 122 to 124 are configured to communicate with host CPU 171 via LP5 interconnects 152 to 154, respectively.

[0056] In some embodiments, accelerator 131 is configured to communicate with memory 126 via inter-die interconnect 161 and external I / O interconnect 168. In some embodiments, accelerator 131 is configured to communicate with memory 126 via inter-die interconnect 161 and external I / O interconnect 169. In some embodiments, accelerator 131 is configured to communicate with host CPU 171 via inter-die interconnect 161. Similarly, accelerators 132 to 134 are configured to communicate with memory 126 via inter-die interconnects 162 to 164 and external I / O interconnect 168, respectively. In some embodiments, accelerators 132 to 134 are configured to communicate with host CPU 171 via inter-die interconnects 162 to 164 and external I / O interconnect 169, respectively. In some embodiments, accelerators 132 to 134 are configured to communicate with host CPU 171 via inter-die interconnects 162 to 164, respectively.

[0057] In some embodiments, accelerator 131 includes a memory controller 114 and an accelerator inter-die interface controller 113. In some embodiments, memory controller 114 is configured to manage memory access and communication between accelerator 131 and a memory (e.g., unified memory 160), such that accelerator 131 can effectively utilize the memory resources available to SoC 120. In some embodiments, accelerator inter-die interface controller 113 is a component in accelerator 131 that is configured to control communication between accelerator 131 and other devices or subsystems within SoC 120, in addition to performing the unified memory access tunneling operations described herein. In some embodiments, accelerator inter-die interface controller 113 includes a unified memory access tunneling system 116. In some embodiments, as described above, unified memory access tunneling system 116 is configured to tunnel a high-level interconnect protocol (e.g., AXI protocol, PCIe protocol, or CXL protocol) into an inter-die interface protocol (e.g., UCIe protocol) of an inter-die interconnect (e.g., inter-die interconnect 161).

[0058] In some embodiments, host CPU interface controller 112 of host CPU 171 includes a unified memory access de-tunneling system 117. In some embodiments, unified memory access de-tunneling system 117 is hardware and / or executable code (further described herein) configured to de-tunnel unified memory access tunnel data packets tunneled by unified memory access tunneling system 116.

[0059] In some embodiments, unified memory access tunneling system 116 and unified memory access de-tunneling system 117 are jointly configured to allow accelerator 131 to directly access unified memory 160, thereby bypassing the need to use accelerator memory for memory access during accelerator processing.

[0060] In some embodiments, in operation, accelerator 131 (via memory controller 114) initiates a memory access request to host CPU 171 to request access to unified memory 160. In some embodiments, as described above, memory controller 114 is a component within accelerator 131 that manages data flow to and from unified memory 160 and is responsible for generating memory access requests to host CPU 171. In some embodiments, unified memory 160 may include memory 126 (e.g., memory associated with host CPU 171) and / or LP memory 121. In some embodiments, the memory access request generated by memory controller 114 includes a virtual memory address associated with the requested data and the type of memory access (e.g., read or write) request.

[0061] In some embodiments, since the accelerator 131 directly and immediately accesses the unified memory 160 to perform processing operations, before sending a memory access request to the host CPU 171, the storage controller 114 notifies the unified memory access tunneling system 116 to perform a unified memory access tunneling operation for the memory access request on the inter-die interface protocol of the inter-die interconnect 161.

[0062] In some embodiments, the unified memory access tunneling system 116 receives the notification and the memory access request, and starts the following process: performing the tunneling operations required to immediately perform the memory access request on the inter-die interconnect 161. In some embodiments, to tunnel the high-level interconnect protocol over the inter-die protocol, the unified memory access tunneling system 116 generates a unified memory access tunneling packet structure that maps to the inter-die interface protocol of the inter-die interconnect 161. In some embodiments, the unified memory access tunneling system 116 generates the unified memory access tunneling packet structure by modifying the inter-die interface protocol packet structure to include additional tunneling fields configured to allow the accelerator 131 to access the unified memory 160 using the high-level interconnect protocol without utilizing the memory associated with the accelerator 131. In some embodiments, after generating the unified memory access tunneling packet structure, the unified memory access tunneling system 116 starts the following process: tunneling the high-level interconnect protocol over the inter-die interface protocol using the modified inter-die interface protocol packet structure to generate a modified inter-die interface packet (e.g., a unified memory access tunneling packet).

[0063] In some embodiments, the unified memory access tunneling system 116 tunnels the high-level interconnect protocol associated with the memory access request over the inter-die interface protocol of the inter-die interconnect 161 by encapsulating the high-level interconnect protocol into a modified inter-die interface packet structure associated with the inter-die interconnect 161. In some embodiments, the unified memory access tunneling system 116 encapsulates the high-level interconnect protocol into the modified inter-die interface packet structure by including the high-level interconnect protocol information in the additional tunneling fields that have been added to the inter-die interface packet structure. In some embodiments, by encapsulating the high-level interconnect protocol (with additional high-level interconnect protocol information) into the modified inter-die interface packet structure associated with the inter-die interconnect 161, the accelerator 131 can forward the unified memory access tunneling packet (e.g., the modified inter-die interface packet) to the unified memory 160 through the inter-die interconnect 161. In some embodiments, the high-level interconnect protocol information can be extracted from the unified memory access tunneling packet and processed according to the high-level interconnect protocol.

[0064] In some embodiments, after tunneling a memory access request into an inter-die interface protocol data packet structure (inter-die interface protocol corresponding to the inter-die interconnect 161), the unified memory access tunneling system 116 provides a modified inter-die interface protocol data packet (e.g., a unified memory access tunneling data packet) to the unified memory access de-tunneling system 117 of the host CPU 171 via the inter-die interconnect 161.

[0065] In some embodiments, the unified memory access de-tunneling system 117 of the host CPU 171 receives a unified memory access tunneling data packet from the unified memory access tunneling system 116. In some embodiments, the unified memory access de-tunneling system 117 is configured to de-tunnel high-level interconnect protocol information from the unified memory access tunneling data packet provided by the accelerator 131 and extract the high-level interconnect protocol information from the unified memory access tunneling data packet provided by the accelerator 131. For example, in some embodiments, the unified memory access de-tunneling system 117 is configured to decode the tunnel fields of the unified memory access tunneling data packet to perform the operations indicated by these fields. For example, the unified memory access de-tunneling system 117 receives a unified memory access tunneling data packet from the unified memory access tunneling system 116 and decodes the write indication tunnel field by evaluating the bits located in this field (e.g., the memory opcode) to determine whether the accelerator 131 is requesting a write operation.

[0066] In some embodiments, when the write indication tunnel field indicates that a write operation is to be performed on the unified memory 160, the associated virtual address is provided to the MMU 111 of the host CPU 171, and the MMU converts the associated virtual address into a physical address of the unified memory 160. In some embodiments, the MMU 111 provides the physical address to the storage controller 114 of the host CPU 171. In some embodiments, the storage controller 118 of the host CPU 171 receives the physical address and determines whether the requested memory associated with this physical address is available in the unified memory 160. In some embodiments, when the storage controller 118 of the host CPU 171 determines that the requested memory is available in the unified memory 160, the host CPU 171 allows the accelerator 131 to write to the requested memory location.

[0067] In some embodiments, the unified memory access de-tunneling system 117 receives unified memory access tunnel data packets from the unified memory access tunneling system 116 and decodes the read indication tunnel field by evaluating the read indication tunnel field to determine whether the bits located in the field (e.g., the memory opcode) indicate that the accelerator 131 is requesting a read operation. In some embodiments, when the read indication tunnel field indicates that a read operation is to be performed by the accelerator 131, the associated virtual address for the read operation is provided to the MMU 111 of the host CPU 171, and the MMU converts the associated virtual address to a physical address of the unified memory 160. In some embodiments, the MMU 111 provides the physical address to the memory controller 118 of the host CPU 171.

[0068] In some embodiments, the memory controller 118 receives the physical address from the MMU 111 and determines whether the requested memory associated with the physical address is available for reading from the unified memory 160. In some embodiments, when the memory controller 114 of the host CPU 171 determines that the requested memory is available in the unified memory 160, the host CPU 171 allows the accelerator 131 to read from the requested memory location. In some embodiments, the memory controller of the host CPU 171 sends the requested data to the memory controller 114 of the accelerator 131 via the inter-die interconnect 161. In some embodiments, the memory controller 114 of the accelerator 131 provides the data to the component in the accelerator 131 that requires the requested data. In some embodiments, with the operations described herein, the accelerator 131 improves existing computer systems because the accelerator 131 can save accelerator power and energy in the accelerator 131 by mainly focusing on accelerator processing operations rather than operations typically associated with locally fetching data from accelerator memory.

[0069] Figure 1C A process flow of the server system 100 is shown in more detail. In some embodiments, Figure 1CThe process flow shown in [Fig.] illustrates that accelerator 131 utilizes the operations described herein to access unified memory 160 (e.g., shared memory 127 and / or memory 126) for direct accelerator processing. In some embodiments, referring to process 149, accelerator 131 is a video die that generates a memory access request for shared memory 127. In some embodiments, unified memory access tunnel system 116 tunnels the AXI protocol onto the UCIe protocol of inter-die interconnect 161 to generate a unified memory access tunnel data packet. In some embodiments, using unified memory access de-tunnel system 117, host CPU 171 de-tunnels the unified memory access tunnel data packet to perform a memory request operation decoded by unified memory access de-tunnel system 117, in which case the unified memory access de-tunnel system accesses data from shared memory 127.

[0070] Similarly, referring to process 148, accelerator 131 generates a memory access request for memory 126. In some embodiments, unified memory access tunnel system 116 tunnels the CXL protocol onto the UCIe protocol of inter-die interconnect 161 to generate a unified memory access tunnel data packet. In some embodiments, using unified memory access de-tunnel system 117, host CPU 171 de-tunnels the unified memory access tunnel data packet to perform a memory request operation decoded by unified memory access de-tunnel system 117. In some embodiments, by utilizing the operations described herein, accelerator 131 is able to immediately and directly perform processing operations using the data accessed at memory 126.

[0071] In some embodiments, processes 165 and 166 illustrate that CPU core 159 accesses memory 126 and shared memory 127 respectively using the shared address space of unified memory 160. As shown, CPU core 159 and accelerator 131 are able to access the same unified memory 160 memory using the shared address space described herein.

[0072] Figure 1DIllustrates a unified memory access method 179 according to some embodiments. The methods, process steps, or stages shown in the figures may be implemented as independent routines or processes, or as part of a larger routine or process. It should be noted that each process step or stage depicted may be implemented as other embodiments such as a device, method, or system of a processor that executes a set of instructions. In some embodiments, at operation 185, the accelerator 131 generates a unified memory access tunnel packet structure that maps to the inter-die interface protocol of the inter-die interconnect 161. In some embodiments, at operation 187, the accelerator 131 uses the unified memory access tunnel packet structure to tunnel the high-level interconnect protocol over the inter-die interface protocol to generate a unified memory access tunnel packet. In some embodiments, at operation 189, the accelerator 131 uses the unified memory access tunnel packet to access the unified memory 160, thereby allowing the accelerator 131 to perform processing operations immediately without an accelerator memory for accelerator processing operations.

[0073] Figure 2 Further illustrates according to some embodiments Figure 1A Block diagram of the SoC 120 in. In some embodiments, the SoC 120 includes a host CPU 171 coupled to the accelerator 131 via the inter-die interconnect 161. In some embodiments, the host CPU 171 includes a host CPU-based coprocessor unit 271, and the accelerator 131 includes an accelerator-based coprocessor unit 272. In some embodiments, the accelerator-based coprocessor unit 272 is the hardware and / or executable code in the accelerator 131 configured to perform the operations of the specified accelerator described herein. In some embodiments, the operations of the specified accelerator are operations to be performed by the accelerator-based coprocessor unit 272 specified by the host CPU 171 and / or the accelerator 131. For example, in some embodiments, the accelerator 131 may be a video codec converter, and the accelerator-based coprocessor unit 272 is configured to perform video codec conversion operations, such as the decoding operations, preprocessing operations, encoding operations, and postprocessing operations described and illustrated by way of example in Figure 3 In.

[0074] In some embodiments, the host CPU-based coprocessor unit 271 is hardware and / or executable code within the host CPU 171 configured to perform the operations of the designated host CPU described herein. In some embodiments, the operations of the designated host CPU are operations to be performed by the host CPU-based coprocessor unit 271 as designated by the host CPU 171 and / or the accelerator 131. For example, in some embodiments, the host CPU-based coprocessor unit 271 may be the following hardware and / or executable code within the host CPU 171: the hardware and / or executable code is configured to perform the demultiplexing operation, decoding operation, preprocessing operation, encoding operation, and multiplexing operation described herein and shown by way of example in Figure 3 In some embodiments, as further described in the text reference Figures 3 to 5 , the host CPU-based coprocessor unit 271 and the accelerator-based coprocessor unit 272 are jointly configured to perform the processing operations routinely performed by the accelerator 131.

[0075] Figure 3 is a block diagram showing the SoC 120 in accordance with some embodiments of Figure 2 As previously described, in some embodiments, the host CPU 171 and the accelerator 131 of the SoC 120 are configured to utilize the host CPU-based coprocessor unit 271 and the accelerator-based coprocessor unit 272 to perform the processing operations routinely performed by the accelerator 131 for the server system 100. In some embodiments, utilizing the host CPU 171 and the accelerator 131 to perform coprocessing operations (e.g., decoding operations, preprocessing operations, and encoding operations) allows the SoC 120 to maximize the efficiency of the server system 100 by performing operations that are more appropriately configured for the hardware and / or software located within the host CPU 171 or the accelerator 131.

[0076] In some embodiments, the host CPU-based coprocessor unit 271 includes a demultiplexer 311, a decoder 312, a preprocessing unit 313, an encoder 314, and a multiplexer 315. In some embodiments, the accelerator-based coprocessor unit 272 includes a decoder 331, a preprocessing unit 332, an encoder 333, and a postprocessing unit 334. In some embodiments, the demultiplexer 311, the decoder 312, the preprocessing unit 313, the encoder 314, and the multiplexer 315 of the host CPU-based coprocessor unit 271 are jointly configured to perform operations together with the decoder 331, the preprocessing unit 332, the encoder 333, and the postprocessing unit 334 of the accelerator-based coprocessor unit 272 to perform the operations described herein.

[0077] In some embodiments, the demultiplexer 311 is based on the following hardware and / or executable code in the coprocessor unit 271 of the host CPU: the hardware and / or executable code is configured to receive the input data stream 340 and divide the input data stream 340 into multiple output data streams defined by the host CPU 171 and / or the accelerator 131. In some embodiments, the input data stream 340 may be, for example, a digital video data stream that has been multiplexed by a video source. In some embodiments, the demultiplexer 311 is configured to divide the input data stream 340 into: (1) a data stream 341 directed to the host CPU decoder and configured to be decoded by the decoder 312 of the host CPU 171; and (2) a data stream 342 directed to the accelerator decoder and configured to be decoded by the decoder 331 of the accelerator 131.

[0078] In some embodiments, the data stream 341 directed to the host CPU decoder is a data stream configured for decoding operations to be performed by the decoder 312 (which may be, for example, a software-based decoder configured to perform software-based decoding operations). In some embodiments, the data stream 341 directed to the host CPU decoder may be a data stream that requires software-based decoding operations that can only be performed by the decoder 312. For example, since the decoder 331 of the accelerator 131 may be a hardware decoder configured to decode a specific type of hardware-specific data stream, when the input data stream (or a portion thereof) is not such a type of input data stream that can be decoded by the decoder 331 (e.g., a hardware-based decoder), the input data stream may be provided by the demultiplexer 311 to the decoder 312 (e.g., a software-based decoder) as the data stream 341 directed to the host CPU decoder. In some embodiments, portions of the input data stream 340 may be specified by the host CPU 171 and / or the accelerator 131 as the data stream 341 directed to the host CPU decoder or the data stream 342 directed to the accelerator decoder. In some embodiments, the host CPU 171 may use a selection signal provided to the demultiplexer 311 to indicate to the demultiplexer 311 the portion of the input data stream 340 that is designated for decoding by the decoder 312 of the host CPU 171 (e.g., the data stream 341 directed to the host CPU decoder) or the portion of the input data stream 340 that is designated for decoding by the decoder 331 of the accelerator 131. In some embodiments, after performing the demultiplexing operation at the demultiplexer 311, the demultiplexer 311 provides the data stream 341 directed to the host CPU decoder to the decoder 312 of the host CPU 171 and provides the data stream 342 directed to the accelerator decoder (e.g., the data stream 342 directed to the accelerator decoder) to the decoder 331 of the accelerator 131.

[0079] In some embodiments, referring to decoder 312 of host CPU 171, decoder 312 receives data stream 341 directed to the host CPU decoder from demultiplexer 311 and begins the process of decoding data stream 341 directed to the host CPU decoder. In some embodiments, decoder 312 is such a software decoder or hardware decoder or a combination thereof: the software decoder or hardware decoder or a combination thereof is configured to perform decoding operations specific to data stream 341 directed to the host CPU decoder provided by demultiplexer 311 (e.g., software-specific data streams that cannot be decoded by decoder 331 due to the hardware configuration of decoder 331). For example, due to the fixed hardware configuration of decoder 331 and the reconfigurable software configuration of decoder 312, decoder 312 can be configured to perform operations specific to data stream 341 directed to the host CPU decoder. In some embodiments, decoder 312 is a decoder configured to perform decoding operations specific to the processing attributes of host CPU 171 and / or decoding operations specific to the non-processing attributes of accelerator 131. In some embodiments, decoder 312 is configured to perform video decoding operations specific to the video data stream provided by demultiplexer 311 of host CPU 171. In some embodiments, after the decoding operation is performed at decoder 312, decoder 312 provides decoded output data stream 344 to preprocessing unit 332 for preprocessing of decoded output data stream 344.

[0080] In some embodiments, referring to decoder 331 of accelerator 131, decoder 331 receives data stream 342 directed to the accelerator decoder from demultiplexer 311 and begins the process of decoding data stream 342 directed to the accelerator decoder. In some embodiments, decoder 331 is a hardware decoder or software decoder or a combination thereof configured to perform decoding operations specific to accelerator 131. In some embodiments, decoder 331 is a decoder configured to perform decoding operations specific to data stream 342 directed to the accelerator decoder provided by demultiplexer 311 of host CPU 171. For example, in some embodiments, due to the fixed hardware configuration of decoder 331, decoder 331 can be configured to decode only data streams mapped to the fixed hardware configuration of decoder 331. In some embodiments, decoder 331 is a video decoder configured to perform video decoding operations specific to the video data stream provided by demultiplexer 311 of host CPU 171 (e.g., data stream 342 directed to the accelerator decoder). In some embodiments, after the decoding operation is performed at decoder 331, decoder 331 provides decoded output data stream 343 to preprocessing unit 332 for preprocessing of decoded output data stream 343.

[0081] In some embodiments, the preprocessing unit 332 receives the decoded output data stream 343 from the decoder 331 and the decoded output data stream 344 from the decoder 312, and starts the process of performing a shared preprocessing operation with the preprocessing unit 313 of the host CPU 171. In some embodiments, the preprocessing unit 332 is hardware and / or executable code located in the accelerator 131, and the hardware and / or executable code is configured to perform the following operations: (1) evaluate the received decoded data stream to determine whether the received input data stream is configured as an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream; (2) perform accelerator-specific preprocessing operations; and (3) share host CPU-specific preprocessing operations with the preprocessing unit 313 of the host CPU-based coprocessor unit 271. In some embodiments, the preprocessing unit 313 is hardware and / or executable code located in the host CPU 171, and the hardware and / or executable code is configured to perform host CPU-specific processing operations on the received host CPU-specific preprocessing data stream 346 from the accelerator 131. In some embodiments, the accelerator-specific preprocessing data stream is a data stream configured to be preprocessed by the preprocessing unit 332 of the accelerator 131. In some embodiments, the host CPU-specific preprocessing data stream is a data stream configured to be preprocessed by the preprocessing unit 313 of the host CPU 171.

[0082] In some embodiments, the preprocessing unit 332 receives the decoded output data stream 343 and the decoded output data stream 344, and determines whether the received decoded data stream (or a portion thereof) is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the preprocessing unit 332 determines whether the received decoded data stream is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream by evaluating a preprocessing operation configuration associated with the received decoded data stream. In some embodiments, the preprocessing operation configuration serves as an indication of whether the received decoded data stream is an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the preprocessing operation configuration can be evaluated by identifying a data stream identification (ID) in the received decoded data stream. In some embodiments, the data stream ID is a unique identifier used to identify and manage data streams, and in this case, the data stream ID is associated with being an accelerator-specific preprocessing data stream or a host CPU-specific preprocessing data stream. In some embodiments, the data stream ID can be assigned by the operating system of the SoC 120 and / or the accelerator 131 when the data stream is created, and is used by the accelerator 131 to identify and manage the data stream. In some embodiments, the accelerator 131 utilizes the data stream ID to schedule the preprocessing of the decoded data stream and determine whether to switch between multiple decoded data streams for preprocessing by the host CPU 171 or the accelerator 131 (and allocate resources such as memory and processing time for each data stream). In some embodiments, the data stream ID is either mapped to an accelerator-specific operation configured to be executed by the preprocessing unit 332 of the accelerator 131 or mapped to a host CPU-specific operation configured to be executed by the preprocessing unit 313.

[0083] In some embodiments, when the decoded data stream is identified by the preprocessing unit 332 as an accelerator-specific preprocessing data stream, the accelerator-specific preprocessing data stream remains at the preprocessing unit 332 for accelerator-specific preprocessing. In some embodiments, the preprocessing unit 332 uses accelerator-specific preprocessing operations to preprocess the accelerator-specific processing data stream to generate an accelerator-specific preprocessed output data stream 345.

[0084] In some embodiments, when the data stream is recognized by the preprocessing unit 332 as a host CPU specific preprocessing data stream after decoding, the preprocessing unit 332 provides this data stream as the host CPU specific preprocessing data stream 346 to the preprocessing unit 313 for host CPU specific preprocessing. In some embodiments, the preprocessing unit 313 preprocesses the host CPU specific preprocessing data stream 346 using host CPU specific preprocessing operations to generate a preprocessed output data stream 347. In some embodiments, the preprocessing unit 313 provides the preprocessed output data stream 347 to the preprocessing unit 332. In some embodiments, the preprocessing unit 332 receives the preprocessed output data stream 347 and provides the preprocessed output data stream 347 and the accelerator specific preprocessed output data stream 345 together as the preprocessed output data stream 348 to the encoder 333.

[0085] In some embodiments, the encoder 333 receives the preprocessed output data stream 348 from the preprocessing unit 332 and begins the process of performing a shared encoding operation with the encoder 314 of the host CPU 171. In some embodiments, the encoder 333 is a hardware and / or software encoder configured to perform the following operations: (1) evaluate the preprocessed output data stream 348 to identify an accelerator specific encoding data stream and a host CPU specific encoding data stream; (2) perform encoding operations specific to the encoder 333 (e.g., accelerator specific encoding operations); and (3) share the host CPU specific encoding operations with the encoder 314 of the co - processing unit 271 based on the host CPU. In some embodiments, the encoder 314 is a software encoder and / or hardware encoder in the host CPU 171 that is configured to perform host specific encoding operations on the host CPU specific encoding data stream 349 provided by the encoder 333. For example, in some embodiments, the encoder 333 is a video encoder configured to perform accelerator specific video encoding operations with a fixed hardware configuration specific to the encoder 333. In some embodiments, the encoder 314 is a software video encoder configured to perform host CPU specific video encoding operations, where the host CPU specific video encoding operations: (1) cannot be performed by the encoder 333, for example, due to the fixed configuration of the encoder 333; or (2) can be performed more efficiently by the encoder 314 using the unique processing capabilities of the host CPU 171.

[0086] In some embodiments, the encoder 333 receives the pre - processed output data stream 348 from the pre - processing unit 332 and evaluates the pre - processed output data stream 348 to identify the accelerator - specific encoded data stream and the host CPU - specific encoded data stream. In some embodiments, the encoder 333 identifies the accelerator - specific encoded data stream or the host CPU - specific encoded data stream in the data stream by searching for specific markers in the pre - processed output data stream 348 that indicate whether a portion of the pre - processed output data stream 348 is an accelerator - specific encoded data stream or a host CPU - specific encoded data stream. In some embodiments, for example, the encoder 333 searches for the accelerator - specific encoded data stream marker and the host CPU - specific encoded data stream marker in the pre - processed output data stream 348.

[0087] In some embodiments, when the encoder 333 identifies the pre - processed output data stream 348 or a portion thereof as an accelerator - specific encoded data stream, the encoder 333 encodes the accelerator - specific encoded data stream at the encoder 333 to generate the accelerator - specific encoded output data stream 336. In some embodiments, when the encoder 333 identifies a portion of the pre - processed output data stream 348 as the host CPU - specific encoded data stream 349, the encoder 333 provides the host CPU - specific encoded data stream 349 to the encoder 314. In some embodiments, the encoder 314 receives the host CPU - specific encoded data stream 349 and encodes the host CPU - specific encoded data stream 349 using the host CPU - specific encoding operation provided by the encoder 314 of the host CPU 171. In some embodiments, after performing the host CPU - specific encoding operation, the encoder 314 provides the encoded output as the host CPU - specific encoded output 361 to the encoder 333. In some embodiments, the encoder 333 receives the host CPU - specific encoded output 361 from the encoder 314 and provides the host CPU - specific encoded output 361 and the accelerator - specific encoded output data stream 336 together as the encoded - pre - processed output data stream 365 to the post - processing unit 334.

[0088] In some embodiments, the post - processing unit 334 receives the encoded - pre - processed output data stream 365 and performs post - processing operations on the encoded - pre - processed output data stream 365. In some embodiments, the post - processing unit 334 is hardware and / or executable code configured to perform post - processing operations on the encoded - pre - processed output data stream 365 of the encoder 333, such as data compression, error correction, or other post - processing operations. In some embodiments, the post - processing unit 334 provides the post - processed data stream 368 to the multiplexer 315 for further processing or storage by the SoC 120.

[0089] Figure 4is a block diagram showing a shared pre - processing flow of the SoC 120 according to some embodiments. In some embodiments, as Figure 4 shown in the pre - processing flow in, the pre - processing unit 332 includes an accelerator - specific pre - processing unit 411 and an accelerator - specific pre - processing unit 413. In some embodiments, the pre - processing unit 313 includes a host CPU - specific pre - processing unit 412. In some embodiments, the accelerator - specific pre - processing unit 411 is hardware and / or executable code within the pre - processing unit 332 configured to perform accelerator - specific pre - processing operations on accelerator - specific pre - processing data streams identified by the pre - processing unit 332. In some embodiments, the host CPU - specific pre - processing unit 412 is hardware and / or executable code within the pre - processing unit 313 configured to perform host CPU - specific pre - processing operations on host CPU - specific pre - processing data streams identified by the pre - processing unit 332. In some embodiments, the accelerator - specific pre - processing unit 413 is hardware and / or executable code within the pre - processing unit 332 that, in addition to being configured to perform accelerator - specific pre - processing operations on accelerator - specific pre - processing data streams identified by the pre - processing unit 332, is also configured to perform additional pre - processing operations on the output of the host CPU - specific pre - processing unit 412 (e.g., the pre - processed output data stream 347) and / or to combine the output of the host CPU - specific pre - processing unit 412 with the accelerator - specific pre - processed output data stream (e.g., the accelerator - specific pre - processed output data stream 345).

[0090] As part of the shared pre - processing flow, the pre - processing unit 332 identifies accelerator - specific pre - processing data streams and pre - processes the accelerator - specific data streams at the accelerator - specific pre - processing unit 411. In some embodiments, the pre - processing unit 332 identifies host CPU - specific pre - processing data streams and provides the host CPU - specific pre - processing data streams for host CPU - specific pre - processing at the host CPU - specific pre - processing unit 412. In some embodiments, the pre - processing unit 332 receives the pre - processed output data stream from the host CPU - specific pre - processing unit 412 and combines the pre - processed output data stream with the accelerator - specific pre - processed output data stream.

[0091] Figure 5 is a block diagram showing a shared encoding flow of the SoC 120 according to some embodiments. In some embodiments, as Figure 5As shown in the shared encoding process, encoder 333 includes an accelerator-specific encoding unit 511 and an accelerator-specific encoding unit 513. In some embodiments, encoder 314 includes a host CPU-specific encoding unit 512. In some embodiments, the accelerator-specific encoding unit 511 is hardware and / or executable code within encoder 333 that is configured to perform accelerator-specific encoding operations on accelerator-specific encoding data streams identified by encoder 333. In some embodiments, the host CPU-specific encoding unit 512 is hardware and / or executable code within encoder 314 that is configured to perform host CPU-specific encoding operations on host CPU-specific encoding data streams identified by encoder 333. In some embodiments, the accelerator-specific encoding unit 513 is hardware and / or executable code within encoder 333 that, in addition to being configured to perform accelerator-specific encoding operations on accelerator-specific encoding data streams identified by encoder 333, is further configured to perform additional encoding operations on the output of the host CPU-specific encoding unit 512 (e.g., the host CPU-specific encoded output 361) and / or combine the output of the host CPU-specific encoding unit 512 with the accelerator-specific encoded output data stream (e.g., the accelerator-specific encoded output data stream 336).

[0092] As part of the shared encoding process, encoder 333 identifies accelerator-specific encoding data streams and encodes the accelerator-specific encoding data streams at the accelerator-specific encoding unit 511. In some embodiments, encoder 333 identifies host CPU-specific encoding data streams and provides the host CPU-specific encoding data streams for host CPU-specific encoding at the host CPU-specific encoding unit 512. In some embodiments, encoder 333 receives the encoded output data stream from the host CPU-specific encoding unit 512 and combines the encoded output data stream with the accelerator-specific encoded output data stream.

[0093] In some embodiments, the existing computer system is improved by the operations described herein because the SoC 120 is capable of dynamically switching between the hardware decoder in the accelerator 131 and the software decoder in the host CPU 171 to avoid hardware codec issues (e.g., new codec configurations, error concealment). In some embodiments, multiple preprocessing operations are dynamically partitioned between the preprocessor of the host CPU 171 and the preprocessor of the accelerator 131 to achieve more flexible algorithms and power usage. In some embodiments, the critical decisions of the encoder (e.g., mode decision or frame parameters of the convex-hull approach) can be executed on the host CPU 171 instead of on the accelerator 131 to obtain better video quality and bitrate. In some embodiments, the fine-grained interaction between the host CPU 171 and the accelerator 130 (e.g., accelerator hardware) enables the framework described herein to be used for improved co-processing.

[0094] Figure 6 which shows a Figure 1A block diagram of the SoC 120 in accordance with some embodiments. In some embodiments, the SoC 120 includes a host CPU 171, a memory 104, accelerators 131 to 134, and a flash device 140. In some embodiments, the flash device 140 is a storage device configured to store accelerator firmware associated with the accelerators 131 to 134 and host CPU firmware associated with the host CPU firmware. In some embodiments, the accelerator firmware can be stored in the form of an accelerator firmware image file, and the host CPU firmware can be stored in the form of a host CPU firmware image file. In some embodiments, the flash device 140 can be, for example, an embedded Multi-Media Controller (EMMC) device or a Universal Flash Storage (UFS) device.

[0095] In some embodiments, as described above, the accelerators 131 to 135 are coupled to the host CPU 171 using the inter-die interconnects 161 to 164. In some embodiments, each of the inter-die interconnects 161 to 164 includes a sideband bus and an inter-die bus (mainband bus) (e.g., Figure 7The inter-die bus 770 and the sideband bus 790 shown in []. In some embodiments, the sideband bus is a set of dedicated communication lines that are configured to transmit control and management information between the accelerators 131 to 134 and the host CPU 171, and are also configured to transmit accelerator firmware components (e.g., critical accelerator firmware components) between the accelerators 131 to 134 and the host CPU 171 based on the evaluation of the accelerator firmware components performed by the root of trust (ROT) 620 of the host CPU 171. In some embodiments, the inter-die bus (or main bus) is a set of communication lines optimized for high-speed data transmission that are configured to transmit data and payloads between the accelerators 131 to 134 and the host CPU 171, and are also configured to transmit accelerator firmware components, such as accelerator firmware components that are considered non-critical firmware components, between the host CPU 171 and the accelerators 131 to 134.

[0096] In some embodiments, the host CPU 171 includes a root of trust (ROT) 620. In some embodiments, the ROT 620 is a secure hardware module and / or executable code or trusted execution environment (TEE) within the host CPU 171 that is configured to instantaneously authenticate accelerator firmware based on the accelerator firmware evaluation of the accelerator firmware, in addition to performing traditional root of trust operations in a trusted computing environment. In some embodiments, the accelerator firmware evaluation performed by the ROT 620 includes, for example, determining whether the various parts of the accelerator firmware associated with the accelerators coupled to the host CPU 171 are critical accelerator firmware components or non-critical accelerator firmware components of the accelerator firmware. In some embodiments, as further described herein Figures 7 to 10 Based on the results of the accelerator firmware evaluation, the host CPU 171 provides the associated accelerators with the certified critical accelerator firmware components for processing via the sideband bus that connects the accelerators to the host CPU 171, and provides the associated accelerators with the certified non-critical accelerator firmware components for processing via the inter-die bus that connects the accelerators to the host CPU 171.

[0097] Figure 7 Shows according to some embodiments Figure 6Block diagram of the SoC 120. In some embodiments, the host CPU 171 is coupled to the flash device 140, the memory 104, and the accelerators 131 to 134. In some embodiments, the host CPU 171 is coupled to the accelerators 131 to 134 via the inter-die bus 770 and the sideband bus 790. In some embodiments, the accelerators 131 to 134 include an accelerator embedded microcontroller and accelerator memory. For example, in some embodiments, the accelerator 131 may include an accelerator embedded microcontroller 740 and an accelerator memory 741. In some embodiments, the accelerator embedded microcontroller 740 is an embedded controller in the accelerator 131, and the embedded controller is configured to coordinate the data flow between the accelerator 131 and other components of the SoC 120. In some embodiments, the accelerator embedded microcontroller 740 is configured to utilize the authentication control unit 781, and the authentication control unit may be configured to control the operations associated with the authentication operations performed by the host CPU 171 within the accelerator 131.

[0098] In some embodiments, the ROT 620 includes a security agent 721, and the security agent is configured to perform the sideband-based accelerator firmware authentication method described herein by utilizing the accelerator firmware identification unit 756, the accelerator firmware authentication unit 752, the accelerator firmware parsing unit 753, and / or the accelerator firmware component size determination unit 754. In some embodiments, the accelerator firmware parsing unit 753 is hardware and / or executable code that is configured to parse or partition the accelerator firmware into multiple accelerator firmware components by examining the code structure of the accelerator firmware to identify the unique functional components or modules of the accelerator firmware and splitting the accelerator firmware into each uniquely identified functional component or module (e.g., accelerator firmware component).

[0099] In some embodiments, the accelerator firmware identification unit 756 is hardware and / or executable code configured to identify accelerator firmware components parsed by the accelerator firmware parsing unit 753 as critical accelerator firmware components (e.g., accelerator firmware components critical for performing processing operations of the accelerator 131) and non-critical accelerator firmware components (e.g., accelerator firmware components non-critical for performing processing operations of the accelerator 131). In some embodiments, the accelerator firmware identification unit 756 identifies critical non-accelerator components and non-critical accelerator firmware components of the accelerator firmware based on accelerator-specific information stored in the accelerator firmware image file. In some embodiments, the accelerator-specific information may be included in the form of headers, sections, symbols, or other metadata that define the structure and organization of the accelerator firmware. For example, in some embodiments, in the Executable and Linkable Format (ELF) or Common Object File Format (COFF) format, the accelerator firmware image file may include sections and symbols that define the various components of the firmware and their roles in the overall system. In some embodiments, the accelerator firmware identification unit 756 may utilize headers and metadata associated with these sections and symbols to identify critical and non-critical accelerator firmware components of the accelerator firmware. In some embodiments, the accelerator firmware identification unit 756 may utilize accelerator-specific information to determine which components are critical and which are not based on the specific requirements of the accelerator 131 and the SoC 120. For example, the accelerator firmware identification unit 756 may identify the bootloader, drivers, and low-level software as critical accelerator firmware components because these accelerator firmware components may be required for the correct operation of the accelerator 131 and the overall system (e.g., the SoC 120). In some embodiments, the host CPU 171 may identify applications, libraries, and high-level software as non-critical accelerator firmware components because these accelerator firmware components provide additional functionality but are not strictly required for the operation of the accelerator.

[0100] In some embodiments, the accelerator firmware component size determination unit 754 is hardware and / or executable code configured to determine the accelerator firmware component size of the accelerator firmware and the size of the accelerator firmware. In some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 is configured to determine the accelerator firmware component size of the accelerator firmware and the size of the accelerator firmware by evaluating size information (such as a table of headers or content) provided by the accelerator firmware itself to determine the size and location of each component. In some embodiments, the size information may be included in the firmware image of the accelerator firmware and may be used by the ROT 620 to divide the firmware into discrete accelerator firmware components.

[0101] In some embodiments, the accelerator firmware authentication unit 752 is hardware and / or executable code configured to perform accelerator firmware authentication operations for accelerators 131 to 134. In some embodiments, the accelerator firmware authentication unit 752 includes an immediate accelerator firmware component authentication unit 761 and a delayed accelerator firmware component authentication unit 762. In some embodiments, the immediate accelerator firmware component authentication unit 761 is hardware and / or executable code configured to receive an accelerator firmware component (e.g., a critical accelerator firmware component) and immediately authenticate the accelerator firmware component. In some embodiments, immediate authentication of performance-critical firmware means that critical accelerator firmware components are authenticated immediately or instantaneously without latency at the host CPU 171 such that the authenticated accelerator firmware can be directly provided to the associated accelerator via a sideband bus (e.g., sideband bus 791 for accelerator 131) coupled to the associated accelerator.

[0102] In some embodiments, the delayed accelerator firmware component authentication unit 762 is hardware and / or executable code configured to authenticate accelerator components (e.g., non-critical accelerator firmware components) at a delayed time indicated or commanded by the host CPU 171. For example, in some embodiments, delayed authentication refers to authentication performed by the delayed accelerator firmware component authentication unit 762, which is delayed by the ROT 620 such that non-critical accelerator firmware components are authenticated only after the critical accelerator firmware components have been authenticated by the immediate accelerator firmware component authentication unit 761.

[0103] In some embodiments, as described above, the security agent 721 is configured to use the accelerator firmware identification unit 756, the accelerator firmware authentication unit 752, the accelerator firmware parsing unit 753, and / or the accelerator firmware component size determination unit 754 to perform the sideband-based accelerator firmware authentication method described herein. In some embodiments, the operation of the SoC 120 is described below with reference to Figures 8 to 10 the operation of the SoC 120 is described.

[0104] Figure 8FIG. 0 is a flowchart showing a sideband-based accelerator firmware authentication method 800 according to some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 800 is configured to authenticate the accelerator firmware immediately (e.g., instantaneously) or non-immediately (e.g., non-instantaneously) based on an accelerator firmware evaluation performed on the accelerator firmware by the security agent 721 of the ROT 620. In some embodiments, as part of the accelerator firmware evaluation, the security agent 721 divides the accelerator firmware into a plurality of accelerator firmware components, and these accelerator firmware components are considered critical accelerator firmware components or non-critical accelerator firmware components. In some embodiments, the accelerator firmware components considered critical are provided to the associated accelerator for processing via a sideband bus of the inter-die interconnect. In some embodiments, the accelerator firmware considered non-critical is provided to the associated accelerator for processing via an inter-die bus of the inter-die interconnect. The methods, process steps, or stages shown in the figures may be implemented as independent routines or processes, or as part of a larger routine or process. It should be noted that each of the depicted process steps or stages may be implemented as other embodiments such as a device, method, or system including a processor that executes a set of instructions.

[0105] In some embodiments, at operation 810, the security agent 721 of the ROT 620 reads the accelerator firmware from the flash device 140 (e.g., non-volatile memory). In some embodiments, as previously described, the accelerator firmware read from the flash device 140 may be associated with a particular accelerator (e.g., accelerator 131, etc.) and may be stored in the flash device 140 in the form of an accelerator firmware image file. In some embodiments, reading the accelerator firmware from the flash device 140 occurs during system bootup of the SoC 120. In some embodiments, once the accelerator firmware is read from the flash device 140, the accelerator firmware is provided to the accelerator firmware parsing unit 753.

[0106] In some embodiments, at operation 815, the accelerator firmware parsing unit 753 of the security agent 721 receives the accelerator firmware from the flash device 140 and parses the accelerator firmware into a plurality of accelerator firmware components. In some embodiments, the accelerator firmware parsing unit 753 parses the accelerator firmware into a plurality of accelerator firmware components by examining the code structure of the accelerator firmware to identify the unique functional components or modules of the accelerator firmware, and splitting the accelerator firmware into each uniquely identified functional component or module. In some embodiments, the accelerator firmware parsing unit 753 identifies each accelerator firmware component by scanning the unique digital signature of the accelerator firmware representing each functional component or module. In some embodiments, after parsing the accelerator firmware into a plurality of accelerator firmware components, the accelerator firmware parsing unit 753 provides these accelerator firmware components to the accelerator firmware identification unit 756, and operation 815 proceeds to operation 820.

[0107] In some embodiments, at operation 820, the accelerator firmware identification unit 756 receives the plurality of accelerator firmware components from the accelerator firmware parsing unit 753 and evaluates the plurality of accelerator firmware components of the accelerator firmware to identify the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware. In some embodiments, the accelerator firmware identification unit 756 identifies the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware by analyzing the accelerator firmware metadata and other accelerator firmware code associated with each accelerator firmware component. For example, in some embodiments, the accelerator firmware identification unit 756 identifies the critical accelerator firmware components and non-critical accelerator firmware components of the accelerator firmware by analyzing the dependencies (e.g., mutual dependencies) of the accelerator firmware components, analyzing the metadata associated with each accelerator firmware component, and analyzing the previous versions of the accelerator firmware and each accelerator firmware component. For example, in some embodiments, the accelerator firmware identification unit 756 examines the dependencies between the various accelerator firmware components in the accelerator firmware by determining which accelerator firmware components are necessary (e.g., critical) for the normal operation of other accelerator firmware components and which components are not necessary (e.g., non-critical) for the normal operation of other components. In some embodiments, the accelerator firmware identification unit 756 identifies the dependencies of different accelerator firmware components by scanning the accelerator firmware code for any inter-component communication mechanisms and examining the inter-component communication mechanisms to determine the type of inter-component communication dependencies (e.g., function calls or shared data structures). In another example, in some embodiments, the accelerator firmware identification unit 756 utilizes the metadata to identify the critical accelerator firmware components and non-critical accelerator firmware components by scanning the metadata to find the version number or annotation associated with each accelerator firmware component that indicates the importance (e.g., critical or non-critical) of the accelerator firmware component.

[0108] In some embodiments, the accelerator firmware identification unit 756 identifies critical and non-critical accelerator firmware components by using a previous version of the accelerator firmware, by comparing the current version of the accelerator firmware components with the previous version of the accelerator firmware components to identify any changes in the accelerator firmware components. In some embodiments, no change from the previous version of the accelerator firmware components to the current version of the accelerator firmware components may indicate that the accelerator firmware components are not critical accelerator firmware components, while a change from the previous version of the accelerator firmware components to the current version of the accelerator firmware components may indicate that the accelerator firmware components are critical accelerator firmware components. In some embodiments, after the accelerator firmware identification unit 756 identifies a component as a non-critical accelerator firmware component or a critical accelerator firmware component, the accelerator firmware identification unit 756 provides the non-critical accelerator firmware components to the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752, and provides the critical accelerator firmware components to the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752. In some embodiments, before providing the non-critical accelerator firmware components to the delayed accelerator firmware component authentication unit 762, the non-critical accelerator firmware components may be stored in a secure area (e.g., a secure boot ROM) of the memory of the ROT 620.

[0109] In some embodiments, at operation 825, the instant accelerator firmware component authentication unit 761 receives a critical accelerator firmware component and instantaneously authenticates the critical accelerator firmware component. In some embodiments, as described above, instantaneously authenticating performance-critical firmware means instantaneously authenticating, without latency, a critical accelerator firmware component at the host CPU 171 such that the authenticated accelerator firmware can be provided directly to the associated accelerator via a sideband bus coupled to the associated accelerator (e.g., sideband bus 791 for accelerator 131). In some embodiments, the accelerator firmware authentication unit 752 is configured to format the packet structure of an accelerator firmware component data packet sent to the associated accelerator via the sideband bus such that the packet structure indicates to the associated accelerator that an accelerator firmware component is being transmitted via the sideband bus. For example, in some embodiments, the accelerator firmware authentication unit 752 is configured to format the packet structure of an accelerator firmware data packet sent to accelerator 131 via sideband bus 791 such that the packet structure indicates to accelerator 131 that a critical accelerator firmware component is being transmitted in the data packet. In some embodiments, bit positions in the packet structure of the accelerator firmware component data packet can indicate to accelerator 131 that a critical accelerator firmware component is being transmitted in the data packet. In some embodiments, bit positions in the packet structure of the accelerator firmware data packet can indicate to the accelerator that a non-critical accelerator firmware component is being transmitted in via an inter-die bus (e.g., inter-die bus 771 associated with accelerator 131). In some embodiments, after instantaneously authenticating the accelerator firmware component, operation 825 proceeds to operation 830.

[0110] In some embodiments, at operation 830, the host CPU 171 provides the authenticated critical accelerator firmware component to accelerator 131 via sideband bus 791. In some embodiments, after providing the authenticated critical accelerator firmware component to accelerator 131 via sideband bus 791, operation 830 proceeds to operation 835.

[0111] In some embodiments, at operation 835, accelerator 131 receives the accelerator firmware component via sideband bus 791 and executes the critical accelerator firmware component. In some embodiments, upon receiving the accelerator firmware component, the authentication control unit 781 of accelerator 131 is configured to scan the data packet for a bit indicator indicating that the received data packet is an accelerator firmware component. In some embodiments, accelerator 131 is configured to scan the data packet for a bit indicator indicating that an associated non-critical accelerator firmware component is being transmitted via inter-die bus 771 for execution by accelerator 131. In some embodiments, after accelerator 131 executes the critical accelerator firmware component, operation 835 proceeds to operation 840.

[0112] In some embodiments, returning to reference operation 820, when the accelerator firmware identification unit 756 regards the accelerator firmware component as a non-critical accelerator firmware component, at operation 850, the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 receives the non-critical accelerator firmware component and authenticates the non-critical accelerator firmware component. In some embodiments, the delayed accelerator firmware component authentication unit 762 uses delayed authentication to authenticate the non-critical accelerator firmware component. In some embodiments, delayed authentication refers to the authentication performed by the delayed accelerator firmware component authentication unit 762, which is delayed by the ROT 620 such that the non-critical accelerator firmware component is authenticated after the critical accelerator firmware component has been authenticated by the immediate accelerator firmware component authentication unit 761. In some embodiments, after authenticating the critical accelerator firmware component at the delayed accelerator firmware component authentication unit 762, operation 850 proceeds to operation 855.

[0113] In some embodiments, at operation 855, after authenticating the critical accelerator firmware component at the delayed accelerator firmware component authentication unit 762, the ROT 620 provides the authenticated non-critical accelerator firmware component to the memory 104 for storage. In some embodiments, after the storage in the memory 104, operation 855 proceeds to operation 860.

[0114] In some embodiments, at operation 860, the authentication control unit 781 of the accelerator embedded u-controller 740 installs the non-performance-critical firmware from the memory 104 into the accelerator memory 741 of the accelerator 131 via the inter-die bus 771. In some embodiments, operation 860 proceeds to operation 840, at which the authenticated non-performance-critical firmware provided via the inter-die bus 771 is executed by the accelerator 131.

[0115] Figure 9 is a block diagram showing an example process flow utilized in the Figure 7 SoC 120 according to some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 800 is utilized in the example process flow in the Figure 9 and is configured to immediately (e.g., instantaneously) authenticate a first accelerator firmware component (e.g., the bootloader accelerator firmware component) based on an accelerator firmware evaluation of the accelerator firmware, and non-immediately (e.g., non-instantaneously) delay the authentication of a second accelerator firmware component (e.g., the body of the accelerator firmware).

[0116] In some embodiments, in step S1, the security agent 721 of ROT 620 reads the accelerator firmware from the flash device 140 at startup. After reading the accelerator firmware from the flash device 140, the accelerator firmware parsing unit 753 and the accelerator firmware identification unit 756 parse the accelerator firmware into a plurality of accelerator firmware components (e.g., a bootloader accelerator firmware component and a main accelerator firmware component), and identify each accelerator firmware component of the accelerator firmware (e.g., a critical accelerator firmware component and a non-critical accelerator firmware component). In some embodiments, after the accelerator firmware parsing unit 753 and the accelerator firmware identification unit 756 have parsed the accelerator firmware into a bootloader accelerator firmware component and a main accelerator firmware component and have identified each accelerator firmware component as a critical accelerator firmware component and a non-critical accelerator firmware component, the accelerator firmware authentication unit 752 immediately authenticates the bootloader accelerator firmware component ("bootloader") at the immediate accelerator firmware component authentication unit 761. In some embodiments, after delaying the authentication of the main accelerator firmware component of the accelerator firmware (e.g., delaying until the bootloader accelerator firmware component is authenticated), the delayed accelerator firmware component authentication unit 762 authenticates the main accelerator firmware component. In some embodiments, an unauthenticated accelerator firmware component (e.g., the main body of the accelerator firmware component) can be temporarily stored in a secure area of the memory of ROT 620, such as a secure boot ROM or a secure enclave within a trusted execution environment (TEE). In some embodiments, the unauthenticated accelerator firmware component is stored in the secure boot memory until the accelerator firmware authentication unit 752 authenticates the unauthenticated accelerator firmware component.

[0117] In some embodiments, in step S2A, immediately following the authentication of the bootloader accelerator firmware component, the security agent 721 provides or pushes the bootloader accelerator firmware component to the accelerator memory 741 of the accelerator 131 via the sideband bus 791. In some embodiments, the security agent 721 directly provides the bootloader accelerator firmware component to the accelerator memory 741 via the sideband bus 791 without the help of a storage controller located in the accelerator 131. In some embodiments, the security agent 721 provides the bootloader to the storage controller in the accelerator 131 via the sideband bus 791 before the bootloader is written to the accelerator memory 741.

[0118] In some embodiments, at step S2B, the security agent 721 writes the main accelerator firmware component of the accelerator firmware into the memory 104 for transmission to the accelerator 131 at step S4. In some embodiments, at step S3, after the inter-die bus 771 becomes operable, the host CPU 171 performs device initialization and releases the reset of the accelerator embedded u-controller 740. In some embodiments, releasing the reset of the accelerator embedded u-controller 740 means that the host CPU 171 sends a reset signal to the accelerator embedded u-controller, and this reset signal enables the accelerator 131 to start executing instructions and control the internal operations of the accelerator 131. In some embodiments, releasing the reset of the accelerator embedded u-controller 740 enables the accelerator 131 to start operating and performing the expected functions of the accelerator 131, and is generally the initial step in the startup process of the entire SoC 120.

[0119] In some embodiments, at step S4, after releasing the reset of the accelerator embedded u-controller, the accelerator embedded u-controller 740 executes the bootloader accelerator firmware component, downloads the authenticated main accelerator firmware component from the memory 104 via the inter-die bus 771, and executes the authenticated main accelerator firmware component at the accelerator 131.

[0120] Figure 10 FIG. is a flowchart showing a sideband-based accelerator firmware authentication method 1000 according to some embodiments. In some embodiments, the sideband-based accelerator firmware authentication method 1000 is a method implemented by the SoC 120, and this method is configured to use the size of the accelerator firmware component to determine whether the accelerator firmware component is to be immediately authenticated and provided to the accelerator 131 via the sideband bus 791. In some embodiments, for example, according to the design of the accelerator or the SoC 120, the accelerator firmware component can be, for example, a bootloader accelerator firmware component, a main accelerator firmware component, or other non-bootloader accelerator firmware components. The methods, process steps, or stages shown in the figures can be implemented as independent routines or processes, or as part of a larger routine or process. It should be noted that each of the depicted process steps or stages can be implemented as other embodiments such as a device, method, or system including a processor that executes a set of instructions.

[0121] In some embodiments, at operation 1010, the security agent 721 of the ROT 620 reads accelerator firmware from the flash device 140 at startup. In some embodiments, at operation 1015, after reading the accelerator firmware from the flash device 140, the accelerator firmware parsing unit 753 of the security agent 721 receives the accelerator firmware and parses the accelerator firmware into a plurality of accelerator firmware components. In some embodiments, at operation 1020, after parsing the accelerator firmware into a plurality of accelerator firmware components, the accelerator firmware component size determination unit 754 of the security agent 721 determines the size of each accelerator firmware component of the accelerator firmware and the size of the accelerator firmware. In some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 determines the size of the accelerator firmware components and the size of the accelerator firmware by evaluating the size information provided by the firmware itself to determine the size and location of each component. In some embodiments, the size information may be included in the firmware image of the accelerator firmware and may be used by the ROT 620 to divide the firmware into discrete multiple accelerator firmware components. In some embodiments, the ROT 620 may also utilize a heuristic or algorithm to determine the size of the accelerator firmware components. For example, in some embodiments, the ROT 620 may estimate the size of the accelerator firmware components based on the storage amount required to execute the expected functions of the accelerator firmware components. In some embodiments, as described above, the accelerator firmware components may be, for example, a bootloader accelerator firmware component or a non-bootloader accelerator firmware component associated with the accelerator in the SoC 120.

[0122] In some embodiments, at operation 1020, the accelerator firmware component size determination unit 754 of the security agent 721 determines whether the size of each individual accelerator firmware component is less than the immediate authentication size threshold 755. For example, in some embodiments, the accelerator firmware component size determination unit 754 of the security agent 721 determines whether the size of the bootloader accelerator firmware component is less than the immediate authentication size threshold 755. In some embodiments, the immediate authentication size threshold 755 is a threshold: the accelerator firmware component size determination unit 754 uses this threshold to determine whether the accelerator firmware component is to be immediately authenticated by the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752 or whether the accelerator firmware component is to be authenticated by the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 at a delayed time. In some embodiments, the immediate authentication size threshold 755 may be a byte size value of 10 gigabytes, 20 gigabytes, or some other byte size value that can be used as an immediate authentication threshold to determine whether the accelerator firmware component is to be immediately authenticated by the immediate accelerator firmware component authentication unit 761 of the accelerator firmware authentication unit 752.

[0123] In some embodiments, at operation 1050, when the accelerator firmware component sizing unit 754 determines that the size of the accelerator firmware component is less than the immediate authentication size threshold 755, the security agent 721 instantaneously authenticates the accelerator firmware component using the immediate accelerator firmware component authentication unit 761. In some embodiments, for example, when the accelerator firmware identification unit 756 determines that the size of the bootloader accelerator firmware component is below the immediate authentication size threshold, the security agent 721 instantaneously authenticates the bootloader accelerator firmware component at the security agent 721 of the host CPU 171 using the immediate accelerator firmware component authentication unit 761 without storing the bootloader accelerator firmware component in the memory 104. In some embodiments, since the size of the bootloader accelerator firmware component can be relatively small compared to the overall size of the accelerator firmware, the bootloader accelerator firmware component is the component in the accelerator firmware that is instantaneously authenticated by the host CPU 171.

[0124] In some embodiments, at operation 1055, after the accelerator firmware component is instantaneously authenticated by the immediate accelerator firmware component authentication unit 761, the accelerator firmware component is pushed into the accelerator memory 741 via the sideband bus 791. For example, in some embodiments, after the bootloader accelerator firmware component is instantaneously authenticated by the immediate accelerator firmware component authentication unit 761, the host CPU 171 pushes the bootloader into the accelerator memory 741 of the accelerator 131 via the sideband bus 791.

[0125] In some embodiments, at operation 1070, the accelerator 131 receives the accelerator firmware component via the sideband bus 791 and executes the critical accelerator firmware component. In some embodiments, since the speed of the data flow in the sideband bus 791 is generally lower than the speed of the data flow in the inter-die bus 771, the use of the sideband bus 791 as described herein improves the performance of the SoC 120 by using the sideband bus 791 for actions that were not previously used for the sideband bus 791, making the SoC 120 more efficient than other SoCs or computer systems.

[0126] In some embodiments, returning to reference operation 1020, when the accelerator firmware component sizing unit 754 determines that the size of the accelerator firmware component is not less than the immediate authentication size threshold 755, at operation 1025, the delayed accelerator firmware component authentication unit 762 of the accelerator firmware authentication unit 752 receives the accelerator firmware component (whose size is not less than the immediate authentication size threshold 755) and authenticates the accelerator firmware component. In some embodiments, after authenticating the accelerator firmware component, operation 1025 proceeds to operation 1030.

[0127] In some embodiments, at operation 1030, the host CPU 171 stores the accelerator firmware component in the memory 104. In some embodiments, at operation 1035, the accelerator firmware component is downloaded from the memory 104 to the accelerator 131 via the inter-die bus 771 and executed by the accelerator 131.

[0128] In some embodiments, when the accelerator firmware component size determination unit 754 determines that all accelerator firmware component sizes are greater than the immediate authentication size threshold 755, the entire authenticated accelerator firmware is provided to the accelerator 131 via the inter-die bus 771 for execution by the accelerator 131.

[0129] In some embodiments, using the embodiments described herein, the efficiency of the SoC 120 is partially improved because the sideband bus (e.g., sideband bus 790), which is not typically used for accelerator firmware component transfer, is used to transfer critical accelerator firmware components while the host CPU 171 is still processing non-critical accelerator firmware components. This allows the accelerator to first process the critical accelerator firmware components until the non-critical accelerator firmware components are provided to the accelerator via the inter-die bus (e.g., inter-die bus 770). Thus, in some embodiments, the systems and methods described herein improve and provide advantages over other methods, such as reducing the time and resources required for firmware authentication during startup or firmware updates, and leveraging resources (e.g., sideband bus 790) that are not fully utilized to transfer accelerator firmware.

Claims

1. A system-on-chip, comprising: A central processing unit CPU; An accelerator, the accelerator being coupled to the CPU via a first inter-die interconnect; And A unified memory, the unified memory being coupled to the CPU via a second inter-die interconnect, wherein, in order to prevent the accelerator from using an accelerator memory for processing operations, the accelerator utilizes a unified memory access tunneling system located in the accelerator to tunnel an advanced interconnect protocol associated with the second inter-die interconnect to an inter-die interconnect protocol associated with the first inter-die interconnect, and the unified memory access tunneling system is configured to allow access to the unified memory using a shared address space.

2. The system-on-chip according to claim 1, wherein: The unified memory access tunneling system tunnels the advanced interconnect protocol to the inter-die interconnect protocol using a unified memory access tunneling data packet; and optionally Wherein, the unified memory access tunneling data packet includes a plurality of tunnel fields, and the plurality of tunnel fields are configured to allow the accelerator and the CPU to access the shared address space mapped to the unified memory.

3. The system-on-chip according to claim 2, wherein: The unified memory access tunneling system provides the unified memory access tunneling data packet to the CPU to access the unified memory.

4. The system-on-chip according to claim 2 or 3, wherein: A unified memory access de-tunneling system in the CPU receives the unified memory access tunneling data packet; and optionally Wherein, the unified memory access de-tunneling system de-tunnels the unified memory access tunneling data packet to access the shared address space of the unified memory.

5. The system-on-chip according to any one of the preceding claims, wherein: The advanced interconnect protocol is an Advanced eXtensible Interface AXI protocol or a Compute Express Link CXL protocol; and / or Wherein, the inter-die interconnect protocol is a Universal Chiplet Interconnect Express (UCIe) protocol.

6. The system-on-chip according to any one of the preceding claims, wherein: Based on the tunneling of the advanced interconnect protocol to the inter-die interconnect protocol, the accelerator does not need to use an accelerator memory.

7. A method, comprising: Generating a unified memory access tunneling data packet structure mapped to an inter-die interconnect at an accelerator; Generating, at the accelerator, a unified memory access tunneling data packet using the unified memory access tunneling data packet structure, the unified memory access tunneling data packet being used to tunnel an advanced interconnect protocol to an inter-die interconnect protocol associated with the inter-die interconnect; and Directly accessing a unified memory using the unified memory access tunneling data packet.

8. The method according to claim 7, wherein: The unified memory access tunneling data packet includes a plurality of tunnel fields, and the plurality of tunnel fields are configured to allow the accelerator and a central processing unit (CPU) to access the shared address space mapped to the unified memory.

9. The method according to claim 7 or 8, further comprising: Provide the unified memory access tunnel data packet to the unified memory access de-tunneling system; Optionally wherein, the unified memory access de-tunneling system in the CPU receives the unified memory access tunnel data packet; And optionally The method further includes: using the unified memory access de-tunneling system to perform de-tunneling transmission on the unified memory access tunnel data packet to access the shared address space of the unified memory.

10. The method according to any one of claims 7 to 9, wherein: The advanced interconnect protocol is an Advanced eXtensible Interface (AXI) protocol or a Compute Express Link (CXL) protocol; and / or wherein, the inter-die interconnect is a Universal Chiplet Interconnect Express (UCIe) interconnect.

11. The method according to any one of claims 7 to 10, wherein: Based on the tunneling transmission from the advanced interconnect protocol to the inter-die interconnect protocol, the accelerator does not need to use local memory.

12. A server system, comprising: An accelerator interface controller; An inter-die interconnect coupled to the accelerator interface controller; And A central processing unit (CPU) interface controller coupled to the inter-die interconnect, wherein the accelerator interface controller is configured to tunnel the advanced interconnect protocol to an inter-die interconnect protocol associated with the inter-die interconnect for accessing a unified memory.

13. The server system according to claim 12, wherein: The inter-die interconnect is a Universal Chiplet Interconnect Express (UCIe) interconnect, and the advanced interconnect protocol is an Advanced eXtensible Interface (AXI) protocol or a Compute Express Link (CXL) protocol.

14. The server system according to claim 12 or 13, wherein: The CPU interface controller is located in a host CPU, and the accelerator interface controller is located in a video accelerator.