Parallel processor, electronic device, data transmission method

CN122733740APending Publication Date: 2026-09-11SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611191509.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-07
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

例如,受半导体制造工艺偏差、制程缺陷等不可控物理因素影响,通用图形处理器(GPGPU)芯片的二级缓存中的部分缓存单元可能存在先天硬件故障,导致对应缓存单元失效

Benefits of technology

[0018]In the parallel processor provided in at least one embodiment of this disclosure, a first crossbar switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second crossbar switch circuit is used as the interconnection network between the L2 cache and off-chip memory. When a cache unit fails, the first crossbar switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second crossbar switch circuit can realize data interaction between normally functioning cache units and all memory of off-chip memory without losing the storage space of off-chip memory, effectively improving the robustness and yield of the chip.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122733740A_ABST
    Figure CN122733740A_ABST
Patent Text Reader

Abstract

A parallel processor, electronic device, and data transmission method are disclosed. This invention relates to the field of electronic digital data processing technology. The parallel processor includes multiple hardware execution units, a secondary cache, and off-chip memory. The secondary cache includes L cache units, and the off-chip memory includes L memory units. m of the L cache units are in a failed and unusable state, while the other L-m cache units operate normally. The parallel processor also includes a first crossbar switch circuit and a second crossbar switch circuit. The first crossbar switch circuit is disposed between the multiple hardware execution units and the secondary cache, configured to perform data interaction between the multiple hardware execution units and the L-m cache units. The second crossbar switch circuit is disposed between the secondary cache and the off-chip memory, configured to perform data interaction between the L-m cache units and the L memory units. This parallel processor can still operate normally even when cache units are faulty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of electrical digital data processing technology, and more specifically to a parallel processor, an electronic device, and a data transmission method. Background Technology

[0002] Due to uncontrollable factors in the manufacturing process, some modules of finished chips may have defects, especially when using advanced manufacturing processes, making it even more difficult to guarantee chip yield. For example, due to uncontrollable physical factors such as semiconductor manufacturing process deviations and defects, some cache units in the L2 cache of a general-purpose graphics processing unit (GPGPU) chip may have inherent hardware failures, causing the corresponding cache units to fail. In this case, since the cache units and memory units in the L2 cache have a one-to-one access and connection relationship, the memory units connected to the failed cache units cannot be used, resulting in storage space loss. Furthermore, the Streaming Processor Cluster (SPC) is unaware of whether the cache units are damaged and continues to send requests normally. Requests sent to the failed cache units will not receive a response, affecting the chip's normal operation and data read / write logic, causing the overall chip function abnormally and rendering it unusable. Summary of the Invention

[0003] This invention provides at least one embodiment of a parallel processor, including multiple hardware execution units, a secondary cache, and off-chip memory. The multiple hardware execution units are configured to execute computing tasks in parallel. The multiple hardware execution units share the secondary cache and the off-chip memory. The secondary cache includes L cache units, and the off-chip memory includes L memory units, where L is a positive integer greater than 1. m of the L cache units are in a failed and unavailable state, while the other Lm cache units are working normally, where m is an integer greater than or equal to 0 and less than L. The parallel processor further includes a first cross-switch circuit and a second cross-switch circuit. The first cross-switch circuit is disposed between the multiple hardware execution units and the secondary cache, and is configured to perform data interaction between the multiple hardware execution units and the Lm cache units. In response to m being greater than 0, data access requests originally destined for the failed cache units are sent to other normally functioning cache units. The second cross-switch circuit is disposed between the secondary cache and the off-chip memory, and is configured to perform data interaction between the Lm cache units and the L memory units.

[0004] For example, in a parallel processor provided in at least one embodiment of this invention, the plurality of hardware execution units and the L cache units have a predetermined routing mapping relationship. The first cross-switch circuit is configured to: when m=0, route the received data access request to the destination cache unit of the received data access request according to the predetermined routing mapping relationship; when m is greater than 0, update the routing mapping relationship between the plurality of hardware execution units and the Lm cache units, and route the received data access request to the destination cache unit of the received data access request according to the updated routing mapping relationship, wherein, in the updated routing mapping relationship, the data access request that originally went to the invalid cache unit is sent to the normally functioning cache unit.

[0005] For example, in the parallel processor provided in at least one embodiment of the present invention, in the updated routing mapping relationship, the data access requests that originally went to the invalid cache unit are evenly distributed to the Lm cache units, and the data access requests that originally went to the Lm cache units change their destination cache unit due to the update of the routing mapping relationship.

[0006] For example, in a parallel processor provided in at least one embodiment of this invention, the parallel processor is configured to receive a first indication signal and a second indication signal, wherein the first indication signal indicates m equals 0 when it is a first value, and indicates m is greater than 0 when it is a second value. The first cross-switch circuit is configured to: in response to receiving the data access request, extract information at a predetermined position in the request address of the data access request as address information; perform a first modulo operation on the address information with L as the modulus to obtain a first result, and perform a second modulo operation with Lm as the modulus to obtain a second result; in response to the first indication signal being the first value, send the data access request to the cache unit indicated by the first result; in response to the first indication signal being the second value, determine the destination cache unit of the data access request based on the second result and the second indication signal, and send the data access request to the destination cache unit, wherein the destination cache unit is one of the Lm cache units, and the second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo operation and the Lm cache units.

[0007] For example, in a parallel processor provided in at least one embodiment of this invention application, for each hardware execution unit, in response to a data access request initiated by the hardware execution unit not being hit in the secondary cache, the second cross switch circuit is configured to send a data request to the memory unit indicated by the request address in the data access request to perform memory data access related to the data access request.

[0008] For example, in a parallel processor provided in at least one embodiment of this invention, the first cross-switch circuit includes a plurality of first multiplexers coupled to the plurality of hardware execution units via an on-chip interconnect network, and L first multiplexers coupled to the L cache units, wherein the plurality of first multiplexers and the L first multiplexers are cross-interconnected; the second cross-switch circuit includes L second multiplexers coupled to the L cache units, and L second multiplexers coupled to the L memory units, wherein the L second multiplexers and the L second multiplexers are cross-interconnected.

[0009] For example, in a parallel processor provided in at least one embodiment of this invention, the first cross-connect circuit includes address processing modules coupled one-to-one with the plurality of first multiplexers. Each address processing module includes a first processing module, a second processing module, a third multiplexer, and a third processing module. The first processing module is configured to perform a first modulo operation on the address information modulo L to obtain a first result, wherein the address information is information at a predetermined position in the request address of a received data access request. The second processing module is configured to perform a second modulo operation on the address information modulo Lm to obtain a second result. The two input terminals of the third multiplexer are coupled to the first processing module and the second processing module, and the output of the third multiplexer... The third processing module is coupled to the third processing module. The third multiplexer is configured to receive a first indication signal as a control signal. When the first indication signal indicates that m equals 0, the first result is selected and output to the third processing module. When the first indication signal indicates that m is greater than 0, the second result is selected and output to the third processing module. The third processing module is also coupled to a corresponding first multiplexer and is configured to output a signal indicating the destination cache unit of the received data access request to the corresponding first multiplexer based on the first indication signal, the second indication signal, and the result output by the third multiplexer. The second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo processing and the Lm cache units.

[0010] For example, in a parallel processor provided in at least one embodiment of this invention application, each first multiplexer is configured to: send a received data access request as an arbitration candidate to a first multiplexer coupled to the destination cache unit to participate in arbitration, and upon obtaining arbitration, send the received data access request to the destination cache unit.

[0011] For example, in the parallel processor provided in at least one embodiment of this invention application, each first multiplexer is configured to use a round-robin arbitration method to determine the order in which data access requests received by the first multiplexer enter the cache unit coupled to the first multiplexer.

[0012] For example, in a parallel processor provided in at least one embodiment of this invention, each second multiplexer is configured to send a received data request as an arbitration candidate to a second multiplexer coupled to the memory cell indicated by the request address in the data request, and, upon obtaining arbitration, send the data request to the memory cell indicated by the request address.

[0013] For example, in the parallel processor provided in at least one embodiment of this application, the parallel processor is a graphics processor or a general-purpose graphics processor, the hardware execution unit is a streaming processor cluster, and the streaming processor cluster is configured with an independently accessible level 1 cache.

[0014] This invention application provides at least one embodiment of a parallel processor, which includes multiple hardware execution units, multiple secondary caches, and multiple off-chip memories. The multiple hardware execution units share the multiple secondary caches and the multiple off-chip memories. Each secondary cache includes L cache units, and each off-chip memory includes L memory units. m of the L cache units are in a failed and unavailable state, while the other Lm cache units are working normally. m is an integer greater than or equal to 0 and less than L. Each secondary cache is equipped with a corresponding first crossbar switch circuit and a second crossbar switch circuit. The multiple secondary caches are paired one-to-one with the multiple off-chip memories. Correspondingly, the first crossbar switch circuit is disposed between the plurality of hardware execution units and the corresponding L2 cache, and is configured to perform data interaction between the plurality of hardware execution units and the Lm cache units that are working normally in the corresponding L2 cache, wherein, in response to m being greater than 0, data access requests that were originally destined for invalid cache units are sent to other working cache units; the second crossbar switch circuit is disposed between the corresponding L2 cache and the off-chip memory connected to the corresponding L2 cache, and is configured to perform data interaction between the Lm cache units that are working normally in the corresponding L2 cache and the L memory units in the off-chip memory connected to the corresponding L2 cache.

[0015] For example, in a parallel processor provided in at least one embodiment of this invention, the parallel processor further includes an on-chip interconnect network. The on-chip interconnect network is disposed between a plurality of first cross-connect circuits corresponding to the plurality of hardware execution units and the plurality of secondary caches. The on-chip interconnect network is configured to send the data access request to the first cross-connect circuit corresponding to the destination secondary cache according to the destination secondary cache indicated by the data access request sent by each hardware execution unit, so that the corresponding first cross-connect circuit sends the data access request to the corresponding destination cache unit in the destination secondary cache.

[0016] This application provides at least one embodiment of an electronic device, including a parallel processor as described in at least one embodiment of this application.

[0017] This invention application provides at least one embodiment of a data transmission method for performing data transmission between multiple hardware execution units and a secondary cache, and for performing data transmission between the secondary cache and off-chip memory. The multiple hardware execution units are configured to execute computing tasks in parallel, and the multiple hardware execution units share the secondary cache and the off-chip memory. The secondary cache includes L cache units, and the off-chip memory includes L memory units. m of the L cache units are in a failed and unavailable state, while the other Lm cache units are working normally. m is an integer greater than or equal to 0 and less than L. The secondary cache is coupled with a first crossbar switch circuit and a second crossbar switch circuit. The first crossbar switch circuit is disposed between the plurality of hardware execution units and the secondary cache, and the second crossbar switch circuit is disposed between the secondary cache and the off-chip memory. The data transmission method includes: data interaction between the plurality of hardware execution units and the Lm cache units is performed by the first crossbar switch circuit, wherein, in response to m being greater than 0, data access requests originally destined for invalid cache units are sent to other normally functioning cache units; and data interaction between the Lm cache units and the L memory units is performed by the second crossbar switch circuit.

[0018] In the parallel processor provided in at least one embodiment of this disclosure, a first crossbar switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second crossbar switch circuit is used as the interconnection network between the L2 cache and off-chip memory. When a cache unit fails, the first crossbar switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second crossbar switch circuit can realize data interaction between normally functioning cache units and all memory of off-chip memory without losing the storage space of off-chip memory, effectively improving the robustness and yield of the chip. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0020] Figure 1 This is a schematic diagram of the architecture of a general-purpose graphics processing unit (GPGPU). Figure 2 It is a diagram of the interconnect architecture inside a general-purpose graphics processor; Figure 3 A schematic diagram illustrating data writing from a multi-master device to a multi-slave device. Figure 4 This is a schematic diagram of the structure of a parallel processor provided in at least one embodiment of the present disclosure; Figure 5 A schematic structural diagram of an address processing module provided in at least one embodiment of this disclosure; Figure 6 A schematic structural diagram of a parallel processor provided in an embodiment of this disclosure; Figure 7 A schematic structural diagram of a parallel processor provided for another embodiment of this disclosure; Figure 8 A schematic structural diagram of an electronic device provided for at least one embodiment of this disclosure; Figure 9 A schematic diagram of the specific structure of an electronic device provided in at least one embodiment of this disclosure; Figure 10 This is a schematic flowchart illustrating a data transmission method provided in at least one embodiment of the present disclosure. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0022] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0023] To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of some known functions and known components have been omitted.

[0024] General-purpose graphics processing units (GPGPUs) are core computing power chips designed for large-scale parallel computing scenarios. Their core computing architecture is based on a parallel computing array composed of massive standardized streaming processor clusters (SPCs), working in conjunction with a multi-level high-speed cache storage system to meet the high-throughput and high-concurrency computing requirements of scenarios such as artificial intelligence, high-performance computing, and parallel simulation.

[0025] Figure 1 This is a schematic diagram of a general-purpose graphics processing unit (GPGPU).

[0026] like Figure 1 As shown, a general-purpose graphics processor is actually an array of programmable multiprocessors. For example, a programmable multiprocessor can be a cluster of streaming processors, such as including... Figure 1 The diagram shows multiple streaming processor clusters. In a general-purpose graphics processor, one streaming processor cluster handles one computational task, or multiple streaming processor clusters handle one computational task. Multiple streaming processor clusters share data through a global cache or global memory.

[0027] like Figure 1 As shown, a streaming processor cluster includes multiple computing units (CUs). Each computing unit (CU) is used to perform arithmetic and logical operations, such as accumulation, reduction, and regular addition, subtraction, multiplication, and division.

[0028] A computing unit comprises multiple cores (also called computing kernels or computing cores). Each computing core includes an arithmetic logic unit (ALU), a floating-point unit, etc., and is used to execute specific computational tasks. In addition, the computing unit also includes registers (e.g., ... Figure 1 The register file and shared memory in a computing unit are used to store source and destination data related to computing tasks in a hierarchical manner. The shared memory in a computing unit is used to share data between the cores of that computing unit.

[0029] like Figure 1 As shown, each computing unit also provides a tensor core for performing tensor-related computations, such as tensor shrinking operations. Tensor cores can accelerate tensor operations such as matrix multiplication. Tensor cores in multiple computing units can be scheduled and controlled uniformly.

[0030] In parallel computing, computational tasks are typically executed by multiple threads. These threads are divided into multiple thread blocks before execution in the graphics processing unit (or parallel computing processor), and then dispatched via a thread block distribution module. Figure 1 (Not shown in the image) Multiple thread blocks are distributed to various computation units. All threads in a thread block must be assigned to the same computation unit for execution. Simultaneously, thread blocks are broken down into minimum execution thread bundles (or simply warps), each containing a fixed number (or less than this fixed number) of threads, for example, 32 threads. Multiple thread blocks can execute in the same computation unit or in different computation units.

[0031] In each computing unit, the thread beam scheduling / distribution module ( Figure 1 (Not shown in the diagram) Thread bundles are scheduled and allocated so that multiple computing cores within the computing unit can run thread bundles. Depending on the number of computing cores in the computing unit, multiple thread bundles within a thread block can execute concurrently or in a time-sharing manner. Multiple threads within each thread bundle execute the same instructions. Memory-executed instructions are issued to shared memory within the computing unit or further issued to intermediate-level caches, global caches, or global memory (e.g., [example cache]). Figure 1 High Bandwidth Memory (HBM) is used for read and write operations.

[0032] In the task execution process of a graphics processing unit (GPU) or general-purpose GPU, various computing tasks of the streaming processor cluster need to frequently interact with the cache. Core interaction requests include read requests for cached data before the execution of computing tasks, write requests for cached data after the generation of computing results, and flush requests to write dirty cache lines back to high-bandwidth memory (HBM). The access efficiency and working stability of the multi-level cache directly determine the overall computing power output performance of the GPU or general-purpose GPU.

[0033] In the hardware architecture design of mainstream or general-purpose graphics processing units (GPUs), the cache system has a clear hierarchy and sharing rules, with each streaming processor cluster (SPC) independently configured with a Level 1 cache (L1 Cache). Figure 1 Shared memory in the cluster ensures low-latency data interaction for single-cluster computing tasks, while multiple streaming processor clusters share a second-level cache (L2 cache). Figure 1 (Global cache in the system) enables efficient reuse of computing resources and cache resources.

[0034] To meet the ultra-high bandwidth demands of multi-threaded concurrent scenarios, secondary caches commonly employ a multi-cache unit grouping architecture, also known as a cache group architecture. This architecture allocates access to target cache groups through a physical address hash mapping mechanism, thereby significantly improving overall throughput. Specifically, the multi-cache unit grouping architecture divides the entire secondary cache into multiple independent cache units and manages them in groups. The system routes and distributes requests based on the physical address of the access request, mapping data access requests from different addresses to the corresponding cache groups for processing. By scheduling multiple cache units in parallel, the access pressure is effectively distributed, significantly improving the overall concurrent throughput of the secondary cache and meeting the high bandwidth demands of massive concurrent access by graphics processing units (GPUs) or general-purpose GPUs.

[0035] Figure 2 It is a diagram of the interconnect architecture inside a general-purpose graphics processor.

[0036] like Figure 2 As shown, multiple SPCs are interconnected with the L2 cache via an on-chip interconnect network, and the L2 cache is interconnected with main memory. The L2 cache comprises multiple cache units, and the main memory also comprises multiple memory units. For a detailed description of the SPC structure, please refer to [reference needed]. Figure 1 The description.

[0037] For example, in a collaborative computing architecture between a central processing unit (CPU) and a general-purpose graphics processing unit (GPU), the CPU, as the master controller, is responsible for task scheduling and instruction issuance. It sends pre-processed computational task instructions and data scheduling instructions in batches to the GPU. After receiving various instructions from the CPU, the GPU uses its internal task distribution mechanism to perform instruction parsing, task decomposition, and load balancing. It then rationally distributes the decomposed fine-grained computational tasks to various streaming processor clusters (SPCs) for parallel execution, fully leveraging the massively parallel computing power of the GPU. Each SPC integrates an independent L1 cache, which can meet its own low-latency data read and write requirements during computation. At the same time, all SPCs can access multiple independent memory partitions within the chip through the GPU's high-speed on-chip interconnect network, enabling shared access to global memory resources.

[0038] In terms of hardware architecture, general-purpose graphics processors divide backend storage resources into multiple physically independent storage partitions, such as... Figure 2 As shown, each storage partition can contain a cache unit and a storage unit of off-chip memory connected to it (e.g., Figure 2 The external memory (such as the memory cells in the main memory) can be, for example, Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM) or high-bandwidth memory. Furthermore, the SPC can access any cache cell via the on-chip interconnect network. Multiple cache cells in the L2 cache are connected one-to-one with multiple memory cells in main memory; a cache cell can only access its corresponding connected memory cell and cannot access other memory cells.

[0039] The hierarchical caching architecture and distributed storage partitioning design can effectively distribute massive concurrent access requests, balance data access latency and overall throughput bandwidth, adapt to high-intensity work scenarios with multiple SPCs performing parallel computing and frequent reading and writing of storage data, and ensure the overall efficient and stable operation of GPGPU.

[0040] Due to uncontrollable factors in the manufacturing process, some modules of the finished chip may have defects. This is especially true when using advanced manufacturing processes, where chip yield is even more difficult to guarantee. Therefore, it is necessary to minimize the impact of hardware defects in some modules on the overall chip. For example, due to uncontrollable physical factors such as semiconductor manufacturing process deviations and defects, some cache units in the L2 cache of a general-purpose graphics processing unit (GPU) chip may have inherent hardware faults, causing the corresponding cache units to fail. Since the cache units and memory units in the L2 cache have a one-to-one access and connection relationship, the memory units connected to the failed cache units cannot be used, resulting in storage space loss. Furthermore, the System Processing Unit (SPC) is unaware of whether a cache unit is damaged and continues to send requests normally. Requests sent to the failed cache units will not receive a response, affecting the chip's normal operation and data read / write logic, causing the overall chip to malfunction and become unusable.

[0041] The crossbar architecture is a hardware interconnect scheme that employs a matrix switching structure. This architecture establishes a fully connected topology between multiple input and output ports, enabling dynamic connections from any input to output through crossbar switches. Its core characteristics lie in supporting concurrent communication between multiple sets of input and output ports, providing non-blocking communication capabilities, high parallelism, and reconfigurable connection flexibility. Within multi-core processors, the crossbar architecture is often used as a high-speed data path between the processor's internal computing units and the memory system, with data flowing between multiple master ports and multiple slave ports via the crossbar switches.

[0042] Figure 3 This is a schematic diagram illustrating the process of writing data from a multi-master device to a multi-slave device.

[0043] like Figure 3 As shown, there are three slave devices (device 0 to slave 2) and three master devices (master devices 0 to master 2). For example, when data is transmitted from a master device to a slave device, master device 0 returns data packets of varying sizes to different slave devices, such as slave device 0, slave device 1, etc., according to the source index of the write request. Master device 1 also returns data packets of varying sizes to different slave devices according to the source index of the write request. In one clock cycle, a slave device will only receive data packets sent by one master device. If slave device 0 receives data packets returned by two master devices (e.g., master device 0 and master device 1) simultaneously, arbitration is required between these two data packets from different sources. The data packet that wins the arbitration (is selected) can continue to be returned to the corresponding slave device.

[0044] This disclosure utilizes a crossbar switch circuit to enable data interaction and transmission between multiple master devices and multiple slave devices. A first crossbar switch circuit architecture is added as an interconnection network between the hardware execution unit and the secondary cache to realize request transmission between the hardware execution unit and the invalidation cache unit. Furthermore, a second crossbar switch circuit is added as an interconnection network between the secondary cache and off-chip memory to realize data access to the memory unit corresponding to the invalidation cache unit.

[0045] This disclosure provides at least one embodiment of a parallel processor, an electronic device, and a data transmission method. The parallel processor provided in at least one embodiment includes multiple hardware execution units, a secondary cache, and off-chip memory. The multiple hardware execution units are configured to execute computational tasks in parallel. The multiple hardware execution units share the secondary cache and off-chip memory. The secondary cache includes L cache units, and the off-chip memory includes L memory units. m of the L cache units are in a failed and unavailable state, while the other Lm cache units are functioning normally. m is an integer greater than or equal to 0 and less than L. The parallel processor also includes a first cross-switch circuit and a second cross-switch circuit. The first cross-switch circuit is disposed between the multiple hardware execution units and the secondary cache, configured to perform data interaction between the multiple hardware execution units and the Lm cache units. In response to m being greater than 0, data access requests originally destined for the failed cache units are sent to other functioning cache units. The second cross-switch circuit is disposed between the secondary cache and the off-chip memory, configured to perform data interaction between the Lm cache units and the L memory units.

[0046] In the parallel processor provided in at least one embodiment of this disclosure, a first crossbar switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second crossbar switch circuit is used as the interconnection network between the L2 cache and off-chip memory. When a cache unit fails, the first crossbar switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second crossbar switch circuit can realize data interaction between normally functioning cache units and all memory of off-chip memory without losing the storage space of off-chip memory, effectively improving the robustness and yield of the chip.

[0047] The circuit structure of the parallel processor provided in at least one embodiment of this disclosure will be described in detail below with reference to the accompanying drawings.

[0048] Figure 4 This is a schematic diagram of the structure of a parallel processor provided in at least one embodiment of the present disclosure.

[0049] like Figure 4As shown, the parallel processor 100 includes multiple hardware execution units 101, a secondary cache 102, and off-chip memory 103.

[0050] The parallel processor 100 can be a processor that integrates a large number of hardware execution units, capable of executing multiple instructions and processing multiple sets of data simultaneously to complete computational tasks in a parallel manner. For example, the parallel processor 100 can specifically be a Graphics Processing Unit (GPU), a General-Purpose Graphics Processing Unit (GPGPU), a Tensor Processing Unit (TPU), a Data Processing Unit (DPU), a Neural Processing Unit (NPU), etc., and this disclosure does not impose specific limitations in this regard.

[0051] Parallel processors can take the form of System on Chip (SOC), which integrates modules such as crossbar switch circuits, processor cores, memory, and peripheral interfaces (such as UART, SPI, and GPIO interfaces) into a single chip, forming a highly integrated hardware architecture.

[0052] In addition, the parallel processor may also include auxiliary circuit modules to ensure the stable operation of the cross switch circuit. These may include a power management module to provide a stable operating voltage for the cross switch circuit; a clock module to provide a precise clock signal for the cross switch circuit; and a reset module to generate a reset signal. Upon receiving the reset signal, the cross switch circuit can first clear its internal cache and configuration parameters, and then cooperate with the processor to complete the device reset and restart.

[0053] This disclosure does not limit the function, form, or use of parallel processors.

[0054] Hardware execution unit 101 is configured as a parallel computing cluster of parallel processors to execute computing tasks in parallel.

[0055] For example, when the parallel processor is a graphics processing unit (GPU) or a general-purpose GPU, the hardware execution unit is a streaming processor cluster, which is configured with an independently accessible L1 cache. A description of streaming processor clusters can be found above. Figure 1 The relevant content will not be repeated here.

[0056] Level 2 cache is a high-speed storage unit built into the parallel processor. Its read and write speeds are much higher than those of off-chip memory. It can temporarily cache operation data and instructions, reducing processor latency. Off-chip memory is a storage device located outside the chip, such as HBM and DDR. It has a larger capacity and is used to store programs and massive amounts of data for a long time. The two work together to achieve a balance between high-speed access and large-capacity storage.

[0057] For example, multiple hardware execution units share the L2 cache and off-chip memory. In other words, multiple hardware execution units can access the L2 cache, which can be considered a global cache from the perspective of the hardware execution units.

[0058] like Figure 4 As shown, the off-chip memory 103 includes L memory units. The L2 cache is divided into L cache units, and the off-chip memory is divided into L memory units, both having the same number of units.

[0059] Of the L cache units, m are in a failed and unavailable state, while the other Lm cache units are functioning normally, where m is an integer greater than or equal to 0 and less than L. Since the chip's state is unknown before tape-out, it is also unknown whether any cache units have failed, and which cache units will fail. This disclosure provides at least one embodiment of a general hardware circuit that supports data access when all cache units are functioning correctly, as well as data access when some cache units have failed, thereby enhancing robustness and flexibility.

[0060] like Figure 4 As shown, the parallel processor 100 also includes a first cross switch circuit 104 and a second cross switch circuit 105.

[0061] The first crossbar switch circuit 104 is disposed between multiple hardware execution units 101 and L2 cache 102, and is configured to perform data interaction between the multiple hardware execution units 101 and Lm cache units. In response to m being greater than 0, data access requests that were originally destined for invalid cache units are sent to other normally functioning cache units.

[0062] For example, the first cross-connect circuit 104 establishes a cross-connection relationship between L cache units and multiple hardware execution units, which can realize data interaction between any hardware execution unit and any cache unit.

[0063] For example, a hardware execution unit is unaware of which cache units are invalid. Therefore, the hardware execution unit can still send data access requests normally, and the request address may very well point to the invalid cache unit. If at least one cache unit is invalid, data access requests that would normally be directed to those invalid cache units can be redirected to other working cache units via the first crossbar switch circuit. This avoids the situation where some cache units fail due to manufacturing processes or other issues, causing the parallel processor to malfunction and become inoperable.

[0064] The second cross switch circuit 105 is located between the L2 cache and the off-chip memory, and is configured to perform data interaction between Lm cache units and L memory units.

[0065] For example, in situations requiring data access to off-chip memory, such as when a data access request sent by the hardware execution unit requests data that is not found in the cache, a data access request needs to be sent to the off-chip memory. The second crossbar switch circuit 105 establishes crossbar interconnections between L cache units and L memory units, enabling data interaction between any cache unit and any memory unit. Therefore, the second crossbar switch circuit allows data interaction between the currently functioning Lm cache units and L memory units, thereby fully utilizing storage space and avoiding waste of off-chip memory space due to failed cache units.

[0066] For example, there is a predetermined routing mapping relationship between multiple hardware execution units and L cache units. For instance, this predetermined routing mapping relationship is the mapping relationship when all cache units are working normally. The hardware execution units will still generate data access requests normally when generating data access requests. According to the predetermined routing mapping relationship, the data access request will be sent to the specified cache unit.

[0067] For example, in some embodiments, the first cross switch circuit is configured to: when m=0, route the received data access request to the destination cache unit of the received data access request according to a predetermined routing mapping relationship; when m is greater than 0, update the routing mapping relationship between the plurality of hardware execution units and Lm cache units, and route the received data access request to the destination cache unit of the received data access request according to the updated routing mapping relationship, wherein, in the updated routing mapping relationship, the data access request that originally went to the invalid cache unit is sent to the normally functioning cache unit.

[0068] In the parallel processor provided in at least one embodiment of this disclosure, the failure and unavailability of cache units are imperceptible to the processing logic of hardware execution units, cache units, and memory units. These circuit structures can still perform request generation, processing, etc., as if all cache units were working normally. When no cache units fail, the first cross-switch circuit routes the received data access request to the destination cache unit of the received data access request according to a predetermined routing mapping relationship. When at least one cache unit fails, the first cross-switch circuit updates the routing mapping relationship between the currently working cache units and multiple hardware execution units. According to the updated routing mapping relationship, the data access request is sent to the corresponding destination cache unit. Furthermore, in the updated routing mapping relationship, the data access request that was originally destined for the failed cache unit is sent to the working cache unit.

[0069] Therefore, the parallel processor in this disclosure can update the routing relationship and transmit data by the first cross-switch circuit without modifying the processing logic of the hardware execution unit and cache unit, thereby minimizing the impact on other circuit structures and achieving normal operation of the parallel processor as a whole even if a cache unit fails, with minimal modification cost.

[0070] For example, in the updated routing mapping, data access requests that originally went to invalid cache units are evenly distributed across Lm cache units, and data access requests that originally went to Lm cache units have changed their destination cache units due to the update of the routing mapping.

[0071] In other words, the updated routing mapping changes the mapping between multiple hardware execution units and the Lm cache units that are currently working normally. This allows data access requests that were originally destined for invalid cache units to be evenly distributed across the Lm cache units. Furthermore, data access requests that were originally destined for the Lm cache units may also change their destination cache unit due to the update of the routing mapping.

[0072] For example, in some embodiments, the parallel processor is configured to receive a first indication signal and a second indication signal.

[0073] When the first indicator signal is a first value, it indicates that m equals 0; when the first indicator signal is a second value, it indicates that m is greater than 0. For example, the first indicator signal can be a mode enable signal. When the first indicator signal is a first value (e.g., 0), it indicates that no cache unit has failed. When the first indicator signal is a second value (e.g., 1), it indicates that a cache unit has failed.

[0074] The second indication signal can be a mapping relationship indication signal. For example, the second indication signal is used to indicate the mapping relationship between Lm modulo results of the second modulo process (modulo process with Lm as the modulus) and Lm cache units. For example, the second indication signal can have multiple values, each value representing the mapping relationship between all possible modulo results (Lm modulo results) of the second modulo process for a given address information and the normally functioning Lm cache units. In different mapping relationships, the normally functioning Lm cache units are not exactly the same.

[0075] The first and second indicator signals can be obtained through registers. For example, after tape-out testing, it can be determined whether there are cache cell failures and the identification information (e.g., number) of the failed cache cells. This allows the determination of the first and second indicator signals, which are then written into the parallel processor through registers. Subsequently, the first crossbar switch circuit rearranges the routing relationships based on the first and second indicator signals.

[0076] Of course, in other embodiments, other forms may be used to indicate whether a cache unit is invalid and the identification information of the invalid cache unit. For example, the functions of the first indicator signal and the second indicator signal may be achieved by simply using different values ​​of an indicator signal. This disclosure does not impose any specific limitations on this.

[0077] The first cross-switch circuit is configured as follows: in response to receiving a data access request, it extracts information from a predetermined position in the request address of the data access request as address information; performs a first modulo operation on the address information with L as the modulus to obtain a first result, and performs a second modulo operation with Lm as the modulus to obtain a second result; in response to a first indication signal being a first value, it sends the data access request to the cache unit indicated by the first result; in response to a first indication signal being a second value, it determines the destination cache unit of the data access request based on the second result and the second indication signal, and sends the data access request to the destination cache unit, wherein the destination cache unit is one of Lm cache units, and the second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo operation and the Lm cache units.

[0078] For example, the data access requests sent by the hardware execution unit to the cache unit all carry a request address, such as a physical address (PA). The first cross switch circuit can extract some bits at specific positions in the physical address as address information, which indicates the destination cache unit specified according to the preset routing mapping relationship.

[0079] Next, the address information is subjected to a first modulo operation with L as the modulus. For example, the address information is subjected to a mod L operation, and the result is used as the first result.

[0080] Furthermore, a second modulo operation (modulo operation with Lm as the modulus) is performed on the address information, and the result is used as the second result. For example, the second result is obtained by performing a mod(Lm) operation (modulo operation with Lm as the modulus) on the address information.

[0081] For example, if the first indication signal is a first value, then the cache unit indicated by the first result is used as the destination cache unit, and the data access request is forwarded to the cache unit indicated by the first result.

[0082] For example, if the first indication signal is the second value, based on the second result and the second indication signal, the destination cache unit for the data access request is determined, and the data access request is sent to the destination cache unit. The cache unit indicated by the second result is one of Lm cache units. That is, the cache unit indicated by the second result is a normally functioning cache unit, and data access requests that would normally be sent to a failed cache unit will also be sent to a normally functioning cache unit.

[0083] For example, in one embodiment, L=4, m=1, and Lm=3. The correspondence between the first indicator signal, the second indicator signal, and the cache unit numbers is shown in Table 1. The four cache units in Table 1 are distinguished as "Cache Unit 1", "Cache Unit 2", "Cache Unit 3", and "Cache Unit 4". Of course, the same principle applies to other embodiments, which will not be elaborated upon here.

[0084] Table 1

[0085] Referring to Table 1, when the first indication signal is 0, the cache unit and the request address have a predetermined mapping relationship; when the first indication signal is 1, the cache unit and the request address have an updated mapping relationship.

[0086] The following section, in conjunction with Table 1, explains in detail how to determine the destination cache unit for a data access request.

[0087] For example, referring to Table 1, when the first indicator signal is 0, it means that m=0, i.e., no cache unit has failed. At this time, the second indicator signal can be any value (represented as "x" in Table 1). After obtaining the address information at the predetermined position in the request address of the data access request, the address information is subjected to the first modulo operation, i.e., mod4 processing. There are four possible results: 00, 01, 10, and 11 (binary representation). For example, referring to Table 1, result 00 can be associated with cache unit 1, i.e., when the first result 0 is obtained after mod4 processing of the address information, it means that the data access request needs to be forwarded to cache unit 1; result 01 can be associated with cache unit 2, i.e., when the first result 1 is obtained after mod4 processing of the address information, it means that the data access request needs to be forwarded to cache unit 2; result 10 can be associated with cache unit 1, i.e., when the first result 2 is obtained after mod4 processing of the address information, it means that the data access request needs to be forwarded to cache unit 3; result 11 can be associated with cache unit 4, i.e., when the first result 3 is obtained after mod4 processing of the address information, it means that the data access request needs to be forwarded to cache unit 4.

[0088] When the first indicator signal is 1, it means that m is greater than 0, that is, at least one cache unit has failed. According to the second indicator signal, different cache unit failure situations can be matched. For example, the second indicator signal is 0 to indicate that cache unit 4 is failed, the second indicator signal is 1 to indicate that cache unit 3 is failed, the second indicator signal is 2 to indicate that cache unit 2 is failed, and the second indicator signal is 3 to indicate that cache unit 1 is failed.

[0089] In the second modulo operation, the acquired address information is processed using mod3 to obtain a second result. Then, based on the second indicator signal, the cache unit corresponding to the second result obtained from the mod3 processing can be determined. There are three possible results from the mod3 processing of the address information: 00, 01, and 10 (binary representation). Referring to Table 1, the cache units mapped to the three mod3 results can be obtained.

[0090] Taking a second indicator signal of 0 as an example, cache units 1 through caching units 3 are operating normally at this time. For example, referring to Table 1, when the second result obtained by receiving address information and performing mod3 processing is 0, the destination cache unit for the data access request is cache unit 1, and the data access request can be sent to cache unit 1; when the intermediate result obtained by receiving address information and performing mod3 processing is 1, the destination cache unit for the data access request is cache unit 2, and the data access request can be sent to cache unit 2; when the intermediate result obtained by receiving address information and performing mod3 processing is 2, the destination cache unit for the data access request is cache unit 3, and the data access request can be sent to cache unit 3.

[0091] Taking the second indicator signal as 2 as an example, cache unit 1, cache unit 3, and cache unit 4 are working normally at this time. For example, when the intermediate result obtained by mod3 processing of the address information is 0, the destination cache unit of the data access request is cache unit 1, and the data access request can be sent to cache unit 1; when the intermediate result obtained by mod3 processing of the address information is 1, the destination cache unit of the data access request is cache unit 3, and the data access request can be sent to cache unit 3; when the intermediate result obtained by mod3 processing of the address information is 2, the destination cache unit of the data access request is cache unit 4, and the data access request can be sent to cache unit 4.

[0092] The same applies to the cases where the second indicator signal is 1 or 3, and will not be elaborated here.

[0093] It should be noted that Table 1 is only an illustrative mapping relationship, and those skilled in the art can adapt the above mapping relationship as needed. This disclosure does not impose any specific restrictions on it.

[0094] For example, in some embodiments, in response to a data access request initiated by a hardware execution unit that misses in the L2 cache, the second cross switch circuit is configured to send a data request to the memory unit indicated by the request address in the data access request to perform data access related to the data access request.

[0095] For example, in some embodiments, taking a read request as an example, after the hardware execution unit sends a read request to a cache unit according to the aforementioned embodiments, the cache unit parses the address, queries the directory (tag), and determines that the data to be read by the read request already exists in the cache unit. If so, it is considered a cache hit, and the data is returned to the hardware execution unit. If it is determined that the data to be read by the read request does not exist in the cache unit, it is considered a cache miss. At this time, a data request is sent to off-chip memory. The address, transaction type (read / write), etc., of this data request are completely inherited from the data access request sent by the hardware execution unit. Specifically, the data request is sent to the memory unit indicated by the request address (e.g., physical address) of the data access request to obtain the data to be read by the data access request.

[0096] As mentioned earlier, the hardware execution unit is unaware of which cache units have failed, therefore... Figure 2The request address of the data access request sent by the hardware execution unit may very well point to the memory unit connected to the invalid cache unit. In this disclosure, a second cross switch circuit is set between the L2 cache and the off-chip memory. Therefore, the data access request sent by the hardware execution unit can be sent to any memory unit of the off-chip memory, so as not to lose the storage space of the off-chip memory and effectively improve the robustness and yield of the chip.

[0097] For example, such as Figure 4 As shown, the first cross switch circuit 104 includes a plurality of first multiplexers (DMUX) coupled to a plurality of hardware execution units via an on-chip interconnect network, and L first multiplexers (MUX) coupled to L cache units, wherein the plurality of first multiplexers and the L first multiplexers are cross-connected.

[0098] like Figure 4 As shown, each first multiplexer includes one input and multiple outputs. The inputs of the multiple first multiplexers are coupled to an on-chip network and, through the on-chip network, to a hardware execution unit. Each first multiplexer includes multiple inputs and one output. The outputs of the multiple first multiplexers are coupled one-to-one to L buffer units. The multiple outputs of each first multiplexer are respectively coupled to one input of each of the multiple first multiplexers.

[0099] The second cross-switch circuit 105 includes L second multiplexers (DMUX) coupled to L cache units and L second multiplexers (MUX) coupled to L memory units, with the L second multiplexers and L second multiplexers cross-connected.

[0100] like Figure 4 As shown, each second multiplexer includes one input and multiple outputs, with the inputs of the L second multiplexers coupled one-to-one with the L cache units. Each second multiplexer includes multiple inputs and one output, with the outputs of the L second multiplexers coupled one-to-one with the L memory units. The multiple outputs of each second multiplexer are each coupled to one input of each of the L second multiplexers. Thus, data requests from cache units can be transmitted to any memory unit connected via a second crossbar switch circuit.

[0101] For example, in some embodiments, the process of sending data access requests to normally functioning cache units as described in the foregoing embodiments can be implemented by software, hardware, or a combination of software and hardware.

[0102] For example, the first cross switch circuit includes an address processing module that is coupled one-to-one with a plurality of first multiplexers.

[0103] Figure 5 This is a schematic structural diagram of an address processing module provided in at least one embodiment of the present disclosure.

[0104] Figure 5 First multiplexers DMUX1 and DMUX2 are shown, along with an address processing module connected to DMUX1 and DMUX2. The address processing modules connected to the other first multiplexers have the same structure and will not be shown individually here.

[0105] like Figure 5 As shown, each address processing module includes a first processing module, a second processing module, a third processing module, and a third multiplexer 1041.

[0106] The first processing module is configured to perform a first modulo operation on the address information, modulo L, to obtain a first result. For example, the first processing module is configured to perform mod L processing on the address information, and the result is sent as the first result to an input terminal of the third multiplexer 1041. The address information is information at a predetermined position in the request address of the received data access request.

[0107] The second processing module is configured to perform a second modulo operation on the address information using Lm as the modulus to obtain a second result. For example, the second processing module is configured to perform mod(Lm) processing on the address information, and the result is sent as the second result to another input of the third multiplexer 1041.

[0108] like Figure 5 As shown, the two input terminals of the third multiplexer are coupled to the first processing module and the second processing module to receive the first result and the second result, and the output terminal of the third multiplexer is coupled to the third processing module.

[0109] The third multiplexer is configured to receive a first indication signal as a control signal, and when the first indication signal indicates that m is equal to 0 (e.g., the first indication signal is a first value), select a first result to output to the third processing module, and when the first indication signal indicates that m is greater than 0 (e.g., the first indication signal is a second value), select a second result to output to the third processing module.

[0110] The third processing module is also coupled to the corresponding first multiplexer and configured to output a signal indicating the destination buffer unit of the received data access request to the corresponding first multiplexer based on the first indication signal, the second indication signal, and the result output by the third multiplexer. The second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo processing and the Lm buffer units.

[0111] For example, the second indication signal includes multiple values, each value indicating the mapping relationship between all possible Lm results obtained by performing the second modulo operation on the address information and Lm normally functioning cache units.

[0112] The Lm possible results include 0, 1, 2, ..., Lm-1. For example, referring to Table 1, when L=4 and m=1, the second indicator signal has 4 values, namely 0, 1, 2, and 3. Each value represents the mapping relationship between all three possible results (0, 1, 2) obtained by mod3 processing of any address information and the three normally functioning cache units.

[0113] For example, in some embodiments, when the first indication signal is a first value, the input received by the third processing module is a first result, and the destination cache unit can be determined based on the first result.

[0114] When the first indication signal is the second value, the input received by the third processing module is the second result. The third processing module can determine the mapping relationship between Lm results and Lm normally functioning cache units based on the second indication signal. From the determined mapping relationship, the identification information of the cache unit corresponding to the second result is determined, and an enable signal for enabling the cache unit indicated by the identification information is output to the corresponding first multiplexer.

[0115] For example, based on the value of the second indication signal, the mapping relationship between Lm possible results of mod(Lm) (0, 1, 2..., Lm-1) and Lm normally functioning cache units is determined. For instance, in a certain mapping relationship, a modulo result of 0 indicates that the destination cache unit is cache unit 1, a modulo result of 1 indicates that the destination cache unit is cache unit 2, and so on. Then, based on the second result, the identification information (e.g., number) of the cache unit corresponding to the second result is determined to identify the destination cache unit.

[0116] For example, each first multiplexer is configured to: send a received data access request as an arbitration candidate to a first multiplexer coupled to the destination cache unit to participate in arbitration, and upon obtaining arbitration, send the received data access request to the destination cache unit.

[0117] For example, each first multiplexer is configured to use a round-robin arbitration method to determine the order in which data access requests received by the first multiplexer enter the buffer unit coupled to the first multiplexer.

[0118] For example, after address processing is complete, the first cross switch circuit will send the data access request to the cache unit indicated by the result, based on the output of the address processing module.

[0119] For example, assuming the address processing module outputs a result indicating cache unit 1, the first crossbar switch circuit is configured to send a data access request to cache unit 1. Specifically, for example, refer to... Figure 5 The data access request is used as an arbitration candidate, for example, by sending a valid signal to cache unit 1 to participate in the arbitration of the first multiplexer coupled to cache unit 1; when the first multiplexer coupled to a cache unit receives multiple data access requests at the same time, it will poll and arbitrate these data access requests to determine the order in which the data access requests enter the second-level cache; when the data access request is arbitrated, the first multiplexer sends the data access request to cache unit 1.

[0120] For example, each second multiplexer is configured to send a request to participate in arbitration to a second multiplexer coupled to the memory cell indicated by the request address in the data access request, with the received data request as an arbitration candidate, and to send the data request to the memory cell indicated by the request address when arbitration is obtained.

[0121] The relationship between data requests and data access requests can be found in the previous description, and will not be repeated here.

[0122] For the second multiplexer, the routing mapping remains unchanged; the memory location to which a data request is destined is still determined by the request address carried in the data access request sent by the hardware execution unit. Although the first crossbar switch circuit may change the destination cache unit of the data access request, the data request used to perform that data access will still go from the destination cache unit to the memory unit that the original data access request would have gone to. Due to the crossbar switch circuit structure, even if a cache unit fails, the data request can still be sent to a normally functioning cache unit, and when it needs to be sent to off-chip memory, it will be sent to the memory unit that the data request was originally destined for.

[0123] For example, the L2 cache and off-chip memory are connected via a second crossbar switch circuit. For example, L cache units correspond one-to-one with L memory units; this one-to-one correspondence means that data returned by a memory unit will be cached in its corresponding cache unit. When m equals 0, for example, when the first indicator signal has a first value, each cache unit still accesses its corresponding memory unit independently. For example, when at least one cache unit fails, for example, when m is greater than 1 and the first indicator signal has a second value, the routing relationship between the hardware execution unit and the cache units is rearranged. Data access requests that should have gone to the failed cache unit are now evenly distributed across the remaining Lm normally functioning cache units, and data access requests that were originally sent to normally functioning cache units also change their destination cache unit due to the change in address routing.

[0124] When the L2 cache is working, a cache miss triggers operations such as read allocate and cache eviction, requiring read and write access to off-chip memory. The L2 cache issues a data request. Upon receiving the request, the second crossbar switch circuit processes the address information representing the off-chip memory within the request address (e.g., the physical address, inherited from the data access request sent by the hardware execution unit), and allocates the data request to the corresponding memory location. When a memory location receives multiple data requests simultaneously, arbitration is performed to determine the processing order.

[0125] Figure 6 This is a schematic structural diagram of a parallel processor provided in one embodiment of the present disclosure. Figure 6 In the example, L=4, m=1. The following section combines... Figure 6 This disclosure provides a detailed description of the processing flow of a parallel processor provided in at least one embodiment.

[0126] like Figure 6 As shown, the parallel processor 100 includes multiple hardware execution units 101, a secondary cache 102, and off-chip memory 103. Each secondary cache includes four cache units, which are illustrated as follows for clarity. Figure 6 The memory consists of cache unit 1, cache unit 2, cache unit 3, and cache unit 4. Each off-chip memory comprises four memory units, which are illustrated as follows for clarity. Figure 6 The memory units are memory unit 1, memory unit 2, memory unit 3 and memory unit 4.

[0127] For a description of the hardware execution unit, L2 cache, and off-chip memory, please refer to the foregoing embodiments; they will not be repeated here.

[0128] The parallel processor includes a first cross switch circuit 104 and a second cross switch circuit 105. The first cross switch circuit 104 includes four first multiplexers coupled to multiple hardware execution units via an on-chip interconnect network, and four first multiplexers coupled to four cache units. The four first multiplexers and the four first multiplexers are cross-connected.

[0129] The second cross-switch circuit includes four second multiplexers coupled to four cache units and four second multiplexers coupled to four memory units, with the four second multiplexers and four second multiplexers cross-connected.

[0130] After tape-out testing, the first and second indicator signals can be configured via registers based on the test results of the cache cells. For example, referring to Table 1, when all cache cells are working normally, the register is set to write the first indicator signal to 0, and the second indicator signal can be any value. When cache cell 4 fails, the register is set to write the first indicator signal to 1, and the second indicator signal to 0. And so on.

[0131] For example, the mapping relationship between the address information of data access requests sent by multiple hardware execution units and Lm cache units can be found in Table 1, which will not be elaborated here.

[0132] For example, the hardware execution unit sends a data access request (REQ) to the secondary cache. The REQ carries the physical address as the request address, uses specific bits of the physical address as address information, processes the address information, and obtains the destination cache unit of the REQ.

[0133] Specifically, after receiving the data access request REQ, the first cross switch circuit takes the bit representing the predetermined position of the destination cache unit in the request address as the address information, performs mod3 and mod4 operations on the address information simultaneously, takes the result of the mod4 operation as the first result, and takes the result of the mod3 operation as the second result.

[0134] For example, in one instance, all cache units are working correctly, the first indicator signal is 0, and the second indicator signal can be any value.

[0135] Since the first indication signal is 0, the data access request REQ is sent to the cache unit indicated by the first result. For example, assuming the first result is 00, the data access request REQ is sent to cache unit 1.

[0136] For example, in another example, if cache unit 4 fails, the first indicator signal is 1 and the second indicator signal is 0.

[0137] Since the first indication signal is 1, the destination cache unit for the data access request (REQ) is determined based on the second result and the second indication signal. For example, assuming the second result is 00, the destination cache unit is determined to be cache unit 1, and the data access request (REQ) is sent to cache unit 1.

[0138] Subsequently, when the data access request REQ initiated by the hardware execution unit misses in the L2 cache, the second cross switch circuit is configured to send a data request to the memory unit indicated by the request address in the data access request REQ in order to perform memory data access related to the data access request.

[0139] For example, when all cache units are working properly, for a data request initiated by cache unit 1, the second crossbar switch circuit sends the data to memory unit 1, and for a data request initiated by cache unit 2, the second crossbar switch circuit sends the data to memory unit 2.

[0140] For example, assuming cache unit 4 fails, data access requests that would normally go to cache unit 4 will be sent to memory unit 4.

[0141] In at least one embodiment of the cross-switch circuit provided in this disclosure, a first cross-switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second cross-switch circuit is used as the interconnection network between the L2 cache and the off-chip memory. When a cache unit fails, the first cross-switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second cross-switch circuit can realize data interaction between the normally functioning cache units and all memory of the off-chip memory without losing the storage space of the off-chip memory, effectively improving the robustness and yield of the chip.

[0142] At least one embodiment of this disclosure also provides a parallel processor. Figure 7 A schematic structural diagram of a parallel processor provided for another embodiment of this disclosure.

[0143] like Figure 7 As shown, the parallel processor includes multiple hardware execution units, multiple L2 caches, and multiple off-chip memories.

[0144] Multiple hardware execution units share multiple L2 caches and multiple off-chip memories.

[0145] Each L2 cache consists of L cache units, and each off-chip memory consists of L memory units. Of the L cache units, m are invalid and unavailable, while the other Lm cache units are functioning normally, where m is an integer greater than or equal to 0 and less than L. The failure scenarios for different L2 caches can vary; for example, some L2 caches may have no failed cache units, while others may have failed cache units.

[0146] For example, each L2 cache may contain multiple cache units; the specific structure can be found in [reference needed]. Figure 4 The description is as follows. Each off-chip memory can also include multiple memory units. For example, the entire off-chip memory space can be divided into multiple partitions, each partition being represented as... Figure 7 The "off-chip memory" in the text is further divided into multiple storage units. For the specific structure, please refer to [reference needed]. Figure 4 The description.

[0147] Each L2 cache is equipped with a corresponding first crossbar switch circuit and a second crossbar switch circuit, and multiple L2 caches are connected to multiple off-chip memories in a one-to-one correspondence.

[0148] The first cross-switch circuit is set between multiple hardware execution units and their corresponding secondary caches, and is configured to perform data interaction between the multiple hardware execution units and the Lm cache units that are working normally in the corresponding secondary caches. In response to m being greater than 0, data access requests that were originally destined for invalid cache units are sent to other working cache units.

[0149] The second cross switch circuit is located between the corresponding L2 cache and the corresponding off-chip memory connected to the L2 cache, and is configured to perform data interaction between the Lm cache units that are working normally in the corresponding L2 cache and the L memory units in the corresponding off-chip memory connected to the L2 cache.

[0150] For example, in some embodiments, a parallel processor may include multiple L2 caches. For instance, due to the large number of hardware execution units and extremely high memory access bandwidth requirements, the L2 cache can be divided into multiple independent slices. Each slice is responsible for a segment of address space, binding a portion of the hardware execution units and the memory controller, enabling parallel processing of memory access requests, improving bandwidth and reducing latency. Logically, multiple L2 caches still belong to a unified global shared cache; structurally, they represent a distributed L2 cache design with physical partitions or physical slices.

[0151] For example, when a parallel processor includes multiple L2 caches, each L2 cache is equipped with a corresponding first crossbar switch circuit and a second crossbar switch circuit.

[0152] For example, such as Figure 7 As shown, the parallel processor also includes an on-chip interconnect network, which is disposed between multiple first crossbar switch circuits corresponding to multiple hardware execution units and multiple secondary caches. The on-chip interconnect network is configured to forward data access requests to the first crossbar switch circuit corresponding to the destination secondary cache based on the destination secondary cache indicated by the data access requests sent by the hardware execution units. Then, the first crossbar switch circuit performs the processing as described above, forwarding the data access request to the destination cache unit in the destination secondary cache.

[0153] For example, multiple hardware execution units have a predetermined routing mapping relationship with L cache units. The first cross switch circuit is configured to: when m=0, route the received data access request to the destination cache unit of the received data access request according to the predetermined routing mapping relationship; when m is greater than 0, update the routing mapping relationship between the multiple hardware execution units and Lm cache units, and route the received data access request to the destination cache unit according to the updated routing mapping relationship. In the updated routing mapping relationship, data access requests that originally went to invalid cache units are sent to normally functioning cache units.

[0154] In response to a data access request initiated by the hardware execution unit not being hit in the secondary cache, the second cross switch circuit is configured to send a data request to the memory unit indicated by the request address in the data access request to perform data access related to the data access request.

[0155] The working process, structure, and connection relationship of each first cross switch circuit and second cross switch circuit can be referred to the relevant description in the foregoing embodiments, and will not be repeated here.

[0156] In at least one embodiment of the cross-switch circuit provided in this disclosure, a first cross-switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second cross-switch circuit is used as the interconnection network between the L2 cache and the off-chip memory. When a cache unit fails, the first cross-switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second cross-switch circuit can realize data interaction between the normally functioning cache units and all memory of the off-chip memory without losing the storage space of the off-chip memory, effectively improving the robustness and yield of the chip.

[0157] At least one embodiment of this disclosure also provides an electronic device. Figure 8 A schematic structural diagram of an electronic device provided for at least one embodiment of this disclosure.

[0158] like Figure 8 As shown, the electronic device 200 includes a parallel processor 100. Further description of the parallel processor 100 can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.

[0159] For example, the electronic device could be a multi-GPU system, which integrates multiple GPU cards and works together to complete large-scale parallel computing tasks. This architecture is commonly used for tasks such as artificial intelligence training, scientific computing, and high-performance graphics rendering.

[0160] For example, in some embodiments, the electronic device 200 provided in at least one embodiment of this disclosure can be a single-machine multi-card form, where multiple GPU cards are integrated in a single server or workstation, and inter-card communication is achieved through a bus (such as PCIe bus, AMD bus, etc.). For example, the electronic device 200 can also be a multi-machine multi-card form, where multiple single-machine multi-card servers are interconnected through a high-speed network to form a cluster, and each card can not only communicate with other cards in the same machine, but also interact with cards in other groups through the network.

[0161] For example, electronic device 200 can also be a TPU cluster, a multi-FPGA accelerator card system, etc., and this disclosure does not impose specific limitations on it.

[0162] For example, this electronic device 200 overcomes the limitations of single-card computing power or memory through parallel computing, and is therefore widely used in scenarios requiring large-scale data processing, high-throughput computing, or low-latency parallel tasks. Examples include deep learning and large model training, high-performance computing, large-scale data processing and AI inference, real-time rendering and visualization, and it can be applied to fields such as large model training, scientific computing, high-concurrency AI services, and real-time rendering.

[0163] Figure 9 This is a schematic diagram of the specific structure of an electronic device provided in at least one embodiment of the present disclosure.

[0164] The following is for reference. Figure 9 The diagram illustrates a specific structural schematic of an electronic device (e.g., a terminal device or a server) 300 suitable for implementing a cross-switch circuit including embodiments of the present disclosure.

[0165] The electronic devices in this disclosure can include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. For example, the electronic device can be in the form of a server, used for various application scenarios such as deep learning and artificial intelligence, scientific computing, graphics rendering and video editing, virtual reality and game development, and cloud services. For example, the electronic device can be a dedicated server such as a data center or cloud computing center that is deployed for tasks such as deep learning training, large-scale data analysis, and high-performance computing.

[0166] Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0167] like Figure 9As shown, the electronic device 300 may include a processing unit 301, such as the aforementioned parallel processor 100, which can execute various appropriate actions and processes according to non-transitory computer-readable instructions stored in memory to achieve various functions. The processing unit 301 may also include devices with instruction optimization capabilities and / or program execution capabilities, such as a central processing unit (CPU) or a tensor processor (TPU). The CPU can be based on x86, ARM, or RISC-V architectures. The GPU can be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.

[0168] like Figure 9 As shown, for example, the memory may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory, etc., for example, computer-readable instructions may be loaded from storage device 308 into RAM 303 to execute computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various applications and various data, such as style images, and various data used and / or generated by the applications, may also be stored in the computer-readable storage medium.

[0169] For example, the processing device 301, the read-only memory (ROM) 302, and the random access memory (RAM) 303 are interconnected via a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0170] Typically, the following devices can be connected to the input / output (I / O) interface 305: input devices 306 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 308 including, for example, magnetic tape, hard disk, flash memory, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 9 An electronic device 300 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 300 may alternatively implement or have more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions.

[0171] At least one embodiment of this disclosure also provides a data transmission method. Figure 10 This is a schematic flowchart illustrating a data transmission method provided in at least one embodiment of the present disclosure.

[0172] The data transmission method provided in at least one embodiment of this disclosure is used for data transmission between multiple hardware execution units and a secondary cache, as well as for data transmission between a secondary cache and off-chip memory.

[0173] Multiple hardware execution units are configured to execute computing tasks in parallel. The multiple hardware execution units share a secondary cache and off-chip memory. The secondary cache includes L cache units, and the off-chip memory includes L memory units. m of the L cache units are in a failed and unavailable state, while the other Lm cache units are working normally. m is an integer greater than or equal to 0 and less than L. The secondary cache is coupled to a first crossbar switch circuit and a second crossbar switch circuit. The first crossbar switch circuit is located between the multiple hardware execution units and the secondary cache, and the second crossbar switch circuit is located between the secondary cache and the off-chip memory.

[0174] For details regarding the structure, connection relationships, and functions of the aforementioned hardware execution unit, L2 cache, and off-chip memory, please refer to the relevant sections on the parallel processor 100 mentioned above; they will not be repeated here.

[0175] like Figure 10 As shown, at least one embodiment of the present disclosure provides a data transmission method including steps S10 and S20.

[0176] In step S10, the first cross-switch circuit performs data interaction between multiple hardware execution units and Lm cache units.

[0177] In response to m being greater than 0, data access requests that were originally destined for invalid cache units are sent to other working cache units.

[0178] In step S20, the second cross switch circuit performs data interaction between Lm cache units and L memory units.

[0179] For example, there is a predetermined routing mapping relationship between multiple hardware execution units and L cache units. In some embodiments, step S10 may include: when m=0, routing the received data access request to the destination cache unit of the received data access request by the first cross switch circuit according to the predetermined routing mapping relationship; when m is greater than 0, updating the routing mapping relationship between the multiple hardware execution units and Lm cache units by the first cross switch circuit, and routing the received data access request to the destination cache unit according to the updated routing mapping relationship, wherein, in the updated routing mapping relationship, data access requests that originally went to invalid cache units are sent to normally functioning cache units.

[0180] For example, in the updated routing mapping, the data access requests that originally went to the invalid cache unit are evenly distributed to Lm cache units, and the data access requests that originally went to Lm cache units have changed their destination cache unit due to the update of the routing mapping.

[0181] For example, in some embodiments, the received first indication signal can be used to determine whether m=0 or m is greater than 0. Furthermore, by combining the first indication signal and the received second indication signal, the specific details of the predetermined routing mapping relationship and the updated routing relationship can also be determined.

[0182] For example, in some embodiments, step S10 may include: in response to receiving the data access request, extracting information from a predetermined position in the request address of the data access request as address information; performing a first modulo operation on the address information with L as the modulus to obtain a first result, and performing a second modulo operation with Lm as the modulus to obtain a second result; in response to the first indication signal being the first value, sending the data access request to the cache unit indicated by the first result by the first cross-connect circuit; in response to the first indication signal being the second value, determining the destination cache unit of the data access request based on the second result and the second indication signal by the first cross-connect circuit, and sending the data access request to the destination cache unit, wherein the destination cache unit is one of the Lm cache units, and the second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo operation and the Lm cache units.

[0183] For example, in some embodiments, step S20 may include: in response to a data access request initiated by the hardware execution unit not being hit in the L2 cache, the second cross switch circuit sends a data request to the memory unit indicated by the request address in the data access request to perform data access related to the data access request.

[0184] For example, the first cross switch circuit includes a plurality of first multiplexers coupled to the plurality of hardware execution units via an on-chip interconnect network, and L first multiplexers coupled to the L cache units, wherein the plurality of first multiplexers and the L first multiplexers are cross-interconnected; the second cross switch circuit includes L second multiplexers coupled to the L cache units, and L second multiplexers coupled to the L memory units, wherein the L second multiplexers and the L second multiplexers are cross-interconnected.

[0185] The specific structure, function, and operation of the first and second cross switch circuits can be found in the descriptions of the foregoing embodiments, and will not be repeated here.

[0186] In the data transmission method provided in at least one embodiment of this disclosure, a first crossbar switch circuit is used as the interconnection network between the hardware execution unit and the L2 cache, and a second crossbar switch circuit is used as the interconnection network between the L2 cache and off-chip memory. When a cache unit fails, the first crossbar switch circuit can send the data access request that was originally destined for the failed cache unit to other normally functioning cache units, so that the chip can still function normally. The second crossbar switch circuit can realize data interaction between normally functioning cache units and all memory of off-chip memory without losing the storage space of off-chip memory, effectively improving the robustness and yield of the chip.

[0187] For more detailed information about steps S10-S20 in the data transmission method provided in at least one embodiment of this disclosure, please refer to the relevant descriptions of the first cross switch circuit and the second cross switch circuit in the foregoing cross switch circuit. Repeated descriptions will not be repeated here.

[0188] The following points should be noted regarding this disclosure: The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0189] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0190] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

[0191] The following points should be noted regarding this disclosure: (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0192] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0193] The above description is only a specific embodiment of this disclosure, but the protection scope of this disclosure is not limited thereto. The protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A parallel processor, characterized in that, Includes multiple hardware execution units, L2 cache, and off-chip memory. The multiple hardware execution units are configured to execute computational tasks in parallel. The multiple hardware execution units share the L2 cache and the off-chip memory. The secondary cache comprises L cache units, and the off-chip memory comprises L memory units, where L is a positive integer greater than 1. Of the L cache units, m are in a failed and unavailable state, while the other Lm cache units are functioning normally, where m is an integer greater than or equal to 0 and less than L. The parallel processor also includes a first cross switch circuit and a second cross switch circuit. The first cross switch circuit is disposed between the plurality of hardware execution units and the secondary cache, and is configured to perform data interaction between the plurality of hardware execution units and the Lm cache units, wherein, in response to m being greater than 0, data access requests that were originally directed to invalid cache units are sent to other normally functioning cache units. The second cross switch circuit is disposed between the L2 cache and the off-chip memory, and is configured to perform data interaction between the Lm cache units and the L memory units.

2. The parallel processor according to claim 1, characterized in that, There is a predetermined routing mapping relationship between the plurality of hardware execution units and the L cache units. The first cross switch circuit is configured as follows: When m=0, the received data access request is routed to the destination cache unit of the received data access request according to the predetermined routing mapping relationship; When m is greater than 0, the routing mapping relationship between the plurality of hardware execution units and the Lm cache units is updated. According to the updated routing mapping relationship, the received data access request is routed to the destination cache unit of the received data access request. In the updated routing mapping relationship, the data access request that originally went to the invalid cache unit is sent to the normally functioning cache unit.

3. The parallel processor according to claim 2, characterized in that, In the updated routing mapping, data access requests that originally went to the invalidated cache unit are evenly distributed across the Lm cache units. The data access requests that originally went to the Lm cache units have changed their destination cache unit due to the update of the routing mapping.

4. The parallel processor according to claim 1, characterized in that, The parallel processor is configured to receive a first indication signal and a second indication signal. When the first indication signal is a first value, it indicates that m equals 0; when the first indication signal is a second value, it indicates that m is greater than 0. The first cross switch circuit is configured as follows: In response to receiving a data access request, information from a predetermined location in the request address of the received data access request is extracted as address information; The address information is subjected to a first modulo operation with L as the modulus to obtain a first result, and a second modulo operation with Lm as the modulus to obtain a second result; In response to the first indication signal being the first value, the received data access request is sent to the cache unit indicated by the first result. In response to the first indication signal being the second value, based on the second result and the second indication signal, the destination cache unit of the received data access request is determined, and the received data access request is sent to the destination cache unit, wherein the destination cache unit is one of the Lm cache units, and the second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo process and the Lm cache units.

5. The parallel processor according to claim 2, characterized in that, For each hardware execution unit, in response to a data access request initiated by the hardware execution unit not being hit in the secondary cache, the second cross switch circuit is configured to send a data request to the memory unit indicated by the request address in the data access request, so as to perform memory data access related to the data access request.

6. The parallel processor according to claim 1, characterized in that, The first cross-switch circuit includes a plurality of first multiplexers coupled to the plurality of hardware execution units via an on-chip interconnect network, and L first multiplexers coupled to the L cache units, wherein the plurality of first multiplexers and the L first multiplexers are cross-interconnected. The second cross-switch circuit includes L second multiplexers coupled to the L cache units and L second multiplexers coupled to the L memory units, wherein the L second multiplexers and the L second multiplexers are cross-connected.

7. The parallel processor according to claim 6, characterized in that, The first crossbar switch circuit includes an address processing module that is coupled one-to-one with each of the plurality of first multiplexers. Each address processing module includes a first processing module, a second processing module, a third multiplexer, and a third processing module. The first processing module is configured to perform a first modulo operation on the address information with L as the modulus to obtain a first result, wherein the address information is information at a predetermined position in the request address of the received data access request; The second processing module is configured to perform a second modulo operation on the address information with Lm as the modulus to obtain a second result; The two input terminals of the third multiplexer are coupled to the first processing module and the second processing module, and the output terminal of the third multiplexer is coupled to the third processing module. The third multiplexer is configured to receive a first indication signal as a control signal, and when the first indication signal indicates that m equals 0, select the first result to be output to the third processing module, and when the first indication signal indicates that m is greater than 0, select the second result to be output to the third processing module. The third processing module is also coupled to the corresponding first multiplexer and configured to output a signal indicating the destination cache unit of the received data access request to the corresponding first multiplexer based on the first indication signal, the second indication signal and the result output by the third multiplexer. The second indication signal is used to indicate the mapping relationship between the Lm modulo results of the second modulo processing and the Lm cache units.

8. The parallel processor according to claim 7, characterized in that, Each first multiplexer is configured as follows: The received data access request is used as an arbitration candidate and sent to the first multiplexer coupled to the destination cache unit to participate in the arbitration request. Upon obtaining arbitration, the received data access request is sent to the destination cache unit.

9. The parallel processor according to claim 6, characterized in that, Each first multiplexer is configured to use a round-robin arbitration method to determine the order in which data access requests received by the first multiplexer enter the buffer unit coupled to the first multiplexer.

10. The parallel processor according to claim 6, characterized in that, Each second multiplexer is configured to send a request to participate in arbitration to a second multiplexer coupled to the memory cell indicated by the request address in the received data request, as an arbitration candidate, and, upon obtaining arbitration, send the received data request to the memory cell indicated by the request address.

11. The parallel processor according to any one of claims 1-10, characterized in that, The parallel processor is a graphics processing unit or a general-purpose graphics processing unit, and the hardware execution unit is a streaming processor cluster. The streaming processor cluster is configured with an independently accessible L1 cache.

12. A parallel processor, characterized in that, The parallel processor includes multiple hardware execution units, multiple L2 caches, and multiple off-chip memories. The plurality of hardware execution units share the plurality of secondary caches and the plurality of off-chip memories. Each L2 cache consists of L cache units, and each off-chip memory consists of L memory units. Of the L cache units, m are in a failed and unavailable state, while the other Lm cache units are functioning normally, where m is an integer greater than or equal to 0 and less than L. Each L2 cache is equipped with a corresponding first crossbar switch circuit and a second crossbar switch circuit, and the plurality of L2 caches are connected to the plurality of off-chip memories in a one-to-one correspondence. The first cross switch circuit is disposed between the plurality of hardware execution units and the corresponding secondary cache, and is configured to perform data interaction between the plurality of hardware execution units and the Lm cache units that are working normally in the corresponding secondary cache, wherein, in response to m being greater than 0, the data access request that was originally sent to the invalid cache unit is sent to other working cache units. The second cross switch circuit is disposed between the corresponding L2 cache and the off-chip memory connected to the corresponding L2 cache, and is configured to perform data interaction between the Lm cache units that are working normally in the corresponding L2 cache and the L memory units in the off-chip memory connected to the corresponding L2 cache.

13. The parallel processor according to claim 12, characterized in that, The parallel processor further includes an on-chip interconnect network, which is disposed between multiple first cross-switch circuits corresponding to the plurality of hardware execution units and the plurality of secondary caches. The on-chip interconnect network is configured to send the data access request to the first crossbar switch circuit corresponding to the destination second-level cache according to the destination second-level cache indicated by the data access request sent by each hardware execution unit, so that the corresponding first crossbar switch circuit sends the data access request to the corresponding destination cache unit in the destination second-level cache.

14. An electronic device, characterized in that, Including the parallel processor as described in any one of claims 1-13.

15. A data transmission method, characterized in that, Used for data transfer between multiple hardware execution units and the L2 cache, as well as data transfer between the L2 cache and off-chip memory. The plurality of hardware execution units are configured to execute computing tasks in parallel. The multiple hardware execution units share the L2 cache and the off-chip memory. The secondary cache comprises L cache units, and the off-chip memory comprises L memory units. Of the L cache units, m are in a failed and unavailable state, while the other Lm cache units are functioning normally, where m is an integer greater than or equal to 0 and less than L. The secondary cache is coupled to a first crossbar switch circuit and a second crossbar switch circuit. The first crossbar switch circuit is disposed between the plurality of hardware execution units and the secondary cache, and the second crossbar switch circuit is disposed between the secondary cache and the off-chip memory. The data transmission method includes: The data interaction between the plurality of hardware execution units and the Lm cache units is performed by the first cross switch circuit, wherein, in response to m being greater than 0, the data access request that was originally sent to the invalid cache unit is sent to other normally functioning cache units. The data exchange between the Lm cache units and the L memory units is performed by the second cross switch circuit.