A deep computation processor system

By setting a detection circuit on the substrate, the depth computing processor is detected and initialized according to the system type, which solves the problem of unsuccessful handshake of computing cards in multi-channel systems and realizes the flexibility and adaptability of multi-channel systems.

CN115269488BActive Publication Date: 2026-08-25HYGON INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210886874.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2026-08-25
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

In multi-channel systems composed of multiple computing cards, handshake failures often occur between the computing cards, a problem that is difficult to solve effectively with existing technologies.

Method used

By setting a detection circuit on the substrate, the system in which each depth computing processor is currently located is detected as either a single-path system or a multi-path system, and the corresponding initialization is performed according to the system type to ensure that all depth computing processors in a multi-path system are running synchronously.

Benefits of technology

It effectively prevents the handshake failure between computing cards in a multi-channel system, improves the flexibility and adaptability of the multi-channel system, and simplifies the setting of the detection circuit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269488B_ABST
    Figure CN115269488B_ABST
Patent Text Reader

Abstract

The application provides a deep computing processor system, which comprises a substrate, a central processor or chipset, a plurality of deep computing processors, each of which is connected with the central processor or chipset, and the plurality of deep computing processors are interconnected through an inter-chip global memory interconnection bus and an inter-chip interconnection control bus. The substrate is further provided with a detection circuit, which detects whether the system in which each deep computing processor is currently located is a single-path system or a multi-path system. A firmware is embedded in each deep computing processor, which initializes the deep computing processor according to the single-path system when the system in which the deep computing processor is currently located is the single-path system, and initializes the deep computing processor according to the multi-path system when the system in which the deep computing processor is currently located is the multi-path system, and synchronizes the running states of all deep computing processors in the multi-path system. The phenomenon that handshaking between the plurality of deep computing processors in the multi-path system is unsuccessful is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a deep computing processor system. Background Technology

[0002] In the field of computer technology, there are numerous application scenarios for compute cards. For example, in some high-end servers, two compute cards are interconnected to enhance computing power. Currently, there are various compute card interconnection technologies, generally using a custom-defined dedicated high-speed interconnect bus. When multiple compute cards form a multi-processor system, the phenomenon of multiple compute cards failing to successfully handshake often occurs. Summary of the Invention

[0003] This invention provides a deep computing processor system to offer a new method for multiplexing deep computing processors, preventing handshake failures between multiple deep computing processors in a multi-way system.

[0004] This invention provides a deep computing processor system, comprising a substrate, a central processing unit (CPU) or chipset disposed on the substrate, and multiple deep computing processors, each of which is connected to the CPU or chipset. The multiple deep computing processors are interconnected via an inter-chip global memory interconnect bus and an inter-chip interconnect control bus. A detection circuit is also disposed on the substrate and connected to each deep computing processor to detect whether the system currently in which each deep computing processor operates is a single-processor system or a multi-processor system. Each deep computing processor is embedded with firmware. This firmware is used to initialize the deep computing processor according to a single-processor system when the current system is a single-processor system; and to initialize the deep computing processor according to a multi-processor system when the current system is a multi-processor system, and to synchronize the operating states of all deep computing processors in the multi-processor system.

[0005] In the above scheme, a detection circuit is set on the substrate to detect whether the system in which each depth computing processor is currently located is a single-processor system or a multi-processor system. This allows each depth computing processor to be initialized using a different initialization method depending on the system type it is in. Specifically, when the system in which the depth computing processor is currently located is a multi-processor system, the processor is initialized according to the multi-processor system's requirements, and the operating status of all depth computing processors in the multi-processor system is synchronized to prevent handshaking failures between multiple depth computing processors in a multi-processor system.

[0006] In one specific implementation, each of the multiple deep learning processors has an inter-chip global memory interconnect (IPE) bus interface and an IPE control bus interface. Among the multiple deep learning processors, the IPE bus interface of any one deep learning processor is interconnected with the IPE bus interfaces of every other deep learning processor; similarly, the IPE control bus interface of any one deep learning processor is interconnected with the IPE control bus interfaces of every other deep learning processor. This interconnection of the IPE bus interface and IPE control bus interface of each deep learning processor with the IPE bus interfaces and IPE control bus interfaces of every other deep learning processor facilitates the formation of multiplexed systems between any two or more deep learning processors, improving the flexibility and adaptability of multiplexed systems.

[0007] In one specific implementation, there are two deep learning processors. Each deep learning processor has a first inter-chip global memory interconnect (IPE) bus interface and a second IPE bus interface; the first IPE bus interfaces of the two deep learning processors are interconnected, and the second IPE bus interfaces of the two deep learning processors are interconnected. Each deep learning processor also has an IPE control bus interface, and the IPE control bus interfaces of the two deep learning processors are interconnected. This simplifies the structure of the deep learning processor system.

[0008] In one specific implementation, each depth computing processor further includes an in-situ state interface and a first general-purpose input / output interface. The in-situ state interface outputs a signal indicating whether the depth computing processor is in an in-situ state, and the first general-purpose input / output interface outputs a signal indicating whether the depth computing processor is in an operational state. A detection circuit is connected to both the in-situ state interface and the first general-purpose input / output interface of each depth computing processor to acquire the in-situ state and operational state of each depth computing processor. The detection circuit is further configured to determine that the system currently occupied by at least two depth computing processors simultaneously in both in-situ and operational states is a multi-channel system when at least two depth computing processors are simultaneously in both in-situ and operational states. Furthermore, the detection circuit is configured to determine that the system currently occupied by each depth computing processor is a single-channel system when at least one depth computing processor is simultaneously in both in-situ and operational states. Each depth computing processor also has a second general-purpose input / output interface connected to the detection circuit, which receives a signal from the detection circuit indicating whether the depth computing processor is in a single-channel or multi-channel system. This is to help determine whether the system in which each depth computing processor is currently located is a single-socket system or a multi-socket system.

[0009] In one specific implementation, there are two depth computing processors. The detection circuit includes: a first OR gate with two inputs and one output, an AND gate with two inputs and one output, a second OR gate with two inputs and one output, and a switch circuit with inputs and an output. The two inputs of the first OR gate are respectively connected to the presence / absence interfaces of the two depth computing processors; a high-level signal output from each presence / absence interface indicates that the depth computing processor is in an absent state, and a low-level signal output from each presence / absence interface indicates that the depth computing processor is in a present state. The two inputs of the AND gate are respectively connected to the first general-purpose input / output interfaces of the two depth computing processors; a high-level signal output from each first general-purpose input / output interface indicates that the depth computing processor is in a non-operating state, and a low-level signal output from each interface indicates that the depth computing processor is in an operating state. The two inputs of the second OR gate are respectively connected to the outputs of the first OR gate and the AND gate; a high-level signal output from the second OR gate indicates that each depth computing processor is in a single-path system, and a low-level signal output from the second OR gate indicates that each depth computing processor is in a multiplexed system. The input of the switching circuit is connected to the output of the second OR gate circuit, and the output of the switching circuit is connected to the second general-purpose input / output interface of both depth computing processors. Through simple combinational logic circuitry, it is possible to determine whether each depth computing processor is currently in a single-path or multi-path system, simplifying the configuration of the detection circuit.

[0010] In one specific implementation, there are four deep learning processors, divided into two pairs. Each deep learning processor has a first inter-chip global memory interconnect (IPI) bus interface and a second IPI bus interface; the two first IPI bus interfaces in each pair are interconnected; and the second IPI bus interfaces of two deep learning processors in one pair are interconnected with the second IPI bus interfaces of two deep learning processors in the other pair. Each deep learning processor also has a first IPI control bus interface and a second IPI control bus interface; the two first IPI control bus interfaces in each pair are interconnected; and the second IPI control bus interfaces of two deep learning processors in one pair are interconnected with the second IPI control bus interfaces of two deep learning processors in the other pair. This improves the computing power of the deep learning processor system while reducing the number of interconnect traces between processors using different deep learning technologies, facilitating manufacturing.

[0011] In one specific implementation, each depth computing processor also has an in-situ state interface and a first general-purpose input / output interface. The in-situ state interface outputs a signal indicating whether the depth computing processor is in an in-situ state, and the first general-purpose input / output interface outputs a signal indicating whether the depth computing processor is in an operational state. The detection circuit is connected to both the in-situ state interface and the first general-purpose input / output interface of each depth computing processor to acquire the in-situ state and operational state of each depth computing processor. The detection circuit is further configured to determine that the system currently occupied by each of the four depth computing processors is a multi-channel system when all four depth computing processors are simultaneously in an in-situ state and an operational state; and to determine that the system currently occupied by each of the four depth computing processors is a single-channel system when at least one depth computing processor is not simultaneously in an in-situ state and an operational state. Each depth computing processor also has a second general-purpose input / output interface connected to the detection circuit, which receives a signal from the detection circuit indicating whether the depth computing processor is in a single-channel system or a multi-channel system. This is to facilitate determining whether the current system of each deep computing processor is a single-socket system or a multi-socket system, so that the deep computing processor system can either operate as a multi-socket system consisting of four deep computing processors or as a single-socket system consisting of each deep computing processor, thus simplifying the combination of multi-socket systems.

[0012] In one specific embodiment, the detection circuit includes: a first OR gate circuit with four inputs and one output, an AND gate circuit with four inputs and one output, a second OR gate circuit with two inputs and one output, and a switching circuit with inputs and an output. The four inputs of the first OR gate circuit are respectively connected to the presence / absence interfaces of four depth computing processors; a high-level signal output from each presence / absence interface indicates that the depth computing processor is in an absent state, and a low-level signal output from each presence / absence interface indicates that the depth computing processor is in a present state. The two inputs of the AND gate circuit are respectively connected to the first general-purpose input / output interfaces of the four depth computing processors; a high-level signal output from each first general-purpose input / output interface indicates that the depth computing processor is in a non-operating state, and a low-level signal output from each interface indicates that the depth computing processor is in an operating state. The two inputs of the second OR gate circuit are respectively connected to the output of the first OR gate circuit and the output of the AND gate circuit; a high-level signal output from the second OR gate circuit indicates that each depth computing processor is in a single-path system, and a low-level signal output from the second OR gate circuit indicates that each depth computing processor is in a multiplexed system. The input of the switching circuit is connected to the output of the second OR gate circuit, and the output of the switching circuit is connected to the second general-purpose input / output interface of each of the four depth computing processors. Through simple combinational logic circuitry, it is possible to determine whether each depth computing processor is currently in a single-path or multi-path system, simplifying the configuration of the detection circuit.

[0013] In one specific implementation, the firmware is further configured to initialize the inter-chip interconnect control bus interface of the deep learning processor according to the multi-processor system when the system in which the deep learning processor currently resides is a multi-processor system. The firmware is also configured to synchronize the operating state of all deep learning processors in the multi-processor system via the inter-chip global memory interconnect bus interface after the inter-chip interconnect control bus interfaces of all deep learning processors in the multi-processor system have been initialized. By first fully initializing the inter-chip interconnect control bus interface of each deep learning processor in the multi-processor system and then synchronizing the operating state of all deep learning processors in the multi-processor system, the occurrence of handshake failures between multiple deep learning processors in the subsequent multi-processor system can be better prevented.

[0014] In one specific implementation, the first and second general-purpose input / output (GPIO) interfaces are further configured to output signals indicating that the inter-chip interconnect control bus (IPC) interface of each deep learning processor has completed initialization. The firmware is also configured to, after the IPC interface of the deep learning processor has been initialized according to the multiplexing system, control the first GPIO interface of the deep learning processor to output a high-level signal, thereby causing the second GPIO interface of the deep learning processor to output a high-level signal. The firmware is also configured to, upon detecting that the second GPIO interfaces of all deep learning processors in the multiplexing system are outputting high-level signals, synchronize the operating state of all deep learning processors in the multiplexing system via the inter-chip global memory interconnect (IGBTE) interface. This allows the firmware to identify that each deep learning processor in the multiplexing system has completed the initialization of its IPC interface, better synchronizing the operating state of all deep learning processors in the multiplexing system, and better preventing subsequent handshake failures between multiple deep learning processors in the multiplexing system.

[0015] In one specific implementation, the detection circuit is a combinational logic circuit or a programmable logic device, so that it can be arbitrarily adjusted by the logic device to adapt to different types of multi-channel system combinations.

[0016] In one specific implementation, the PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) interface of the central processing unit or chipset is connected to the PCIe interface of each depth computing processor. A clock buffer is also provided on the substrate, connected to the PCIe reference clock interface of the central processing unit or chipset, and also connected to the PCIe reference clock interface of each depth computing processor. This design ensures that multiple depth computing processors share the same clock signal as the central processing unit or chipset, enabling synchronized clock reference signals across multiple systems.

[0017] In one specific implementation, the clock buffer is also connected to the inter-chip memory interconnect bus reference clock interface of each depth computing processor. The clock reference signal used for communication between each depth computing processor and the central processing unit or chipset, and the clock reference signal used for communication between multiple depth computing processors, are clock signals of the same origin.

[0018] In one specific embodiment, the deep computing processor system further includes: an oscillator or resonator disposed on the substrate, and a clock generator disposed on the substrate and connected to the oscillator or resonator. The clock transmitter is also connected to the inter-chip memory interconnect bus reference clock interface of each deep computing processor. The clock reference signal used for communication between each deep computing processor and the central processing unit or chipset, and the clock reference signal used for communication between multiple deep computing processors, are clock signals from different sources. Attached Figure Description

[0019] Figure 1 A structural block diagram of a deep computing processor system provided in an embodiment of the present invention;

[0020] Figure 2 for Figure 1 The diagram shows the interconnects of a deep computing processor system.

[0021] Figure 3 for Figure 1 The diagram shows a schematic of the structure of a detection circuit in a depth computing processor system;

[0022] Figure 4 for Figure 3 The flowchart shown illustrates the operation of the detection circuit.

[0023] Figure 5 A structural block diagram of another deep computing processor system provided in an embodiment of the present invention;

[0024] Figure 6 for Figure 5 The diagram shows the interconnects of a deep computing processor system.

[0025] Figure 7 for Figure 5 The diagram shows a schematic of the structure of a detection circuit in a depth computing processor system;

[0026] Figure 8 for Figure 7 The flowchart shown illustrates the operation of the detection circuit.

[0027] Figure label:

[0028] 10 - Central Processing Unit or Chipset 20 - Deep Computing Processor 30 - Detection Circuit

[0029] 41- Clock buffer 42- Oscillator or resonator 43- Clock generator

[0030] 51 - First OR gate circuit 52 - Second OR gate circuit 53 - AND gate circuit 54 - Switching circuit Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] To facilitate understanding of the deep computing processor system provided in this embodiment of the invention, the application scenario of the deep computing processor system provided in this embodiment of the invention will be described first. This deep computing processor system is applied in a computer to perform tasks such as logical operations. The deep computing processor system will then be described in detail below with reference to the accompanying drawings.

[0033] refer to Figure 1 and Figure 5 The deep computing processor system provided in this embodiment of the invention includes a substrate, on which a central processing unit or chipset 10 is disposed, and a plurality of deep computing processors 20 are also disposed. Each of the plurality of deep computing processors 20 is connected to the central processing unit or chipset 10, and the plurality of deep computing processors 20 are interconnected through an inter-chip global memory interconnect bus and an inter-chip interconnect control bus. A detection circuit 30 is also disposed on the substrate, which is connected to each deep computing processor 20 to detect whether the system in which each deep computing processor 20 is currently located is a single-processor system or a multi-processor system. Each deep computing processor 20 is embedded with firmware, which is used to initialize the deep computing processor 20 according to the single-processor system when the system in which the deep computing processor 20 is currently located is a single-processor system; the firmware is also used to initialize the deep computing processor 20 according to the multi-processor system when the system in which the deep computing processor 20 is currently located is a multi-processor system, and to synchronize the operating status of all deep computing processors 20 in the multi-processor system.

[0034] In the above scheme, a detection circuit 30 is set on the substrate to detect whether the system in which each depth computing processor 20 is currently located is a single-path system or a multi-path system. This allows each depth computing processor 20 to be initialized using a different initialization method depending on the system type it is in. Specifically, when the system in which the depth computing processor 20 is currently located is a multi-path system, the depth computing processor 20 is initialized according to the multi-path system requirements, and the operating status of all depth computing processors 20 in the multi-path system is synchronized to prevent handshaking failures between multiple depth computing processors 20 in a multi-path system. The following is a detailed description of each of the above structures with reference to the accompanying drawings.

[0035] When setting up the substrate, the substrate serves as a carrier for mounting components such as, but not limited to, a central processing unit, chipset, depth computing processor 20, and detection circuitry 30. A circuit board, such as, but not limited to, a printed circuit board, can be used as the substrate. In practical applications, a motherboard can be used as the substrate. (Reference) Figure 1 and Figure 5 The substrate also includes a central processing unit or chipset 10 to run an operating system, read and write memory such as, but not limited to, hard disks and disks, and other functions.

[0036] Continue to refer to Figure 1 and Figure 5 Multiple depth computing processors 20 are also disposed on the substrate, and each of the multiple depth computing processors 20 is connected to the central processing unit or chipset 10. Specifically, the PCIe interface of the central processing unit or chipset 10 can be connected to the PCIe interface of each depth computing processor 20 to realize communication and interaction between each depth computing processor 20 and the central processing unit or chipset 10. The number of depth computing processors 20 disposed on the substrate can be as follows: Figure 1 The two shown can also be as follows: Figure 5 The four shown are examples of the four types of depth computing processors 20. Of course, the number of depth computing processors 20 on the substrate can also be any value of 3, 5, 6, 8, etc., with no fewer than 2.

[0037] like Figure 1 and Figure 5 As shown, a clock buffer 41 can also be disposed on the substrate. The clock buffer 41 is connected to the PCIe reference clock interface of the central processing unit or chipset 10 to receive the reference clock signal generated by the central processing unit or chipset 10. Furthermore, the clock buffer 41 is also connected to the PCIe reference clock interface of each depth computing processor 20 to transmit the reference clock signal received from the central processing unit or chipset 10 to the PCIe reference clock interface of each depth computing processor 20, providing a clock reference signal for communication between the central processing unit or chipset 10 and each depth computing processor 20. This configuration, by designing multiple depth computing processors 20 to use the same source clock signal as the central processing unit or chipset 10, enables synchronized clock reference signals between multiple systems.

[0038] refer to Figure 1 , Figure 2 , Figure 5 and Figure 6Multiple deep learning processors 20 are interconnected via an inter-chip global memory interconnect bus and an inter-chip interconnect control bus. The inter-chip global memory interconnect bus is used to transmit specific data signals between different deep learning processors 20, and the inter-chip interconnect control bus is used to transmit specific control signals between different deep learning processors 20. This enables the multiple deep learning processors 20 to transmit control and data signals to each other, facilitating the formation of multiplexed systems with at least two deep learning processors 20 to improve overall computing power. Figure 1 and Figure 5 As shown, each depth computing processor 20 has an inter-chip global memory interconnect bus interface (e.g., Figures 1 to 8 In this context, the xHMI bus interface represents the inter-chip global memory interconnect bus interface. When different deep computing processors 20 interconnect via the inter-chip global memory interconnect bus, the inter-chip global memory interconnect bus connects between the inter-chip global memory interconnect bus interfaces of the different deep computing processors 20 to achieve inter-chip global memory interconnection between the different deep computing processors 20. Each deep computing processor 20 has an inter-chip interconnect control bus interface (such as...). Figures 1 to 8 The CFIP interface in the text represents the chip interconnect control bus interface. When different depth computing processors 20 interconnect with each other via the chip interconnect control bus, the chip interconnect control bus is connected between the chip interconnect control bus interfaces of different depth computing processors 20 to realize the interconnection of the chip interconnect control bus between different depth computing processors 20.

[0039] It is important to note that the number of paths in a multi-path system can be dual-path, quad-path, or octa-path, specifically equal to the number of depth computing processors 20 included in the system. However, the number of depth computing processors 20 in a multi-path system can be equal to or less than the number of depth computing processors 20 on the substrate. When the number of depth computing processors 20 in a multi-path system is equal to the number of depth computing processors 20 on the substrate, all depth computing processors 20 on the substrate together form a multi-path system. When the number of depth computing processors 20 in a multi-path system is less than the number of depth computing processors 20 on the substrate, the substrate contains one multi-path system and at least one single-path system.

[0040] refer to Figure 1 and Figure 5 Furthermore, an oscillator or resonator 42 and a clock generator 43 can be disposed on the substrate, wherein the clock generator 43 is connected to the oscillator or resonator 42 to generate a clock reference signal. The clock transmitter is also connected to the inter-chip memory interconnect bus reference clock interface (e.g., ...) of each depth computing processor 20. Figures 1 to 8The xHMI reference clock interface (representing the inter-chip memory interconnect bus reference clock interface) is connected to transmit the generated clock reference signal to each depth computing processor 20, providing a reference signal for communication and interaction between different depth computing processors 20, and also providing a reference signal for the calculation of each depth computing processor 20. This ensures that the clock reference signal used for communication between each depth computing processor 20 and the central processing unit or chipset 10, and the clock reference signal used for communication between multiple depth computing processors 20, are clock signals from different sources. Additionally, the reference... Figure 2 and Figure 6 Furthermore, the clock buffer 41 on the substrate can also be connected to the inter-chip memory interconnect bus reference clock interface of each depth computing processor 20, so as to synchronously provide the clock reference signal generated by the central processing unit or chipset 10 to the inter-chip memory interconnect bus reference clock interface of each depth computing processor 20, so that when different depth computing processors 20 communicate with each other, they use the clock reference signal originally generated by the central processing unit or chipset 10, so that the clock reference signal used by each depth computing processor 20 to communicate with the central processing unit or chipset 10 and the clock reference signal used to communicate between multiple depth computing processors 20 are the same source clock signal.

[0041] refer to Figure 1 and Figure 5 A detection circuit 30 is also provided on the substrate. This detection circuit 30 is connected to each depth computing processor 20 to detect whether the system currently occupied by each depth computing processor 20 is a single-path system or a multi-path system, thereby identifying whether each depth computing processor 20 is a single-path system composed of its own components or a multi-path system composed of at least one other depth computing processor 20. In the specific configuration of the detection circuit 30, combinational logic circuits or programmable logic devices can be used. The programmable logic devices can be, but are not limited to, CPLDs (Complex Programmable Logic Devices) and FPGAs (Field Programmable Gate Arrays), allowing for arbitrary adjustments to the logic devices to adapt to different combinations of multi-path systems.

[0042] Furthermore, each depth computing processor 20 is embedded with firmware. This firmware is used to initialize the depth computing processor 20 according to the single-processor system when the current system is a single-processor system, thus eliminating the need to configure registers and other devices for synchronizing the operating states of different depth computing processors 20. The firmware is also used to initialize the depth computing processor 20 according to the multi-processor system when the current system is a multi-processor system, configuring registers and other devices for subsequent operating state synchronization. After each depth computing processor 20 in a multi-processor system has completed initialization, the firmware can also synchronize the operating states of all depth computing processors 20 in the multi-processor system, facilitating handshaking between different depth computing processors 20 in the multi-processor system. By setting a detection circuit 30 on the substrate to detect whether the current system of each depth computing processor 20 is a single-processor system or a multi-processor system, the firmware embedded in each depth computing processor 20 initializes the depth computing processor 20 using different initialization methods according to the different system types in which the depth computing processor 20 is currently located. Especially when the system in which the depth computing processor 20 is currently located is a multi-way system, the depth computing processor 20 is initialized according to the multi-way system, and the running status of all depth computing processors 20 in the multi-way system is synchronized to prevent the phenomenon of unsuccessful handshake between multiple depth computing processors 20 in the multi-way system.

[0043] It should be noted that the interconnection methods of the inter-chip global memory interconnect bus and inter-chip interconnect control bus among the multiple deep computing processors 20 are somewhat related to the configuration of the multiplexed system composed of the multiple deep computing processors 20. The following examples illustrate different configurations of the detection circuit 30 for different interconnection methods.

[0044] For example, among multiple deep learning processors 20, the inter-chip global memory interconnect (IPE) bus interface of any one deep learning processor 20 can be interconnected with the IPE bus interfaces of every other deep learning processor 20. That is, each deep learning processor 20 has an IPE bus interconnection with every other deep learning processor 20. Furthermore, the IPE control bus interface of any one deep learning processor 20 is interconnected with the IPE control bus interfaces of every other deep learning processor 20. This ensures that the IPE bus interface and IPE control bus interface of each deep learning processor 20 are interconnected with the IPE bus interfaces and IPE control bus interfaces of every other deep learning processor 20. At this time, even if some of the depth computing processors 20 are out of place or not working, it does not affect the interconnection between the remaining depth computing processors 20. Therefore, a multiplex system composed of multiple depth computing processors 20 can be composed of any two or more depth computing processors 20, which facilitates the formation of a multiplex system between any two or more depth computing processors 20, and improves the flexibility and adaptability of forming a multiplex system.

[0045] At this point, the number of depth calculation processors 20 on the substrate can be any value, such as two, three, four, or more than two. For example... Figure 1 and Figure 2 The number of depth computing processors 20 shown is two. Each of the two depth computing processors 20 has a first inter-chip global memory interconnect bus interface (e.g., Figures 1 to 8 The xHMI0 bus interface in the text represents the first inter-chip global memory interconnect bus interface and the second inter-chip global memory interconnect bus interface (e.g., ...). Figures 1 to 8 The xHMI1 bus interface in the diagram represents the first inter-chip global memory interconnect bus interface. The first inter-chip global memory interconnect bus interfaces of the two deep learning processors 20 are interconnected, and the second inter-chip global memory interconnect bus interfaces of the two deep learning processors 20 are interconnected to achieve inter-chip global memory interconnection between the two deep learning processors 20. (Continue to refer to...) Figure 1 and Figure 2 Each depth computing processor 20 also has an inter-chip interconnect control bus interface, and the inter-chip interconnect control bus interfaces of two depth computing processors 20 are interconnected. This simplifies the structure of the depth computing processor system.

[0046] When using the above interconnection method, the detection circuit 30 can confirm that at least two depth computing processors 20 are in the same multiplexing system as long as it detects that at least two depth computing processors 20 are simultaneously in both the in-place and working states. When specifically detecting the in-place and working states of each depth computing processor 20, refer to... Figure 1 Each depth computing processor 20 may also have an in-situ state interface and a first general-purpose input / output interface (such as...). Figures 1 to 8 The first GPIO in the diagram represents the first general-purpose input / output interface. The presence / inventory interface outputs a signal indicating whether the depth computing processor 20 is in a present / inventory state, and the first general-purpose input / output interface outputs a signal indicating whether the depth computing processor 20 is in an operational state. The detection circuit 30 is connected to both the presence / inventory interface and the first general-purpose input / output interface of each depth computing processor 20 to collect the presence / inventory and operational states of each depth computing processor 20, thereby obtaining the number of depth computing processors 20 currently simultaneously in both present and operational states. When the detection circuit 30 collects at least two depth computing processors 20 simultaneously in both present and operational states, it directly determines that the system currently occupied by these at least two depth computing processors 20 is a multi-channel system. Furthermore, when the detection circuit 30 collects no more than one depth computing processor 20 simultaneously in both present and operational states (specifically, one or zero), it determines that the system currently occupied by each depth computing processor 20 is a single-channel system. (Reference) Figure 1 Each depth computing processor 20 may also have a second general-purpose input / output interface (e.g., connected to the detection circuit 30) that is connected to the detection circuit 30. Figures 1 to 8 The second GPIO in the figure represents the second general purpose input / output interface. The second general purpose input / output interface is used to receive the signal determined by the detection circuit 30 as to whether the depth computing processor 20 is in a single-path system or a multi-path system, so as to determine whether the current system of each depth computing processor 20 is a single-path system or a multi-path system.

[0047] The following is based on Figure 1 and Figure 2 The number of depth computing processors 20 shown is two (the two depth computing processors are respectively) Figures 1-4 Taking the first and second depth processing processors in the image as examples, an exemplary configuration of the detection circuit 30 is shown. (See reference...) Figure 3The detection circuit 30 includes a first OR gate 51, an AND gate 53, a second OR gate 52, and a switching circuit 54. The first OR gate 51 has two inputs and one output. The two inputs of the first OR gate 51 are respectively connected to the presence / absence interfaces of the two depth computing processors 20. A high-level signal output from each presence / absence interface indicates that the depth computing processor 20 is in an absent state, and a low-level signal output from each presence / absence interface indicates that the depth computing processor 20 is in a present state. The AND gate 53 has two inputs and one output. The two inputs of the AND gate 53 are respectively connected to the first general-purpose input / output interfaces of the two depth computing processors 20. A high-level signal output from each first general-purpose input / output interface indicates that the depth computing processor 20 is in a non-operating state, and a low-level signal output from each interface indicates that the depth computing processor 20 is in an operating state. The second OR gate circuit 52 has two input terminals and one output terminal. The two input terminals of the second OR gate circuit 52 are respectively connected to the output terminals of the first OR gate circuit 51 and the AND gate circuit 53. A high-level signal output by the second OR gate circuit 52 indicates that each depth computing processor 20 is in a single-path system, and a low-level signal output by the second OR gate circuit 52 indicates that each depth computing processor 20 is in a multiplexed system. The switch circuit 54 has one input terminal and one output terminal. The input terminal of the switch circuit 54 is connected to the output terminal of the second OR gate circuit 52, and the output terminal of the switch circuit 54 is connected to the second general-purpose input / output interface of both depth computing processors 20. This configuration, using simple combinational logic circuits, can determine whether each depth computing processor 20 is currently in a single-path or multiplexed system, simplifying the configuration of the detection circuit 30. Additionally, refer to... Figure 3 A diode can be provided between each of the first general-purpose input / output interfaces and the input terminal of the AND gate circuit 53, so that the first general-purpose input / output interface can only be input to the input terminal of the AND gate circuit 53 in one direction, and the input terminal of the AND gate circuit 53 cannot be input to the first general-purpose input / output interface in reverse, so that the first general-purpose input / output interface and the detection circuit 30 can be isolated to prevent mutual interference.

[0048] In Adoption Figure 3 When the detection circuit 30 shown performs detection, it first closes the switch circuit 54 to power on the depth computing processor system. Figure 1When only one of the two depth computing processors 20 is in the active state, and the other is not, one input of the first OR gate 51 is a high-level signal and the other is a low-level signal. After logical operation, the output of the first OR gate 51 becomes a high-level signal. At this time, regardless of whether the output of the AND gate 53 is low or high, since the output of the first OR gate 51 serves as an input of the second OR gate 52, the output of the second OR gate 52 always outputs a high-level signal. This causes the second general-purpose input / output interface of the depth computing processor 20 to receive a high-level signal. When the firmware in each depth computing processor 20 detects that the second general-purpose input / output interface is high, it initializes each depth computing processor 20 according to a single-path system. Simultaneously, by disconnecting the switch circuit 54, the second general-purpose input / output interface of each depth computing processor 20 can also be made high, thus initializing each depth computing processor 20 according to a single-path system regardless of whether it is in the active or active state.

[0049] Of course, after the switch circuit 54 is closed and the depth computing processor system is powered on, when both depth computing processors 20 are in the active state, both inputs of the first OR gate circuit 51 are low-level signals, making the output of the first OR gate circuit 51 also a low-level signal. At this time, the output of the second OR gate circuit 52 is equal to the output of the AND gate circuit 53. If both depth computing processors 20 are in the working state after the depth computing processor system is powered on, both inputs of the AND gate circuit 53 are low-level signals, making the output of the AND gate circuit 53 also a low-level signal. Then the output of the second OR gate circuit 52 is also a low-level signal. The firmware in each depth computing processor 20 can then determine that each depth computing processor 20 is in a multi-path system (specifically a dual-path system), and initialize each depth computing processor 20 according to the multi-path system.

[0050] In addition to the interconnection methods shown above, other interconnection methods can also be used between the multiple depth computing processors 20. For example, it is not necessary for each depth computing processor 20 to be interconnected with every other depth computing processor 20, but rather with some of the other depth computing processors 20. Each depth computing processor 20 only needs to be interconnected with at least one of the other depth computing processors 20 via an inter-chip global memory interconnect bus and an inter-chip control bus. However, since each depth computing processor 20 is not interconnected with every other depth computing processor 20, interconnection between other depth computing processors 20 may not be possible when some depth computing processors 20 are in an inactive or non-operating state. Therefore, the detection circuit 30 can confirm that all depth computing processors 20 together form a multiplexed system only when it detects that all depth computing processors 20 are simultaneously in an inactive and operating state; otherwise, it is confirmed that each depth computing processor 20 forms a single-path system. The following uses... Figure 5 and Figure 6 The number of depth computing processors 20 shown is four (e.g., Figures 5-8 Taking the first, second, third, and fourth deep computing processors in the system as an example, an interconnection method is exemplified.

[0051] refer to Figure 5 and Figure 6 For ease of description, the four depth computing processors 20 are divided into two pairs of depth computing processors 20, each pair of depth computing processors 20 containing two depth computing processors 20, such as... Figure 5 and Figure 6 The two adjacent deep learning processors 20 on the left side are shown as a pair, and the two adjacent deep learning processors 20 on the right side are also shown as a pair. Similarly, each deep learning processor 20 has a first inter-chip global memory interconnect (IPE) bus interface and a second IPE bus interface. During interconnection, the two first IPE bus interfaces in each pair of deep learning processors 20 are interconnected. Simultaneously, the second IPE bus interfaces of two deep learning processors 20 in one pair of deep learning processors 20 are interconnected with the second IPE bus interfaces of two deep learning processors 20 in the other pair of deep learning processors 20, respectively. This ensures that each deep learning processor 20 has at least one IPE bus interface interconnected with the IPE bus interfaces of at least one other deep learning processor 20. Figure 5 and Figure 6As shown, each depth computing processor 20 also has a first inter-chip interconnect control bus interface (such as...). Figures 5-8 In this context, CFIP0 represents the first inter-chip interconnect control bus interface and the second inter-chip interconnect control bus interface (e.g., ...). Figures 5-8 In this context, CFIP1 represents the second inter-chip interconnect control bus interface. The two first inter-chip interconnect control bus interfaces in each pair of deep computing processors 20 are interconnected. Simultaneously, the second inter-chip interconnect control bus interfaces of two deep computing processors 20 in one pair are interconnected with the second inter-chip interconnect control bus interfaces of two deep computing processors 20 in the other pair. This ensures that each deep computing processor 20 has at least one inter-chip interconnect control bus interface, interconnected with at least one inter-chip interconnect control bus interface of at least one other deep computing processor 20. This not only improves the computing power of the deep computing processor system but also reduces the number of interconnect traces between processors using different deep computing technologies, facilitating manufacturing.

[0052] At this time, in response to such Figure 5 and Figure 6 When the interconnection method shown is configured with the detection circuit 30, it can still collect the presence and operation status information of each depth computing processor 20, and determine whether all depth computing processors 20 are simultaneously in the presence and operation state, thereby determining whether it is a single-path system or a multi-path system. Similarly, refer to... Figure 5 and Figure 6 Each depth computing processor 20 may also have an in-situ state interface and a first general-purpose input / output interface. The in-situ state interface outputs a signal indicating whether the depth computing processor 20 is in an in-situ state, and the first general-purpose input / output interface outputs a signal indicating whether the depth computing processor 20 is in an operational state. The detection circuit 30 is connected to both the in-situ state interface and the first general-purpose input / output interface of each depth computing processor 20 to collect the in-situ and operational states of each depth computing processor 20, thereby obtaining the number of depth computing processors 20 currently simultaneously in an in-situ and operational state. When the detection circuit 30 collects data showing that all four depth computing processors 20 are simultaneously in an in-situ and operational state, it determines that the system currently in which each of the four depth computing processors 20 is located is a multi-channel system. Furthermore, when the detection circuit 30 collects data showing that at least one (specifically one, two, three, or four) depth computing processors 20 are not simultaneously in an in-situ and operational state, it determines that the system currently in which each of the four depth computing processors 20 is located is a single-channel system. (Reference) Figure 5Each depth computing processor 20 also has a second general-purpose input / output interface connected to the detection circuit 30. The second general-purpose input / output interface is used to receive a signal determined by the detection circuit 30 as to whether the depth computing processor 20 is in a single-path system or a multi-path system, so as to determine whether the current system of each depth computing processor 20 is a single-path system or a multi-path system, so that the depth computing processor system can work either as a multi-path system composed of four depth computing processors 20 or as a single-path system composed of each depth computing processor 20, simplifying the combination of multi-path systems.

[0053] The following is based on Figure 5 and Figure 6 Taking a four-unit depth computing processor 20 as an example, this illustration demonstrates an exemplary configuration of the detection circuit 30. (See reference...) Figure 7 The detection circuit 30 includes a first OR gate 51, an AND gate 53, a second OR gate 52, and a switching circuit 54. The first OR gate 51 has four inputs and one output. The four inputs of the first OR gate 51 are respectively connected to the presence / absence interfaces of the four depth computing processors 20. A high-level signal output from each presence / absence interface indicates that the depth computing processor 20 is in an absent state, and a low-level signal output from each presence / absence interface indicates that the depth computing processor 20 is in a present state. The AND gate 53 has four inputs and one output. The two inputs of the AND gate 53 are respectively connected to the first general-purpose input / output interfaces of the four depth computing processors 20. A high-level signal output from each first general-purpose input / output interface indicates that the depth computing processor 20 is in a non-operating state, and a low-level signal output from each interface indicates that the depth computing processor 20 is in an operating state. The second OR gate circuit 52 has two input terminals and one output terminal. The two input terminals of the second OR gate circuit 52 are respectively connected to the output terminals of the first OR gate circuit 51 and the AND gate circuit 53. A high-level signal output by the second OR gate circuit 52 indicates that each depth computing processor 20 is in a single-path system, and a low-level signal output by the second OR gate circuit 52 indicates that each depth computing processor 20 is in a multiplexed system. The switch circuit 54 has one input terminal and one output terminal. The input terminal of the switch circuit 54 is connected to the output terminal of the second OR gate circuit 52, and the output terminal of the switch circuit 54 is connected to the second general-purpose input / output interface of all four depth computing processors 20. Through a simple combinational logic circuit, it is possible to determine whether the system currently occupied by each depth computing processor 20 is a single-path system or a multiplexed system, simplifying the configuration of the detection circuit 30. Additionally, refer to... Figure 5A diode can be provided between each of the first general-purpose input / output interfaces and the input terminal of the AND gate circuit 53, so that the first general-purpose input / output interface can only be input to the input terminal of the AND gate circuit 53 in one direction, and the input terminal of the AND gate circuit 53 cannot be input to the first general-purpose input / output interface in reverse, so that the first general-purpose input / output interface and the detection circuit 30 can be isolated to prevent mutual interference.

[0054] In Adoption Figure 7 When the detection circuit 30 shown performs detection, it first closes the switch circuit 54 to power on the depth computing processor system. Figure 5 In this system, not all four depth processing processors 20 are in the active state; at least one is in the inactive state. When at least one input of the first OR gate 51 is high, its output becomes high after logical operation. At this time, regardless of whether the output of the AND gate 53 is low or high, the output of the first OR gate 51 serves as an input to the second OR gate 52, causing the second OR gate 52 to output a high-level signal. This results in all second general-purpose input / output interfaces of the depth processing processor 20 receiving a high-level signal. When the firmware in each depth processing processor 20 detects a high-level signal at the second general-purpose input / output interface, it initializes each depth processing processor 20 according to a single-path system. Alternatively, by disconnecting the switch circuit 54, the second general-purpose input / output interface of each depth processing processor 20 can be made high, thus initializing each depth processing processor 20 according to a single-path system regardless of whether it is in the active or active state.

[0055] Of course, after the switch circuit 54 is closed and the depth computing processor system is powered on, when all four depth computing processors 20 are in the active state, both inputs of the first OR gate circuit 51 are low-level signals, making the output of the first OR gate circuit 51 also a low-level signal. At this time, the output of the second OR gate circuit 52 is equal to the output of the AND gate circuit 53. If all four depth computing processors 20 are in the working state after the depth computing processor system is powered on, both inputs of the AND gate circuit 53 are low-level signals, making the output of the AND gate circuit 53 also a low-level signal. Then the output of the second OR gate circuit 52 is also a low-level signal. The firmware in each depth computing processor 20 can then determine that each depth computing processor 20 is in a multi-way system (specifically a four-way system), and initialize each depth computing processor 20 according to the multi-way system.

[0056] When the detection circuit 30 detects that at least two of the multiple depth computing processors 20 are currently in a multi-processor system, the firmware of each depth computing processor 20 in the multi-processor system initializes the inter-chip interconnect control bus interface of each depth computing processor 20 in the multi-processor system according to the multi-processor system initialization. After the inter-chip interconnect control bus interfaces of all depth computing processors 20 in the multi-processor system are initialized, the firmware can also synchronize the running state of all depth computing processors 20 in the multi-processor system through the inter-chip global memory interconnect bus interface. By first fully initializing the inter-chip interconnect control bus interface of each depth computing processor 20 in the multi-processor system and then synchronizing the running state of all depth computing processors 20 in the multi-processor system, the phenomenon of unsuccessful handshake between multiple depth computing processors 20 in the multi-processor system can be better prevented.

[0057] Additionally, in the detection circuit 30, for example... Figure 3 or Figure 7 In the illustrated structure, the first and second general-purpose input / output interfaces can also be used to output signals indicating that the inter-chip interconnect control bus interface of each depth computing processor 20 has completed initialization. This allows each depth computing processor 20 in the multiplexing system to confirm whether its inter-chip interconnect control bus interface has completed initialization by recognizing the level signal of the second general-purpose input / output interface. Specifically, refer to... Figure 4 and Figure 8 First, the firmware in each depth computing processor 20 sets both the first and second general-purpose input / output interfaces to input mode to collect whether the signal received by the second general-purpose input / output interface is high or low. If it is determined that the second general-purpose input / output interface receives a high-level signal, each depth computing processor 20 is initialized directly according to the single-path system initialization. If it is determined that the second general-purpose input / output interface receives a low-level signal, the inter-chip interconnection control bus interface of each depth computing processor 20 is initialized according to the multi-path system initialization, and the system waits for the inter-chip interconnection control bus interface of each depth computing processor 20 to complete initialization and be in the ready state.

[0058] After the firmware in each depth computing processor 20 of the multiplexing system initializes the inter-chip interconnect control bus interface of the depth computing processor 20 according to the multiplexing system, the firmware can also control the first general-purpose input / output interface of the depth computing processor 20 to output a high-level signal and adjust the first general-purpose input / output interface to the output state, through, for example... Figure 3 and Figure 7The detection circuit 30 shown enables the second general-purpose input / output interface of the depth computing processor 20 to output a high-level signal. The firmware of each depth computing processor 20 in the multiplexing system is also used to determine that the inter-chip interconnect control bus interfaces of all depth computing processors 20 in the multiplexing system are ready when a high-level signal is detected from the second general-purpose input / output interfaces of all depth computing processors 20 in the multiplexing system. Then, the operating state of all depth computing processors 20 in the multiplexing system is synchronized through the inter-chip global memory interconnect bus interface. Specifically, the registers of the peer depth computing processor 20 interconnected with each depth computing processor 20 are read, and inter-chip global memory interconnect bus status queries and state machine transitions are performed to establish an inter-chip global memory interconnect link, synchronizing the operating state of all depth computing processors 20 in the multiplexing system. This avoids subsequent handshake failures caused by asynchronous initialization of the inter-chip global memory interconnect bus interfaces of all depth computing processors 20 in the multiplexing system. In this way, the firmware can identify that each depth computing processor 20 in the multi-path system has completed the initialization of the chip interconnection control bus interface of the depth computing processor 20, so as to better synchronize the running status of all depth computing processors 20 in the multi-path system and better prevent the phenomenon of unsuccessful handshake between multiple depth computing processors 20 in the subsequent multi-path system.

[0059] By setting a detection circuit 30 on the substrate, the system in which each depth computing processor 20 is currently located is detected as either a single-processor system or a multi-processor system. This allows each depth computing processor 20 to be initialized using a different initialization method depending on the system type it is in. Specifically, when the system in which the depth computing processor 20 is currently located is a multi-processor system, the depth computing processor 20 is initialized according to the multi-processor system's requirements, and the operating status of all depth computing processors 20 in the multi-processor system is synchronized to prevent handshake failures between multiple depth computing processors 20 in a multi-processor system.

[0060] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A deep computing processor system, characterized in that, include: substrate; A central processing unit or chipset is disposed on the substrate; Multiple depth computing processors are disposed on the substrate, each of the multiple depth computing processors is connected to the central processing unit or chipset, and the multiple depth computing processors are interconnected through an inter-chip global memory interconnect bus and an inter-chip interconnect control bus. A detection circuit is disposed on the substrate and is connected to each depth computing processor to detect whether the system currently in which each depth computing processor is located is a single-path system or a multi-path system. Each deep computing processor is embedded with firmware. The firmware is used to initialize the deep computing processor according to the single-path system when the system in which the deep computing processor is currently located is a single-path system. The firmware is also used to initialize the deep computing processor according to the multi-path system when the system in which the deep computing processor is currently located is a multi-path system, and to synchronize the running status of all deep computing processors in the multi-path system. Each depth computing processor also has an in-situ state interface and a first general-purpose input / output interface. The in-situ state interface is used to output a signal indicating whether the depth computing processor is in an in-situ state, and the first general-purpose input / output interface is used to output a signal indicating whether the depth computing processor is in a working state. The detection circuit is connected to the in-situ state interface and the first general-purpose input / output interface of each depth computing processor to collect the in-situ state and working state of each depth computing processor; the detection circuit is also used to determine that the system in which the at least two depth computing processors that are simultaneously in the in-situ state and working state are currently in a multi-channel system when at least two depth computing processors that are simultaneously in the in-situ state and working state are collected.

2. The deep computing processor system as described in claim 1, characterized in that, Each of the plurality of deep computing processors has an inter-chip global memory interconnect bus interface and an inter-chip interconnect control bus interface. Among the plurality of deep computing processors, the inter-chip global memory interconnect bus interface of any one deep computing processor is interconnected with the inter-chip global memory interconnect bus interface of each of the other deep computing processors; the inter-chip interconnect control bus interface of any one deep computing processor is interconnected with the inter-chip interconnect control bus interface of each of the other deep computing processors.

3. The deep computing processor system as described in claim 2, characterized in that, The number of depth computing processors is two; Each deep computing processor has a first inter-chip global memory interconnect bus interface and a second inter-chip global memory interconnect bus interface; the first inter-chip global memory interconnect bus interfaces of the two deep computing processors are interconnected, and the second inter-chip global memory interconnect bus interfaces of the two deep computing processors are interconnected. Each deep computing processor also has an inter-chip interconnect control bus interface, and the inter-chip interconnect control bus interfaces of the two deep computing processors are interconnected.

4. The deep computing processor system as described in claim 2 or 3, characterized in that, The detection circuit is also used to determine that each deep computing processor is currently in a single-path system when the number of deep computing processors that are simultaneously in the in-place and working states is no more than one. Each depth computing processor also has a second general-purpose input / output interface connected to the detection circuit, the second general-purpose input / output interface being used to receive a signal determined by the detection circuit as to whether the depth computing processor is in a single-path system or a multi-path system.

5. The deep computing processor system as described in claim 4, characterized in that, The number of depth computing processors is two; The detection circuit includes: A first OR gate circuit with two input terminals and one output terminal is provided. The two input terminals of the first OR gate circuit are respectively connected to the in-situ state interfaces of the two depth computing processors. A high-level signal output by each in-situ state interface indicates that the depth computing processor is in an out-of-situ state, and a low-level signal output by each in-situ state interface indicates that the depth computing processor is in an in-situ state. An AND gate circuit with two input terminals and one output terminal is provided. The two input terminals of the AND gate circuit are respectively connected to the first general-purpose input / output interface of the two depth computing processors. A high-level signal output by each first general-purpose input / output interface indicates that the depth computing processor is in a non-working state, and a low-level signal output by each state interface indicates that the depth computing processor is in a working state. A second OR gate circuit has two input terminals and one output terminal. The two input terminals of the second OR gate circuit are respectively connected to the output terminal of the first OR gate circuit and the output terminal of the AND gate circuit. The high-level signal output by the second OR gate circuit indicates that each depth computing processor is in a single-path system, and the low-level signal output by the second OR gate circuit indicates that each depth computing processor is in a multi-path system. A switching circuit having an input terminal and an output terminal, wherein the input terminal of the switching circuit is connected to the output terminal of the second OR gate circuit, and the output terminal of the switching circuit is connected to the second general-purpose input / output interface of the two depth computing processors.

6. The deep computing processor system as described in claim 1, characterized in that, The number of depth computing processors is four, and the four depth computing processors are divided into two pairs of depth computing processors; Each deep computing processor has a first inter-chip global memory interconnect bus interface and a second inter-chip global memory interconnect bus interface; the two first inter-chip global memory interconnect bus interfaces in each pair of deep computing processors are interconnected; the second inter-chip global memory interconnect bus interfaces of the two deep computing processors in one pair of deep computing processors are interconnected with the second inter-chip global memory interconnect bus interfaces of the two deep computing processors in another pair of deep computing processors respectively. Each deep computing processor also has a first inter-chip interconnect control bus interface and a second inter-chip interconnect control bus interface; the two first inter-chip interconnect control bus interfaces in each pair of deep computing processors are interconnected; the second inter-chip interconnect control bus interfaces of the two deep computing processors in one pair of deep computing processors are interconnected with the second inter-chip interconnect control bus interfaces of the two deep computing processors in another pair of deep computing processors, respectively.

7. The deep computing processor system as described in claim 6, characterized in that, Each depth computing processor also has an in-situ state interface and a first general-purpose input / output interface, wherein the in-situ state interface is used to output a signal indicating whether the depth computing processor is in an in-situ state, and the first general-purpose input / output interface is used to output a signal indicating whether the depth computing processor is in a working state. The detection circuit is connected to both the presence status interface and the first general-purpose input / output interface of each depth computing processor to collect the presence status and operating status of each depth computing processor. The detection circuit is also used to determine that the system currently in which each of the four depth computing processors is located is a multi-processor system when all four depth computing processors are simultaneously in both presence and operating states. Furthermore, the detection circuit is used to determine that the system currently in which each of the four depth computing processors is located is a single-processor system when at least one depth computing processor is not simultaneously in both presence and operating states. Each depth computing processor also has a second general-purpose input / output interface connected to the detection circuit, the second general-purpose input / output interface being used to receive a signal determined by the detection circuit as to whether the depth computing processor is in a single-path system or a multi-path system.

8. The deep computing processor system as described in claim 7, characterized in that, The detection circuit includes: A first OR gate circuit with four input terminals and one output terminal is provided. The four input terminals of the first OR gate circuit are respectively connected to the in-situ state interfaces of the four depth computing processors. A high-level signal output by each in-situ state interface indicates that the depth computing processor is in an out-of-situ state, and a low-level signal output by each in-situ state interface indicates that the depth computing processor is in an in-situ state. An AND gate circuit with four input terminals and one output terminal is provided. The two input terminals of the AND gate circuit are respectively connected to the first general-purpose input / output interface of the four depth computing processors. A high-level signal output by each first general-purpose input / output interface indicates that the depth computing processor is in a non-working state, and a low-level signal output by each state interface indicates that the depth computing processor is in a working state. A second OR gate circuit has two input terminals and one output terminal. The two input terminals of the second OR gate circuit are respectively connected to the output terminal of the first OR gate circuit and the output terminal of the AND gate circuit. The high-level signal output by the second OR gate circuit indicates that each depth computing processor is in a single-path system, and the low-level signal output by the second OR gate circuit indicates that each depth computing processor is in a multi-path system. A switching circuit having an input terminal and an output terminal, wherein the input terminal of the switching circuit is connected to the output terminal of the second OR gate circuit, and the output terminal of the switching circuit is connected to the second general-purpose input / output interface of the four depth computing processors.

9. The deep computing processor system as described in claim 5 or 8, characterized in that, The firmware is also used to initialize the chip interconnect control bus interface of the deep computing processor in accordance with the multi-path system when the system in which the deep computing processor is currently located is a multi-path system. The firmware is also used to synchronize the operating status of all deep computing processors in the multi-path system after the inter-chip interconnect control bus interface of all deep computing processors in the multi-path system has been initialized, through the inter-chip global memory interconnect bus interface and the inter-chip global memory interconnect bus.

10. The deep computing processor system as described in claim 9, characterized in that, The first general-purpose input / output interface and the second general-purpose input / output interface are also used to output a signal indicating that the inter-chip interconnect control bus interface of each depth computing processor has completed initialization; The firmware is also used to control the first general-purpose input / output interface of the deep computing processor to output a high-level signal after the chip interconnect control bus interface of the deep computing processor is initialized according to the multiplexing system, so that the second general-purpose input / output interface of the deep computing processor outputs a high-level signal. The firmware is also used to synchronize the operating status of all deep computing processors in the multi-path system through the inter-chip global memory interconnect bus interface and the inter-chip global memory interconnect bus when a high-level signal is detected from the second general-purpose input / output interface of all deep computing processors in the multi-path system.

11. The deep computing processor system as described in claim 1, characterized in that, The detection circuit is a combinational logic circuit or a programmable logic device.

12. The deep computing processor system as described in claim 1, characterized in that, The PCIe interface of the central processing unit or chipset is connected to the PCIe interface of each deep computing processor. The substrate is also provided with a clock buffer, which is connected to the PCIe reference clock interface of the central processing unit or chipset, and is also connected to the PCIe reference clock interface of each depth computing processor.

13. The deep computing processor system as described in claim 12, characterized in that, The clock buffer is also connected to the inter-chip memory interconnect bus reference clock interface of each depth computing processor.

14. The deep computing processor system as described in claim 1, characterized in that, Also includes: An oscillator or resonator disposed on the substrate; A clock generator is disposed on the substrate and connected to the oscillator or resonator, and the clock generator is also connected to the inter-chip memory interconnect bus reference clock interface of each depth computing processor.

Citation Information

Patent Citations

  • Processor system and method for operating computer processor

    CN103377169A

  • Device and method of detecting running state of processor

    CN112416665A

  • Multi-chip interconnection system and neural network acceleration processing method

    CN113902111A