Data Transmission Processing Method and Related Devices in a Chip System

By configuring the stream output table and stream input table in advance in the chip system, the problem of low data transmission efficiency between sub-chips is solved, and efficient data transmission and improved processing performance is achieved.

CN114328623BActive Publication Date: 2025-05-30SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111633371.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-05-30
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In chip systems, the data transmission efficiency between sub-chips is low, which affects the overall processing performance.

Method used

By configuring the stream output table and stream input table in the source subchip and destination subchip in advance, the subchip can quickly generate and receive data packets, achieving efficient data transmission.

Benefits of technology

It improves the data transmission efficiency between sub-chips and improves the processing performance of the chip system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328623B_ABST
    Figure CN114328623B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a data transmission processing method and related device in a chip system. The method includes: a source sub-chip receives configuration parameters and configures a preset flow output table according to the configuration parameters; the flow output table includes an identifier of a data stream and an identifier of a destination sub-chip of the data stream; the source sub-chip and the destination sub-chip are sub-chips among a plurality of sub-chips included in the chip system, and the plurality of sub-chips are connected in a preset topological structure; the source sub-chip performs a first data processing task to obtain output data; the source sub-chip generates a plurality of data packets based on the flow output table and sends the data packets, and the data packets include the identifier of the data stream and the identifier of the destination sub-chip. The present application can achieve efficient data transmission between sub-chips in the chip system and improve the processing performance of the chip system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a data transmission processing method and related devices in a chip system. Background Art

[0002] A chip system may include multiple sub-chips, each sub-chip having the function of independently processing data. The multiple sub-chips are connected in a certain topology to achieve communication with each other. Moreover, the multiple sub-chips can cooperate to process a single large computing task in a model parallel manner to improve the processing efficiency of the task. During the process of cooperating to process the task, data needs to be frequently exchanged and transmitted between the multiple sub-chips, and the efficiency of this data transmission affects the processing performance of the entire chip system. Summary of the Invention

[0003] Embodiments of the present application disclose a data transmission processing method and related devices in a chip system, which can achieve efficient data transmission between sub-chips in the chip system and improve the processing performance of the chip system.

[0004] In a first aspect, the present application provides a data transmission processing method in a chip system, the method including:

[0005] A source sub-chip receives configuration parameters and configures a preset flow output table according to the foregoing configuration parameters; the foregoing flow output table includes an identifier of a data stream and an identifier of a destination sub-chip of the foregoing data stream; the foregoing source sub-chip and the foregoing destination sub-chip are sub-chips among the multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology;

[0006] The foregoing source sub-chip executes a first data processing task to obtain output data;

[0007] The foregoing source sub-chip generates multiple data packets based on the foregoing flow output table and sends the data packets, and the data packets include the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip.

[0008] In the present application, by pre-configuring the flow output table in the source sub-chip, the source sub-chip can quickly generate and send data packets by querying the configured flow output table, improving the data sending efficiency, achieving efficient data transmission between sub-chips in the chip system, and improving the processing performance of the chip system.

[0009] In a possible implementation manner, the foregoing flow output table further includes one or more of an identifier of the foregoing first data processing task, indication information of the number of data packets of the foregoing data stream that have been sent, and a starting address in the foregoing source sub-chip for storing data included in the foregoing data stream; the foregoing data packet further includes the identifier of the foregoing first data processing task.

[0010] In the present application, the flow output table includes information that can be used to distinguish different tasks by the identifiers of the tasks, and enables the generated data packets to also include the corresponding task identifiers to indicate the tasks to which the data in the data packets belongs. The indication information of the number of the above-mentioned data packets included in the flow output table and the above-mentioned starting address can quickly calculate the storage address of the data to be sent, so that the data can be quickly obtained, packed and sent.

[0011] In a possible implementation manner, before the foregoing source sub-chip generates and sends a plurality of data packets based on the foregoing flow output table, it further includes: the foregoing source sub-chip receives an unlock data packet; wherein, the foregoing unlock data packet includes the identifier of the foregoing data stream, and the foregoing unlock data packet is used to indicate to the foregoing source sub-chip that the foregoing destination sub-chip is ready to receive the foregoing data stream.

[0012] In the present application, after the destination sub-chip is ready to receive data, it sends an unlock data packet to the source sub-chip, which can avoid data loss and other situations caused by the destination sub-chip receiving data without being ready.

[0013] In a possible implementation manner, the foregoing source sub-chip sends the foregoing plurality of data packets, including: the foregoing source sub-chip sends the foregoing plurality of data packets based on a port forwarding mapping table; wherein, the foregoing port forwarding mapping table includes the mapping relationship between the identifier of the foregoing destination sub-chip and the sending port.

[0014] In the present application, the sending port of the data packet is determined by querying the port forwarding mapping table to achieve fast transmission of the data packet.

[0015] In a possible implementation manner, when there are multiple destination sub-chips for the foregoing data stream, the foregoing data packet includes the identifiers of the foregoing multiple destination sub-chips.

[0016] In the present application, the data packet can carry the identifiers of multiple destination sub-chips. Compared with the existing situation where a data packet is sent to each destination, the number of data packets to be sent can be reduced, saving the transmission bandwidth.

[0017] In a possible implementation manner, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, and the foregoing subsystem is configured with a subsystem identifier; when the foregoing destination sub-chip is a sub-chip of the foregoing subsystem, the foregoing data packet further includes the identifier of the foregoing subsystem.

[0018] In the present application, the destination of the data packet can be quickly located through the identifier of the chip subsystem, improving the efficiency of data transmission.

[0019] In a second aspect, the present application provides a data transmission processing method in a chip system, and the method includes:

[0020] The destination sub-chip receives configuration parameters and configures a preset flow input table according to the foregoing configuration parameters. The foregoing flow input table includes at least one data stream identifier to be received. The foregoing destination sub-chip is a sub-chip among multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topological structure.

[0021] When the foregoing destination sub-chip receives a data packet, it determines whether the data stream identifier in the foregoing data packet is included in the foregoing flow input table.

[0022] When the foregoing data stream identifier in the foregoing flow input table is included in the foregoing data packet, the data in the foregoing data packet is stored.

[0023] In this application, by pre-configuring the flow input table in the destination sub-chip, the destination sub-chip can quickly receive and process data packets by querying the configured flow input table, improving the data reception efficiency, realizing efficient data transmission between sub-chips in the chip system, and improving the processing performance of the chip system.

[0024] In a possible implementation manner, the foregoing flow input table further includes one or more of the identifier of the foregoing first data processing task, indication information of the number of data packets of the foregoing data stream that have been received, and the starting address in the foregoing destination sub-chip for storing the data included in the foregoing data stream. The foregoing data packet further includes the identifier of the foregoing first data processing task.

[0025] In this application, the identifier of the task included in the flow input table can be used to distinguish data packets of different tasks. The indication information of the number of the foregoing data packets and the foregoing starting address included in the flow output table can quickly calculate the storage address of the data in the received data packet, so that data storage can be quickly realized.

[0026] In a possible implementation manner, before the foregoing destination sub-chip receives the foregoing data packet, it further includes: the foregoing destination sub-chip sends an unlock data packet. Wherein, the foregoing unlock data packet includes the identifier of the foregoing data stream and the identifier of the source sub-chip, and the foregoing unlock data packet is used to indicate to the foregoing source sub-chip that the foregoing destination sub-chip is ready to receive the foregoing data stream. The foregoing source sub-chip is the sub-chip that sends the foregoing data packet in the foregoing chip system.

[0027] In this application, after the destination sub-chip is ready to receive data, it sends an unlock data packet to the source sub-chip, which can avoid data loss and other situations caused by the destination sub-chip receiving data without being ready.

[0028] In a possible implementation manner, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, and the foregoing subsystem is configured with a subsystem identifier. The foregoing data packet further includes the identifier of the foregoing subsystem.

[0029] In this application, the destination of the data packet can be quickly located through the identification of the chip subsystem, improving the efficiency of data transmission.

[0030] In a third aspect, this application provides a data transmission processing method in a chip system, and the method includes:

[0031] The controller assigns a first data processing task to the source sub-chip, and the data obtained after the execution of the foregoing first data processing task is sent to the destination sub-chip in the form of a data stream; the foregoing controller, the foregoing source sub-chip, and the foregoing destination sub-chip are sub-chips among multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology structure;

[0032] The foregoing controller configures an identifier for the foregoing data stream;

[0033] The foregoing controller sends the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip to the foregoing source sub-chip; wherein, the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip are used for associated storage in the flow output table of the foregoing source sub-chip; the foregoing flow output table is the basis for the foregoing source sub-chip to send data.

[0034] In this application, the controller assigns tasks to the source sub-chip and configures the flow output table of the source sub-chip, so that the source sub-chip can quickly generate and send data packets by querying the configured flow output table, improving the data sending efficiency, realizing efficient data transmission between sub-chips in the chip system, and improving the processing performance of the chip system.

[0035] In a possible implementation manner, the foregoing method further includes:

[0036] The foregoing controller assigns a second data processing task to the foregoing destination sub-chip, and the foregoing second data processing task is executed based on the data obtained after the execution of the foregoing first data processing task;

[0037] The foregoing controller sends the identifier of the foregoing data stream to the foregoing destination sub-chip; wherein, the identifier of the foregoing data stream is used for storage in the flow input table of the foregoing destination sub-chip, and the foregoing flow input table is the basis for the foregoing destination sub-chip to receive data.

[0038] In this application, the controller assigns tasks to the destination sub-chip and configures the flow input table of the destination sub-chip, so that the destination sub-chip can quickly receive and process data packets by querying the configured flow input table, improving the data receiving efficiency, realizing efficient data transmission between sub-chips in the chip system, and improving the processing performance of the chip system.

[0039] In a possible implementation manner, the foregoing method further includes:

[0040] The foregoing controller obtains the data transmission situation between sub-chips in the foregoing chip system;

[0041] The foregoing controller generates scheduling information for the foregoing source sub-chip based on the foregoing data transmission situation, and the foregoing scheduling information indicates a transmission port for sending the foregoing data stream to the foregoing destination sub-chip in the foregoing source sub-chip;

[0042] The foregoing controller sends the foregoing scheduling information to the foregoing source sub-chip.

[0043] In this application, the controller realizes the scheduling of the data packet sending port based on the data transmission situation of the entire chip system, thereby avoiding the transmission of data packets from congested paths and improving the efficiency of data packet transmission.

[0044] In a fourth aspect, this application provides a source sub-chip, which includes:

[0045] A receiving unit, configured to receive configuration parameters;

[0046] A configuration unit, configured to configure a preset flow output table according to the foregoing configuration parameters; the foregoing flow output table includes an identifier of a data stream and an identifier of the foregoing destination sub-chip of the foregoing data stream; the foregoing source sub-chip and the foregoing destination sub-chip are sub-chips among a plurality of sub-chips included in the chip system, and the foregoing plurality of sub-chips are connected in a preset topological structure;

[0047] An execution unit, configured to execute a first data processing task to obtain output data;

[0048] A generating unit, configured to generate a plurality of data packets based on the foregoing flow output table from the foregoing output data, where the foregoing data packets include an identifier of the foregoing data stream and an identifier of the foregoing destination sub-chip;

[0049] A sending unit, configured to send the foregoing plurality of data packets.

[0050] In a possible implementation manner, the foregoing flow output table further includes one or more of an identifier of the foregoing first data processing task, indication information of the number of data packets of the foregoing data stream that have been sent, and a starting address in the foregoing source sub-chip for storing data included in the foregoing data stream;

[0051] The foregoing data packet further includes an identifier of the foregoing first data processing task.

[0052] In a possible implementation manner, the foregoing receiving unit is further configured to, before the foregoing generating unit generates a plurality of data packets based on the foregoing flow output table from the foregoing output data,

[0053] receive an unlock data packet; wherein, the foregoing unlock data packet includes an identifier of the foregoing data stream, and the foregoing unlock data packet is used to indicate to the foregoing source sub-chip that the foregoing destination sub-chip is ready to receive the foregoing data stream.

[0054] In a possible implementation manner, the foregoing sending unit is specifically configured to:

[0055] Send the foregoing multiple data packets based on a port forwarding mapping table; wherein, the foregoing port forwarding mapping table includes a mapping relationship between the identifier of the foregoing destination sub-chip and the sending port.

[0056] In a possible implementation manner, when there are multiple destination sub-chips for the foregoing data stream, the foregoing data packet includes the identifiers of the foregoing multiple destination sub-chips.

[0057] In a possible implementation manner, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, and the foregoing subsystem is configured with a subsystem identifier;

[0058] When the foregoing destination sub-chip is a sub-chip of the foregoing subsystem, the foregoing data packet further includes the identifier of the foregoing subsystem.

[0059] In a fifth aspect, the present application provides a destination sub-chip, and the destination sub-chip includes:

[0060] A receiving unit, configured to receive configuration parameters;

[0061] A configuration unit, configured to configure a preset flow input table according to the foregoing configuration parameters, the foregoing flow input table includes at least one data stream identifier to be received; the foregoing destination sub-chip is a sub-chip among multiple sub-chips included in a chip system, and the foregoing multiple sub-chips are connected in a preset topological structure;

[0062] A judgment unit, configured to judge whether the foregoing flow input table includes the data stream identifier in the foregoing data packet when receiving the data packet;

[0063] A storage unit, configured to store the data in the foregoing data packet when the foregoing flow input table includes the data stream identifier in the foregoing data packet.

[0064] In a possible implementation manner, the foregoing flow input table further includes one or more of the identifier of the foregoing first data processing task, indication information of the number of data packets of the foregoing received data stream, and the starting address in the foregoing destination sub-chip for storing the data included in the foregoing data stream;

[0065] The foregoing data packet further includes the identifier of the foregoing first data processing task.

[0066] In a possible implementation manner, the foregoing destination sub-chip further includes a sending unit, configured to before the foregoing receiving unit receives the foregoing data packet,

[0067] Send an unlock data packet; wherein, the unlock data packet includes the identifier of the foregoing data stream and the identifier of the source sub-chip, and the unlock data packet is used to indicate to the foregoing source sub-chip that the foregoing destination sub-chip is ready to receive the foregoing data stream; the foregoing source sub-chip is the sub-chip that sends the foregoing data packet in the foregoing chip system.

[0068] In a possible implementation, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, and the foregoing subsystem is configured with a subsystem identifier;

[0069] The foregoing data packet further includes the identifier of the foregoing subsystem.

[0070] In a sixth aspect, the present application provides a controller, which includes:

[0071] An allocation unit, configured to allocate a first data processing task to a source sub-chip, and the data obtained after the execution of the foregoing first data processing task is sent to a destination sub-chip in the form of a data stream; the foregoing controller, the foregoing source sub-chip, and the foregoing destination sub-chip are sub-chips among a plurality of sub-chips included in the chip system, and the foregoing plurality of sub-chips are connected in a preset topology;

[0072] A configuration unit, configured to configure an identifier for the foregoing data stream;

[0073] A sending unit, configured to send the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip to the foregoing source sub-chip; wherein, the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip are used to be associated and stored in the flow output table of the foregoing source sub-chip; the foregoing flow output table is the basis for the foregoing source sub-chip to send data.

[0074] In a possible implementation, the foregoing allocation unit is further configured to allocate a second data processing task to the foregoing destination sub-chip, and the foregoing second data processing task is executed based on the data obtained after the execution of the foregoing first data processing task;

[0075] The foregoing sending unit is further configured to send the identifier of the foregoing data stream to the foregoing destination sub-chip; wherein, the identifier of the foregoing data stream is used to be stored in the flow input table of the foregoing destination sub-chip, and the foregoing flow input table is the basis for the foregoing destination sub-chip to receive data.

[0076] In a possible implementation, the foregoing controller further includes:

[0077] An acquisition unit, configured to acquire the data transmission situation between sub-chips in the foregoing chip system;

[0078] A generation unit, configured to generate scheduling information for the foregoing source sub-chip based on the foregoing data transmission situation, and the foregoing scheduling information indicates the sending port in the foregoing source sub-chip for sending the foregoing data stream to the foregoing destination sub-chip;

[0079] The foregoing sending unit is further configured to send the foregoing scheduling information to the foregoing source sub-chip.

[0080] In a seventh aspect, the present application provides a sub-chip, which includes a processor, a memory, and a communication port; wherein, the foregoing memory and communication port are coupled to the foregoing processor, the foregoing communication port is used for receiving and sending data, the foregoing memory is used for storing a computer program, and the foregoing processor is used for calling the foregoing computer program so that the foregoing sub-chip executes the method according to any one of the first aspect;

[0081] The foregoing sub-chip is a sub-chip among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure.

[0082] In an eighth aspect, the present application provides a sub-chip, which includes a processor, a memory, and a communication port; wherein, the foregoing memory and communication port are coupled to the foregoing processor, the foregoing communication port is used for receiving and sending data, the foregoing memory is used for storing a computer program, and the foregoing processor is used for calling the foregoing computer program so that the foregoing sub-chip executes the method according to any one of the second aspect;

[0083] The foregoing sub-chip is a sub-chip among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure.

[0084] In a ninth aspect, the present application provides a sub-chip, which includes a processor, a memory, and a communication port; wherein, the foregoing memory and communication port are coupled to the foregoing processor, the foregoing communication port is used for receiving and sending data, the foregoing memory is used for storing a computer program, and the foregoing processor is used for calling the foregoing computer program so that the foregoing sub-chip executes the method according to any one of the third aspect;

[0085] The foregoing sub-chip is a sub-chip among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure.

[0086] In a tenth aspect, the present application provides a chip system, which includes a source sub-chip, a destination sub-chip, and a controller; wherein, the foregoing source sub-chip is the source sub-chip according to any one of the fourth aspect above, the foregoing destination sub-chip is the destination sub-chip according to any one of the fifth aspect above, and the foregoing controller is the controller according to any one of the sixth aspect above; or,

[0087] The foregoing source sub-chip is the sub-chip according to the seventh aspect above, the foregoing destination sub-chip is the sub-chip according to the eighth aspect above, and the foregoing controller is the sub-chip according to the ninth aspect above.

[0088] In an eleventh aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method according to any one of the first aspect; or,

[0089] which, when executed by a processor, implements the method according to any one of the second aspect; or,

[0090] which, when executed by a processor, implements the method according to any one of the third aspect.

[0091] In a twelfth aspect, the present application provides a computer program product including a computer program, which, when executed by a processor, implements the method according to any one of the first aspect; or,

[0092] which, when executed by a processor, implements the method according to any one of the second aspect; or,

[0093] which, when executed by a processor, implements the method according to any one of the third aspect.

[0094] It can be understood that the above fourth aspect to twelfth aspect are all applicable to execute the method provided by any one of the first aspect to third aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] The following will introduce the drawings required to be used in the embodiments of the present application.

[0096] Figure 1 is a schematic diagram of a chip system provided by the present application;

[0097] Figure 2 is a schematic structural diagram of a sub-chip provided by the present application;

[0098] Figures 3 to 6 is a schematic diagram of a chip system provided by the present application;

[0099] Figure 7 is a schematic diagram of sub-chip group division provided by the present application;

[0100] Figure 8 is a schematic flowchart of a data transmission and processing method in a chip system provided by the present application;

[0101] Figure 9A and Figure 9B is a schematic diagram of a data processing flow provided by the present application;

[0102] Figure 10 is a schematic diagram of a data packet structure provided by the present application;

[0103] Figure 11 Schematic diagram of the ports of the sub-chips in the chip system provided by this application;

[0104] Figure 12 Schematic diagram of the structure of the data packet provided by this application;

[0105] Figure 13 Schematic diagram of the process of the routing implementation solution provided by this application;

[0106] Figure 14 Schematic diagram of the direction coordinate system provided by this application;

[0107] Figure 15 Schematic diagram of constructing the direction coordinate system based on the sub-chips provided by this application;

[0108] Figure 16 Schematic diagram of the ports of the sub-chips in the chip system provided by this application;

[0109] Figures 17 to 19 Schematic diagram of the structure of the virtual device provided by this application;

[0110] Figure 20 Schematic diagram of the structure of the physical device provided by this application. Detailed implementation manners

[0111] The embodiments of this application will be described below with reference to the accompanying drawings.

[0112] Figure 1 Shown is a schematic diagram of the structure of a chip system provided by an embodiment of this application. The chip system 110 includes multiple sub-chips ( Figure 1 16 sub-chips are exemplarily shown in it), and the multiple sub-chips are connected according to a preset topological connection relationship. For example, Figure 1 the 16 sub-chips in it can be arranged in a matrix form, and then, each individual sub-chip is respectively connected to two, three or four surrounding sub-chips.

[0113] Each sub-chip has its own memory, Figure 1 and the memories of some sub-chips are exemplarily drawn in it. The memory can be, for example, synchronous dynamic random access memory (SDRAM) or Double Data Rate SDRAM (DDRSDRAM), and DDRSDRAM can be abbreviated as DDR. Each sub-chip in the chip system 100 has complete processing capabilities and can execute tasks independently. Of course, the multiple sub-chips in the chip system 100 can cooperate with each other to execute large-scale processing tasks.

[0114] See Figure 2 , Figure 2 An exemplary structural schematic diagram of the sub-chip in the above chip system 110 is shown. The structure of the sub-chip can be presented in the form of a network-on-chip (NoC). It can be seen that the sub-chip may include a processing module, a routing module, a static memory, a memory controller, and four ports (d0, d1, d2, and d3).

[0115] The above processing module is the control unit (CU) in the sub-chip, responsible for the management of each processing flow in the sub-chip.

[0116] The above routing module is responsible for data synchronization within the sub-chip, data synchronization between sub-chips, data broadcasting, and data transmission. Among them, the routing module also includes a control unit, which is used to manage the routing process in the routing module. The routing module also includes a local buffer, which can be used to temporarily store data to be processed. The routing module also includes a forwarding-port mapper (FPM). The FPM can be a hardware module or a software module. A port forwarding mapping table is stored in the FPM. The port forwarding mapping table includes the mapping relationship between the destination sub-chip and the sending port, and can be used to map data packets to the corresponding ports for sending. A stream in table (SIT) and a stream out table (SOT) are stored in the routing module. The SIT and SOT are used for data transmission between sub-chips, which will be introduced in detail later and will not be elaborated here for the time being.

[0117] The above static memory can be a static random-access memory (SRAM), etc., for storing data in the sub-chip.

[0118] The above memory controller is connected to the memory corresponding to the sub-chip. The memory controller can be, for example, a DDR controller, etc.

[0119] The above four ports (d0, d1, d2, and d3) are the network interfaces of the sub-chip, which can realize data transmission between the above sub-chips. The connection between the above sub-chips is realized through these four ports. Optionally, the sub-chip can also include two such ports, such as Figure 1 sub-chip 0, sub-chip 3, sub-chip 12, and sub-chip 15 in Figure 1 Or, optionally, the sub-chip can also include three such ports, such as

[0120] The chip system provided by the embodiments of the present application is not limited to the aboveFigure 1 The structure shown may also be other structures. For example, refer to Figure 3 . Figure 3 Exemplarily shown is a chip system 120, which also includes a plurality of sub-chips ( Figure 2 8 sub-chips are exemplarily shown in it). The plurality of sub-chips of the chip system 120 are arranged and connected in the form of a cuboid. This connection method can minimize the data transmission path between the sub-chips. The structure of the sub-chips in the chip system 120 can refer to the corresponding description in the above Figure 2 , and will not be elaborated here.

[0121] In a possible implementation, for a super-large task, more sub-chips are required to process together to improve the task processing efficiency. Then, the above chip system 110 or chip system 120 can be used as a sub-system of a chip, and a larger chip system is composed of a plurality of such sub-systems. Exemplarily, refer to Figure 4 the shown chip system 130.

[0122] The chip system 130 may include a plurality of sub-systems of the above chips, Figure 4 taking 8 sub-systems as an example in it. Each sub-system of the plurality of sub-systems may be the above chip system 110 or chip system 120. Each sub-system can be regarded as a whole. Then, the plurality of sub-systems can be connected through a preset topological connection relationship. For example, they can be arranged and connected in the form of a cuboid, as Figure 4 shown. To facilitate understanding of the connection method of each sub-system in the chip system 130, exemplarily, taking the sub-system as the above chip system 110 as an example, the connection structure schematic diagram of each sub-system in the chip system 130 can be referred to Figure 5 .

[0123] In Figure 5 , it can be seen that the chip system 130 includes 8 sub-systems. Each sub-system of the 8 sub-systems includes 16 sub-chips, and the 16 sub-chips can be arranged and connected in the form of a matrix. Between two adjacent sub-systems, the connection between the two sub-systems can be achieved by connecting any one sub-chip in one sub-system to any one sub-chip in the other sub-system. Figure 5 Exemplarily, the sub-chips arranged at the corners of the matrix in each sub-system are used as the sub-chips connected to another sub-system. For example, in sub-system 0, the connection between sub-system 0 and sub-system 1 is established through sub-chip 3 in sub-system 0 and sub-chip 0 in sub-system 1. The connection between sub-system 0 and sub-system 2 is established through sub-chip 12 in sub-system 0 and sub-chip 0 in sub-system 2. The connection between sub-system 0 and sub-system 4 is established through sub-chip 0 in sub-system 0 and sub-chip 0 in sub-system 4.

[0124] In a possible implementation, a control bus may also be included between the above-mentioned subsystems. For example, reference may be made to Figure 6 the chip system 140 shown. The chip system 140 includes a central controller (the central controller may be a sub-chip or a control module in the chip system 140, etc.). The central controller can manage the task processing flow in the chip system 140. The control bus is connected to the central controller for each subsystem to receive control instructions from the central controller. In a specific implementation, each subsystem may be connected to the control line by a sub-chip. After receiving the control instruction, the sub-chip can forward it on behalf of the corresponding sub-chip within the same subsystem. Alternatively, each sub-chip in each subsystem is connected to the control line to directly receive control instructions. This application does not limit the specific connection of the controller.

[0125] In addition to the above chip system 140, the chip systems provided in the embodiments of this application (such as the above chip systems 110, 120, and 130, etc.) also include a central controller for managing the task processing flow of the entire chip system. Similarly, the central controller may be a sub-chip or a control module in the chip system, etc. The central controller can obtain the load conditions and resource usage conditions of all sub-chips in the chip system, etc., so as to allocate tasks to each sub-chip based on these conditions. Exemplarily, a task scheduler in the central controller can allocate tasks to these sub-chips based on information such as the load conditions and resource usage conditions of each sub-chip.

[0126] In the chip system provided in the embodiments of this application, each sub-chip has the ability to independently process data. However, in the case of large data tasks, the processing efficiency of a single sub-chip is relatively low. To improve the task processing efficiency, multiple sub-chips included in the above chip system can be divided into multiple sub-chip groups, and each sub-chip group includes at least one sub-chip. In this way, data tasks can be processed with the sub-chip group as the processing unit, thereby improving the processing efficiency. To facilitate understanding of the chip group, reference may be made to Figure 7 .

[0127] Figure 7 Taking the above chip system 110 as an example, the 16 sub-chips in the chip system are divided into 9 sub-chip groups. For the specific division situation, reference may be made to Figure 7 the division situation shown. Each sub-chip group includes at least one sub-chip. In addition, multiple sub-chips in the same sub-chip group may be adjacent sub-chips, such as sub-chip group 3, sub-chip group 4, sub-chip group 7, and sub-chip group 8. Alternatively, multiple sub-chips in the same sub-chip group may be non-adjacent sub-chips, such as sub-chip group 2, which is composed of non-adjacent sub-chips 1 and 12.

[0128] The chip system provided by the embodiments of the present application can implement data task processing in a data parallel, model parallel, or model parallel plus data parallel manner. Among them:

[0129] Data parallelism means dividing the data to be processed into several data blocks, and distributing the several data blocks to different sub-chip groups respectively. Each chip group runs the same processing program to process the assigned data. For example, assuming that the data to be processed is divided into 3 data blocks, and there are 3 sub-chip groups that can run the same processing program to process the data blocks, then, the first data block among the 3 data blocks can be sent to the first sub-chip group among the 3 sub-chip groups for processing, the second data block among the 3 data blocks can be sent to the second sub-chip group among the 3 sub-chip groups for processing, and the third data block among the 3 data blocks can be sent to the third sub-chip group among the 3 sub-chip groups for processing.

[0130] Model parallelism means that multiple sub-chip groups jointly complete a data processing task, and each sub-chip group in the multiple sub-chip groups only executes part of the steps of the entire data processing task (the part of the steps can be one or more processing steps). For example, assuming that a data processing task needs to go through 3 steps to complete the processing, then, two sub-chip groups can be configured to jointly complete the task. Among them, the first sub-chip group completes the processing of the first 2 steps among the 3 steps, and the second sub-chip group obtains the processed data from the first sub-chip group to complete the processing of the third step. Or, three sub-chip groups can be configured to jointly complete the task. Among them, the first sub-chip group completes the processing of the first step among the 3 steps, the second sub-chip group obtains the processed data from the first sub-chip group to complete the processing of the second step, and the third sub-chip group obtains the processed data from the second sub-chip group to complete the processing of the third step. That is, the steps completed by each chip group can be one or more, and can be specifically determined according to the load situation and resource usage situation of the chip group.

[0131] The model parallel plus data parallel approach combines the above-mentioned data parallel and model parallel methods to process data. For example, if a data processing task requires three steps to complete, then three sub-chip groups can be configured to jointly complete this task. Among them, the first sub-chip group completes the processing of the first step among these three steps, the second sub-chip group obtains the processed data from the first sub-chip group to complete the processing of the second step, and the third sub-chip group obtains the processed data from the second sub-chip group to complete the processing of the third step. However, since the processing of the first step is relatively complex and takes more time to complete, in order to improve the processing efficiency, one or more additional sub-chip groups can be configured to jointly execute the processing task of the first step. For example, a fourth sub-chip group can be configured to execute the processing task of the first step together with the aforementioned first sub-chip group. Specifically, the data used for the first step of processing can be divided into two parts, one part is sent to the first sub-chip group for processing, and the other part is sent to the fourth sub-chip group for processing. Then, the data processed by the first sub-chip group and the fourth sub-chip group is sent together to the second sub-chip group for the second step of processing.

[0132] It should be noted that in the above model parallel plus data parallel approach, each processing step can adopt the data parallel processing method for processing, or some processing steps can adopt the data parallel processing method for processing. Specifically, it can be determined according to the specific implementation, and the present application does not limit this.

[0133] In specific implementation, the data processing task can be assigned to each sub-chip group through the central controller of the chip system. To implement the processing of data tasks using the model parallel or model parallel plus data parallel approach, data transmission is required between sub-chips. Data transmission will cause time delay, resulting in reduced processing efficiency. To achieve efficient data transmission between sub-chips in the chip system and improve the processing performance of the chip system, the embodiments of the present application provide a data transmission processing method in the chip system.

[0134] See Figure 8 The data transmission processing method provided by the embodiments of the present application includes but is not limited to the following steps:

[0135] S801. The source sub-chip receives configuration parameters and configures a preset flow output table according to the configuration parameters; the flow output table includes the identifier of the data stream and the identifier of the destination sub-chip of the data stream; the source sub-chip executes the first data processing task to obtain output data.

[0136] In a specific implementation, when processing data tasks in a model parallel or model parallel plus data parallel manner, the foregoing source sub-chip may be any sub-chip in the first sub-chip group, and the first sub-chip group is a sub-chip group in the chip system for executing the processing tasks of the first step in the target task. The first step may include one or more processing steps of the target task. The chip system may be the chip system 110, chip system 120, chip system 130, or chip system 140 described above, etc.

[0137] The foregoing flow output table is pre-initialized and stored in the source sub-chip for the source sub-chip to transmit corresponding data. The flow output table is the basis for the source sub-chip to send data. Specifically, the flow output table includes the identifier of the data stream and the identifier of the destination sub-chip to which the data stream goes. In a possible implementation manner, the flow output table may further include one or more of the identifier of the task to which the data included in the data stream belongs, the indication information of the number of data packets of the data stream that have been sent, and the starting address in the source sub-chip for storing the data included in the data stream.

[0138] Specifically, regarding the data stream, one sub-chip sends data to another sub-chip. These data are encapsulated into multiple data packets, and these data packets are numbered in sequence and sent. The continuously sent data packets form a data stream. In a possible implementation, each data packet may carry 1 kb of data. If the total size of the data to be transmitted is 64 kb, then these data can be split and encapsulated into 64 data packets for sending, and the 64 data packets can form a data stream.

[0139] In a specific implementation, the flow output table can be initialized by the central controller of the chip system. Specifically, based on the foregoing description, it is the central controller that distributes each data processing task to each sub-chip group. Then, the central controller can configure a corresponding task identifier for each data processing task to distinguish different tasks.

[0140] In addition, for a data processing task, it is also the central controller that configures the sub-chip group to execute each processing step of the data processing task. Therefore, the central controller can know the flow direction of the data stream corresponding to the data processing task and configure a data stream identifier for the corresponding data stream to distinguish different data streams. For the convenience of understanding, an example is given below.

[0141] Exemplarily, see Figure 9A Suppose a certain data processing task includes eight processing steps. The central controller can allocate the processing task based on the data volume size of the data processing task, the processing complexity of the eight steps, and the load and resource usage conditions of each sub-chip group in the chip system. For example, as Figure 9AAs shown, the central controller configures the sub-chip group 1 to execute step 1 of the data processing task, the sub-chip group 2 to execute steps 2 and 3 of the data processing task, the sub-chip group 3 to execute step 4 of the data processing task, the sub-chip group 4 to execute steps 5, 6 and 7 of the data processing task, and the sub-chip group 5 to execute step 8 of the data processing task. Among them, the sub-chip group 1 includes sub-chip 4, the sub-chip group 2 includes sub-chip 5, the sub-chip group 3 includes sub-chip 6 and sub-chip 7, the sub-chip group 4 includes sub-chip 8, and the sub-chip group 5 includes sub-chip 0 and sub-chip 9.

[0142] In the above Figure 9A If it is assumed that the data after being processed in step 1 is required for the execution of steps 2, 5 and 8, then the sub-chip group 1 can send the data after being processed in step 1 to the sub-chip group 2, the sub-chip group 4 and the sub-chip group 5 respectively. Since the data sent by the sub-chip group 1 to the sub-chip group 2, the sub-chip group 4 and the sub-chip group 5 is the same, the identifiers of the data streams from the sub-chip group 1 to the sub-chip group 2, from the sub-chip group 1 to the sub-chip group 4, and from the sub-chip group 1 to the sub-chip group 5 are the same, for example, all are 17. In addition, the data obtained after the sub-chip group 2 completes the processing of steps 2 and 3 is sent to the sub-chip group 3 for the processing of step 4, and the identifier of the data stream from the sub-chip group 2 to the sub-chip group 3 can be 18. The data obtained after the sub-chip group 3 completes the processing of step 4 is sent to the sub-chip group 5 for the processing of step 8, and the identifier of the data stream from the sub-chip group 3 to the sub-chip group 5 can be 19. The data obtained after the sub-chip group 4 completes the processing of steps 5, 6 and 7 is sent to the sub-chip group 5 for the processing of step 8, and the identifier of the data stream from the sub-chip group 4 to the sub-chip group 5 can be 20.

[0143] In specific implementation, the data transmission between sub-chip groups is actually the data transmission between the sub-chips in the sub-chip groups. For example, when the sub-chip group 1 sends the data after being processed in step 1 to the sub-chip group 5, it is actually the sub-chip 4 in the sub-chip group 1 that sends the data after being processed in step 1 to the sub-chip 0 and the sub-chip 9 in the sub-chip group 5. The same applies to others and will not be elaborated here.

[0144] In another possible implementation manner, refer to Figure 9BIn this case, it is assumed that the execution of step 2, step 5, and step 8 all requires the data processed by step 1, but the required data is different or not completely the same. For example, it is assumed that the data 1 in the data processed by step 1 is required for the execution of step 2, the data 2 in the data processed by step 1 is required for the execution of step 5, and the data 3 in the data processed by step 1 is required for the execution of step 8. Then, the sub-chip group can send the data 1 to sub-chip group 2, send the data 2 to sub-chip group 4, and send the data 3 to sub-chip group 5. Since the data sent to the three sub-chip groups, namely sub-chip group 2, sub-chip group 4, and sub-chip group 5, is different, the identifiers of the corresponding data streams are also different. For example, the identifier of the data stream from sub-chip group 1 to sub-chip group 2 can be 15, the identifier of the data stream from sub-chip group 1 to sub-chip group 4 can be 16, and the identifier of the data stream from sub-chip group 1 to sub-chip group 5 can be 17. The identifiers of the data streams between other sub-chip groups can be referred to the previous description about Figure 9A and will not be elaborated here.

[0145] In a specific implementation, the data transmission between the sub-chip processing steps within the sub-chip group does not require the configuration of the identifier of the data stream. For example, in sub-chip group 2, the data obtained after step 2 is processed is sent to the corresponding module in sub-chip 5 for the processing of step 3, and the data transmission within sub-chip 5 does not require the configuration of the identifier of the data stream. Similarly, the data transmission within sub-chip 8 in sub-chip group 4 does not require the configuration of the identifier of the data stream.

[0146] Based on the above description, the central controller can know the flow direction of the corresponding data stream, that is, it can know the identifier of the destination sub-chip of the data stream. The central controller can associate and store information such as the identifier of the above data processing task, the identifier of the data stream corresponding to the data processing task, the identifier of the source sub-chip of the data stream, the identifier of the destination sub-chip of the data stream, and the steps corresponding to the processing of the source sub-chip. This associated storage can be stored in the form of a table, and this table can be called a data stream table (stream table, ST). For the convenience of understanding this data stream table, for example, refer to Table 1.

[0147] Table 1

[0148]

[0149]

[0150] Exemplarily, based on the above Figure 9B, Table 1 above exemplarily shows the information associated with step 1 of the data processing task. The information associated with other steps is the same and will not be elaborated here. As shown in Table 1, the data flow table includes the identity of the task (task identity document, TID), the identity of the source sub-chip (source mask, S_mask), the steps executed by the source sub-chip, the identity of the destination sub-chip (destination mask, D_mask), the steps executed by the destination sub-chip, and the identity of the data stream (stream identity document, SID).

[0151] The identity of the above task refers to the identity of the corresponding data processing task. The identity of this task can be, for example, 1 or other identification symbols, and this application does not limit this.

[0152] The identity of the above source sub-chip and the steps executed by the source sub-chip indicate the identity of the sub-chip that executes step 1. The identity of the above destination sub-chip and the steps executed by the destination sub-chip indicate the destination where the data processed in step 1 goes, and the corresponding processing steps executed at this destination. The identity of the data stream indicates the identity of the data stream formed by the source sub-chip sending the data processed in step 1 to each corresponding destination sub-chip. See the specific identities shown in Table 1, but the identities shown in Table 1 are only examples and can be replaced by other identification symbols. In the program executed by the computer, these identities can be represented in binary or hexadecimal, etc., and this application does not limit the representation of various identities.

[0153] It should be noted that in the embodiments of this application, the source sub-chip and the destination sub-chip mentioned are for the data stream, and the source sub-chip and the destination sub-chip corresponding to different data streams may be different.

[0154] In addition, optionally, the content included in the above data flow table is not limited to the items shown in Table 1 above. In specific implementations, the content included in the data flow table can be some or all of the items shown in Table 1 above. Or it can also include other content, such as other steps executed by the destination sub-chip (for example, if the above sub-chip 8 processes steps 6 and 7 in addition to step 5, then the above data flow table can include the information of steps 6 and 7, etc.).

[0155] Based on the above description, since the central controller stores the above data flow table, the central controller can initialize the flow output table in the above source sub-chip based on this data flow table. Specifically, the central controller can find the associated information corresponding to the above source sub-chip in the data flow table, that is, find information such as the corresponding task identifier, the identifier of the destination sub-chip, and the identifier of the data flow. Then, send this associated information to the source sub-chip. After receiving this associated information, the source sub-chip fills this information into the flow output table. One or more of the associated information corresponding to the source sub-chip in the data flow table are the above configuration parameters.

[0156] In a possible implementation manner, based on the previous description, the flow output table in the source sub-chip may further include indication information on the number of data packets of the data flow that have been sent and the starting address in the source sub-chip for storing the data included in the data flow. These two items of content can be filled into the flow output table by the source sub-chip based on its own information.

[0157] Exemplarily, the source sub-chip receives a data processing task assigned by the central controller (this data processing task may be the first data processing task in S801 above). Then, after completing the preset processing steps, it obtains the processed data (this data is the output data obtained by executing the first data processing task in S801 above). The processed data is stored in the memory buffer of the source sub-chip. Then, the source sub-chip can obtain the size and the starting storage address of the processed data. The processed data is the data to be sent, and the size of the data packet to be sent is preset. Then, by obtaining the size of the processed data, the number of data packets to be sent can be known. Thus, the source sub-chip can fill the number of data packets to be sent and the starting storage address of the data to be sent into the above flow output table. Thus, the initialization of the flow output table in the source sub-chip is completed.

[0158] The first data processing task received by the above source sub-chip may be specific task data and / or task execution instructions, etc. After receiving the first data processing task, the source sub-chip can cache the task in a preset storage space for subsequent execution.

[0159] To facilitate the understanding of the above flow output table, Table 2 can be referred to exemplarily.

[0160] Table 2

[0161]

[0162] The information corresponding to Task 1 in the flow output table is exemplarily shown in Table 2 above. It can be seen that the TID, SID, and D_mask corresponding to Task 1 in Table 2 are obtained by the central controller from Table 1 above and sent to the source sub-chip, so they are the same as Table 1 above. In addition, the start address (S_addr) in Table 2 refers to the start address in the above source sub-chip for storing the data included in the data stream to be sent. The count packet (C_packet) in Table 2 refers to the indication information of the number of data packets of the data stream that have been sent. This packet count can be a countdown. For example, at initialization, the packet count is the number of all data packets included in the data stream. Then, each time the source sub-chip sends a data packet, the packet count is decremented by one.

[0163] In addition, the flow output table may include information corresponding to multiple tasks, which are distinguished by the identifiers of the tasks. The identifiers of the data streams of different tasks may be the same or different, specifically determined according to the data stream table of the central controller.

[0164] Alternatively, exemplarily, the source sub-chip may complete the initialization of the flow output table before executing the first data processing task assigned by the central controller. Specifically, the source sub-chip may initialize the number of data packets of the data stream that has been sent in the flow output table to zero, and initialize the start address for storing the data included in the data stream to the start address of a specified storage space. Then, the data obtained after the source sub-chip executes the data processing task assigned by the central controller can be stored in the specified storage space. And, during the process of the source sub-chip sending the processed data, each time a data packet is sent, the number of data packets of the data stream that has been sent in the flow output table is incremented by one.

[0165] S802. The source sub-chip generates a plurality of data packets based on the flow output table, and the data packets include the identifier of the data stream and the identifier of the destination sub-chip.

[0166] After the above source sub-chip initializes the completion flow output table and executes the above first data processing task to obtain the output data, it can generate data packets corresponding to the data stream based on the flow output table. The data packet can include the type of the data packet, the identifier of the task, the identifier of the data stream, the identifier of the destination sub-chip, the data packet number, and the data. The identifier of the destination sub-chip in the data packet can be the identifier of one or more destination sub-chips. If the destination sub-chip corresponding to the data in the data packet is one, the identifier of the destination sub-chip in the data packet is the identifier of the one destination sub-chip. If the destination sub-chips corresponding to the data in the data packet are multiple, the identifier of the destination sub-chip in the data packet is the identifiers of the multiple destination sub-chips. For example, in the above Table 2, there are two destination sub-chips corresponding to the data stream 17, namely sub-chip 0 and sub-chip 9. Then, the data packets in the data stream 17 include the identifiers of sub-chip 0 and sub-chip 9.

[0167] In a specific implementation, the source sub-chip obtains the starting address of the data to be stored and sent (i.e., the above output data) based on the identifier of the task and the identifier of the data stream in the above flow output table, and reads the data to be sent based on the starting address to generate the above data packet.

[0168] In a possible implementation manner, the generated data packet may also carry sideband information, and the sideband information may include one or more pieces of information such as the identifier of the task, the identifier of the data stream, or the identifier of the destination sub-chip. These sideband information may not be encapsulated in the data packet, but are sent together with the data packet. In a specific implementation, only the routing module of the sub-chip can know the information included inside the data packet, and other modules such as the ports in the sub-chip are not aware of it. Therefore, in order to facilitate the fast forwarding of data packets, the data packet can be configured to carry the above sideband information. To facilitate the understanding of the format of the data packet and the format of the sideband information, an exemplary reference can be made to Figure 10 . Figure 10 The format of the data packet shown and the corresponding sideband information are only examples. In a specific implementation, the data packet may also include other information, and the sideband information may also include more information. This application does not limit this.

[0169] S803. The source sub-chip sends the above multiple data packets to the destination sub-chip.

[0170] After the source sub-chip generates a data packet based on the above flow output table, it can send the data packet to the destination sub-chip. Specifically, the source sub-chip can send the data packet based on the port forwarding mapping table. Based on the previous introduction about the chip system, the routing module of the source sub-chip also includes a port forwarding mapping module FPM, and a port forwarding mapping table is stored in the FPM. The port forwarding mapping table includes the mapping relationship between the identifier of the destination sub-chip and the sending port.

[0171] Specifically, the source sub-chip can query the port forwarding mapping table based on the identifier of the destination sub-chip to which the data packet is sent, and can find the corresponding sending port. Then, the data packet is sent out from this sending port. For easy understanding, please refer to Figure 11 . In Figure 11 , assume that sub-chip 0 is the source sub-chip and sub-chip 1 is the destination sub-chip. The port forwarding mapping table in sub-chip 0 stores the association relationship between the identifier of sub-chip 1 and port d1 of sub-chip 0. Then, when sub-chip 0 has a data packet to send to sub-chip 1, sub-chip 0 queries the port forwarding mapping table based on the identifier of sub-chip 1, and learns that the sending port of the data packet is d1. Then, sub-chip 0 sends the data packet from port d1.

[0172] S804. The destination sub-chip receives the above data packet.

[0173] After the source sub-chip sends out the data packet, the data packet reaches the destination sub-chip after a series of transmissions. The destination sub-chip receives the data packet. For example, taking the above Figure 11 as an example, sub-chip 0 sends a data packet from port d1, and the data packet reaches the destination sub-chip through port d3 of sub-chip 1. Then, sub-chip 1 receives the data packet through port d3.

[0174] S805. The destination sub-chip determines whether the flow input table contains the data flow identifier in the received data packet. The flow input table includes at least one data flow identifier to be received; when the flow input table contains the data flow identifier in the received data packet, the data in the received data packet is stored.

[0175] In specific implementation, after the destination sub-chip receives a data packet destined for itself, it can first store the data in the data packet in a buffer for subsequent processing. Specifically, the destination sub-chip can store the data in the data packet based on the flow input table. The flow input table is the basis for the destination sub-chip to receive data.

[0176] The flow input table may include the identification of tasks and the identification of data flows. In a specific implementation, the flow input table can be initialized by the central controller of the chip system. Based on the previous description, it is known that the data flow table is stored in the central controller. The central controller can query the data flow table based on the identification of the destination sub-chip, and send the information associated with the destination sub-chip to the destination sub-chip. The information associated with the destination sub-chip includes information such as the identification of the corresponding task and the identification of the data flow. After receiving this information, the destination sub-chip can write this information into its own flow input table. One or more of the associated information corresponding to the destination sub-chip in the data flow table are the configuration parameters for configuring the flow input table. The identification of the data flow in the flow input table is the data flow identification to be received by the destination sub-chip, that is, only when the data flow identification in the data packet belongs to the data flow identification in the flow input table, the destination sub-chip will obtain the data in the data packet for storage for subsequent processing.

[0177] In addition, the above flow input table may further include indication information of the number of data packets of the received data flow and information such as the starting address in the destination sub-chip for storing the data included in the data flow. These information can be filled into the flow input table by the destination sub-chip based on its own information. Exemplarily, the destination sub-chip can configure a specified storage space for the data flow to be received, and then initialize the starting address of the specified storage space into the flow input table. And exemplarily, the destination sub-chip can initialize the number of data packets of the received data flow in the flow input table to zero, and then during the process of the destination sub-chip receiving the data packets of the data flow, each time a data packet is received, the number of data packets of the corresponding received data flow in the flow input table is incremented by one. Or, exemplarily, the destination sub-chip can learn the total number of data packets included in the data flow to be received from the source sub-chip or the controller, and initialize the number of data packets of the corresponding received data flow in the flow input table to the total number. Then during the process of the destination sub-chip receiving the data packets of the data flow, each time a data packet is received, the number of data packets of the corresponding received data flow in the flow input table is decremented by one.

[0178] To facilitate the understanding of the flow input table, Table 3 can be referred to exemplarily.

[0179] Table 3

[0180]

[0181] The information corresponding to Task 1 in the flow input table is exemplarily shown in Table 3 above. It can be seen that the TID and SID corresponding to Task 1 in Table 3 are obtained by the central controller from Table 1 above and sent to the source sub-chip, so they are the same as Table 1 above. In addition, the start address (S_addr) in Table 3 refers to the start address in the aforementioned destination sub-chip for storing the data included in Data Stream 15. The count packet (C_packet) in this Table 3 refers to the indication information of the number of data packets of Data Stream 15 that have been received. This packet count can be a countdown. For example, at initialization, the packet count is the total number of all data packets included in this data stream. Then, each time the destination sub-chip receives a data packet, the packet count is decremented by one.

[0182] In addition, the flow input table may include information corresponding to multiple tasks, which are distinguished by the identifiers of the tasks. The identifiers of the data streams of different tasks may be the same or different, specifically determined according to the data stream table of the central controller.

[0183] In another possible implementation, the destination sub-chip can initialize the above flow input table by receiving a header packet from the source sub-chip. This header packet may include the type of the data packet, the identifier of the task, the identifier of the data stream, the identifier of the destination sub-chip, the total number of data packets included in the data stream, and the start address in the destination sub-chip for storing the data of this data stream. Optionally, this header packet may also carry sideband information, and the content of this sideband information may be the same as the content of the aforementioned sideband information, which will not be elaborated here. For ease of understanding the format of the header packet, reference can be made exemplarily to Figure 12 .

[0184] Figure 12 where N_PK represents the total number of data packets included in the data stream, Address represents the start address in the aforementioned destination sub-chip for storing the data of this data stream, and other identifiers refer to the previous description, which will not be elaborated here. Figure 12 The format of the data packet and the corresponding sideband information shown are only examples. In specific implementations, the data packet may also include other information, and the sideband information may also include more information. This application does not limit this.

[0185] In a possible implementation, after the flow input table of the destination sub-chip is initialized, it indicates that the destination sub-chip is ready to receive the corresponding data stream. Therefore, the destination sub-chip can send an unlock data packet to the source sub-chip, and this unlock data packet is used to indicate to the source sub-chip that the destination sub-chip is ready to receive the corresponding data stream.

[0186] The unlock packet may include a packet type, an identifier of a task, an identifier of a data stream, and an identifier of a sub-chip that receives the packet. The sub-chip that receives the packet is the aforementioned source sub-chip. Optionally, the unlock packet may also carry sideband information, and the content of the sideband information may be the same as the content of the aforementioned sideband information, which will not be elaborated here.

[0187] Based on the above description, after the destination sub-chip initializes the flow input table, for the packet received from the source sub-chip, it can first obtain the data stream identifier in the packet and compare it with the data stream identifier in its own flow input table. If the flow input table includes the data stream identifier in the received packet, it can query the flow input table based on the task identifier and data stream identifier in the packet to obtain the starting address of the corresponding stored data, and then calculate the corresponding storage address based on the starting address, and store the data of the packet in the corresponding storage address.

[0188] In a possible implementation, the central controller assigns a second data processing task to the destination sub-chip, and the second data processing task is executed based on the data obtained after the execution of the above first data processing task. Then, the source sub-chip executes the above first data processing task to obtain output data, and sends the output data to the destination sub-chip in the form of a data stream. The destination sub-chip receives and stores the output data based on its own flow input table. Then, the destination sub-chip can read the output data from the storage space based on the known storage address for executing its own second data processing task.

[0189] In a possible implementation manner, based on the relevant description of the foregoing chip system, the chip system may include multiple subsystems, and each subsystem is connected according to a preset topological connection relationship. If the above source sub-chip and destination sub-chip are not sub-chips in the same subsystem, then, the packet transmitted between the two sub-chips further includes an identifier of the subsystem where the sub-chip that receives the packet is located. For example, if the source sub-chip sends a packet to the destination sub-chip (such as the packet included in the above data stream, etc.), then, the packet further includes an identifier of the subsystem where the destination sub-chip is located. If the destination sub-chip sends a packet to the source sub-chip (such as the above unlock packet, etc.), then, the packet further includes an identifier of the subsystem where the source sub-chip is located. Through the identifier of the subsystem in the packet, data transmission between sub-chips across subsystems can be realized, which is beneficial to the processing of large data tasks and improves the processing efficiency.

[0190] In summary, the embodiments of the present application can achieve efficient data transmission within the chip system based on the above initialized flow output table and flow input table, and improve the data processing performance of the chip system.

[0191] In a possible implementation, in order to further improve the data transmission efficiency within the above chip system, the embodiments of the present application may provide a routing implementation method to enable flexible scheduling and efficient transmission between the sub-chips of the data chip system. The following takes the first sub-chip as an example for introduction.

[0192] See Figure 13 , the routing implementation method provided by the embodiments of the present application includes but is not limited to the following steps:

[0193] S1301. The first sub-chip obtains a first data packet; wherein, the first data packet includes the identifier of the destination sub-chip; the first sub-chip and the destination sub-chip are sub-chips included in the chip system, and the multiple sub-chips included in the chip system are connected in a preset topological structure.

[0194] The chip system may be the chip system 110, chip system 120, chip system 130, or chip system 140 introduced above, etc. The first sub-chip may be any sub-chip in any one of these chip systems.

[0195] Exemplarily, the above first sub-chip may be the source sub-chip described above Figure 13 , then, the above first sub-chip obtaining the first data packet may be the first sub-chip generating the first data packet. The specific implementation of generating the first data packet may refer to the corresponding description in step S802 above, which will not be elaborated here.

[0196] Or, exemplarily, the above first sub-chip may be a sub-chip passed through during the process of the data packet being transmitted from the source sub-chip to the destination sub-chip, then, the above first sub-chip obtaining the first data packet may be the first sub-chip receiving the first data packet.

[0197] For the convenience of subsequent description, the chip system where the first sub-chip is located is referred to as the first chip system. The first sub-chip receives the above first data packet from another sub-chip in the first chip system.

[0198] In a specific implementation, the first data packet may include one or more of the following information: the type of the packet, the identifier of the task, the identifier of the data stream, the identifier of the destination sub-chip, the packet number, and the data. In a possible implementation manner, the above first data packet may also carry sideband information, and these sideband information may include one or more of the identifier of the task, the identifier of the data stream, or the identifier of the destination sub-chip. Regarding the content included in the first data packet, reference may be made to the description of the data packet in step S802 above, which will not be elaborated here.

[0199] S1302. The first sub-chip sends the data in the first data packet based on the data transmission situation among the sub-chips in the chip system.

[0200] In specific implementation, the data transmission situations among the sub-chips in the chip system include various types. Hereinafter, several possible implementation manners are introduced by way of example.

[0201] In the first possible implementation manner, the first sub-chip may send the data in the first data packet based on the congestion situation of its own ports.

[0202] Specifically, based on the foregoing introduction to the chip system, each sub-chip includes multiple ports for communicating with other sub-chips. Among them, each port is configured with a corresponding transmission buffer for storing data to be transmitted.

[0203] After receiving the first data packet, the first sub-chip parses the first data packet to obtain the identifier of the destination sub-chip in the first data packet. If the identifier of the destination sub-chip indicates that the first sub-chip is the destination sub-chip, then the first sub-chip extracts and stores the data in the first data packet for subsequent processing. Otherwise, the first sub-chip uses the identifier of the destination sub-chip as an index to search for the transmission port of the first data packet in its own forwarding mapping table. For the introduction to the forwarding mapping table, reference can be made to the corresponding description in the foregoing description of Figure 2 If multiple found transmission ports are included, then the specific transmission port can be determined based on the congestion situations of the transmission buffers of the multiple transmission ports. Specifically, in order to improve the data transmission efficiency, the port with the least amount of data to be transmitted in the transmission buffers of the multiple transmission ports can be selected to send the first data packet.

[0204] In a possible implementation manner, if the first data packet includes identifiers of multiple destination sub-chips and the first sub-chip is one of the destination sub-chips, then the first sub-chip extracts and stores the data in the first data packet for subsequent processing. And the first sub-chip uses the identifiers of the remaining destination sub-chips as an index to search for the transmission port of the data in the first data packet in its own forwarding mapping table.

[0205] If the identifier of the remaining destination sub-chip is one, then similarly, after finding the corresponding transmission port, the port with the least amount of data to be transmitted in the transmission buffer of the transmission port is selected to send the data in the first data packet. Specifically, the data will be re-encapsulated into a data packet for transmission, and the identifier of the destination sub-chip in the re-encapsulated data packet no longer includes the identifier of the first sub-chip, but only includes the identifier of the remaining destination sub-chip.

[0206] If there are still multiple identifiers of the remaining destination sub-chips as described above, then the first sub-chip searches for the corresponding sending ports in its own forwarding mapping table respectively. If the found sending ports are the same, then a data packet can be regenerated by copying the data included in the above first data packet, and the newly generated data packet includes the identifiers of the multiple remaining destination sub-chips. And send the newly generated data packet from the same found sending port. Similarly, the sending port can be the port with the least amount of data to be sent in the sending buffer among the found sending ports.

[0207] Or, if there are still multiple identifiers of the remaining destination sub-chips as described above, taking two as an example, assume that the remaining destination sub-chips are sub-chip A and sub-chip B. The above first sub-chip searches for the sending port mapped by the identifier of sub-chip A in its own forwarding mapping table, and searches for the sending port mapped by the identifier of sub-chip B. Assume that the found sending ports are different, then the first sub-chip can regenerate two data packets: data packet A and data packet B. Both data packets include the data included in the above first data packet, where the identifier of the destination sub-chip included in data packet A is the identifier of sub-chip A, and the identifier of the destination sub-chip included in data packet B is the identifier of sub-chip B. Then, send data packet A and data packet B through their respective found sending ports. Similarly, the sending port can be the port with the least amount of data to be sent in the sending buffer among the found sending ports.

[0208] The second possible implementation method is that the above first sub-chip can send the data in the above first data packet to the destination sub-chip based on the principle of minimum bandwidth consumption. The principle of minimum bandwidth consumption refers to the principle of delivering data to the destination sub-chip with the minimum transmission bandwidth.

[0209] To facilitate the understanding of this implementation method, first introduce the direction coordinate system constructed with the first sub-chip as the center. Figure 14 An exemplary schematic diagram of the direction coordinate system constructed with the first sub-chip as the center is shown. It can be seen that the direction coordinate system includes four direction axes: the first direction axis, the second direction axis, the third direction axis, and the fourth direction axis. The four direction axes all diverge outward with the first sub-chip as the center. Among them, the first direction axis and the second direction axis are collinear and in opposite directions; the third direction axis and the fourth direction axis are collinear and in opposite directions. The direction coordinate system also includes four regions: the first region, the second region, the third region, and the fourth region. Among them, the first region is bounded by the first direction axis and the third direction axis; the second region is bounded by the second direction axis and the third direction axis; the third region is bounded by the second direction axis and the fourth direction axis, and the fourth region is bounded by the first direction axis and the fourth direction axis.

[0210] In a chip system, the row where the first sub-chip is located is on at least one of the first direction axis and the second direction axis, and the column where the first sub-chip is located is on at least one of the third direction axis and the fourth direction axis. Exemplarily, reference can be made to Figure 15 . Assume that sub-chip 5 in the chip system is the first sub-chip. Then, a direction coordinate system is established with sub-chip 5 as the center. In this direction coordinate system, the second row of sub-chip 5 is on the first direction axis and the second direction axis, and the second column of sub-chip 5 is on the third direction axis and the fourth direction axis. Then, sub-chip 2 and sub-chip 3 are in the first region of this direction coordinate system. Sub-chip 0 is in the second region of this direction coordinate system. Sub-chip 8 and sub-chip 12 are in the third region of this direction coordinate system. Sub-chip 10, sub-chip 11, sub-chip 14, and sub-chip 15 are in the fourth region of this direction coordinate system.

[0211] In a possible implementation, if sub-chip 0 in the above Figure 15 chip system is the first sub-chip, then a direction coordinate system is established with sub-chip 0 as the center. In this direction coordinate system, the first row of sub-chip 0 is on the first direction axis, and the first column of sub-chip 0 is on the fourth direction axis. Then, except for the sub-chips in the row and column where sub-chip 0 is located, the remaining sub-chips are all in the fourth region of this direction coordinate system.

[0212] In a possible implementation, the above minimum bandwidth consumption principle includes: when the destination sub-chip included in the above first data packet is on the target direction axis, the first sub-chip sends the data of the first data packet along the direction of the target direction axis; the target direction axis is the first direction axis, the second direction axis, the third direction axis, or the fourth direction axis. For the sake of easy understanding, taking the above Figure 15 as an example for illustration.

[0213] In Figure 15 , assume that sub-chip 5 is the above first sub-chip, and it receives a data packet. The identifier of the destination sub-chip in the data packet indicates that the destination sub-chip is sub-chip 7. If there is only the identifier of one destination sub-chip in the data packet, then, since sub-chip 7 is on the first direction axis, sub-chip 5 sends the data packet along the direction of the first direction axis. That is, sub-chip 5 first sends the data packet to sub-chip 6, and then sub-chip 6 forwards it to sub-chip 7. If there are multiple identifiers of destination sub-chips in the data packet, and sub-chip 7 as one of the destination sub-chips is on the first direction axis. Therefore, sub-chip 5 copies the data in the data packet to generate a new data packet and sends the new data packet along the direction of the first direction axis. That is, sub-chip 5 first sends the new data packet to sub-chip 6, and then sub-chip 6 forwards it to sub-chip 7. The newly generated data packet includes the identifier of sub-chip 7.

[0214] In a possible implementation manner, the above-mentioned first data packet includes the identifiers of the first destination sub-chip and the second destination sub-chip. The above-mentioned minimum bandwidth consumption principle further includes: in the direction coordinate system established with the above-mentioned first sub-chip as the center, when the first destination sub-chip and the second destination sub-chip are respectively in two adjacent regions among the first region, the second region, the third region, and the fourth region of the coordinate system, the first sub-chip sends a second data packet along the direction of the common direction axis. The second data packet includes the data, the identifiers of the first destination sub-chip and the second destination sub-chip. The common direction axis is the direction axis of the common boundary of the two adjacent regions. For the sake of easy understanding, in combination with the above Figure 15 as an example for illustration.

[0215] In Figure 15 , assume that sub-chip 5 is the above-mentioned first sub-chip, and it receives a data packet. The identifier of the destination sub-chip in the data packet indicates that the destination sub-chips are sub-chip 8 and sub-chip 14. Sub-chip 8 is located in the third region, and sub-chip 14 is located in the fourth region. These two regions are adjacent regions, and the common boundary is the fourth direction axis. Therefore, sub-chip 5 sends the data packet along the direction of the fourth direction axis. That is, sub-chip 5 first sends the data packet to sub-chip 9, and then sub-chip 9 further forwards it. Specifically, sub-chip 9 can also be regarded as the above-mentioned first sub-chip, a direction coordinate system is established with this sub-chip 9 as the center, and then the data is forwarded based on the above-mentioned minimum bandwidth consumption principle.

[0216] In a possible implementation manner, the above-mentioned first data packet includes the identifiers of the first destination sub-chip and the second destination sub-chip. The above-mentioned minimum bandwidth consumption principle further includes: in the direction coordinate system established with the above-mentioned first sub-chip as the center, when the first destination sub-chip is in the first region of the coordinate system and the second destination sub-chip is in the third region, the first sub-chip sends a third data packet along one of the direction axes of the two boundaries of the first region, and sends a fourth data packet along one of the direction axes of the two boundary direction axes of the third region. The third data packet includes the data and the identifier of the first destination sub-chip. The fourth data packet includes the data and the identifier of the second destination sub-chip. For the sake of easy understanding, in combination with the above Figure 15 as an example for illustration.

[0217] In Figure 15In this case, assume that sub-chip 5 is the first sub-chip mentioned above. It receives a data packet, and the identifier of the destination sub-chip in this data packet indicates that the destination sub-chips are sub-chip 2 and sub-chip 12. Sub-chip 2 is located in the first area, and sub-chip 12 is located in the third area. Then, sub-chip 5 can regenerate two data packets based on the data in the received data packet: data packet A and data packet B. Data packet A includes the data and the identifier of sub-chip 2, and data packet B includes the data and the identifier of sub-chip 12. Then, send data packet A along the direction of the first direction axis or the third direction axis. For example, send data packet A along the direction of the first direction axis, that is, first send data packet A to sub-chip 6, and then sub-chip 6 forwards data packet A to sub-chip 2. In addition, sub-chip 5 sends data packet B along the direction of the second direction axis or the fourth direction axis. For example, send data packet B along the direction of the fourth direction axis, that is, first send data packet B to sub-chip 9, and then sub-chip 9 continues to forward it further.

[0218] In a possible implementation manner, the above first data packet includes the identifiers of the first destination sub-chip and the second destination sub-chip. The above minimum bandwidth consumption principle further includes: in the direction coordinate system established with the above first sub-chip as the center, when the first destination sub-chip is in the second area of the coordinate system and the second destination sub-chip is in the fourth area, the first sub-chip sends a third data packet along one of the direction axes of the two boundaries of the second area, and sends a fourth data packet along one of the direction axes of the two boundaries of the fourth area. The third data packet includes the data and the identifier of the first destination sub-chip. The fourth data packet includes the data and the identifier of the second destination sub-chip. For the sake of easy understanding, combined with the above Figure 15 as an example for illustration.

[0219] In Figure 15 In this case, assume that sub-chip 5 is the first sub-chip mentioned above. It receives a data packet, and the identifier of the destination sub-chip in this data packet indicates that the destination sub-chips are sub-chip 0 and sub-chip 10. Sub-chip 0 is located in the second area, and sub-chip 10 is located in the fourth area. Then, sub-chip 5 can regenerate two data packets based on the data in the received data packet: data packet C and data packet D. Data packet C includes the data and the identifier of sub-chip 0, and data packet D includes the data and the identifier of sub-chip 10. Then, send data packet C along the direction of the second direction axis or the third direction axis. For example, send data packet C along the direction of the second direction axis, that is, first send data packet C to sub-chip 4, and then sub-chip 4 forwards data packet C to sub-chip 0. In addition, sub-chip 5 sends data packet D along the direction of the first direction axis or the fourth direction axis. For example, send data packet D along the direction of the first direction axis, that is, first send data packet D to sub-chip 6, and then sub-chip 6 forwards it to sub-chip 10.

[0220] In a possible implementation, the above first data packet includes the identifiers of the first destination sub-chip and the second destination sub-chip. The above minimum bandwidth consumption principle further includes: in the direction coordinate system established with the above first sub-chip as the center, when the first destination sub-chip is in the target area and the second destination sub-chip is on the direction axis of the boundary of the target area, the first sub-chip sends a fifth data packet along the direction of the direction axis of the boundary of the target area. The fifth data packet includes the data and the identifiers of the first destination sub-chip and the second destination sub-chip. The target area is the first area, the second area, the third area, or the fourth area. For the sake of easy understanding, combined with the above Figure 15 as an example for illustration.

[0221] In Figure 15 , assume that sub-chip 5 is the above first sub-chip, and it receives a data packet. The identifier of the destination sub-chip in the data packet indicates that the destination sub-chips are sub-chip 14 and sub-chip 9. Sub-chip 14 is located in the fourth area, and sub-chip 9 is on the fourth direction axis. The fourth direction axis is the boundary direction axis of the fourth area. Then, sub-chip 5 can send the received data packet along the direction of the fourth direction axis, that is, send it to sub-chip 9. After receiving the data packet, sub-chip 9 stores the data in the data packet. And copies the data to regenerate a new data packet. The new data packet includes the identifier of sub-chip 14, and then sends the new data packet to sub-chip 13 or sub-chip 10, and then sub-chip 13 or sub-chip 10 forwards it to sub-chip 14.

[0222] In a third possible implementation, the above first sub-chip can receive scheduling information from the central controller in the chip system and send data correspondingly based on the scheduling information.

[0223] Specifically, based on the previous introduction to the chip system, it can be known that the central controller of the chip system can also be responsible for data scheduling in the chip system. Specifically, the central controller obtains the data transmission status of each sub-chip through the control bus, and can learn about the congestion status of each transmission path and / or the port congestion status of each sub-chip through the analysis of these data transmission statuses. Thus, a data transmission strategy can be formulated based on these situations and sent to each sub-chip in the form of scheduling information. Each sub-chip sends data correspondingly based on the scheduling information sent by the controller, thereby reducing the probability of congestion and improving the data transmission efficiency.

[0224] Exemplarily, the scheduling information sent by the central controller to a sub-chip may include the identifiers of one or more sub-chips and the identifiers of the ports corresponding to the one or more sub-chips. After receiving the scheduling information, the sub-chip updates these information to its own forwarding mapping table for subsequent data forwarding. For the sake of easy understanding, it can be exemplarily referred to Figure 16 .

[0225] Figure 16 An exemplary structural schematic diagram of a chip system is shown. Assuming that sub-chip 0 serves as the central controller of the chip system, sub-chip 0 can communicate with other sub-chips in the chip system through a control bus ( Figure 16 not shown in the figure). Specifically, sub-chip 0 can collect information such as the amount of data to be sent at the ports of each sub-chip through the control bus. Based on this information, the congestion situation of each port can be analyzed, and then the congestion situation of each transmission path can be analyzed. Based on the situations obtained from these analyses, sub-chip 0 can comprehensively formulate a data transmission strategy in the chip system and send it to each sub-chip in the form of scheduling information.

[0226] For example, for the data transmitted from sub-chip 1 to sub-chip 6, sub-chip 0 analyzes and learns that port d2 of sub-chip 1 is relatively idle, and port d1 of sub-chip 5 is also relatively idle. Then, sub-chip 0 sends a scheduling information to sub-chip 1, and the scheduling information includes the identifier of sub-chip 6 and the identifier of port d2. After receiving the scheduling information, sub-chip 1 updates the information that the sending port corresponding to the destination sub-chip being sub-chip 6 is port d2 to its own forwarding mapping table. In addition, sub-chip 0 sends a scheduling information to sub-chip 5, and the scheduling information includes the identifier of sub-chip 6 and the identifier of port d1. After receiving the scheduling information, sub-chip 5 updates the information that the sending port corresponding to the destination sub-chip being sub-chip 6 is port d1 to its own forwarding mapping table. Then, when sub-chip 1 receives a data packet destined for sub-chip 6, it queries its own forwarding mapping table and learns that its sending port is d2. Therefore, the data packet is sent out from port d2. After the data packet arrives at sub-chip 5, sub-chip 5 queries its own forwarding mapping table and learns that its sending port is d1. Therefore, the data packet is sent out from port d1 and the data packet is delivered to sub-chip 6.

[0227] In summary, the embodiments of the present application transmit the received data based on the data transmission situation in the chip system, so as to flexibly schedule the sending of data, improve the data transmission efficiency, and further improve the processing performance of the chip system.

[0228] The above mainly introduces the data transmission processing method in the chip system provided by the embodiments of the present application. It can be understood that, in order to implement the corresponding functions, each device includes the corresponding hardware structure and / or software module for executing each function. Combining the units and steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0229] The embodiments of the present application can divide the functions of the device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. There may be other division methods in actual implementation.

[0230] In the case of dividing each function module corresponding to each function, Figure 17 Fig. shows a specific schematic diagram of the logical structure of the device, and the device may be the above-mentioned source sub-chip. The device 1700 includes:

[0231] A receiving unit 1701, configured to receive configuration parameters;

[0232] A configuration unit 1702, configured to configure a preset flow output table according to the foregoing configuration parameters; the foregoing flow output table includes the identifier of the data stream and the identifier of the destination sub-chip of the foregoing data stream; the foregoing device 1700 and the foregoing destination sub-chip are sub-chips among multiple sub-chips included in the chip system, and the foregoing multiple sub-chips are connected in a preset topology structure;

[0233] An execution unit 1703, configured to execute a first data processing task to obtain output data;

[0234] A generating unit 1704, configured to generate multiple data packets based on the foregoing flow output table, and the foregoing data packets include the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip;

[0235] A sending unit 1705, configured to send the foregoing multiple data packets.

[0236] In a possible implementation manner, the foregoing flow output table further includes one or more of the identifier of the foregoing first data processing task, the indication information of the number of data packets of the foregoing data stream that have been sent, and the starting address in the foregoing device 1700 for storing the data included in the foregoing data stream; the foregoing data packet further includes the identifier of the foregoing first data processing task.

[0237] In a possible implementation manner, the foregoing receiving unit 1701 is further configured to receive an unlock data packet before the foregoing generating unit generates a plurality of data packets based on the foregoing flow output table; wherein, the foregoing unlock data packet includes the identifier of the foregoing data stream, and the foregoing unlock data packet is used to indicate to the foregoing device 1700 that the foregoing destination sub-chip is ready to receive the foregoing data stream.

[0238] In a possible implementation manner, the foregoing sending unit 1705 is specifically configured to: send the foregoing plurality of data packets based on a port forwarding mapping table; wherein, the foregoing port forwarding mapping table includes the mapping relationship between the identifier of the foregoing destination sub-chip and the sending port.

[0239] In a possible implementation manner, when there are multiple destination sub-chips of the foregoing data stream, the foregoing data packet includes the identifiers of the foregoing multiple destination sub-chips.

[0240] In a possible implementation manner, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, and the foregoing subsystem is configured with a subsystem identifier; when the foregoing destination sub-chip is a sub-chip of the foregoing subsystem, the foregoing data packet further includes the identifier of the foregoing subsystem.

[0241] Figure 17 For the specific operations and beneficial effects of each unit in the foregoing device 1700, reference may be made to the corresponding descriptions in the foregoing Figure 8 and its possible method embodiments, which will not be elaborated here.

[0242] In the case of dividing each functional module according to the corresponding functions, Figure 18 FIG. shows a specific logical structure diagram of a device, and this device may be the foregoing destination sub-chip. The device 1800 includes:

[0243] A receiving unit 1801, configured to receive configuration parameters;

[0244] A configuration unit 1802, configured to configure a preset flow input table according to the foregoing configuration parameters, the foregoing flow input table includes at least one data stream identifier to be received; the foregoing device 1800 is a sub-chip among a plurality of sub-chips included in a chip system, and the foregoing plurality of sub-chips are connected in a preset topology structure;

[0245] A determination unit 1803, configured to determine whether the data stream identifier in the foregoing data packet is included in the foregoing flow input table when receiving the data packet;

[0246] A storage unit 1804, configured to store the data in the foregoing data packet when the foregoing data packet includes the foregoing data stream identifier to be received.

[0247] In a possible implementation, the foregoing flow input table further includes one or more of the identifier of the foregoing first data processing task, the indication information of the number of data packets of the foregoing received data stream, and the starting address for storing the data included in the foregoing data stream in the foregoing apparatus 1800; the foregoing data packet further includes the identifier of the foregoing first data processing task.

[0248] In a possible implementation, the foregoing apparatus 1800 further includes a sending unit, configured to send an unlock data packet before the foregoing receiving unit 1801 receives the foregoing data packet; wherein, the foregoing unlock data packet includes the identifier of the foregoing data stream and the identifier of the source sub-chip, and the foregoing unlock data packet is used to indicate to the foregoing source sub-chip that the foregoing apparatus 1800 is ready to receive the foregoing data stream; the foregoing source sub-chip is the sub-chip that sends the foregoing data packet in the foregoing chip system.

[0249] In a possible implementation, the foregoing chip system includes a subsystem, the foregoing subsystem includes at least two sub-chips, the foregoing subsystem is configured with a subsystem identifier; the foregoing data packet further includes the identifier of the foregoing subsystem.

[0250] Figure 18 For the specific operations and beneficial effects of each unit in the foregoing apparatus 1800, reference may be made to the corresponding descriptions in the foregoing Figure 8 and its possible method embodiments, which will not be elaborated herein.

[0251] In the case of dividing each functional module according to the corresponding functions, Figure 19 shows a specific logical structure diagram of an apparatus, and the apparatus may be the foregoing central controller. The apparatus 1900 includes:

[0252] An allocation unit 1901, configured to allocate a first data processing task to a source sub-chip, and the data obtained after the execution of the foregoing first data processing task is sent to a destination sub-chip in the form of a data stream; the foregoing apparatus 1900, the foregoing source sub-chip, and the foregoing destination sub-chip are sub-chips among multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology structure;

[0253] A configuration unit 1902, configured to configure an identifier for the foregoing data stream;

[0254] A sending unit 1903, configured to send the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip to the foregoing source sub-chip; wherein, the identifier of the foregoing data stream and the identifier of the foregoing destination sub-chip are used for associated storage in the stream output table of the foregoing source sub-chip; the foregoing stream output table is the basis for the foregoing source sub-chip to send data.

[0255] In a possible implementation manner, the foregoing allocation unit 1901 is further configured to allocate a second data processing task to the foregoing destination sub-chip, and the foregoing second data processing task is executed based on the data obtained after the execution of the foregoing first data processing task;

[0256] The foregoing sending unit 1903 is further configured to send the identifier of the foregoing data stream to the foregoing destination sub-chip; wherein, the identifier of the foregoing data stream is used for storage in the stream input table of the foregoing destination sub-chip, and the foregoing stream input table is the basis for the foregoing destination sub-chip to receive data.

[0257] In a possible implementation manner, the foregoing apparatus 1900 further includes:

[0258] An obtaining unit, configured to obtain the data transmission situation between sub-chips in the foregoing chip system;

[0259] A generating unit, configured to generate scheduling information for the foregoing source sub-chip based on the foregoing data transmission situation, where the foregoing scheduling information indicates a sending port for sending the foregoing data stream to the foregoing destination sub-chip in the foregoing source sub-chip;

[0260] The foregoing sending unit 1903 is further configured to send the foregoing scheduling information to the foregoing source sub-chip.

[0261] Figure 19 For the specific operations and beneficial effects of each unit in the foregoing apparatus 1900, reference may be made to the corresponding descriptions in the foregoing Figure 8 and Figure 13 and its possible method embodiments, which will not be elaborated herein.

[0262] Figure 20 Shown is a specific hardware structure diagram of the apparatus provided in this application. The apparatus 2000 includes: a processor 2001, a memory 2002, and a communication port 2003. The processor 2001, the communication port 2003, and the memory 2002 may be connected to each other or connected to each other through a bus 2004.

[0263] Exemplarily, the memory 2002 is used to store the computer programs and data of the device 2000. The memory 2002 may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM), etc. Exemplarily, the memory 2002 may be the static memory shown in the above Figure 2 as shown.

[0264] The communication port 2003 includes a sending port and a receiving port. The number of communication ports 2003 may be multiple, and is used to support the device 2000 to communicate, such as receiving or sending data or messages, etc. Exemplarily, the communication port 2003 may be the ports d0, d1, d2, and d3 shown in the above Figure 2 as shown.

[0265] Exemplarily, the processor 2001 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. Exemplarily, the processor 2001 may be the processing module shown in the above Figure 2 as shown.

[0266] In a possible implementation manner, the above device 2000 is the source sub-chip in the above Figure 8 and its possible implementation manners. Then, the processor 2001 in the device 2000 may be used to read the program stored in the above memory 2002, so that the device 2000 executes the operations performed by the source sub-chip as described in the above Figure 8 and its specific embodiments.

[0267] In a possible implementation manner, the above device 2000 is the destination sub-chip in the above Figure 8 and its possible implementation manners. Then, the processor 2001 in the device 2000 may be used to read the program stored in the above memory 2002, so that the device 2000 executes the operations performed by the destination sub-chip as described in the above Figure 8 and its specific embodiments.

[0268] In a possible implementation manner, the above device 2000 is the above Figure 8and the central controller in its possible embodiments. Then, the processor 2001 in the device 2000 can be used to read the program stored in the above-mentioned memory 2002, so that the device 2000 performs the operations performed by the central controller as described above Figure 8 and the operations performed by the central controller described in its specific embodiments.

[0269] Figure 20 For the specific operations and beneficial effects of each unit in the device 2000 shown, reference can be made to the above Figure 8 and the corresponding descriptions in its specific method embodiments, which will not be elaborated here.

[0270] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the operations performed by the source sub-chip described in any one of the above Figure 8 and its possible method embodiments.

[0271] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the operations performed by the destination sub-chip described in any one of the above Figure 8 and its possible method embodiments.

[0272] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the operations performed by the central controller described in any one of the above Figure 8 and its possible method embodiments.

[0273] The embodiments of the present application also provide a computer program product. When the computer program product is read and executed by a computer, the operations performed by the source sub-chip described in any one of the above Figure 8 and its possible method embodiments will be realized.

[0274] The embodiments of the present application also provide a computer program product. When the computer program product is read and executed by a computer, the operations performed by the destination sub-chip described in any one of the above Figure 8 and its possible method embodiments will be realized.

[0275] The embodiments of the present application also provide a computer program product. When the computer program product is read and executed by a computer, the operations performed by the central controller described in any one of the above Figure 8 and its possible method embodiments will be realized.

[0276] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data transmission and processing method in a chip system, characterized in that, the method includes: The source sub-chip receives configuration parameters and configures a preset flow output table according to the configuration parameters; the flow output table includes the identifier of the data stream and the identifier of the destination sub-chip of the data stream; the source sub-chip and the destination sub-chip are sub-chips among multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology; The source sub-chip executes a first data processing task to obtain output data; The source sub-chip generates multiple data packets based on the flow output table and sends them, and the data packets include the identifier of the data stream and the identifier of the destination sub-chip.

2. The method according to claim 1, characterized in that, the flow output table further includes one or more of the identifier of the first data processing task, indication information of the number of data packets of the data stream that have been sent, and the starting address in the source sub-chip for storing the data included in the data stream; The data packet further includes the identifier of the first data processing task.

3. The method according to claim 1, characterized in that, before the source sub-chip generates multiple data packets based on the flow output table and sends them, it further includes: The source sub-chip receives an unlock data packet; wherein, the unlock data packet includes the identifier of the data stream, and the unlock data packet is used to indicate to the source sub-chip that the destination sub-chip is ready to receive the data stream.

4. The method according to claim 1, characterized in that, the source sub-chip sends the multiple data packets, including: The source sub-chip sends the multiple data packets based on a port forwarding mapping table; wherein, the port forwarding mapping table includes the mapping relationship between the identifier of the destination sub-chip and the sending port.

5. The method according to claim 1, characterized in that, when there are multiple destination sub-chips for the data stream, the data packet includes the identifiers of the multiple destination sub-chips.

6. The method according to any one of claims 1-5, characterized in that, the chip system includes a subsystem, the subsystem includes at least two sub-chips, and the subsystem is configured with a subsystem identifier; when the destination sub-chip is a sub-chip of the subsystem, the data packet further includes the identifier of the subsystem.

7. A data transmission and processing method in a chip system, characterized in that, the method includes: The destination sub-chip receives configuration parameters and configures a preset flow input table according to the configuration parameters, the flow input table includes at least one data stream identifier to be received; the destination sub-chip is a sub-chip among multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology; When the destination sub-chip receives a data packet, it determines whether the data stream identifier in the data packet is included in the flow input table; When the data stream identifier in the data packet is included in the flow input table, store the data in the data packet.

8. The method according to claim 7, characterized in that, The flow input table further includes one or more of the identifier of the first data processing task, indication information of the number of data packets of the data stream that have been received, and the starting address in the destination sub-chip for storing the data included in the data stream; The data packet further includes the identifier of the first data processing task.

9. The method according to claim 7, wherein, Before the destination sub-chip receives the data packet, it further includes: The destination sub-chip sends an unlock data packet; wherein, the unlock data packet includes the identifier of the data stream and the identifier of the source sub-chip, and the unlock data packet is used to indicate to the source sub-chip that the destination sub-chip is ready to receive the data stream; the source sub-chip is the sub-chip that sends the data packet in the chip system.

10. The method according to any one of claims 7-9, wherein, The chip system includes a subsystem, the subsystem includes at least two sub-chips, and the subsystem is configured with a subsystem identifier; The data packet further includes the identifier of the subsystem.

11. A data transmission and processing method in a chip system, wherein, The method includes: The controller assigns a first data processing task to the source sub-chip, and the data obtained after the execution of the first data processing task is sent to the destination sub-chip in the form of a data stream; the controller, the source sub-chip, and the destination sub-chip are sub-chips among the multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology; The controller configures an identifier for the data stream; The controller sends the identifier of the data stream and the identifier of the destination sub-chip to the source sub-chip; wherein, the identifier of the data stream and the identifier of the destination sub-chip are used to be associated and stored in the flow output table of the source sub-chip; the flow output table is the basis for the source sub-chip to send data.

12. The method according to claim 11, wherein, The method further includes: The controller assigns a second data processing task to the destination sub-chip, and the second data processing task is executed based on the data obtained after the execution of the first data processing task; The controller sends the identifier of the data stream to the destination sub-chip; wherein, the identifier of the data stream is used to be stored in the flow input table of the destination sub-chip, and the flow input table is the basis for the destination sub-chip to receive data.

13. The method according to claim 11 or 12, wherein, The method further includes: The controller obtains the data transmission situation between the sub-chips in the chip system; The controller generates scheduling information for the source sub-chip based on the data transmission situation, and the scheduling information indicates the sending port in the source sub-chip for sending the data stream to the destination sub-chip; The controller sends the scheduling information to the source sub-chip.

14. A source sub-chip, wherein, The source sub-chip includes: A receiving unit for receiving configuration parameters; A configuration unit, configured to configure a preset flow output table according to the configuration parameters; the flow output table includes an identifier of a data stream and an identifier of a destination sub-chip of the data stream; the source sub-chip and the destination sub-chip are sub-chips among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure; An execution unit, configured to execute a first data processing task to obtain output data; A generation unit, configured to generate a plurality of data packets based on the flow output table, where the data packets include the identifier of the data stream and the identifier of the destination sub-chip; A sending unit, configured to send the plurality of data packets.

15. A destination sub-chip, characterized in that, the destination sub-chip includes: A receiving unit, configured to receive configuration parameters; A configuration unit, configured to configure a preset flow input table according to the configuration parameters, the flow input table includes at least one data stream identifier to be received; the destination sub-chip is a sub-chip among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure; A judging unit, configured to judge whether the data stream identifier in the data packet is included in the flow input table when receiving the data packet; A storage unit, configured to store the data in the data packet when the data stream identifier in the data packet is included in the flow input table.

16. A controller, characterized in that, the controller includes: An allocation unit, configured to allocate a first data processing task to a source sub-chip, and the data obtained after the execution of the first data processing task is sent to a destination sub-chip in the form of a data stream; the controller, the source sub-chip and the destination sub-chip are sub-chips among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure; A configuration unit, configured to configure an identifier for the data stream; A sending unit, configured to send the identifier of the data stream and the identifier of the destination sub-chip to the source sub-chip; wherein, the identifier of the data stream and the identifier of the destination sub-chip are used to be associated and stored in the flow output table of the source sub-chip; the flow output table is the basis for the source sub-chip to send data.

17. A sub-chip, characterized in that, it includes a processor, a memory and a communication port; wherein, the memory and the communication port are coupled to the processor, the communication port is used for receiving and sending data, the memory is used for storing a computer program, and the processor is used for calling the computer program so that the sub-chip executes the method according to any one of claims 1-6; the sub-chip is a sub-chip among a plurality of sub-chips included in a chip system, and the plurality of sub-chips are connected in a preset topological structure.

18. A sub-chip, characterized in that, it includes a processor, a memory and a communication port; wherein, the memory and the communication port are coupled to the processor, the communication port is used for receiving and sending data, the memory is used for storing a computer program, and the processor is used for calling the computer program so that the sub-chip executes the method according to any one of claims 7-10; The sub-chip is one of the multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology structure.

19. A sub-chip Characterized in that it includes a processor, a memory, and a communication port; wherein, the memory and the communication port are coupled to the processor, the communication port is used for receiving and transmitting data, the memory is used for storing computer programs, and the processor is used for calling the computer programs so that the sub-chip executes the method according to any one of claims 11-13; The sub-chip is one of the multiple sub-chips included in the chip system, and the multiple sub-chips are connected in a preset topology structure.

20. A chip system Characterized in that the chip system includes a source sub-chip, a destination sub-chip, and a controller; wherein, the source sub-chip is the source sub-chip according to claim 14, the destination sub-chip is the destination sub-chip according to claim 15, and the controller is the controller according to claim 16; or the source sub-chip is the sub-chip according to claim 17, the destination sub-chip is the sub-chip according to claim 18, and the controller is the sub-chip according to claim 19.

21. A computer-readable storage medium Characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-6; or when the computer program is executed by a processor, it implements the method according to any one of claims 7-10; or when the computer program is executed by a processor, it implements the method according to any one of claims 11-13.

Citation Information

Patent Citations

  • Method and apparatus for transmitting and receiving data packets in multi-chip communication

    CN103516627A

  • Data processing chip and system, data storage, forwarding and reading processing method

    CN107451075A

  • Data processing device and method

    CN111274193A

  • Task processing method and device and cloud analysis system

    CN111381956A

  • Data handling method and device

    CN112463680A