Data transmission network, data processing method, device and chip
Patent Information
- Application Number
- CN202211459104.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2042-11-21
AI Technical Summary
[0013]根据本公开的一个或多个实施例,能够在硬件实现中简化布线,降低功耗损失。
Smart Images

Figure CN115952828B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to the field of artificial intelligence and chip technology, specifically to a data transmission network, data processing method, apparatus, chip, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] With the popularization of neural networks and graphics processing applications, more and more companies are starting to develop their own domain-specific architecture (DSA) chips, such as neural network processing units (NPUs) or graphics processing units (GPUs), in order to accelerate specific tasks in parallel and thus accelerate application business.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a data transmission network, a data processing method, an apparatus, a chip, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of this disclosure, a data transmission network is provided, comprising: a first node layer, each node in the first node layer being used to read and write data to any storage unit in at least one storage unit corresponding to the node, the first node layer including a first node and a second node; and a second node layer, the second node layer including a third node and a fourth node, each of the first node and the second node being connected to the third node and the fourth node respectively to form a first-level network, so that target data can be transferred from the source storage unit to the target storage unit via a first data path in the first-level network, wherein the source storage unit and the target storage unit are located in multiple storage units corresponding to the first node and the second node.
[0007] According to another aspect of this disclosure, a data processing method is provided, the method comprising: in response to receiving a first instruction, transferring target data from a source storage unit to a target storage unit via the aforementioned data transmission network, wherein the first instruction includes address information of the source storage unit and address information of the target storage unit.
[0008] According to another aspect of this disclosure, a data processing apparatus is provided, the apparatus comprising: a transfer unit configured to, in response to receiving a first instruction, transfer target data from a source storage unit to a target storage unit via the aforementioned data transmission network, wherein the first instruction includes address information of the source storage unit and address information of the target storage unit.
[0009] According to another aspect of this disclosure, a chip is provided that includes at least one of the above-described data transmission network and the above-described data processing device.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the data processing method described above.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described data processing method.
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described data processing method when executed by a processor.
[0013] According to one or more embodiments of this disclosure, wiring can be simplified and power consumption loss can be reduced in hardware implementation.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0017] Figure 2 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0018] Figure 3 This diagram illustrates a data shaping operation in the related technology.
[0019] Figure 4 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0020] Figure 5 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0021] Figure 6 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0022] Figure 7 A schematic diagram of the structure of a node in a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0023] Figure 8 A schematic diagram of the structure of a node in a data transmission network according to an exemplary embodiment of the present disclosure is shown;
[0024] Figure 9 A flowchart of a data processing method 900 according to an exemplary embodiment of the present disclosure is shown;
[0025] Figure 10 A flowchart of a data processing method 1000 according to an exemplary embodiment of the present disclosure is shown;
[0026] Figure 11 A schematic diagram of the structure of a data processing apparatus according to an exemplary embodiment of the present disclosure is shown;
[0027] Figure 12A A schematic diagram illustrating data reading and writing according to exemplary embodiments of the present disclosure is shown;
[0028] Figure 12B A schematic diagram of a data transmission route according to an exemplary embodiment of the present disclosure is shown;
[0029] Figure 13 A structural block diagram of a data processing apparatus 1300 according to an exemplary embodiment of the present disclosure is shown;
[0030] Figure 14 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0031] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0032] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0033] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0034] In related technologies, since each network node in a data transmission network is only connected to its corresponding upstream node, each path in the data transmission network can only support data reading and writing of its corresponding storage block. For data shaping operations, an additional data shaping module must be used to read the data into that module, perform the shaping operation, and then transmit the data to the corresponding storage address through the corresponding data path. The application of the data shaping module results in data shaping operations consuming more bandwidth. Furthermore, it increases the complexity of wiring in the hardware implementation, and the complex wiring further increases power consumption.
[0035] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0036] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0037] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the data processing methods described above.
[0038] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0039] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0040] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to issue data processing commands. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0041] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0042] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0043] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0044] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0045] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0046] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0047] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0048] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0049] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0050] Figure 2 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown.
[0051] According to some embodiments, such as Figure 2 As shown, a data transmission network 200 is provided, including: a first node layer 210, each node in the first node layer 210 being used to read and write data to any storage unit in at least one storage unit corresponding to that node, the first node layer 210 including a first node 211 and a second node 212; and a second node layer 220, the second node layer 220 including a third node 221 and a fourth node 222, each of the first node 211 and the second node 212 being connected to the third node 221 and the fourth node 222 respectively to form a first-level network, so that target data can be transferred from the source storage unit to the target storage unit via a first data path in the first-level network, wherein the source storage unit and the target storage unit are located in the corresponding multiple storage units of the first node 211 and the second node 212.
[0052] This simplifies wiring and reduces power consumption in hardware implementation.
[0053] Data shaping operations are used to transform multi-dimensional data, such as matrix transpose or matrix rotation by 90°. They are a type of data preprocessing operation typically performed before inputting data into a computational unit. The following description uses matrix transpose as an example to illustrate a specific implementation.
[0054] Figure 3 A schematic diagram of a data shaping operation in related technologies is shown.
[0055] In static storage space, each dimension of a matrix is often stored contiguously, such as... Figure 3 As shown, the data (1-16) of a 4×4 matrix are stored in different storage units in multiple storage blocks. The only difference between the different data is their storage address; therefore, data reshaping operations are essentially data transfers from one location to another in static storage space.
[0056] In related technologies, such as Figure 3 As shown, to transpose a 4×4 matrix, each data point in the matrix must first be read into the data shaping module. The data shaping module will concatenate the data of different dimensions into continuous data according to certain output rules, and then store it back into the corresponding storage unit in the static storage space. Then, the calculation unit will read the shaped data from the static storage space to perform the corresponding calculations.
[0057] In other words, in related technologies, multi-level data transmission networks are used only for data reading and writing. Each node in the data transmission network is connected to only one superior node and one subordinate node to complete unidirectional data transmission, while data shaping operations are entirely concentrated in the data shaping module. The amount of data that needs to be shaped is often very large, and the on-chip static memory is also often very large, which results in high transmission latency during data reading and writing. At the same time, it consumes a lot of area for chip design, and the power consumption caused by data movement is not conducive to low-power design.
[0058] In some embodiments, such as Figure 2 As shown, taking the first node 211 as an example, it is designed to connect to the third node 221 and the fourth node 222 respectively, thereby enabling data transmission with the two nodes. Each of the aforementioned nodes includes at least one buffer register for storing the data received by that node.
[0059] In some embodiments, each node in the first node layer 210 is used to read and write data to any storage unit in at least one storage unit corresponding to that node. Each node's at least one storage unit can constitute a storage block, and all storage blocks can constitute a storage array. Each storage block contains at least one storage unit with the same column address and distinct row addresses.
[0060] like Figure 2 As shown, the first node 211 and the second node 212 are used to read and write data to each storage unit in storage block 230 and storage block 240, respectively. Each node in the first node 211 and the second node 212 is connected to the third node 221 and the fourth node 222 to form a first-level network. The first-level network includes multiple data paths, and any data stored in storage block 230 or storage block 240 can be stored in any storage unit in storage block 230 or storage block 240 through one of the data paths in the first-level network.
[0061] In one example, the source storage unit 231 of the target data is located in storage block 230, and the target storage unit 241 is located in storage block 240. In response to receiving an instruction to transfer the target data, the target data can be read from the source storage unit 231 by the first node 211, and stored in the target storage unit 241 via a first data path of "first node 211 → fourth node 222 → second node 212". The instruction may include the address information of the source and target storage units of the target data, and the data transfer direction of the first data path and each node in the path is determined based on the address information of the source and target storage units.
[0062] Therefore, the aforementioned data transmission network enables data shaping operations on two storage blocks. The relevant controller in the chip only needs to obtain the target address and source address of each target data based on the data shaping instructions to perform the corresponding data shaping operation based on the aforementioned data transmission network. This eliminates the need for additional data shaping units, saving chip area, reducing hardware design complexity, and reducing power consumption.
[0063] Figure 4 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown.
[0064] In some embodiments, such as Figure 4 As shown, to enable data transfer between more storage blocks, a data transfer network 400 is provided, wherein a first node layer 410 and a second node layer 420 may include multiple first-level networks. Each node in the first node layer 410 and the second node layer 420 is used for one of the multiple first-level networks (that is, each node is only used to form one first-level network and is not reused). The multiple first-level networks include a first network 401 and a second network 402.
[0065] The data transmission network 400 may further include a third node layer 430, which includes a first node group 431 and a second node group 432. The first node group 431 and the second node group 432 are used to cross-connect the first network 401 and the second network 402 to form a second-level network. The first node group 431 includes a plurality of fifth nodes corresponding to a plurality of nodes in the second node layer 420 of the first network 401, and the second node group 432 includes a plurality of sixth nodes corresponding to a plurality of nodes in the second node layer 420 of the second network 402. The cross-connection includes: connecting a plurality of nodes in the second node layer 420 of the first network 401 to a plurality of fifth nodes and to a plurality of sixth nodes respectively; and connecting a plurality of nodes in the second node layer 420 of the second network 402 to a plurality of fifth nodes and to a plurality of sixth nodes respectively.
[0066] That is, for the two first-level networks described above, the first network 401 and the second network 402 are cross-connected through corresponding nodes in the third node layer 430. Specifically, multiple nodes (third node 421 and fourth node 422) in the second node layer 420 of the first network 401 can be connected to multiple fifth nodes (node 431-1 and node 431-2), and at the same time, the third node 421 and the fourth node 422 can be connected to multiple sixth nodes (node 432-1 and node 432-2), respectively; similarly, multiple nodes (third node 423 and fourth node 424) in the second node layer 420 of the second network 402 can be connected to the aforementioned fifth and sixth nodes, thereby forming a second-level network.
[0067] The second-level network includes multiple data paths. Any data stored in storage blocks 440, 450, 460, and 470 can be stored in any storage unit of any storage block in storage blocks 440-470 via one of the data paths in the second-level network 404.
[0068] In one example, the source storage unit 471 of the target data is located in storage block 470, and the target storage unit 444 is located in storage block 440. In response to receiving an instruction to transfer the target data, the target data can be read from the source storage unit 471 by the second node 414, and stored in the target storage unit 444 via a second data path: "Second node 414 → Third node 423 → Node 431-1 → Third node 421 → First node 411". The instruction may include the address information of the source and target storage units of the target data, and the data transfer direction of the first data path and each node in the path is determined based on the address information of the source and target storage units.
[0069] Therefore, when the source storage unit and the target storage unit are storage units corresponding to different first-level networks, data interconnection between the two first-level networks is achieved by utilizing nodes at the next higher level of the first-level network, thereby enabling data to be transmitted between more storage blocks and allowing the network to support data shaping operations of higher-dimensional matrices.
[0070] In some embodiments, the data transmission network may further include: multiple node layers, including a fourth node layer, the fourth node layer including a third node group and a fourth node group, the third node group and the fourth node group being used to cross-connect the third network and the fourth network, the third network and the fourth network being the subordinate networks corresponding to the third node group and the fourth node group respectively.
[0071] Therefore, by further expanding the data transmission network, it can be used for higher-dimensional data shaping.
[0072] Figure 5 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown.
[0073] In some exemplary embodiments, to further extend the data transmission network described above, a node layer 520 can be added above the third node layer. The two second-level networks (network 511 and network 512) described above then become the lower-level networks of this node layer 520. This node layer 520 may include a third node group 521 and a fourth node group 522. Multiple nodes in each group are cross-connected with the highest-level nodes in networks 511 and 512 (i.e., the nodes in the third node layer of each second-level network) in a similar manner to the above, thereby obtaining a data transmission matrix capable of being used for higher-dimensional data shaping.
[0074] Data shaping operations often handle high-dimensional data, such as transposing a matrix of several hundred dimensions. Based on methods similar to the cross-connects described above, the scale of the data transmission network can be further expanded, enabling it to meet the demands of data processing. Figure 6 A schematic diagram of the structure of a data transmission network according to an exemplary embodiment of the present disclosure is shown. Figure 6 As shown, the data transmission network 600 can, for example, perform integer operations such as transposing a 32×32 dimensional matrix.
[0075] In the network described above, since most nodes have two inputs and two outputs, data collisions will occur when two nodes connected to a node transmit data to that node simultaneously, resulting in performance loss.
[0076] In some embodiments, each node in the data transmission network may include an arbitrator and multiple buffer registers, the arbitrator being used to determine the data to be output by the node at the next moment, provided that data is stored in at least two of the multiple buffer registers.
[0077] Therefore, by setting up an arbitrator in each node, the problem of data conflicts can be resolved to some extent.
[0078] Figure 7 A schematic diagram of the structure of a node in a data transmission network according to an exemplary embodiment of the present disclosure is shown.
[0079] like Figure 7 As shown, the node internally may include buffer registers 701 and 702, and an arbitrator 703. When two upstream nodes transmit data to this node at the same time, the two data entries can be stored in buffer registers 701 and 702 respectively. The data to be output by the node at the next moment is determined based on the arbitrator 703. The arbitrator 703 has two outputs, corresponding to the two data output directions of this node. The arbitrator can be, for example, a multiplexer.
[0080] To further reduce transmission losses, the arbitration window can be further expanded. In some embodiments, each node in the data transmission network may include multiple primary arbitrators and two secondary arbitrators. Each of the multiple primary arbitrators is connected to multiple buffer registers, and each primary arbitrator is connected to two secondary arbitrators. The data output directions of the two secondary arbitrators correspond to the two data output directions of the node, respectively. Where data is registered in at least two of the corresponding buffer registers of the multiple primary arbitrators, the multiple primary arbitrators determine, based on the data transmission direction of each data, the data to be output by each of the multiple primary arbitrators at the next moment and the output direction of that data, wherein the data transmission direction is determined based on the address information of the target storage unit of the data. Furthermore, each of the two secondary arbitrators performs the following operations: when the secondary arbitrator receives one data at the next moment, it transmits the data to the next node according to the corresponding data output direction; and when the secondary arbitrator receives multiple data at the next moment, it selects a first data from the multiple data and transmits the first data to the next node according to the corresponding data output direction.
[0081] Therefore, by further expanding the arbitration window, the probability of data conflicts can be further reduced, thereby improving the quality and efficiency of data transmission.
[0082] Figure 8A schematic diagram of the structure of a node in a data transmission network according to an exemplary embodiment of the present disclosure is shown.
[0083] In some exemplary embodiments, the node may include, for example, a primary arbitrator 821 and a primary arbitrator 822. Primary arbitrator 821 is connected to buffer registers 811 and 812, respectively, and primary arbitrator 822 is connected to buffer registers 813 and 814, respectively. The node also includes secondary arbitrators 831 and 832. Each of primary arbitrators 821 and 822 is connected to secondary arbitrators 831 and 832, respectively. Furthermore, the data output direction of secondary arbitrators 831 and 832 corresponds to the data output direction of the node; that is, the two secondary arbitrators can output data to the two downstream nodes connected to the node.
[0084] In actual data transmission, there may be a situation where all four buffer registers contain data at a certain moment. In the next moment, based on the data transmission direction of each data (determined by the target storage address of the data), the first-level arbitrator determines which data to transmit to the corresponding second-level arbitrator. For example, if the data transmission directions in buffer registers 811, 812, 813, and 814 are direction 801, direction 802, direction 801, and direction 801 respectively, then the first-level arbitrator 822 can select buffer registers 813 and 814... Any data in buffer register 814 is transmitted to secondary arbitrator 831, which then forwards the data to the corresponding downstream node. Primary arbitrator 821 can select data to transmit based on the transmission direction of each data, avoiding data conflicts (i.e., both primary arbitrators sending data to the same secondary arbitrator simultaneously). In this example, primary arbitrator 821 can select data in buffer register 812 (transmission direction 802) to transmit to secondary arbitrator 832, which then forwards the data to the corresponding downstream node. This avoids data conflicts and improves the overall performance of the data transmission network.
[0085] In some cases, when two primary arbitrators inevitably send data to the same secondary arbitrator simultaneously, the secondary arbitrator can select one of the data to continue the transmission. The data that is not successfully transmitted can continue to be stored in the corresponding register, thereby avoiding data errors and loss.
[0086] It can be seen that for the network node in the above exemplary embodiment, a certain performance loss will only occur when the data transmission direction in the four buffer registers is the same, and the probability of this occurrence is only 1 / 16 of all cases. Furthermore, in actual application scenarios, each node will only receive 1-2 data at the same time in most cases. Therefore, the network node structure in the above embodiment can effectively avoid the occurrence of data conflicts, thereby improving the overall performance of the data transmission network.
[0087] Figure 9 A flowchart of a data processing method 900 according to an exemplary embodiment of the present disclosure is shown.
[0088] In some embodiments, such as Figure 9 As shown, method 900 includes:
[0089] Step 901: In response to receiving the first instruction, the target data is transferred from the source storage unit to the target storage unit via the aforementioned data transmission network, wherein the first instruction includes the address information of the source storage unit and the address information of the target storage unit.
[0090] Therefore, by applying the data transmission network described above, it is only necessary to determine the source address of each data and the destination address after data shaping. Data shaping can be completed based on this data transmission network, thereby avoiding excessive bandwidth consumption during data shaping. At the same time, it simplifies wiring and reduces power consumption in hardware implementation.
[0091] Figure 10 A flowchart of a data processing method 1000 according to an exemplary embodiment of the present disclosure is shown.
[0092] In some embodiments, such as Figure 10 As shown, method 1000 includes:
[0093] Step 1001: In response to receiving the second instruction, each target data in at least one target data is transmitted from the corresponding source storage unit to the corresponding register in the computing unit via the data transmission network, so that the computing unit performs the corresponding calculation and stores the corresponding calculation result of the target data in the corresponding target storage unit via the data transmission network. The second instruction includes the address information of the source storage unit and the address information of the target storage unit for each target data in at least one target data.
[0094] In related technologies, data shaping and parallel computing are executed serially. That is, the data shaping module reads data from static storage, then stores the shaped data back into static storage. Subsequently, the computing unit reads the shaped data and stores the calculation results back into static storage. This pipelined operation is time-consuming, and the access to static storage by the data shaping module and computing unit consumes significant memory bandwidth, affecting the access bandwidth of other modules and consequently impacting the overall chip performance.
[0095] Figure 11 A schematic diagram of the structure of a data processing apparatus according to an exemplary embodiment of the present disclosure is shown.
[0096] According to some exemplary embodiments, such as Figure 11 As shown, the aforementioned data transmission network can be set between the static memory and the computing unit, so that when the computing unit obtains each target data in the computing instruction (second instruction) and its corresponding source memory unit and target memory unit address information, it can read the data in the corresponding address into the corresponding register in the computing unit through the corresponding data path in the data transmission network, perform calculations based on the data, and directly save the calculation results to the corresponding target memory unit.
[0097] Therefore, by applying the aforementioned data transmission network to the computation process, the computing unit obtains computation instructions, thereby acquiring the source and destination addresses of the data required for each computation. This enables the computing unit to directly obtain the shaped data and store the computation results in the corresponding storage address, thereby reducing the bandwidth occupied by data shaping in memory access, saving module synchronization time, and improving the overall efficiency of data processing.
[0098] To further ensure data transmission performance and avoid data conflicts, in some embodiments, the above data processing method may further include: receiving a third instruction, the third instruction being used to perform data shaping on multiple target data, the third instruction including the address information of the source storage unit of each target data in the multiple target data, the multiple target data being stored in a first storage block and a second storage block respectively; and parsing the third instruction to determine the target data reading order, the target data reading order being used to enable the first node and the second node to read data from the corresponding storage block according to a preset rule, the preset rule including that at the same time, the data read by the first node and the data read by the second node have different row addresses.
[0099] The data transmission network includes a first node and a second node, which correspond to a first storage block and a second storage block in the storage array, respectively. Each storage block in the first and second storage blocks includes at least one storage unit. Each storage unit in each storage block has the same column address but different row addresses. Each node in the first and second nodes is used to read and write data to any storage unit in the storage block corresponding to that node.
[0100] Figure 12A A schematic diagram illustrating data reading and writing according to an exemplary embodiment of the present disclosure is shown. Figure 12B A schematic diagram of a data transmission route according to an exemplary embodiment of the present disclosure is shown.
[0101] In some exemplary embodiments, such as Figure 12A As shown, for a 4×4 matrix transpose operation, each time a node in the first node layer reads data from its corresponding storage block at the same time, it selects data from different row addresses to read. For example, it reads data from storage cell 1212 (stores data 5) in storage block 1210, storage cell 1221 (stores data 2) in storage block 1220, storage cell 1234 (stores data 15) in storage block 1230, and storage cell 1243 (stores data 12) in storage block 1240. The data transmission path for these four data points is as follows: Figure 12B As shown, none of the four data transmission routes will experience data conflicts at any node. Therefore, by determining the order in which data is read, data conflicts can be avoided.
[0102] Figure 13 A structural block diagram of a data processing apparatus 1300 according to an exemplary embodiment of the present disclosure is shown.
[0103] In some embodiments, such as Figure 13 As shown, the data processing apparatus 1300 includes a transfer unit 1310 configured to transfer target data from a source storage unit to a target storage unit via the aforementioned data transmission network in response to receiving a first instruction, wherein the first instruction includes address information of the source storage unit and address information of the target storage unit.
[0104] The operation of unit 1310 of data processing device 1300 is similar to the operation of step S901 in method 900 described above, and will not be repeated here.
[0105] In some embodiments, the data processing apparatus may further include: a transmission unit configured to, in response to receiving a second instruction, transmit each of the at least one target data from a corresponding source storage unit to a corresponding register in a computing unit via a data transmission network, so that the computing unit performs a corresponding calculation and stores the corresponding calculation result of the target data in a corresponding target storage unit via the data transmission network, wherein the second instruction includes address information of the source storage unit corresponding to each of the at least one target data and address information of the target storage unit.
[0106] In some embodiments, the data transmission network includes a first node and a second node, which correspond to a first storage block and a second storage block in the storage array, respectively. Each storage block in the first and second storage blocks includes at least one storage unit. Each storage unit in each storage block has the same column address but different row addresses. Each node in the first and second nodes is used to read and write data to any storage unit in the storage block corresponding to that node.
[0107] The data processing apparatus may further include: a receiving unit configured to receive a third instruction for data shaping of multiple target data, the third instruction including address information of the source storage unit of each target data in the multiple target data, the multiple target data being stored in a first storage block and a second storage block respectively; and a parsing unit configured to parse the third instruction so that a first node and a second node read data from the corresponding storage block at the same time according to a preset rule, the preset rule including that at the same time, the data read by the first node and the data read by the second node have different row addresses.
[0108] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0109] refer to Figure 14 The present invention describes a structural block diagram of an electronic device 1400 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0110] like Figure 14 As shown, the electronic device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. The RAM 1403 may also store various programs and data required for the operation of the electronic device 1400. The computing unit 1401, ROM 1402, and RAM 1403 are interconnected via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0111] Multiple components in electronic device 1400 are connected to I / O interface 1405, including: input unit 1406, output unit 1407, storage unit 1408, and communication unit 1409. Input unit 1406 can be any type of device capable of inputting information to electronic device 1400. Input unit 1406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1407 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1408 may include, but is not limited to, a hard disk and an optical disk. The communication unit 1409 allows the electronic device 1400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0112] The computing unit 1401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1401 performs the various methods and processes described above, such as method 900. For example, in some embodiments, method 900 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by the computing unit 1401, one or more steps of method 900 described above may be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured to execute method 900 by any other suitable means (e.g., by means of firmware).
[0113] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0114] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0117] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0118] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0119] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0120] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A data transmission network for data shaping, comprising: The first node layer, each node in the first node layer is used to read and write data to any storage unit in the storage block corresponding to the node. The first node layer includes a first node and a second node. The storage blocks corresponding to each node in the first node layer belong to the same storage array. The first node and the second node correspond to the first storage block and the second storage block in the storage array, respectively. Each storage block in the first storage block and the second storage block includes at least one storage unit. The column address of each storage unit in each storage block is the same, but the row address is different. as well as The second node layer includes a third node and a fourth node. Each of the first and second nodes is connected to the third and fourth nodes respectively to form a first-level network. This allows target data to be read from the source storage unit via nodes in the first node layer, and written to the target storage unit via a first data path in the first-level network. The source storage unit and the target storage unit are located in the first storage block and the second storage block, respectively. The first node and the second node are configured to, upon receiving an instruction for data shaping of multiple target data, read data from corresponding storage blocks according to preset rules. These preset rules include that, at the same time, the data read by the first node and the data read by the second node have different row addresses, and wherein... Each node in the data transmission network includes multiple primary arbitrators and two secondary arbitrators. Each primary arbitrator is connected to multiple buffer registers, and each primary arbitrator is connected to the two secondary arbitrators. The data output directions of the two secondary arbitrators correspond to the two data output directions of the corresponding node. The plurality of primary arbitrators are used to determine, when data is stored in at least two of the corresponding buffer registers of the plurality of primary arbitrators, the data to be output by each of the plurality of primary arbitrators at the next time step and the output direction of that data based on the data transmission direction, wherein the data transmission direction is determined based on the address information of the target storage unit of the data; and, Each of the two secondary arbitrators is used to perform the following operations: If the secondary arbitrator receives data at the next moment, it will transmit the data to the next node according to the corresponding data output direction; and If the secondary arbitrator receives multiple data at the next moment, it selects the first data from the multiple data and transmits the first data to the next node according to the corresponding data output direction.
2. The network according to claim 1, wherein, Between the first node layer and the second node layer, there are multiple first-level networks. Each node in the first node layer and the second node layer is used in one of the multiple first-level networks. The multiple first-level networks include a first network and a second network. The network further includes: The third node layer includes a first node group and a second node group. The first node group and the second node group are used to cross-connect the first network and the second network to form a second-level network. The first node group includes multiple fifth nodes corresponding to multiple nodes in the second node layer of the first network, and the second node group includes multiple sixth nodes corresponding to multiple nodes in the second node layer of the second network. The cross-connection includes: Connect multiple nodes in the second node layer of the first network to the multiple fifth nodes, and connect them to the multiple sixth nodes; and Multiple nodes in the second node layer of the second network are respectively connected to the multiple fifth nodes, and respectively connected to the multiple sixth nodes.
3. The network according to claim 2, further comprising: Multiple node layers, including a fourth node layer, the fourth node layer including a third node group and a fourth node group, the third node group and the fourth node group being used to perform the cross-connection of the third network and the fourth network, the third network and the fourth network being the subordinate networks corresponding to the third node group and the fourth node group respectively.
4. A data processing method, the method comprising: In response to receiving a first instruction, target data is transferred from a source storage unit to a target storage unit via a data transmission network as described in any one of claims 1-3, wherein the first instruction includes address information of the source storage unit and address information of the target storage unit.
5. The method according to claim 4, further comprising: In response to receiving a second instruction, each of the at least one target data is transmitted from its corresponding source storage unit to its corresponding register in the computing unit via the data transmission network, so that the computing unit performs the corresponding calculation and stores the calculation result of the target data in its corresponding target storage unit via the data transmission network. The second instruction includes the address information of the source storage unit and the address information of the target storage unit for each of the at least one target data.
6. The method according to claim 4 or 5, wherein the data transmission network includes a first node and a second node, the first node and the second node respectively corresponding to a first storage block and a second storage block in the storage array, each storage block in the first storage block and the second storage block includes at least one storage unit, each storage unit in each storage block has the same column address but different row addresses, each of the first node and the second node is used to read data and write data to any storage unit in the storage block corresponding to that node, the method further comprising: A third instruction is received, the third instruction being used to perform data shaping on multiple target data, the third instruction including the address information of the source storage unit of each of the multiple target data, the multiple target data being stored in the first storage block and the second storage block respectively; as well as The third instruction is parsed to determine the target data reading order, which is used to enable the first node and the second node to read data from the corresponding storage blocks according to a preset rule. The preset rule includes that at the same time, the data read by the first node and the data read by the second node have different row addresses.
7. A data processing apparatus, the apparatus comprising: The transfer unit is configured to, in response to receiving a first instruction, transfer target data from a source storage unit to a target storage unit via a data transmission network as described in any one of claims 1-3, wherein the first instruction includes address information of the source storage unit and address information of the target storage unit.
8. The apparatus according to claim 7, further comprising: The transmission unit is configured to, in response to receiving a second instruction, transmit each target data in at least one target data from its corresponding source storage unit to a corresponding register in the computing unit via the data transmission network, so that the computing unit performs a corresponding calculation and then stores the corresponding calculation result of the target data in the corresponding target storage unit via the data transmission network, wherein the second instruction includes address information of the source storage unit and address information of the target storage unit for each target data in the at least one target data.
9. The apparatus according to claim 7 or 8, wherein the data transmission network includes a first node and a second node, the first node and the second node respectively corresponding to a first storage block and a second storage block in the storage array, each storage block in the first storage block and the second storage block includes at least one storage unit, each storage unit in each storage block has the same column address but different row addresses, each of the first node and the second node is used to read data and write data to any storage unit in the storage block corresponding to that node, the apparatus further comprising: The receiving unit is configured to receive a third instruction, which is used to perform data shaping on multiple target data. The third instruction includes the address information of the source storage unit of each of the multiple target data, and the multiple target data are stored in the first storage block and the second storage block respectively. as well as The parsing unit is configured to parse the third instruction so that the first node and the second node read data from the corresponding storage blocks at the same time according to a preset rule, wherein the preset rule includes that at the same time, the data read by the first node and the data read by the second node have different row addresses.
10. A chip comprising at least one of the following: The data transmission network as described in any one of claims 1-3; and The apparatus as described in any one of claims 7-9.
11. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 4-6.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 4-6.
13. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 4-6.
Citation Information
Patent Citations
Data transmission method and device, electronic device and readable storage medium
WO2021022441A1