Data Processing Method, Apparatus, Chip, Device and Medium
By combining the operation unit reading instructions in the neural network processor, the DRAM bandwidth bottleneck problem is solved, and the dynamic allocation and performance improvement of the operation unit is achieved.
Patent Information
- Application Number
- CN202211689557.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-27
AI Technical Summary
In the prior art, neural network processors have DRAM bandwidth bottlenecks when accessing data, especially in the dynamic allocation scenario of computing units, which cannot effectively reduce the transmission of duplicate data, resulting in performance bottlenecks.
By monitoring the read instructions of the calculation unit, the calculation units that read the same instructions within a preset time are merged into a group, and the combined data read instructions are sent to the on-chip network to avoid repeated data transmission to unnecessary calculation units.
When reading duplicate data, the neural network processor saves data access to DRAM, supports dynamic allocation of computing units, and improves overall processing performance.
Smart Images

Figure CN115796254B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, in particular to the field of artificial intelligence and chip technology, and specifically to a data processing method, device, chip, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0003] Neural networks are widely used in various AI scenarios, such as speech recognition, image recognition, and natural language processing. Because neural network operations involve numerous matrix multiplications and convolutions, neural network processors (NPUs), specifically designed to accelerate these operations, can significantly increase the processing speed of various AI applications and are gradually gaining wider application.
[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0006] According to one aspect of the present disclosure, a data processing method for a processor including a plurality of arithmetic units is provided, including: in response to receiving a plurality of program read instructions that meet a first preset condition, determining a plurality of first arithmetic units among the plurality of arithmetic units, wherein the plurality of program read instructions respectively come from the plurality of first arithmetic units, and the first preset condition includes that the time difference between the first instruction and the last instruction among the plurality of program read instructions is less than a first preset time and the plurality of program read instructions have the same program read address; forwarding the plurality of program read instructions so that each of the plurality of first arithmetic units among the plurality of first arithmetic units obtains and executes the corresponding program, and the program includes a data read instruction; in response to receiving a plurality of first data read instructions respectively from the plurality of first arithmetic units, merging the plurality of first data read instructions into a second data read instruction, wherein the plurality of first data read instructions and the second data read instruction have the same data read address; and based on the second data read instruction, obtaining the data in the data read address to send the data to the plurality of first arithmetic units respectively.
[0007] According to another aspect of the present disclosure, a data processing device for a processor including a plurality of arithmetic units is provided, including: a determination unit configured to, in response to receiving a plurality of program read instructions that meet a first preset condition, determine a plurality of first arithmetic units among the plurality of arithmetic units, wherein the plurality of program read instructions respectively come from the plurality of first arithmetic units, and the first preset condition includes that the time difference between the first instruction and the last instruction among the plurality of program read instructions is less than a first preset time and the plurality of program read instructions have the same program read address; a forwarding unit configured to forward the plurality of program read instructions so that each of the plurality of first arithmetic units among the plurality of first arithmetic units obtains and executes the corresponding program, and the program includes a data read instruction; a merging unit configured to, in response to receiving a plurality of first data read instructions respectively from the plurality of first arithmetic units, merge the plurality of first data read instructions into a second data read instruction, wherein the plurality of first data read instructions and the second data read instruction have the same data read address; and an obtaining unit configured to, based on the second data read instruction, obtain the data in the data read address to send the data to the plurality of first arithmetic units respectively.
[0008] According to another aspect of the present disclosure, a chip is provided, including the above data processing device.
[0009] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above data processing method.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above data processing method.
[0011] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program implements the above data processing method when executed by a processor.
[0012] According to one or more embodiments of the present disclosure, when reading duplicate data, it is possible to save the data access amount of a chip (such as a neural network processor) to a DRAM (Dynamic Random Access Memory), and at the same time support the dynamic allocation of arithmetic units, avoiding sending duplicate data to arithmetic units that do not require such data.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure is shown;
[0016] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;
[0017] Figure 3 A schematic diagram showing matrix multiplication calculation performed based on a neural network processor is shown;
[0018] Figure 4 A structural block diagram of a processor according to an exemplary embodiment of the present disclosure is shown;
[0019] Figure 5 A structural block diagram of a memory access merging unit according to an embodiment of the present disclosure is shown;
[0020] Figure 6 A flowchart of a method for determining a plurality of first arithmetic units according to an embodiment of the present disclosure is shown;
[0021] Figure 7 The flowchart of the determination method of a plurality of first operation units according to an embodiment of the present disclosure is shown;
[0022] Figure 8 The flowchart of the merging method of the second data reading instruction according to an embodiment of the present disclosure is shown;
[0023] Figure 9 The structural block diagram of a data processing device according to an embodiment of the present disclosure is shown;
[0024] Fig.10 The structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure is shown. Detailed implementation manners
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0026] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the description of the context, they may also refer to different instances.
[0027] The terms used in the description of various examples in the present disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0028] A large number of operation units that can work concurrently are integrated inside the NPU, and various calculations for matrices and vectors can be completed quickly. In many actual business scenarios, the operation units access the memory controller through the network on chip to obtain data and complete the calculations. The network on chip (NOC) sends the requests of each operation unit to the corresponding memory controller according to the requested address, and then the memory controller issues commands to the DRAM storage to complete the data access.
[0029] Since the DRAM storage is located outside the NPU chip, the data bandwidth is limited by the traces on the printed circuit board and is much lower than the bandwidth achievable between the arithmetic units and the NOC. Therefore, transferring data from the DRAM often becomes the most time-consuming part of the entire computing process and also becomes the bottleneck of the overall system performance.
[0030] In related technologies, in order to improve the bottleneck problem of the memory access bandwidth, various methods for reducing the amount of memory-accessed data have been tried in terms of software and hardware, including placing data that may be repeatedly accessed in the cache, and completing the broadcast of memory-accessed data through a combination of software and hardware.
[0031] Among them, the method of improving memory access through on-chip cache, although no additional adjustment is required, requires a relatively large cache to store the repeated matrix, resulting in a large hardware overhead, and the space occupied when reading the repeated matrix may also affect the residency of other data in the cache, reducing the hit rate of accessing other data; the method of improving memory access through the broadcast module and the method of improving memory access through the multicast synchronization module cannot solve the scenario when only some arithmetic units need to broadcast (considering the resource partitioning on the NPU, some arithmetic units on the NPU will be dynamically allocated during software execution, and this information cannot be obtained when writing the program).
[0032] According to an embodiment of the present disclosure, a data processing method is provided. By monitoring the requests of arithmetic units to read instructions, and in response to multiple arithmetic units reading the same instruction within a preset time, these multiple arithmetic units are combined into a group; subsequently, further monitor the instructions of each arithmetic unit in the group to read data. After receiving the instructions of all units in the group to read repeated data, the read instructions are combined and sent to the on-chip network, and the read data is sent to each arithmetic unit in the group respectively. Thereby, while saving the data access amount of the chip (such as a neural network processor) to the DRAM when reading repeated data, it supports the dynamic allocation of arithmetic units and avoids sending repeated data to arithmetic units that do not need this data.
[0033] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0034] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Referring Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more application programs.
[0035] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of the above-described data processing method.
[0036] In some embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software-as-a-service (SaaS) model.
[0037] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may differ from the system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0038] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to issue tasks to the processor. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0039] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0040] Network 110 can be any type of network known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0041] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0042] The computing units in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0043] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0044] In some embodiments, server 120 can be a server of a distributed system, or a server combined with a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0045] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0046] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0047] Figure 1The system 100 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with the present disclosure.
[0048] Figure 2 FIG. 2 is a flow chart of a data processing method 200 according to an embodiment of the present disclosure.
[0049] According to an embodiment of the present disclosure, the method 200 may be used in a processor including a plurality of operation units, and as shown in FIG. Figure 2 As shown, the method 200 includes:
[0050] Step S201: In response to receiving a plurality of program read instructions that meet a first preset condition, determining a plurality of first operation units from a plurality of operation units, wherein the plurality of program read instructions are respectively from the plurality of first operation units, the first preset condition including that a time difference between receiving a first instruction and a last instruction of the plurality of program read instructions is less than a first preset time and the plurality of program read instructions have the same program read address;
[0051] Step S202: forwarding a plurality of program read instructions so that each of the plurality of first operation units acquires and executes a corresponding program, wherein the program includes a data read instruction;
[0052] Step S203: In response to receiving a plurality of first data read instructions respectively from a plurality of first operation units, merging the plurality of first data read instructions into a second data read instruction, wherein the plurality of first data read instructions and the second data read instruction have the same data read address; and
[0053] Step S204 : Based on the second data read instruction, obtain data in the data read address to send the data to the plurality of first operation units respectively.
[0054] Thus, by monitoring the request for the operation unit to read the instruction, and in response to multiple operation units reading the same instruction within a preset time, the multiple operation units are merged into a group; then, the instruction for each operation unit in the group to read data is further monitored, and in response to receiving the instruction for all units in the group to read duplicate data, the read instructions are merged and sent to the on-chip network, and the read data is sent to each operation unit in the group. In this way, when reading duplicate data, the chip (such as a neural network processor) can save the amount of data access to DRAM, while supporting the dynamic allocation of operation units, avoiding sending duplicate data to operation units that do not need the data.
[0055] In some embodiments, the processor may be a Neural Processing Unit (NPU), or any processing chip with multiple arithmetic units, which is not limited herein. Hereinafter, the NPU will be taken as an example to elaborate on the method provided by the embodiments of the present disclosure.
[0056] Figure 3 A schematic diagram showing matrix multiplication calculation performed based on a neural network processor is shown.
[0057] Taking Figure 3 as an example, when the NPU performs matrix multiplication, the entire operation will be split onto, for example, four arithmetic units. Among them, for matrix A, each of the four arithmetic units reads 1 / 4; for matrix C, each of the four arithmetic units writes 1 / 4; while matrix B needs to be read by all arithmetic units separately once.
[0058] It can be seen that in the above operation process, each data of matrix B has to be read 1 time by each arithmetic unit, and in this example, it has to be read 4 times. If matrix B is only read once from the DRAM storage and distributed to 4 arithmetic units on the chip, the data volume accessed from the DRAM during the operation can be significantly reduced. The above matrix B can be used as a repeated matrix or repeated data, and the DRAM bandwidth bottleneck can be improved by improving the reading operation of the repeated matrix.
[0059] In the related art, the above problems are mainly improved by placing the data that may be repeatedly accessed in the cache and by the method of combining software and hardware to complete the broadcast of the accessed data. Among them, the method of improving memory access through the on-chip cache, although no additional adjustment is required, requires a relatively large cache to store the repeated matrix, resulting in a large hardware overhead, and the space occupied when reading the repeated matrix may also affect the residence of other data in the cache, reducing the hit rate of accessing other data; while the methods of improving memory access through the broadcast module and the multicast synchronization module cannot be applied in the scenario of dynamic allocation of arithmetic units.
[0060] Specifically, when writing an NPU program to process a computing task, generally speaking, the number of arithmetic units in the program block, the number of arithmetic units called by the NPU during program execution, and the number of arithmetic units included in the NPU are not always the same. Because the NPU often processes multiple tasks in parallel, when actually issuing tasks, the NPU will decide which arithmetic units the new task will be assigned to for execution according to the status of the current arithmetic units.
[0061] In some exemplary embodiments, taking Figure 3Taking the matrix multiplication task shown in the figure as an example, the program divides the task into 4 blocks for calculation. At the same time, there are generally 4 computing units inside the NPU. However, when the task is actually issued, only 2 of the computing units may be idle. At this time, the task will only be issued to these 2 computing units. In other words, each of the 2 computing units actually needs to process 2 blocks in sequence to complete the matrix calculation. The detailed calculation process can be:
[0062] 1) The host sends the matrix multiplication program to the NPU and marks cu_num = 4 to indicate that the program is divided into 4 blocks of calculation;
[0063] 2) The task scheduling module of the NPU detects that the operation unit 1 and the operation unit 2 among the four operation units are in an idle state, and determines to assign the task to these two units;
[0064] 3) The task scheduling module first sends the block parameters cu_idx = 0, cu_num = 4, and the address of the program code in DRAM to operation unit 1, so that operation unit 1 performs the calculation of the first block; at the same time, it sends the block parameters cu_idx = 1, cu_num = 4, and the address of the program code in DRAM to operation unit 2, so that operation unit 2 performs the calculation of the second block;
[0065] 4) Operation unit 1 obtains the program code according to the address and performs the calculation of block A[0]*B according to the configured block parameter cu_idx=0 to obtain C[0];
[0066] 5) Operation unit 2 obtains the program code according to the address and performs the calculation of block A[1]*B according to the configured block parameter cu_idx=1 to obtain C[1];
[0067] 6) After receiving the signal that the calculation unit 1 has completed the calculation, the task scheduling module sends the block parameters cu_idx = 2, cu_num = 4, and the address of the program code in the DRAM to the calculation unit 1, so that the calculation unit 1 performs the calculation of the third block;
[0068] 7) Operation unit 1 obtains the program code according to the address and performs the calculation of block A[2]*B according to the configured block parameter cu_idx=2 to obtain C[2];
[0069] 8) The task scheduling module receives the signal that the calculation unit 2 has completed the calculation, and sends the block parameters cu_idx=3, cu_num=4, and the address of the program code in the DRAM to the calculation unit 2, so that the calculation unit 2 performs the calculation of the fourth block;
[0070] 9) Operation unit 2 obtains the program code according to the address and performs the calculation of block A[3]*B according to the configured block parameter cu_idx=3 to obtain C[3];
[0071] 10) After receiving the signal that the calculation unit 1 and the calculation unit 2 have completed the calculation, the task scheduling module can confirm that the calculation of the four blocks has been completed and notify the host side that the task is completed.
[0072] For the scenario where only some computing units execute the same task, the method of improving memory access by broadcast module and the method of improving memory access by multicast synchronization module cannot be confirmed.
[0073] Figure 4 A structural block diagram of a processor according to an exemplary embodiment of the present disclosure is shown.
[0074] like Figure 4 As shown, the processor 400 can be, for example, a neural network processor, including a task scheduling unit 410, four computing units 420, a memory access merging unit 430, an on-chip network 440 and three memory controllers 450, and the processor 400 is connected to multiple DRAM memories outside the chip through the three memory controllers 450 to read the data stored therein.
[0075] In some embodiments, the data processing method of the present disclosure can be performed based on the memory access merging unit 430. Specifically, the memory access merging unit 430 can be set between the operation unit and the on-chip network, and its application monitors the request for the operation unit to read the instruction, and in response to multiple operation units reading the same instruction within a preset time, the multiple operation units are merged into a group; then, the memory access merging unit 430 can further monitor the instruction for each operation unit in the group to read data, and in response to receiving the instruction for all units in the group to read duplicate data, merge the read instructions and send them to the on-chip network, and send the read data to each operation unit in the group respectively.
[0076] Figure 5 A structural block diagram of a memory access merging unit according to an embodiment of the present disclosure is shown.
[0077] like Figure 5 As shown, the memory access merging unit 430 may include a monitoring module 431 , a recording module 432 and a merging module 433 .
[0078] In some embodiments, when the processor 400 starts to execute the first task after startup, the program read instructions sent by each operation unit can be detected based on the monitoring module 431, and in response to the program read instructions sent by multiple operation units within the first preset time are all used to read the program at the same address (that is, multiple program read instructions all contain the same program read address), the monitoring module 431 can judge the multiple operation units that send the above-mentioned program read instructions as multiple first operation units, and record them in the recording module 432.
[0079] In some embodiments, a plurality of first operation units may be recorded in the processor in the form of an operation group.
[0080] In some embodiments, the monitoring module 431 may determine whether the read instruction is a program read instruction or a data read instruction by identifying a special tag possessed by the read instruction.
[0081] In some embodiments, different special tags may be set for different types of instructions during the program code writing stage, so that the monitoring module can determine whether the current instruction is a program read instruction by detecting the special tags.
[0082] In some embodiments, for a processor that has executed at least one task, the monitoring module 431 can, while detecting the program read instruction, further determine whether the operation unit that sends the instruction belongs to an existing operation group, or whether the program read address corresponding to the program read instruction is the same as the program read address corresponding to the existing operation group.
[0083] Figure 6 FIG. 6 is a flow chart illustrating a method 600 for determining a plurality of first operation units according to an embodiment of the present disclosure. According to some embodiments, the method 600 includes:
[0084] Step S601: receiving a first program read instruction sent by a second operation unit among a plurality of operation units, where the first program read instruction includes a first program read address;
[0085] Step S602: In response to the second operation unit not belonging to any preceding operation group and the first program read instruction meeting a second preset condition, creating a new first operation group, wherein the second operation unit is determined to be a first operation unit in the first operation group, the first operation group corresponds to a first program read address, and the second preset condition includes at least one of the following: the first program read address is different from a program read address corresponding to any preceding operation group, and a time when the first program read instruction is received is greater than a second preset time between a time when the preceding program read instruction is received;
[0086] Step S603: Receive a second program read instruction sent by a third arithmetic unit among multiple arithmetic units. The second program read instruction includes a second program read address; and
[0087] Step S604: In response to the third arithmetic unit not belonging to any previous arithmetic group and the second program read instruction meeting the third preset condition, add the third arithmetic unit to the first arithmetic group as a first arithmetic unit in the first arithmetic group, where the third preset condition includes that the time difference between receiving the second program read instruction and the first program read instruction is less than a first preset time and the second program read address is the same as the first program read address.
[0088] Continue to refer to Figure 5 , in some embodiments, after receiving a program read instruction, the monitoring module 431 can first determine whether its corresponding arithmetic unit belongs to a previously created arithmetic group (i.e., the arithmetic group already created in the recording module 432). If it does not belong to the previous arithmetic group, and the address corresponding to this program read instruction is different from the address corresponding to the previous arithmetic group and / or the time difference between the time of receiving this instruction and the reception time corresponding to the last program read instruction in the previous arithmetic group exceeds a second preset time, then at this time, the monitoring module 431 can determine that the current program read instruction meets the second preset condition, create a new arithmetic group in the recording module 432, and record the arithmetic unit and the program read instruction corresponding to this instruction in the associated information of this arithmetic group.
[0089] In some embodiments, in response to receiving, within the first preset time, a program read instruction for reading the same program sent by another or multiple arithmetic units, and these one or more arithmetic units also not belonging to the previous arithmetic group, the one or more arithmetic units can be added to the newly created arithmetic group mentioned above.
[0090] Thus, it is possible to avoid data transmission chaos caused by the same arithmetic unit appearing repeatedly in different arithmetic groups.
[0091] Figure 7 Shows a flowchart of a method 700 for determining multiple first arithmetic units according to an embodiment of the present disclosure.
[0092] According to some embodiments, as Figure 7 shown, the method 700 includes:
[0093] Step S701: Receive a third program read instruction sent by a fourth arithmetic unit among multiple arithmetic units. The third program read instruction includes a third program read address;
[0094] Step S702: In response to the fourth arithmetic unit belonging to the first previous arithmetic group, delete the record regarding the first previous arithmetic group; and
[0095] Step S703: Create a second operation group, wherein the second operation group includes a fourth operation unit, and the second operation group corresponds to the third program read address.
[0096] Continue to refer Figure 4 and Figure 5 In some embodiments, when monitoring module 431 determines that the operation unit corresponding to the currently received program read instruction belongs to a previous operation group, it can first delete the record of the previous operation group from recording module 432 and then create a new operation group. In this way, after the task scheduling unit issues a new round of tasks to the operation units, the groups can be re-established based on the new tasks, achieving dynamic adjustment of the operation groups.
[0097] After making the above judgment and creating and adding the operation group, the monitoring module 431 can forward the program reading instruction directly to the on-chip network 440, so that the on-chip network 440 sends the instruction to the corresponding memory controller 450 to read the program from the corresponding address and send it to the corresponding operation unit. After receiving the program, the operation unit can start executing the program and read the corresponding data.
[0098] Figure 8 FIG. 8 is a flowchart of a method 800 for merging a second data read instruction according to an embodiment of the present disclosure.
[0099] According to some embodiments, Figure 8 As shown, the method 800 includes:
[0100] Step S801: receiving a data read instruction from a first operation group;
[0101] Step S802: In response to detecting that the data read instruction includes a repeat data flag, determining the data read instruction as a first data read instruction, wherein the repeat data flag indicates that the data read instruction is used to obtain data that needs to be obtained by multiple first operation units;
[0102] Step S803: Create a first record, where the first record includes a data read address corresponding to the first data read instruction and a first operation unit;
[0103] Step S804: in response to receiving the first data read instruction including the data read address sent by the remaining first operation units in the first operation group, add the corresponding first operation units to the first record until the first record includes all the first operation units in the first operation group; and
[0104] Step S805 : Generate a second data read instruction, where the second data read instruction is used to obtain data in the data read address.
[0105] Thus, by adding a first record and determining, through this record, the first arithmetic unit that has sent a duplicate data reading instruction (the first data reading instruction), when it is determined that all first arithmetic units within the group have been received, a merged data reading instruction is generated, thereby enabling, through this first record, the determination of the timing for generating and sending the merged reading instruction.
[0106] Continue to refer to Figure 4 and Figure 5 In some embodiments, the merging module 433 may receive data reading instructions from each arithmetic unit. The merging module 433 may detect whether the data reading instruction is an instruction for reading duplicate data. When it is determined that it is a duplicate data reading instruction, it may check, in the recording module 432, the arithmetic group to which the arithmetic unit corresponding to this instruction belongs. When it belongs to a certain arithmetic group, it may check, in the associated information of this arithmetic group, whether other arithmetic units in this arithmetic group have already sent the same duplicate data reading instruction.
[0107] When other arithmetic units in this arithmetic group have not sent this instruction yet, a record may be created in the associated information of this arithmetic group, and the data reading address corresponding to this duplicate data reading instruction and the arithmetic unit for reading this address may be recorded in this record; subsequently, when receiving this duplicate data reading instruction sent by another arithmetic unit in this arithmetic group, this arithmetic unit can be directly added to this record.
[0108] When the merging module 433 detects, through this record, that all arithmetic units in this arithmetic group have sent this duplicate data reading instruction, an instruction for reading this duplicate data may be generated and sent to the on-chip network 440, so that the on-chip network 440 sends the instruction to the corresponding memory controller 450 to read this duplicate data from the corresponding address and return it to the merging module 433; based on the above record, the merging module 433 can send this data to each arithmetic unit in the record (i.e., each arithmetic unit in this arithmetic group).
[0109] In some embodiments, the merging module 433 may determine whether the data reading instruction is for reading duplicate data by identifying a special mark possessed by the data reading instruction.
[0110] In some embodiments, during the stage of writing program code, a duplicate data mark can be added to the data that needs to be repeatedly obtained, so that the merging module 433 can determine whether the current instruction is a duplicate data reading instruction by detecting this mark.
[0111] According to some embodiments, the above data processing method may further include: deleting the first record in response to sending data to multiple first arithmetic units respectively.
[0112] Thus, when the merging module 433 obtains duplicate data and distributes it to each first arithmetic unit, the first record can be deleted. Thereby, when other duplicate data needs to be obtained subsequently, re-recording can be performed to re-determine the timing of generating and sending the merged read instruction.
[0113] In some embodiments, when the merging module 433 receives a normal data request instruction sent by an arithmetic unit (i.e., a data request instruction without a duplicate data mark), it can directly forward it to read the corresponding data.
[0114] In some exemplary embodiments, continuing with the Figure 3 matrix multiplication task shown above, for example, the process of merging the duplicate data read requests may include:
[0115] 1) The host side issues the matrix multiplication program to the NPU and marks cu_num = 4, indicating that the program is divided into 4 blocks for calculation;
[0116] 2) The task scheduling module of the NPU finds that arithmetic unit 1 and arithmetic unit 2 are idle and decides to assign the task to these two units;
[0117] 3) The task scheduling module issues the block parameters cu_idx = 0, cu_num = 4, and the address of the program code in the DRAM to arithmetic unit 1, and issues the block parameters cu_idx = 1, cu_num = 4, and the address of the program code in the DRAM to arithmetic unit 2;
[0118] 4) Arithmetic unit 1 reads address a to fetch the program code and attaches a read instruction mark to the request. After the memory access merging module monitors it, it newly creates an operation group (arithmetic unit 1, address a) and forwards the request;
[0119] 5) Arithmetic unit 1 obtains the program code and starts to execute, and calculates A[0]*B to get C[0] according to the configured block parameters;
[0120] 6) Arithmetic unit 2 reads address a to fetch the program code and attaches a read instruction mark to the request. After the memory access merging module monitors it, it adjusts the operation group (arithmetic unit 1, arithmetic unit 2, address a) and forwards the request;
[0121] 7) Arithmetic unit 2 obtains the program code according to the address and starts to execute, and calculates A[1]*B to get C[1] according to the configured block parameters;
[0122] 8) Arithmetic unit 1 reads address b to duplicate the matrix and attaches a duplicate matrix mark to the request. After the memory access merging module sees it, it checks the mergeable arithmetic units, confirms that arithmetic unit 1 belongs to an operation group, and adds a record (address b, arithmetic unit 1) to the group;
[0123] 9) Operation unit 2 reads the deduplication matrix for address b and adds a duplicate matrix mark to the request. After seeing this, the memory access merge module checks the mergeable operation units and confirms that operation unit 2 belongs to an operation group and that there is already a record for address b in the group. Therefore, the address is adjusted to (address b, operation unit 1, operation unit 2).
[0124] 10) Having received access requests for address b from all the operation units in the operation group (operation units 1 and 2), the memory access merge module sends a request for accessing address b to the on-chip network;
[0125] 11) The memory access and merging module receives the requested data, deletes the record at address b, and returns the data to operation units 1 and 2;
[0126] 12) The task scheduling module receives the signal that the calculation unit 1 has completed the calculation, and sends the block parameters cu_idx=2, cu_num=4, and the address of the program code in DRAM to the calculation unit 1;
[0127] 13) Operation unit 1 reads address a to fetch the program code and adds a read instruction tag to the request. The memory access merge module monitors this and deletes the previous operation group, creates a new operation group (operation unit 1, address a), and forwards the request.
[0128] 14) Operation unit 1 obtains the program code and starts executing it, and calculates A[2]*B according to the configured block parameters to obtain C[2];
[0129] 15) Subsequently, a method similar to steps 6 to 11 is executed until the computing units 1 and 2 complete the assigned 3rd and 4th block tasks, and feedback the corresponding signal to the task scheduling module, which then notifies the host side that the task is completed.
[0130] Through one or more embodiments of the present disclosure, compared with the solutions in the related art, the above-mentioned data processing method can only read from DRAM once when multiple computing units access the same repeated matrix, saving DRAM data access and improving the memory access bottleneck problem most of the time; and the above-mentioned method does not require additional hardware cache resources to store data retrieved from DRAM, and the hardware overhead is small; further, the above-mentioned method only needs to mark two special read requests accordingly, and the program executed by the computing unit itself does not need to be modified, and the hardware changes of the computing unit are also very small; in addition, the above-mentioned method supports the scenario of dynamic scheduling of computing units by the NPU system, and can determine in real time which computing units can optimize memory access by merging requests of repeated matrices without the need for software intervention.
[0131] According to another aspect of the present disclosure, a data processing device is provided. Figure 9The structural block diagram of a data processing device 900 according to an exemplary embodiment of the present disclosure is shown. As Figure 9 shown, the data processing device 900 includes:
[0132] A determination unit 910, configured to determine a plurality of first arithmetic units among a plurality of arithmetic units in response to receiving a plurality of program read instructions that meet a first preset condition, where the plurality of program read instructions respectively come from a plurality of first arithmetic units, and the first preset condition includes that the time difference between the first instruction and the last instruction among the plurality of program read instructions is less than a first preset time and the plurality of program read instructions have the same program read address;
[0133] A forwarding unit 920, configured to forward the plurality of program read instructions so that each of the plurality of first arithmetic units in the plurality of first arithmetic units obtains and executes the corresponding program, and the program includes a data read instruction;
[0134] A merging unit 930, configured to merge the plurality of first data read instructions received from the plurality of first arithmetic units into a second data read instruction, where the plurality of first data read instructions and the second data read instruction have the same data read address; and
[0135] An obtaining unit 940, configured to obtain the data in the data read address based on the second data read instruction to send the data to the plurality of first arithmetic units respectively.
[0136] Among them, the operations of the units 910 - 940 of the data processing device 900 are similar to the operations of steps S201 - S204 in the method 200, and will not be elaborated here.
[0137] According to some embodiments, multiple first arithmetic units are recorded in the processor in the form of arithmetic groups. The determination unit 910 may include: a first receiving subunit, configured to receive a first program read instruction sent by a second arithmetic unit among the multiple arithmetic units, the first program read instruction including a first program read address; a first creating subunit, configured to create a first arithmetic group in response to the second arithmetic unit not belonging to any previous arithmetic group and the first program read instruction meeting a second preset condition, wherein the second arithmetic unit is determined as a first arithmetic unit in the first arithmetic group, the first arithmetic group corresponds to the first program read address, and the second preset condition includes at least one of the following: the first program read address is different from the program read addresses corresponding to any previous arithmetic groups, and the time when the first program read instruction is received is greater than a second preset time compared to the time when a previous program read instruction is received; a second receiving subunit, configured to receive a second program read instruction sent by a third arithmetic unit among the multiple arithmetic units, the second program read instruction including a second program read address; and a first adding subunit, configured to add the third arithmetic unit to the first arithmetic group as a first arithmetic unit in the first arithmetic group in response to the third arithmetic unit not belonging to any previous arithmetic group and the second program read instruction meeting a third preset condition, wherein the third preset condition includes that the time difference between receiving the second program read instruction and the first program read instruction is less than a first preset time and the second program read address is the same as the first program read address.
[0138] According to some embodiments, the determination unit 910 may further include: a third receiving subunit, configured to receive a third program read instruction sent by a fourth arithmetic unit among the multiple arithmetic units, the third program read instruction including a third program read address; a deleting subunit, configured to delete the record regarding the first previous arithmetic group in response to the fourth arithmetic unit belonging to the first previous arithmetic group; and a second creating subunit, configured to create a second arithmetic group, wherein the second arithmetic group includes the fourth arithmetic unit and the second arithmetic group corresponds to the third program read address.
[0139] According to some embodiments, a plurality of first arithmetic units form a first arithmetic group. The merging unit 930 may include: a fourth receiving subunit, configured to receive a data reading instruction from the first arithmetic group; a determining subunit, configured to, in response to detecting that the data reading instruction includes a duplicate data flag, determine the data reading instruction as a first data reading instruction, where the duplicate data flag indicates that the data to be obtained by the data reading instruction is data that all the plurality of first arithmetic units need to obtain; a creating subunit, configured to create a first record, where the first record includes the data reading address corresponding to the first data reading instruction and the first arithmetic unit; a second adding subunit, configured to, in response to receiving a first data reading instruction including a data reading address sent by the remaining first arithmetic units in the first arithmetic group, add the corresponding first arithmetic unit to the first record until all the first arithmetic units in the first arithmetic group are included in the first record; and a generating subunit, configured to generate a second data reading instruction for obtaining the data in the data reading address.
[0140] According to some embodiments, the above data processing device may further include: a deleting unit, configured to delete the first record in response to sending data to the plurality of first arithmetic units respectively.
[0141] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.
[0142] Referring to Fig.10 , a block diagram of an electronic device 1000 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0143] As Fig.10As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0144] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 can be any type of device that can input information into the electronic device 1000. The input unit 1006 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1007 can be any type of device that can present information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 can include but are not limited to a magnetic disk and an optical disk. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0145] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method 200 described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).
[0146] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server combined with a blockchain.
[0152] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0153] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A data processing method for a processor including a plurality of arithmetic units, the method comprising: In response to receiving a plurality of program read instructions that meet a first preset condition, determining a plurality of first arithmetic units among the plurality of arithmetic units, wherein the plurality of program read instructions respectively come from the plurality of first arithmetic units, and the first preset condition includes that the time difference between the first instruction and the last instruction among the plurality of program read instructions is less than a first preset time and the plurality of program read instructions have the same program read address; Forwarding the plurality of program read instructions so that each of the plurality of first arithmetic units among the plurality of first arithmetic units obtains and executes a corresponding program, and the program includes a data read instruction; In response to receiving a plurality of first data read instructions respectively from the plurality of first arithmetic units, combining the plurality of first data read instructions into a second data read instruction, wherein the plurality of first data read instructions and the second data read instruction have the same data read address; and Based on the second data read instruction, obtaining the data in the data read address to respectively send the data to the plurality of first arithmetic units.
2. The method according to claim 1, wherein The plurality of first arithmetic units are recorded in the processor in the form of arithmetic groups, and the determining a plurality of first arithmetic units among the plurality of arithmetic units in response to receiving a plurality of program read instructions that meet a first preset condition includes: Receiving a first program read instruction sent by a second arithmetic unit among the plurality of arithmetic units, the first program read instruction including a first program read address; In response to the second arithmetic unit not belonging to any previous arithmetic group and the first program read instruction meeting a second preset condition, creating a first arithmetic group, wherein the second arithmetic unit is determined as one of the first arithmetic units in the first arithmetic group, and the first arithmetic group corresponds to the first program read address, and the second preset condition includes at least one of the following: The first program read address is different from the program read addresses corresponding to any previous arithmetic group, and The time when the first program read instruction is received is greater than a second preset time from the time when a previous program read instruction is received; Receiving a second program read instruction sent by a third arithmetic unit among the plurality of arithmetic units, the second program read instruction including a second program read address; and In response to the third arithmetic unit not belonging to any previous arithmetic group and the second program read instruction meeting a third preset condition, adding the third arithmetic unit to the first arithmetic group to be one of the first arithmetic units in the first arithmetic group, wherein the third preset condition includes that the time difference between the second program read instruction and the first program read instruction is less than the first preset time and the second program read address is the same as the first program read address.
3. The method according to claim 2, wherein, The determining a plurality of first arithmetic units among the plurality of arithmetic units in response to receiving a plurality of program read instructions that meet a first preset condition further includes: Receive a third program read instruction sent by a fourth arithmetic unit among the multiple arithmetic units, where the third program read instruction includes a third program read address; In response to the fourth arithmetic unit belonging to a first pre-order arithmetic group, delete the record regarding the first pre-order arithmetic group; and Create a second arithmetic group, where the second arithmetic group includes the fourth arithmetic unit, and the second arithmetic group corresponds to the third program read address.
4. The method according to any one of claims 1 to 3, wherein, The multiple first arithmetic units form a first arithmetic group. The step of merging the multiple first data read instructions into a second data read instruction in response to receiving the multiple first data read instructions respectively from the multiple first arithmetic units includes: Receive a data read instruction from the first arithmetic group; In response to detecting that the data read instruction includes a duplicate data flag, determine the data read instruction as a first data read instruction, where the duplicate data flag indicates that the data to be obtained by the data read instruction is the data that all the multiple first arithmetic units need to obtain; Create a first record, where the first record includes the data read address corresponding to the first data read instruction and the first arithmetic unit; In response to receiving the first data read instructions including the data read address sent by the remaining first arithmetic units in the first arithmetic group, add the corresponding first arithmetic units to the first record until all the first arithmetic units in the first arithmetic group are included in the first record; and Generate the second data read instruction, where the second data read instruction is used to obtain the data in the data read address.
5. The method according to claim 4, further comprising: In response to sending the data to the multiple first arithmetic units respectively, delete the first record.
6. A data processing device for a processor including multiple arithmetic units, the device includes: A determination unit configured to determine multiple first arithmetic units among the multiple arithmetic units in response to receiving multiple program read instructions that meet a first preset condition, where the multiple program read instructions are respectively from the multiple first arithmetic units, and the first preset condition includes that the time difference between receiving the first instruction and the last instruction among the multiple program read instructions is less than a first preset time and the multiple program read instructions have the same program read address; A forwarding unit configured to forward the multiple program read instructions so that each first arithmetic unit among the multiple first arithmetic units obtains and executes the corresponding program, where the program includes a data read instruction; A merging unit configured to merge the multiple first data read instructions into a second data read instruction in response to receiving the multiple first data read instructions respectively from the multiple first arithmetic units, where the multiple first data read instructions and the second data read instruction have the same data read address; and An obtaining unit configured to obtain the data in the data read address based on the second data read instruction to send the data to the multiple first arithmetic units respectively.
7. The apparatus according to claim 6, wherein The multiple first arithmetic units are recorded in the processor in the form of arithmetic groups, and the determining unit includes: A first receiving subunit, configured to receive a first program reading instruction sent by a second arithmetic unit among the multiple arithmetic units, where the first program reading instruction includes a first program reading address; A first creating subunit, configured to create a first arithmetic group in response to the second arithmetic unit not belonging to any previous arithmetic group and the first program reading instruction meeting a second preset condition, where the second arithmetic unit is determined as a first arithmetic unit in the first arithmetic group, the first arithmetic group corresponds to the first program reading address, and the second preset condition includes at least one of the following: The first program reading address is different from the program reading addresses corresponding to any previous arithmetic groups, and The time when the first program reading instruction is received is greater than a second preset time from the time when a previous program reading instruction is received; A second receiving subunit, configured to receive a second program reading instruction sent by a third arithmetic unit among the multiple arithmetic units, where the second program reading instruction includes a second program reading address; and A first adding subunit, configured to add the third arithmetic unit to the first arithmetic group as a first arithmetic unit in the first arithmetic group in response to the third arithmetic unit not belonging to any previous arithmetic group and the second program reading instruction meeting a third preset condition, where the third preset condition includes that the time difference between receiving the second program reading instruction and the first program reading instruction is less than the first preset time and the second program reading address is the same as the first program reading address.
8. The apparatus according to claim 7, wherein, The determining unit further includes: A third receiving subunit, configured to receive a third program reading instruction sent by a fourth arithmetic unit among the multiple arithmetic units, where the third program reading instruction includes a third program reading address; A deleting subunit, configured to delete the record regarding the first previous arithmetic group in response to the fourth arithmetic unit belonging to the first previous arithmetic group; and A second creating subunit, configured to create a second arithmetic group, where the second arithmetic group includes the fourth arithmetic unit and the second arithmetic group corresponds to the third program reading address.
9. The device according to any one of claims 6 - 8, wherein The multiple first arithmetic units form a first arithmetic group, and the merging unit includes: A fourth receiving subunit, configured to receive a data reading instruction from the first arithmetic group; A determining subunit, configured to determine the data reading instruction as a first data reading instruction in response to detecting that the data reading instruction includes a duplicate data flag, where the duplicate data flag indicates that the data to be obtained by the data reading instruction is data that all the multiple first arithmetic units need to obtain; A creating subunit, configured to create a first record, where the first record includes the data reading address corresponding to the first data reading instruction and the first arithmetic unit; A second adding subunit, configured to add corresponding first arithmetic units to the first record until all the first arithmetic units in the first arithmetic group are included in the first record in response to receiving the first data reading instruction including the data reading address sent by the remaining first arithmetic units in the first arithmetic group; and A generating subunit, configured to generate the second data reading instruction for obtaining the data in the data reading address.
10. The apparatus according to claim 9, further comprising: A deleting unit, configured to delete the first record in response to sending the data to the plurality of first arithmetic units respectively.
11. A chip, comprising the apparatus according to any one of claims 6-10.
12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
14. A computer program product comprising a computer program, wherein, The computer program implements the method according to any one of claims 1-5 when executed by a processor.
Citation Information
Patent Citations
Data reading method and device
CN108536473A
Data processing method and device, chip, equipment and medium
CN115033286A