An asynchronous chip, data transfer method, device, equipment and medium
Through the coordinated work of the scheduling module and the data handling module in the asynchronous chip, asynchronous data handling is realized, the problems of long handling time and low efficiency in the existing technology are solved, the data handling efficiency is improved and the impact of delay is reduced.
Patent Information
- Application Number
- CN202310209796.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-06
AI Technical Summary
In the prior art, how to asynchronously transport data through chips to save handling time and improve handling efficiency is an urgent problem.
The asynchronous chip is adopted, including a scheduling module and at least two data handling modules. The scheduling module obtains the operating status of each data handling module, determines the target data handling module and issues data scheduling instructions. The data handling module receives instructions for data handling, and issues data handling instructions when the reference data handling module is idle to transport the data.
It realizes asynchronous data handling, improves handling efficiency, saves handling time, and avoids normal working abnormalities caused by excessive data in the storage unit, reduces the delay in reporting statistical events on critical paths, and reduces the impact of application load.
Smart Images

Figure CN116088769B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of neural networks, and in particular, to an asynchronous chip, a data transfer method, a device, equipment, and a medium. Background Art
[0002] With the continuous development of neural network technology, high-performance chips applied to neural networks have also been mass-produced and applied. These chips can perform tasks such as performance analysis of computational loads or data transfer.
[0003] At present, how to asynchronously transfer data to be transferred through a chip, thereby saving transfer time and improving transfer efficiency is a key issue in the industry research. Summary of the Invention
[0004] Embodiments of the present invention provide an asynchronous chip, a data transfer method, a device, equipment, and a medium, which can asynchronously transfer data to be transferred, save transfer time, and improve transfer efficiency.
[0005] According to one aspect of the embodiments of the present invention, an asynchronous chip is provided, which is applied to a deep neural network and includes: a scheduling module and at least two data transfer modules; wherein, the scheduling module is communicatively connected to each of the data transfer modules; each of the data transfer modules includes a corresponding storage unit;
[0006] The scheduling module is configured to obtain the operating states of the data transfer modules, determine a target data transfer module according to the operating states, and issue a data scheduling instruction to the target data transfer module;
[0007] The data transfer module is configured to receive the data scheduling instruction issued by the scheduling module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit;
[0008] The scheduling module is further configured to, when determining that the working state of the reference data transfer module is idle, issue a data transfer instruction to the reference data transfer module, so that the target data transfer module transfers the stored data stored in the corresponding storage unit to an external storage space.
[0009] According to another aspect of the embodiments of the present invention, a data transfer method is provided, which is executed by any asynchronous chip provided by the embodiments of the present invention and includes:
[0010] Obtain the operating states of the data transfer modules through the scheduling module, determine a target data transfer module according to the operating states, and issue a data scheduling instruction to the target data transfer module;
[0011] Receive the data scheduling instruction issued by the scheduling module through the data transfer module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit;
[0012] When it is determined through the scheduling module that the working state of the reference data transfer module is idle, issue a data transfer instruction to the reference data transfer module, so that the reference data transfer module transfers the stored data stored in the corresponding storage unit to the external storage space.
[0013] According to another aspect of the embodiments of the present invention, there is provided a data transfer device, including:
[0014] A data scheduling instruction issuing module, configured to obtain the operating status of each data transfer module through the scheduling module, determine the target data transfer module according to each of the operating statuses, and issue a data scheduling instruction to the target data transfer module;
[0015] A first data transfer module, configured to receive the data scheduling instruction issued by the scheduling module through the data transfer module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit;
[0016] A second data transfer module, configured to issue a data transfer instruction to the reference data transfer module when it is determined through the scheduling module that the working state of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data stored in the corresponding storage unit to the external storage space.
[0017] According to another aspect of the embodiments of the present invention, there is provided an electronic device, the electronic device includes:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the data transfer method according to any one of the embodiments of the present invention.
[0021] According to another aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the data transfer method according to any one of the embodiments of the present invention when executed.
[0022] The technical solution of the embodiment of the present invention provides an asynchronous chip, which can be applied to a neural network, including: a scheduling module and at least two data transfer modules; wherein, the scheduling module is respectively communicatively connected to each of the data transfer modules; each of the data transfer modules includes a corresponding storage unit; the scheduling module is configured to obtain the operating states of each of the data transfer modules, determine a target data transfer module according to each of the operating states, and issue a data scheduling instruction to the target data transfer module; the data transfer module is configured to receive the data scheduling instruction issued by the scheduling module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit; an idle data transfer module can be determined, and the data to be transferred is transferred through the idle transfer model; the scheduling module is further configured to issue a data transfer instruction to the reference data transfer module when it is determined that the working state of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data in the corresponding storage unit to an external storage space, and can also transfer the data stored in the storage unit when the data transfer module is idle, realizing asynchronous data transfer, improving the transfer efficiency, and saving the transfer time.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the embodiments of the present invention. Other features of the embodiments of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 is a schematic structural diagram of an asynchronous chip provided in Embodiment 1 of the present invention;
[0026] Figure 2 is a schematic structural diagram of an asynchronous chip provided in Embodiment 2 of the present invention;
[0027] Figure 3 is a flowchart of a data transfer method provided in Embodiment 3 of the present invention;
[0028] Figure 4 is a schematic diagram of a data asynchronous processing method provided in Embodiment 3 of the present invention;
[0029] Figure 5It is a schematic structural diagram of a data transfer device provided according to Embodiment 4 of the present invention;
[0030] Figure 6 It is a schematic structural diagram of an electronic device for implementing the data transfer method of the embodiments of the present invention. Detailed implementation manners
[0031] In order to enable those skilled in the art to better understand the solutions of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the embodiments of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] Embodiment 1
[0034] Figure 1 It is a schematic structural diagram of an asynchronous chip provided according to Embodiment 1 of the present invention. The asynchronous chip can be applied to a neural network. For example, it can load or run a neural network model, process or transfer the data generated by the neural network model, etc.; optionally, referring to Figure 1 , the asynchronous chip may include: a scheduling module 110 and at least two data transfer modules 120; wherein, the scheduling module 110 is respectively communicatively connected to each of the data transfer modules 120; each of the data transfer modules 120 includes a corresponding storage unit ( Figure 1 not shown in the figure). It should be noted that Figure 1 only one data transfer module is shown in the figure, which is not a limitation of this embodiment.
[0035] In an alternative implementation of this embodiment, the scheduling module 110 can be used to obtain the operating status of each of the data transfer modules 120, determine a target data transfer module based on each of the operating statuses, and issue a data scheduling instruction to the target data transfer module. The operating status of the data transfer module can be idle or non-idle, and this embodiment does not limit it. The number of target data transfer modules can be one or multiple, and this embodiment does not limit it.
[0036] Exemplarily, in this embodiment, the asynchronous chip includes a total of ten data transfer modules. The scheduling module can obtain the operating status of each data transfer module and determine the target data transfer module based on the obtained operating statuses of each data transfer module. If the operating status of the first data transfer module is idle and the operating statuses of the other data transfer modules are all non-idle, the first data transfer module can be determined as the target data transfer module. Further, the scheduling module can also issue a data scheduling instruction to the first data transfer module. The data scheduling instruction can include specific information (size, location, or type, etc.) of the data to be transferred, and can also include the transfer destination of the data to be transferred, etc., and this embodiment does not limit it.
[0037] In an alternative implementation of this embodiment, the data transfer module can be used to receive the data scheduling instruction issued by the scheduling module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the storage unit corresponding to it.
[0038] Optionally, in this embodiment, after receiving the data scheduling instruction issued by the scheduling module, the target data transfer module can perform data transfer according to the received data scheduling instruction. For example, it can transfer the image data stored in the storage unit corresponding to it to other modules of the asynchronous chip or to the external storage space of the asynchronous chip, etc. The data transfer module can also store the data transfer time in the storage unit corresponding to it during the data transfer process. The advantage of this setting is that information such as the processes of each transfer time and the data content being transferred can be obtained by querying each storage unit.
[0039] In another alternative implementation of this embodiment, the scheduling module 110 can also be used to issue a data transfer instruction to the reference data transfer module when determining that the working status of the reference data transfer module is idle, so that the target data transfer module transfers the stored data in the storage unit corresponding to it to the external storage space. The reference data transfer module can be other data transfer modules except the above-mentioned target data transfer modules.
[0040] Optionally, in this embodiment, after the scheduling module determines the target data transfer module according to the operating states of the obtained data transfer modules and sends a data scheduling instruction to the target data transfer module, it can further determine the reference data transfer modules. For example, it can determine the data transfer modules with an idle working state among the data transfer modules other than the target data transfer module, and determine these data transfer modules as the reference data transfer modules. Further, it can issue a data transfer instruction to the reference data transfer modules, so that the reference data transfer modules transfer the stored data in the storage units corresponding to them to the external storage space.
[0041] The technical solution of this embodiment provides an asynchronous chip, which can be applied to neural networks and includes: a scheduling module and at least two data transfer modules; wherein, the scheduling module is communicatively connected to each of the data transfer modules; each of the data transfer modules includes a storage unit corresponding to it; the scheduling module is configured to obtain the operating states of the data transfer modules, determine the target data transfer module according to the operating states, and issue a data scheduling instruction to the target data transfer module; the data transfer module is configured to receive the data scheduling instruction issued by the scheduling module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the storage unit corresponding to it; it can determine the idle data transfer modules and transfer the data to be transferred through the idle transfer model; the scheduling module is further configured to issue a data transfer instruction to the reference data transfer module when it determines that the working state of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data in the storage unit corresponding to it to the external storage space, and can also transfer the data stored in the storage unit, realizing asynchronous data transfer, improving the transfer efficiency, saving the transfer time, and avoiding the phenomenon that the data stored in the storage unit is too much and cannot work properly. At the same time, the solution of this embodiment can collect data asynchronously and write the data into the external storage space asynchronously, eliminating the delay of statistical event reporting on the critical path and minimizing the impact on the application load.
[0042] Based on the above technical solution, the data transfer module can also be configured to compress the data stored in the storage unit corresponding to it to obtain compressed data, and after receiving the data transfer instruction issued by the scheduling module, transfer the compressed data to the external storage space.
[0043] In an alternative implementation of this embodiment, when the data transfer module is in an idle state without receiving a data scheduling instruction sent by the scheduling module, it can compress the data stored in the storage unit corresponding to it to obtain compressed data; further, when receiving a data transfer instruction issued by the scheduling module, it can transfer the compressed data to an external storage space.
[0044] The advantage of this setting is that it can further reduce the possibility of storage space overflow and also expand the space for long-term performance data analysis.
[0045] Embodiment 2
[0046] Figure 2 FIG. 10 is a schematic structural diagram of an asynchronous chip according to Embodiment 2 of the present invention. This embodiment further refines the above technical solutions, and the technical solutions in this embodiment can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the asynchronous chip may further include: at least two data calculation modules 230, each of the data calculation modules is communicatively connected to each of the data transfer modules; each of the data calculation modules includes a storage unit corresponding to it.
[0047] In an alternative implementation of this embodiment, the data transfer module may further be configured to determine target data to be transmitted according to the data scheduling instruction and transmit the target data to the data calculation module.
[0048] Optionally, in this embodiment, after receiving a data scheduling instruction issued by the scheduling module, the data transfer module may parse the received data scheduling instruction to determine the target data to be transmitted and transmit the target data to the corresponding data calculation module; where the target data may be image data, voice data, or log data, etc., and this embodiment does not limit it.
[0049] In an alternative implementation of this embodiment, the data calculation module may be configured to receive the target data transmitted by the data transfer module, perform local calculations according to the target data, and store the intermediate data and result data generated during the calculation in the storage unit corresponding to the data calculation module.
[0050] Optionally, in this embodiment, after receiving the target data transmitted by the data transfer module, the data calculation module may perform local calculations based on the target data, and store the intermediate data and result data generated during the calculation process in the storage unit of the data calculation module; it should be noted that when the data transfer module transmits the target data, it may also transmit the processing task corresponding to the target data to the data calculation module, so that the data calculation module can directly process the received target data according to the received processing task.
[0051] In the solution of this embodiment, the asynchronous chip may further include at least two data calculation modules 230, and each of the data calculation modules is communicatively connected to each of the data transfer modules; each of the data calculation modules includes a corresponding storage unit, and the data calculation module can be used to receive the target data transmitted by the data transfer module, perform local calculations based on the target data, and store the intermediate data and result data generated during the calculation process in the storage unit corresponding to the data calculation module, and can process the received data, improving the calculation ability of the asynchronous chip, and at the same time can also locally store the intermediate data and calculation results generated during the calculation process, facilitating the subsequent rapid search of relevant data.
[0052] Embodiment III
[0053] Figure 3 is a flowchart of a data transfer method provided according to Embodiment III of the present invention. This embodiment is applicable to the situation of transferring neural network data. This method can be executed by a data transfer device, and the data transfer device can be implemented in the form of hardware and / or software. The data transfer device can be configured in the asynchronous chip involved in the embodiments of the present invention, or in an electronic device such as a computer, a server, or a tablet computer. Specifically, referring to Figure 3 , the method specifically includes the following steps:
[0054] Step 310, obtain the running states of the data transfer modules through the scheduling module, determine the target data transfer module according to the running states, and issue a data scheduling instruction to the target data transfer module.
[0055] In an optional implementation manner of this embodiment, the running states of the data transfer modules can be obtained through the scheduling module, and the target data transfer module can be determined according to the obtained running states; further, a data scheduling instruction can be issued to the determined target data transfer module.
[0056] Step 320, receive the data scheduling instruction issued by the scheduling module through the data transfer module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit.
[0057] In an alternative implementation of this embodiment, after the scheduling module sends a data scheduling instruction to the target data transfer module, the target data transfer module can perform data transfer according to the received data scheduling instruction and record the data transfer event in the storage unit corresponding thereto.
[0058] Before performing data transfer according to the received data scheduling instruction, the target data transfer module can determine the target data to be transmitted according to the received data scheduling instruction and transmit the target data to the data calculation module.
[0059] Optionally, in this embodiment, after receiving the target data, the data calculation module can perform local calculation according to the target data and store the intermediate data and result data generated during the calculation in the storage unit corresponding to the data calculation module.
[0060] Step 330: When it is determined by the scheduling module that the working state of the reference data transfer module is idle, issue a data transfer instruction to the reference data transfer module so that the reference data transfer module transfers the stored data stored in the storage unit corresponding thereto to an external storage space.
[0061] In an alternative implementation of this embodiment, when the scheduling module determines that the working state of the reference data transfer module other than the target data transfer module is idle, it can further issue a data transfer instruction to the reference data transfer module; after receiving the data transfer instruction, the reference data transfer module can transfer the stored data stored in the storage unit corresponding thereto to an external space.
[0062] Optionally, in this embodiment, before transferring the stored data stored in the storage unit corresponding thereto to an external space, the reference data transfer module may further include: compressing the data stored in the storage unit corresponding thereto to obtain compressed data, and after receiving the data transfer instruction issued by the scheduling module, transferring the compressed data to an external storage space.
[0063] In the solution of this embodiment, the scheduling module obtains the running status of each data transfer module, determines the target data transfer module according to each running status, and issues a data scheduling instruction to the target data transfer module; the data transfer module receives the data scheduling instruction issued by the scheduling module, performs data transfer according to the data scheduling instruction, and records the data transfer event in the storage unit corresponding thereto; when it is determined through the scheduling module that the working status of the reference data transfer module is idle, a data transfer instruction is issued to the reference data transfer module, so that the reference data transfer module transfers the stored data stored in the storage unit corresponding thereto to the external storage space, and the data stored in the storage unit can be transferred, realizing asynchronous data transfer, improving the transfer efficiency, saving the transfer time, and avoiding the phenomenon that the data stored in the storage unit is too much to work properly. At the same time, the solution of this embodiment can collect data asynchronously and write the data into the external storage space asynchronously, eliminating the delay of statistical event reporting on the critical path and minimizing the impact on the application load.
[0064] Based on the application scenario of load performance analysis of a neural network high-performance computing chip, the embodiment of the present invention combines the data transfer and data calculation characteristics of deep learning, and proposes a non-centralized asynchronous load performance data hardware module design method. Compared with the existing design methods, it maximally reduces the impact on the performance of the original load, and at the same time brings extreme performance data accuracy. Finally, it also brings convenience to the software maintenance and development of the upper layer.
[0065] The performance data collection algorithm proposed by the embodiment of the present invention is executed locally at each statistical node in the form of firmware or a hardware unit, and the collected performance data is also stored in the local storage unit. The entire process does not pass through the external bus, eliminating the uncertain delay caused by external bus network congestion, and further eliminating the out-of-order problem of load performance data.
[0066] Combined with the deep neural network load performance analysis scenario, the embodiment of the present invention uses the time-sharing characteristics of data calculation and data transfer to asynchronously collect scattered statistical data from each statistical node and asynchronously write it into the external storage space. The entire process introduces asynchronous operations, eliminating the delay of statistical event reporting on the critical path and minimizing the impact on the application load.
[0067] The embodiment of the present invention can reprocess block-structured data, so that when the data transfer module writes to the external storage space, a compression algorithm is introduced, further reducing the possibility of storage space overflow, and also expanding the space for long-term performance data analysis. The embodiment of the present invention stores the performance data in the local node, and the statistical time of the load performance event is not affected by the delay introduced by the bus, so the absolute time obtained by statistics is more accurate.
[0068] The performance data of the embodiment of the present invention is stored in the local node to form structured data, which provides more flexibility for the later performance data organization. The method of the embodiment of the present invention involves that the load performance statistics object is a data handling hardware module, a data computing hardware module, and a control scheduling hardware module in a deep learning computing scenario.
[0069] The method of the embodiment of the present invention designs a storage unit for each submodule to store local load performance events; at the same time, each unit has a load performance event collection subunit, which exists in the form of firmware or hardware subunit and has programmable capabilities.
[0070] The control scheduling hardware module included in the method of the embodiment of the present invention is responsible for scheduling the data transfer of the data module, and prepares the data and operator codes required for the calculation for the data calculation hardware unit.
[0071] In the embodiment of the present invention, for the timing of collecting load performance events, the code block for collecting events can exist in the data handling unit and the computing unit in the form of firmware or hardware submodule solidification, with the focus on programability. This part can be embedded into the program by the operator development or deep learning operator library user in the AOT (ahead of time) model compilation period through high-level language, inserted into the final ELF (executable file) as a section, and in the runtimeprogram to the load performance event collection subunit module or directly run, determine the timing of data collection from the code level, and write it to the local storage. The final unified collection is through the control scheduling hardware module, which fully perceives the current model scheduling status, and asynchronously calls the idle data handling unit to move the local load performance data of each load hardware module to a larger storage space.
[0072] Figure 4 is a schematic diagram of a data asynchronous processing method provided according to Embodiment 3 of the present invention, Figure 4 It can be seen that the data handling module handles data for the data calculation module, and usually records the performance analysis time of load handling and calculation. Here, it is simple to record the starting load recording event of data handling and calculation, and save it to the local storage space. Each module performs its own load performance time storage. At the same time, the global scheduling module starts asynchronous load event handling in stages A and B, that is, there is no data handling event stage, and transfers the load analysis events of each hardware module to the external storage space of the device. As shown in the figure below, in the global device storage space, the high-performance load analysis tool can obtain complete load performance analysis data, and then perform post-processing to show the user the complete and sequence-preserving load performance analysis results. Example 4
[0073] Figure 5It is a schematic structural diagram of a data transfer device provided in Embodiment 4 of the present invention. As Figure 5 shown, the device includes: a data scheduling instruction issuing module 510, a first data transfer module 520, and a second data transfer module 530.
[0074] The data scheduling instruction issuing module 510 is configured to obtain the operating status of each data transfer module through the scheduling module, determine the target data transfer module according to each of the operating statuses, and issue a data scheduling instruction to the target data transfer module;
[0075] The first data transfer module 520 is configured to receive the data scheduling instruction issued by the scheduling module through the data transfer module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the storage unit corresponding thereto;
[0076] The second data transfer module 530 is configured to issue a data transfer instruction to the reference data transfer module when it is determined through the scheduling module that the working status of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data in the storage unit corresponding thereto to an external storage space.
[0077] In the solution of this embodiment, the data scheduling instruction issuing module obtains the operating status of each data transfer module through the scheduling module, determines the target data transfer module according to each of the operating statuses, and issues a data scheduling instruction to the target data transfer module; the first data transfer module receives the data scheduling instruction issued by the scheduling module through the data transfer module, performs data transfer according to the data scheduling instruction, and records the data transfer event in the storage unit corresponding thereto; the second data transfer module issues a data transfer instruction to the reference data transfer module when it is determined through the scheduling module that the working status of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data in the storage unit corresponding thereto to an external storage space, and can transfer the data stored in the storage unit, realizing asynchronous data transfer, improving the transfer efficiency, saving the transfer time, and avoiding the phenomenon that the data stored in the storage unit is too much to work properly. At the same time, the solution of this embodiment can collect data asynchronously and write the data into the external storage space asynchronously, eliminating the delay of the statistical event reporting on the critical path and minimizing the impact on the application load.
[0078] In an optional implementation manner of this embodiment, the second data transfer module 530 is further configured to compress the data stored in the storage unit corresponding thereto through the data transfer module to obtain compressed data, and transfer the compressed data to the external storage space after receiving the data transfer instruction issued by the scheduling module.
[0079] In an alternative implementation of this embodiment, the first data transfer module 520 is further configured to determine target data to be transmitted according to the data scheduling instruction, and transmit the target data to the data calculation module; receive the target data transmitted by the data transfer module through the data calculation module, perform local calculation according to the target data, and store the intermediate data and result data generated during the calculation in the storage unit corresponding to the data calculation module.
[0080] The data transfer device provided by the embodiments of the present invention can execute the data transfer method provided by any embodiment of the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0081] Embodiment Five
[0082] Figure 6 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the embodiments of the present invention described and / or claimed herein.
[0083] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0084] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0085] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data transfer method.
[0086] In some embodiments, the data transfer method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data transfer method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data transfer method by any other suitable means (e.g., by means of firmware).
[0087] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0089] In the context of the embodiments of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0091] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0092] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0093] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the embodiments of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the embodiments of the present invention can be achieved, and no limitations are imposed herein.
[0094] The above specific embodiments do not constitute a limitation on the protection scope of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.
Claims
1. An asynchronous chip, applied to neural networks, characterized in that, Comprising: A scheduling module and at least two data transfer modules; wherein, the scheduling module is communicatively connected to each of the data transfer modules; each of the data transfer modules includes a corresponding storage unit; The scheduling module is configured to obtain the operating states of each of the data transfer modules, determine a target data transfer module according to each of the operating states, and issue a data scheduling instruction to the target data transfer module; The data transfer module is configured to receive the data scheduling instruction issued by the scheduling module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the corresponding storage unit; The scheduling module is further configured to, when determining that the working state of a reference data transfer module is idle, issue a data transfer instruction to the reference data transfer module, so that the reference data transfer module transfers the stored data stored in the corresponding storage unit to an external storage space.
2. The chip according to claim 1, wherein: The data transfer module is further configured to compress the data stored in the corresponding storage unit to obtain compressed data, and after receiving the data transfer instruction issued by the scheduling module, transfer the compressed data to an external storage space.
3. The chip according to claim 1, wherein: The data transfer module is further configured to determine target data to be transmitted according to the data scheduling instruction, and transmit the target data to a data calculation module.
4. The chip according to claim 3, characterized in that, The chip further includes at least two data calculation modules, each of the data calculation modules is communicatively connected to each of the data transfer modules; each of the data calculation modules includes a corresponding storage unit; The data calculation module is configured to receive the target data transmitted by the data transfer module, perform local calculation according to the target data, and store the intermediate data and result data generated during the calculation in the storage unit corresponding to the data calculation module.
5. A data transfer method, which is executed by the asynchronous chip described in any one of claims 1-4, characterized in that, Comprising: Obtaining the operating states of each data transfer module through a scheduling module, determining a target data transfer module according to each of the operating states, and issuing a data scheduling instruction to the target data transfer module; Receiving the data scheduling instruction issued by the scheduling module through a data transfer module, performing data transfer according to the data scheduling instruction, and recording the data transfer event in the corresponding storage unit; When determining that the working state of a reference data transfer module is idle through the scheduling module, issuing a data transfer instruction to the reference data transfer module, so that the reference data transfer module transfers the stored data stored in the corresponding storage unit to an external storage space.
6. The method according to claim 5, wherein Further comprising: Compressing the data stored in the corresponding storage unit through a data transfer module to obtain compressed data, and after receiving the data transfer instruction issued by the scheduling module, transferring the compressed data to an external storage space.
7. The method according to claim 5, wherein Further comprising: Determining target data to be transmitted according to the data scheduling instruction, and transmitting the target data to a data calculation module; Receive the target data transmitted by the data transfer module through the data calculation module, perform local calculations based on the target data, and store the intermediate data and result data generated during the calculations in the storage unit corresponding to the data calculation module.
8. A data transfer device, characterized in that, It includes: A data scheduling instruction issuing module, configured to obtain the operating status of each data transfer module through a scheduling module, determine a target data transfer module according to each of the operating statuses, and issue a data scheduling instruction to the target data transfer module; A first data transfer module, configured to receive the data scheduling instruction issued by the scheduling module through the data transfer module, perform data transfer according to the data scheduling instruction, and record the data transfer event in the storage unit corresponding to it; A second data transfer module, configured to issue a data transfer instruction to the reference data transfer module when it is determined through the scheduling module that the working status of the reference data transfer module is idle, so that the reference data transfer module transfers the stored data stored in the storage unit corresponding to it to an external storage space.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the data transfer method according to any one of claims 5-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the data transfer method according to any one of claims 5-7 when executed by a processor.
Citation Information
Patent Citations
Chip, neural network training system, memory management method and device, and equipment
CN112819145A
Server and sorting equipment thereof
CN113672530A