Data Processing Method, Apparatus, Electronic Device, and Storage Medium
By introducing a second computing device (on-chip CPU) into a heterogeneous computing system to configure and control the accelerator, the stability problem of the host CPU under the multi-accelerator configuration and control is solved, and more efficient and stable data processing is achieved.
Patent Information
- Application Number
- CN202010671509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-07-13
AI Technical Summary
In existing heterogeneous computing systems, the host CPU needs to configure and control multiple accelerators, resulting in poor system stability, especially when the data processing volume is large, the control pressure of the host CPU increases.
By establishing a connection between the first computing device (host CPU) and the second computing device (on-chip CPU), the first computing device transmits configuration parameters to the second computing device, the second computing device configures the accelerator, and causes the configured accelerator to perform data processing through message interaction, reducing the control pressure of the host CPU.
This method can improve the efficiency and stability of data processing, reduce the control pressure of the host CPU, and ensure the stable operation of the system under high load conditions.
Smart Images

Figure CN113934677B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a data processing method, a data processing device, an electronic device, and a storage medium. Background Art
[0002] Heterogeneous computing technology is a parallel and distributed computing technology that matches the type of parallelism of computing tasks with the type of computing that a machine can effectively support and can make full use of various computing resources.
[0003] An existing heterogeneous computing system includes a host CPU and an accelerator (or an acceleration chip). Before data processing, a handshake is performed between the host CPU and the accelerator to configure the accelerator. Among them, a handshake can also be called a handover. A handshake is used to reach parameters, such as an information transfer rate, an alphabet, a parity check, an interrupt process, and other protocol characteristics, after a communication circuit is established and before information transmission starts. The accelerator is configured through the handshake between the host CPU and the accelerator. The host CPU can also control the configured accelerator to complete the acceleration processing of data.
[0004] However, adopting this solution, the host CPU needs to configure multiple accelerators and control multiple accelerators to process data. In the case of a large amount of data processing, the host CPU needs to complete complex control, and the stability of the system is poor. Summary of the Invention
[0005] Embodiments of this application provide a data processing method to improve processing efficiency.
[0006] Correspondingly, embodiments of this application also provide a data processing device, an electronic device, and a storage medium to ensure the implementation and application of the above system.
[0007] To solve the above problems, embodiments of this application disclose a data processing method, and the method includes: establishing a connection between a first computing device and a second computing device; the second computing device receives configuration parameters transmitted by the first computing device and configures an accelerator according to the second computing device; the first computing device sends data to be processed to a data cache of an access memory and sends a data write message to the second computing device so that the configured accelerator processes the data to be processed; after the accelerator sends a processing result to a result cache of the access memory, the second computing device sends a result write message to the first computing device so that the first computing device obtains the processing result according to the result write message.
[0008] An embodiment of the present application also discloses a data processing device, which includes: a first computing device and a second computing device, wherein the first computing device is connected to the second computing device; the second computing device receives configuration parameters transmitted by the first computing device to configure an accelerator, and receives a data write message sent by the first computing device, so that the configured accelerator processes the data to be processed cached in the access memory by the first computing device, and sends the processing result to the result cache of the access memory, and sends a result write message to the first computing device, so that the first computing device obtains the processing result according to the result write message.
[0009] An embodiment of the present application also discloses an electronic device, including: a processor, and a memory storing executable code thereon, and when the executable code is executed, the processor executes one or more of the methods in the embodiments of the present application.
[0010] An embodiment of the present application also discloses one or more machine-readable media storing executable code thereon, and when the executable code is executed, a processor executes one or more of the methods in the embodiments of the present application.
[0011] Compared with the prior art, the embodiments of the present application have the following advantages:
[0012] In the embodiments of the present application, a connection between a first computing device and a second computing device is established, the first computing device transmits configuration parameters to the second computing device, the accelerator is configured by the second computing device, then the first computing device sends the data to be processed to the data cache of the access memory, and sends a data write message to the second computing device, the second computing device sends the data write message to the accelerator, so that the accelerator obtains the data to be processed in the data cache of the access memory according to the data write message, and processes the data to be processed, then the accelerator sends the processing result to the result cache of the access memory, and sends a result write message to the first computing device through the second computing device, so that the first computing device obtains the processing result according to the result write message; in the embodiments of the present application, the accelerator can be configured by the second computing device, and according to the message interaction between the first computing device and the second computing device, the configured accelerator performs data processing through the second computing device, reducing the control pressure of the first computing device and improving the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic flowchart of a data processing method according to an embodiment of the present application;
[0014] Figure 2A is a schematic structural diagram of a data processing system according to an embodiment of the present application;
[0015] Figure 2B is a schematic structural diagram of a data processing system according to another embodiment of the present application;
[0016] Figure 3 is a schematic flowchart of a data processing method according to another embodiment of the present application;
[0017] Figure 4 is a schematic flowchart of a data processing method according to still another embodiment of the present application;
[0018] Figure 5 is a schematic structural diagram of a data processing device according to an embodiment of the present application;
[0019] Figure 6 is a schematic structural diagram of an exemplary device according to an embodiment of the present application. Detailed Embodiments
[0020] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] The embodiments of the present application can be applied to the field of data processing based on heterogeneous computing. Heterogeneous computing mainly refers to a system computing method composed of computing units with different types of instruction sets and architectures. The field of data processing based on heterogeneous computing can include fields such as High Efficiency Video Coding (HEVC), artificial intelligence, and speech synthesis.
[0022] The embodiments of the present application disclose a data processing system, such as Figure 2AAs shown, the system includes: a first computing device, a second computing device, and accelerators. The first computing device is used to allocate acceleration tasks to the second computing device; the second computing device is used to communicate with the first computing device to receive acceleration tasks and schedule them to the accelerators. The accelerators are used to perform acceleration processing on data. Among them, the first computing device may include a host CPU, and the second computing device may include a System on Chip (SoC) CPU. Hereinafter, taking the first computing device as the host CPU and the second computing device as the on-chip CPU as an example, the data processing system will be described. The host CPU is used to allocate acceleration tasks to the on-chip CPU; the on-chip CPU is used to communicate with the host CPU to receive acceleration tasks and schedule them to the accelerators. The accelerators can also be referred to as acceleration chips, and the accelerators are used to perform acceleration processing on data. The accelerators may include General-Purpose computing on Graphics Processing Units (GPGPU) and General-Purpose computing Digital Signal Processing (GPDSP) devices, etc.
[0023] Specifically, the first computing device, the second computing device, and the accelerators can all be integrated into chips, plugged on the motherboard, and connected for communication through a PCIe bridge (Peripheral Component Interconnect express Bridge). For example, the on-chip CPU can be a system-level chip including a CPU and a controller. The chips of the host CPU and the on-chip CPU can be plugged on the motherboard, so that the host CPU can be connected to the on-chip CPU through the PCIe bridge to communicate between the host CPU and the accelerators. The host CPU can communicate with the Dynamic Random Access Memory (DRAM) through the PCIE bridge, Direct Memory Access (DMA) connection network, and the path of Chip Interconnect to write the data to be processed and obtain the processing results. The on-chip CPU can control and manage the accelerators through the on-chip interconnect network. The accelerators can communicate with the DRAM through the on-chip interconnect network to obtain the data to be processed and write the processing results.
[0024] In addition, in an optional embodiment, when the data volume of the acceleration task is small, the DRAM can also be a Static Random-Access Memory (SRAM).
[0025] In the embodiments of the present application, the host CPU may schedule the data to be processed to the DRAM and notify the on-chip CPU. According to the notification, the on-chip CPU enables the accelerator to obtain the data to be processed and perform processing. Then, the accelerator may return the processing result to the host CPU through the DRAM. In the embodiments of the present application, the host CPU does not directly handshake with the accelerator, but handshakes with the accelerator through the on-chip CPU, which can configure and control the accelerator by using the on-chip CPU, perform data processing more stably, effectively reduce the control delay of the accelerator, and process data more efficiently.
[0026] Specifically, during the data processing, in addition to communicating data through the DRAM, the host CPU and the on-chip CPU also need to communicate small-format messages such as instructions, parameters, and messages, as Figure 2B shown, the host CPU and the on-chip CPU can communicate instructions, parameters, and messages through the memory. Among them, by communicating instructions on the memory, the startup of the accelerator can be completed and the configuration status of the accelerator can be determined. By communicating configuration parameters on the memory, the accelerator can be configured. By communicating messages on the memory, the write message of the data to be processed of the host CPU can be transmitted to the on-chip CPU, and the write message of the processing result of the accelerator can be returned to the host CPU. Among them, the data receiver can obtain the corresponding data (data to be processed or processing result) from the DRAM according to the write message.
[0027] In the actual processing process, instructions are usually small-scale data, configuration parameters are usually large-scale static data, which can be understood as data that generally does not change during the operation of an acceleration task, and messages are usually large-scale dynamic data, which can be understood as data that changes during an acceleration task. Therefore, in order to transmit small-scale instructions, large-scale static data, and large-scale dynamic data, in an alternative embodiment, the memory may include a first memory, a second memory, and a third memory. The first memory may also be referred to as an instruction memory, and the first memory is used for the communication of instructions. Through the communication of instructions, the startup of the accelerator and the determination of the configuration state of the accelerator can be completed. In an alternative example, when the data volume of the instructions is small, the first memory may be a register; in another alternative example, when the data volume of the instructions is large, the first memory may be implemented based on the SRAM of the on-chip CPU. The second memory may also be referred to as a parameter memory, a shared cache, etc. The second memory is used for the communication of static configuration parameters, and the static configuration parameters are used to configure the accelerator. In an alternative example, when the scale of the static parameters is large, the second memory may be based on an external DRAM; when the scale of the static parameters is small, the second memory may be implemented based on the SRAM of the on-chip CPU. The third memory may also be referred to as a message memory, a message buffer, etc. The third memory is used for the communication of messages, and mainly carries the dynamic information in a single acceleration task, such as the write message (descriptor) of a single input data (such as data to be processed) / output data (such as processing result).
[0028] The embodiment of the present application discloses that the acceleration processing process of the accelerator can be divided into four stages based on the communication between the host CPU and the on-chip CPU, including: a task establishment stage, a task execution stage, an accelerator clearing stage, and a task termination stage. Specifically, in the task establishment stage, the host CPU shakes hands with the accelerator through the on-chip CPU to complete the configuration of the accelerator. In the task execution stage, the accelerator performs acceleration processing on the data to be processed and returns the processing result to the host CPU. In the accelerator clearing stage, when the host CPU has finished transmitting the data to be processed, the on-chip CPU notifies the accelerator to clear the transmitted data to be processed. In the task termination stage, after the host CPU has received the processing result, the on-chip CPU clears and shuts down the accelerator.
[0029] Task establishment stage:
[0030] Step 102: Establish a connection between the first computing device and the second computing device.
[0031] After the host CPU determines an acceleration task, the host CPU can filter out the on-chip CPUs in the idle state and establish a connection between the host CPU and the on-chip CPUs. Specifically, after the accelerator is started, the on-chip CPUs can write an idle instruction (IDLE) in the first memory. The idle instruction indicates that the on-chip CPUs are in the idle state and are ready to receive acceleration tasks. In an optional embodiment, the first computing device determines the idle state of the second computing device; and establishes a connection between the first computing device and the idle second computing device. The host CPU can monitor the idle instructions in the first memory. When the idle instructions are monitored, the host CPU determines that the on-chip CPUs are in the idle state, and the host CPU establishes a connection with the idle on-chip CPUs.
[0032] After establishing a connection between the host CPU and the idle on-chip CPUs, the host CPU can write configuration parameters in the second memory.
[0033] Step 104: The on-chip CPUs receive the configuration parameters transmitted by the host CPU and configure the accelerator according to the second computing device.
[0034] After the host CPU writes the configuration parameters in the second memory, the host CPU can write a start instruction (START) in the first memory. The start instruction is used to notify the on-chip CPUs that the configuration of the accelerator can start. In an optional embodiment, the second computing device can monitor the start instruction written by the first computing device in the first memory; when the start instruction is monitored, the second computing device can obtain the configuration parameters written by the first computing device from the second memory and configure the accelerator.
[0035] In the task execution phase:
[0036] Step 106: The first computing device sends the data to be processed to the data cache of the access memory and sends a data write message to the second computing device so that the configured accelerator can process the data to be processed.
[0037] The host CPU and the on-chip CPU can complete the transmission of data to be processed by accessing a memory (such as DRAM). In an optional embodiment, when the accelerator configuration is completed, the first computing device transmits the data to be processed to the data cache of the access memory. After the accelerator is configured, the on-chip CPU can write a configuration completion instruction (INIT) to the first memory. The configuration completion instruction is used to notify the host CPU that the configuration is completed and data processing can be performed. The host CPU can monitor the configuration completion instruction in the first memory to determine whether the accelerator is configured. When the configuration completion instruction is monitored, the host CPU can transmit the data to be processed to the accelerator through the data cache of the access memory. Then the accelerator processes the data to be processed.
[0038] After the host CPU writes the data to be processed into the data cache of the access memory, it can transmit a data write message to the on-chip CPU through the third memory so that the accelerator can obtain the data to be processed corresponding to the data write message. Specifically, in an optional embodiment, the step of sending a data write message to the second computing device so that the configured accelerator processes the data to be processed includes: the first computing device writes a data write message to the third memory; the second computing device obtains the data write message from the third memory, so that the configured accelerator obtains the data to be processed from the data cache in the access memory and processes it. The access memory includes multiple caches, such as multiple data caches and multiple result caches. The data cache is used to store the data to be processed, and the result cache is used to store the processing results. When storing the data to be processed and the processing results, the corresponding cache can be called to store the data according to the cache status of the data cache and the result cache. Among them, during the data storage process, one cache can be applied for at a time to store the data; multiple caches can also be applied for at a time to store the data. Before the task starts, a data cache manager can be set on the side of the host CPU. The data cache manager is used to manage the cache status of the data cache in the access memory for storing the data to be processed. The host CPU can determine the target data cache capable of storing the data to be processed according to the cache status of the data cache. In an optional embodiment, the first computing device obtains the cache status of the data cache of the access memory and determines the target data cache according to the cache status to store the data to be processed; the first computing device obtains the cache status of the data cache of the access memory and determines the target data cache according to the cache status to store the data to be processed. After the host CPU stores the data to be processed in the target data cache, it can add the address of the target data cache to the data write message (RESERVE IN) in the form of a parameter, and then write the data write message (RESERVE IN) into the third memory. The on-chip CPU reads the data write message (RESERVE IN) from the third memory and notifies the accelerator, so that the accelerator reads the data to be processed from the target data cache according to the data write message (RESERVE IN).
[0039] Accelerator clearing phase:
[0040] After the host CPU writes the data to be processed for the acceleration task into the access memory, it can write a data transfer completion message (FLUSH) to the third memory to notify the on-chip CPU to complete the processing of the data in the access memory. In an optional embodiment, when the first computing device completes the data transfer, the configured accelerator processes the remaining data in the data cache of the access memory. After the on-chip CPU monitors the data transfer completion message (FLUSH) in the third memory, it notifies the accelerator to complete the processing of the remaining data to be processed in the access memory and no longer waits for new data to be processed.
[0041] Step 108: After the accelerator sends the processing result to the result cache of the access memory, the second computing device sends a result write message to the first computing device so that the first computing device can obtain the processing result according to the result write message.
[0042] After the configured accelerator processes the data to be processed, the on-chip CPU can apply for the result cache of the access memory and inform the accelerator of the allocation information of the result cache, so that the accelerator transmits the processing result to the result cache of the access memory and sends a result write message to the first computing device. The first computing device can obtain the processing result from the result cache of the access memory according to the result write message. Among them, the result write message can be sent through the third memory. In an optional embodiment, the sending the result write message to the first computing device so that the first computing device can obtain the processing result according to the result write message includes: the second computing device writes the result write message in the third memory; the first computing device obtains the result write message from the third memory to obtain the processing result from the result cache of the access memory. A result cache manager can be set on the side of the on-chip CPU. The result cache manager is used to manage the cache status of the result cache in the access memory for storing the processing result. The on-chip CPU can determine the target data cache capable of storing the processing result according to the cache status of the result cache. In an optional embodiment, the second computing device obtains the cache status of the result cache of the access memory and determines the target result cache according to the cache status to store the processing result. After the accelerator writes the processing result into the target result cache, the on-chip CPU can write a result write message (RESERVEOUT) to the third memory. The result write message includes the address of the target result cache. The host CPU reads the result write message from the third memory to obtain the corresponding processing result from the target result cache.
[0043] During the task execution phase, in order to release the data cached in the access memory, the data in the access memory can be released after it is consumed. In an alternative embodiment, when the target data to be processed in the data to be processed is processed, the first computing device releases the data cache corresponding to the target data to be processed; when the first computing device finishes receiving the target processing result in the processing result, the second computing device releases the result cache corresponding to the target processing result. After the data to be processed is processed by the accelerator, the on-chip CPU can write a data release message (RELEASEIN) to the third memory, and the host CPU can read the data release message. Based on the data release message, the host CPU can modify the cache status of the data cache in the data cache manager, so as to release the data cache corresponding to the data to be processed, enabling the data cache to store subsequent data to be processed. In one example, the host CPU can overwrite the data to be processed with new data to be processed to release the data cache. After the host CPU receives the processing result, the host CPU can write a result release message (RELEASE OUT) to the third memory, and the on-chip CPU can read the result release message from the third memory. The on-chip CPU can modify the cache status of the result cache in the result cache manager according to the result release message, so as to release the result cache corresponding to the processing result. In one example, the on-chip CPU can overwrite the processing result in the result cache with a new processing result.
[0044] In the above embodiment, the sender of the task actively applies for a cache to store data. In addition to writing data to the access memory in this way, the embodiment of the present application can also apply for caches for storing the data to be processed and the processing result through the host CPU. Specifically, in an alternative embodiment, the first computing device determines the cache status of the data cache and the result cache of the target memory, so as to determine the target data cache and the target result cache according to the cache status. The target data cache is used to store the data to be processed, and the target result cache is used to store the processing result. Before the data to be processed is sent, the host CPU can apply for the data cache corresponding to the data to be processed and the result cache corresponding to the processing result to determine the data cache allocation message and the result cache allocation message. The host CPU can store the data to be processed according to the data cache allocation message. The host CPU can use the third memory to send the result cache allocation message to the on-chip CPU, and the on-chip CPU can store the processing result according to the result cache allocation message.
[0045] In the method of determining the data cache and result cache for accessing the memory by the host CPU, the cache occupied by the data to be processed and the processing result can be released after the host CPU receives the processing result. In an optional embodiment, when the first computing device finishes receiving the processing result, the first computing device releases the result cache corresponding to the target processing result and releases the data cache of the data to be processed corresponding to the target processing result. In this embodiment, the host CPU sends the data to be processed to the accelerator through the data cache of the access memory. After the data to be processed is accelerated by the accelerator, a processing result is obtained. Then the accelerator returns the processing result to the host CPU through the result cache of the access memory. After receiving the processing result, the host CPU releases the result cache corresponding to the target processing result and the data cache of the data to be processed corresponding to the target processing result. Compared with the method of releasing after the data to be processed is consumed, the embodiment of the present application can reduce the number of communications between the host CPU and the on-chip CPU and can release the cache more simply.
[0046] In the task termination phase:
[0047] After the host CPU finishes receiving the processing result corresponding to the acceleration task, the host CPU can also send a notification message to the on-chip CPU to turn off the accelerator through the on-chip CPU. Specifically, in an optional embodiment, when the first computing device finishes receiving the processing result, the second computing device turns off the accelerator and writes an idle instruction (IDLE) to the first memory to receive a setup instruction for the next data processing.
[0048] The host CPU can send a result received instruction (CLOSE) to the on-chip CPU to notify the reception of the processing result. Specifically, in an optional embodiment, the first computing device writes a result received instruction to the first memory, and the second computing device monitors the result received instruction written by the first computing device in the first memory to determine the reception situation of the processing result by the first computing device. When the host CPU finishes receiving the processing result, the on-chip CPU can clear the data of the accelerator and turn off the accelerator. At this time, the accelerator has finished processing the data to be processed in the task and is in an idle state. The on-chip CPU can write an idle instruction to the first memory for the next task processing.
[0049] In addition, to enable the accelerator to reprocess data in the event of an exception occurring in the host CPU, in an optional embodiment, when a preset condition is met, the first computing device writes a setup instruction (START) to the first memory, and the preset condition includes at least one of data write exception of the first computing device, data reception exception of the first computing device, and abnormal exit of the first computing device. During the data processing, the host CPU may accidentally cause an abnormal exit, a data write exception, or a data reception exception. In this case, the host CPU can re-write the setup instruction in the first memory to cause the on-chip CPU to clear the data related to the accelerator and reconfigure the accelerator to re-complete the data processing process, facilitating exception recovery and improving the robustness of the entire system.
[0050] In the embodiment of the present application, the connection between the host CPU and the on-chip CPU is established, the host CPU transmits configuration parameters to the on-chip CPU, configures the accelerator through the on-chip CPU, and then the host CPU sends the data to be processed to the data cache of the access memory and sends a data write message to the on-chip CPU. The accelerator obtains the data to be processed in the data cache of the access memory according to the data write message, processes the data to be processed, and then the accelerator sends the processing result to the result cache of the access memory and sends a result write message to the host CPU so that the host CPU can obtain the processing result according to the result write message. In the embodiment of the present application, the accelerator can be configured by the on-chip CPU, and the configured accelerator can perform data processing through the message interaction between the host CPU and the on-chip CPU, reducing the control pressure of the host CPU and enabling more stable and faster data processing.
[0051] Based on the above embodiments, the embodiment of the present application further provides a data processing method, which can divide the processing of the acceleration task into the following four stages: task establishment stage, task execution stage, accelerator clearing stage, and task termination stage. In this embodiment, the first computing device may include a host CPU, the second computing device may include an on-chip CPU, and the host CPU and the on-chip CPU communicate through a memory and an access memory. The memory is used for communicating instructions, configuration parameters, and messages. The memory includes a first memory, a second memory, and a third memory. The access memory is used for communicating the data to be processed and the processing result.
[0052] The task establishment stage includes the process of establishing the connection between the host CPU and the on-chip CPU and the process of configuring the accelerator by the on-chip CPU. Specifically, as Figure 3 shown, the task establishment stage includes:
[0053] Step 302: The on-chip CPU writes an idle instruction (IDLE) to the first memory. After the accelerator is started or after the accelerator finishes processing the data to be processed, the on-chip CPU can write an idle instruction to the first memory to indicate that the accelerator is capable of processing data.
[0054] Step 304: The host CPU can periodically monitor the idle instruction (IDLE) in the first memory to determine whether the accelerator is capable of processing data. If there is an idle instruction in the first memory, it is determined that the accelerator can process data; if there is no idle instruction in the first memory, it is determined that the accelerator cannot process data.
[0055] Step 306: The host CPU writes configuration parameters to the second memory, and the configuration parameters are used to configure the accelerator.
[0056] Step 308: After the host CPU writes the configuration parameters to the second memory, the host CPU writes a start instruction (START) to the first memory.
[0057] Step 310: The on-chip CPU can periodically monitor the start instruction (START) in the first memory.
[0058] Step 312: When the start instruction is monitored, the on-chip CPU obtains the configuration parameters from the second memory to configure the accelerator.
[0059] Step 314: When the accelerator configuration is completed, the on-chip CPU writes a configuration completed instruction (INIT) to the first memory.
[0060] Step 316: The host CPU periodically monitors the configuration completed instruction (INIT) in the first memory. When the configuration completed instruction is monitored, the host CPU completes the handshake with the accelerator through the on-chip CPU and can send the data to be processed to the accelerator for data processing.
[0061] The task execution stage includes: the process of the configured accelerator processing the data to be processed and the process of returning the processing result of the accelerator to the host CPU. Specifically, the task execution stage includes:
[0062] Step 318: When the host CPU monitors the configuration completed instruction in the first memory, it writes the data to be processed to the access memory and writes a data write message (RESERVEIN) to the third memory. The host CPU can determine the cache status of the data cache according to the data cache manager and determine the data cache capable of storing the data to be processed.
[0063] Step 320: The on-chip CPU obtains a data write message from the third memory so that the accelerator can obtain the data to be processed corresponding to the data write message and process it. The host CPU can apply for one data cache at a time to store the data to be processed, or apply for multiple caches at a time to store multiple data to be processed. Therefore, the on-chip CPU can, upon receiving a data write message, notify the accelerator to process the data; the on-chip CPU can also, upon receiving multiple data write messages, notify the accelerator to process the data.
[0064] Step 322: The on-chip CPU writes a data release message (RELEASE IN) to the third memory.
[0065] Step 324: The host CPU obtains the data release message from the third memory to release the data cache corresponding to the data to be processed in the storage memory.
[0066] Step 326: The on-chip CPU writes the processing result of the data to be processed to the access memory and writes a result write message to the third memory. The on-chip CPU can determine the cache status of the result cache according to the result cache manager and determine the result cache capable of storing the processing result.
[0067] Step 328: The host CPU obtains the result write message from the third memory to obtain the processing result from the access memory.
[0068] Step 330: The host CPU writes a result release message (RELEASE OUT) to the third memory.
[0069] Step 332: The on-chip CPU obtains the result release message from the third memory to release the result cache corresponding to the processing result.
[0070] The accelerator emptying phase includes the process of the host CPU sending all the data to be processed and the process of the accelerator processing the remaining data. Specifically, the accelerator emptying phase includes:
[0071] Step 334: After the host CPU writes all the data to be processed of the task to the third memory, it writes a data transfer complete message (FLUSH) to the third memory.
[0072] Step 336: The on-chip CPU obtains the data transfer complete message. The on-chip CPU processes the remaining data in the access memory according to the data transfer complete message and no longer waits for new data.
[0073] Step 338: The on-chip CPU writes the processing result of the remaining data to the access memory and writes a result write message to the third memory.
[0074] Step 340: The host CPU obtains the result write message from the third accessor to obtain the processing result from the access memory.
[0075] Step 342: The host CPU writes a result release message (RELEASE OUT) to the third memory.
[0076] Step 344: The on-chip CPU obtains the result release message from the third memory to release the result cache corresponding to the processing result.
[0077] The task termination phase may include: the process of the host CPU receiving the processing message completely and the process of shutting down the accelerator. Specifically, the task termination phase includes:
[0078] Step 346: After the host CPU receives the processing result of the task, it writes a result received instruction (CLOSE) to the first memory to shut down the accelerator.
[0079] Step 348: The on-chip CPU obtains the result received instruction in the first memory, clears the data in the accelerator, and shuts down the accelerator.
[0080] Step 350: After the on-chip CPU shuts down the accelerator, it writes an idle instruction to the first memory for the next data processing.
[0081] In the embodiment of the present application, when the on-chip CPU is in an idle state of the accelerator, it writes an idle instruction (IDLE) to the first memory. The host CPU writes configuration parameters to the second memory according to the idle instruction and writes a start instruction (START) to the first memory. The on-chip CPU obtains the configuration parameters from the second memory according to the start instruction to configure the accelerator. When the accelerator configuration is completed, the on-chip CPU writes a configuration completed instruction (INIT) to the first memory. The host CPU writes the data to be processed to the access memory according to the configuration completed instruction. The accelerator obtains the data to be processed from the access memory and performs data processing to obtain a processing result. The on-chip CPU writes the processing result to the access memory. The host CPU obtains the processing result from the access memory. After the host CPU receives the processing result of the task, the host CPU can write a result received instruction (CLOSE) to the first memory. The on-chip CPU shuts down the accelerator according to the result received instruction and writes an idle instruction to the first memory for the next data processing. In the embodiment of the present application, the accelerator can be configured by the on-chip CPU and the configured accelerator can perform data processing, so that the data processing can be more stable.
[0082] Based on the above embodiments, the present application further provides a data processing method. The difference between this method and the data processing method of the above embodiments lies in the different task execution phases. In the task execution phase of the above embodiments, after the data is consumed, through the message interaction between the host CPU and the on-chip CPU, the data can be released. However, using this method, the message volume between the host CPU and the on-chip CPU is relatively large. Specifically, as Figure 4 shown,
[0083] The task establishment phase includes:
[0084] Step 402: The on-chip CPU writes an idle instruction (IDLE) to the first memory in the exchange layer. After the accelerator is started or the accelerator finishes processing the data to be processed, the on-chip CPU can write an idle instruction to the first memory to indicate that the accelerator can process data.
[0085] Step 404: The host CPU can periodically monitor the idle instruction (IDLE) in the first memory to determine whether the accelerator can process data. If there is an idle instruction in the first memory, it is determined that the accelerator can process data; if there is no idle instruction in the first memory, it is determined that the accelerator cannot process data.
[0086] Step 406: The host CPU writes configuration parameters to the second memory, and the configuration parameters are used to configure the accelerator.
[0087] Step 408: After the host CPU writes the configuration parameters to the second memory, the host CPU writes a start instruction (START) to the first memory.
[0088] Step 410: The on-chip CPU can periodically monitor the start instruction (START) in the first memory.
[0089] Step 412: When the start instruction is monitored, the on-chip CPU obtains the configuration parameters from the second memory to configure the accelerator.
[0090] Step 414: When the accelerator configuration is completed, the on-chip CPU writes a configuration completed instruction (INIT) to the first memory.
[0091] Step 416: The host CPU periodically monitors the configuration completed instruction (INIT) in the first memory. When the configuration completed instruction is monitored, the host CPU completes the handshake with the accelerator through the on-chip CPU and can send the data to be processed to the on-chip CPU for data processing.
[0092] The task execution phase includes:
[0093] Step 418: When the configuration completion instruction is monitored, the host CPU writes the data to be processed into the access memory and writes a data write and result cache allocation message into the third memory. Among them, the host CPU determines the data cache capable of storing the data to be processed and the result cache capable of storing the processing result according to the cache status of the data cache and the result cache. The data write and result cache allocation message may include a data write message and a result cache allocation message. The on-chip CPU can read the data to be processed in the access memory according to the data write message, and the on-chip CPU can store the processing result in the access memory according to the result cache allocation message.
[0094] Step 420: The on-chip CPU reads the data write and result cache allocation message from the third memory; the on-chip CPU sends the data write message to the accelerator, so that the accelerator obtains the data to be processed from the access memory and processes the data to be processed.
[0095] Step 422: The on-chip CPU writes the processing result into the access memory according to the result cache allocation message and writes a result write message into the third memory.
[0096] Step 424: The host CPU reads the result write message from the third memory to obtain the processing result from the access memory. And releases the processed result and related data to be processed cached in the access memory.
[0097] The accelerator emptying stage includes:
[0098] Step 426: After the host CPU writes all the data to be processed of the task into the third memory, it writes a data transfer complete message (FLUSH) into the third memory.
[0099] Step 428: The on-chip CPU obtains the data transfer complete message. The on-chip CPU processes the remaining data in the access memory according to the data transfer complete message and no longer waits for new data.
[0100] Step 430: The on-chip CPU writes the processing result of the remaining data into the access memory and writes a result write message into the third memory.
[0101] Step 432: The host CPU obtains the result write message from the third accessor to obtain the processing result from the access memory.
[0102] Step 434: The host CPU writes a result release message (RELEASE OUT) into the third memory.
[0103] Step 436: The on-chip CPU obtains the result release message from the third memory to release the corresponding processing result.
[0104] The task termination phase may include:
[0105] Step 438, after the host CPU receives the processing result of the task, it writes a result received instruction (CLOSE) to the first memory to turn off the accelerator.
[0106] Step 440, the on-chip CPU obtains the result received instruction in the first memory, clears the data in the accelerator, and turns off the accelerator.
[0107] Step 442, after the on-chip CPU turns off the accelerator, it writes an idle instruction to the first memory for the next data processing.
[0108] In the embodiment of the present application, after the host CPU receives the processing result, the host CPU releases the data to be processed and the processing result cached in the access memory, which can reduce the amount of messages for interaction between the host CPU and the on-chip CPU. The other phases in the embodiment of the present application are similar to the processing methods in the above embodiments and will not be elaborated here.
[0109] In the embodiment of the present application, the first computing device can implement the configuration of the accelerator and the transfer of processing tasks based on the interaction with the second computing device. Among them, in some scenarios, the cache required for data can be applied during the execution process, such as the Figure 3 processing method described above, while in some scenarios, the required caches can be applied in advance, such as the Figure 4 processing method described above. Therefore, in some alternative embodiments, tags can be set based on the above different cache processing methods, so that different tags can be selected according to the application scenario, and data can be processed based on different cache methods. For example, in some scenarios, the computing device has good computing power and can adopt the Figure 3 method described above, so as not to occupy too much cache. In some other scenarios, if the memory such as DRAM has good storage capacity, the Figure 4 method described above can be selected to bind the cache with the processing of the computing device and the accelerator, and the processing process is simpler and more reliable.
[0110] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.
[0111] Based on the above embodiments, this embodiment further provides a data processing device, such as Figure 5As shown in the figure, it may specifically include: a first computing device and a second computing device. The first computing device may include a host CPU, and the second computing device may include an on-chip CPU. Among them, the first computing device is connected to the second computing device;
[0112] The second computing device receives the configuration parameters transmitted by the first computing device to configure the accelerator, and receives the data write message sent by the first computing device, so that the configured accelerator processes the data to be processed in the data cache of the access memory sent by the first computing device, and sends the processing result to the result cache of the access memory, and sends a result write message to the first computing device, so that the first computing device obtains the processing result according to the result write message.
[0113] In summary, in the embodiment of the present application, the structures of the first computing device and the second computing device can be similar to Figure 2A the structure of the data processing system described above, which will not be elaborated here. In the embodiment of the present application, the connection between the first computing device and the second computing device is established, the first computing device transmits configuration parameters to the second computing device, the accelerator is configured through the second computing device, and the configured accelerator is used to process the data to be processed transmitted by the first computing device, and then the accelerator returns the processing result to the first computing device; in the embodiment of the present application, the accelerator can be configured through the second computing device and the configured accelerator can perform data processing, so that the data processing can be more stable.
[0114] Based on the above embodiment, this embodiment further provides a data processing device, which may specifically include: a first computing device and a second computing device. The first computing device may include a host CPU, and the second computing device may include an on-chip CPU. Among them, the first computing device is connected to the second computing device;
[0115] The second computing device receives the configuration parameters transmitted by the first computing device to configure the accelerator, and makes the configured accelerator process the data to be processed transmitted by the first computing device, and sends the processing result to the first computing device.
[0116] The first memory receives the idle instruction written by the second computing device, so that the first computing device determines the idle state of the second computing device according to the idle instruction, and establishes a connection between the first computing device and the idle second computing device.
[0117] A second memory that receives configuration parameters written by a first computing device, so that a second computing device configures an accelerator according to the configuration parameters. Among them, the first memory can receive a setup instruction written by the first computing device, so that the second computing device, when monitoring the setup instruction, obtains the configuration parameters from the second memory.
[0118] An access memory that receives data to be processed transmitted by a first computing device to transmit the data to be processed to the accelerator for data processing; receives the processing result transmitted by the accelerator to transmit the processing result to the first computing device.
[0119] A third memory that receives a data write message corresponding to the data to be processed written by the first computing device, so that the accelerator obtains the data to be processed in the data cache of the access memory according to the data write message; receives a result write message corresponding to the processing result written by the second computing device, so that the host CPU obtains the processing result in the result cache of the access memory according to the result write message. In addition, the third memory can also receive a data transmission completion message written by the first computing device, so that the configured accelerator processes the remaining data in the data cache of the access memory. The first memory can also receive a result reception completion instruction written by the first computing device, so that the second computing device closes the accelerator according to the result reception completion instruction and writes an idle instruction to the first memory.
[0120] In the embodiments of the present application, the device structure of the embodiments of the present application is similar to Figure 2A the structure of the data processing system shown, which will not be elaborated here. When the accelerator is in an idle state, the second computing device writes an idle instruction to the first memory. The first computing device writes configuration parameters to the second memory according to the idle instruction and writes a setup instruction to the first memory. The second computing device obtains the configuration parameters from the second memory according to the setup instruction to configure the accelerator. When the accelerator configuration is completed, the second computing device writes a configuration completion instruction to the first memory. The first computing device writes the data to be processed to the access memory according to the configuration completion instruction, so that the accelerator obtains the data to be processed from the access memory and performs data processing to obtain a processing result. The second computing device writes the processing result to the access memory. The first computing device obtains the processing result from the access memory. After the first computing device receives the processing result of the task, the first computing device can write a result reception completion instruction to the first memory. The second computing device closes the accelerator according to the result reception completion instruction and writes an idle instruction to the first memory for the next data processing. In the embodiments of the present application, the accelerator can be configured by the second computing device and the configured accelerator can perform data processing, enabling more stable data processing.
[0121] Optionally, as an embodiment, the data processing device further includes:
[0122] A data cache manager that stores the cache status of the data cache in the access memory, so that the first computing device determines a target data cache according to the cache status of the data cache to store the data to be processed.
[0123] A result cache manager that stores the cache status of the result cache in the access memory, so that the second computing device determines a target result cache according to the cache status of the result cache to store the processing result.
[0124] The third memory can also receive a data release message corresponding to the target data to be processed transmitted by the second computing device, so as to release the data cache corresponding to the target data to be processed when the target data to be processed in the data to be processed is processed.
[0125] The third memory can also receive a result release message corresponding to the target processing result transmitted by the first computing device, so as to release the result cache corresponding to the target processing result when the target processing result in the processing result is received.
[0126] Optionally, as an embodiment, the data processing device further includes:
[0127] A data cache manager that stores the cache status of the data cache in the access memory, so that the first computing device determines a target data cache according to the cache status of the data cache to store the data to be processed.
[0128] A result cache manager that stores the cache status of the result cache in the access memory, so that the first computing device determines a target result cache according to the cache status of the result cache to store the processing result.
[0129] The third memory can receive a result cache allocation message transmitted by the first computing device, so that the second computing device transmits the processing result to the result cache of the access memory according to the result cache allocation message.
[0130] In this embodiment, the first computing device sends the data to be processed to the second computing device through the access memory. After the data to be processed is accelerated by the accelerator, a processing result is obtained. Then the accelerator returns the processing result to the first computing device through the access memory. After receiving the processing result, the first computing device releases the data to be processed and the processing result in the access memory. Compared with the method of releasing after the data to be processed is consumed, the embodiment of the present application can reduce the communication times between the first computing device and the second computing device and can release the cache more simply.
[0131] The embodiments of the present application further provide a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can be caused to execute instructions for each method step in the embodiments of the present application.
[0132] The embodiments of the present application provide one or more machine-readable media, on which instructions are stored. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments. In the embodiments of the present application, the electronic device includes devices such as servers and terminal devices.
[0133] The embodiments of the present disclosure can be implemented as a device configured with any suitable hardware, firmware, software, or any combination thereof. The device may include electronic devices such as servers (clusters) and terminals. Figure 6 Exemplary device 600 that can be used to implement the various embodiments described in the present application is schematically shown.
[0134] For one embodiment, Figure 6 Exemplary device 600 is shown, which has one or more processors 602, a control module (chipset) 604 coupled to at least one of the (one or more) processors 602, a memory 606 coupled to the control module 604, a non-volatile memory (NVM) / storage device 608 coupled to the control module 604, one or more input / output devices 610 coupled to the control module 604, and a network interface 612 coupled to the control module 604.
[0135] Processor 602 may include one or more single-core or multi-core processors. Processor 602 may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). Processor 602 may include a first computing device and a second computing device, such as a host CPU and an on-chip CPU. In some embodiments, device 600 is capable of acting as the server, terminal, etc. devices described in the embodiments of the present application.
[0136] In some embodiments, device 600 may include one or more computer-readable media (such as memory 606 or NVM / storage device 608) having instructions 614 and one or more processors 602 combined with the one or more computer-readable media and configured to execute instructions 614 to implement modules and thereby perform the actions described in the present disclosure.
[0137] For one embodiment, the control module 604 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 602 and / or any suitable device or component communicating with the control module 604.
[0138] The control module 604 may include a memory controller module to provide an interface to the memory 606. The memory controller module can be a hardware module, a software module, and / or a firmware module.
[0139] The memory 606 may be used to load and store data and / or instructions 614 for the device 600, for example. For one embodiment, the memory 606 may include any suitable volatile memory, e.g., suitable DRAM. In some embodiments, the memory 606 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0140] For one embodiment, the control module 604 may include one or more input / output controllers to provide an interface to the NVM / storage device 608 and the input / output device(s) 610.
[0141] For example, the NVM / storage device 608 may be used to store data and / or instructions 614. The NVM / storage device 608 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact discs (CDs) drives, and / or one or more digital versatile discs (DVDs) drives).
[0142] The NVM / storage device 608 may include storage resources that are part of a device on which the device 600 is installed, or it may be accessible by the device without being part of the device. For example, the NVM / storage device 608 may be accessed via the input / output device(s) 610 over a network.
[0143] The input / output device(s) 610 may provide an interface for the device 600 to communicate with any other suitable device. The input / output device 610 may include communication components, audio components, sensor components, etc. The network interface 612 may provide an interface for the device 600 to communicate over one or more networks. The device 600 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, e.g., access a wireless network based on a communication standard such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.
[0144] For one embodiment, at least one of the (one or more) processors 602 may be logically encapsulated with one or more controllers (e.g., memory controller modules) of the control module 604. For one embodiment, at least one of the (one or more) processors 602 may be logically encapsulated with one or more controllers of the control module 604 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 602 may be logically integrated with one or more controllers of the control module 604 on the same die. For one embodiment, at least one of the (one or more) processors 602 may be logically integrated with one or more controllers of the control module 604 on the same die to form a system-on-chip (SoC).
[0145] In various embodiments, the device 600 may be, but is not limited to: a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.) and other terminal devices. In various embodiments, the device 600 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 600 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.
[0146] Among them, a main control chip may be used as the processor or the control module in the detection device, sensor data, location information, etc. are stored in the memory or the NVM / storage device, the sensor group may be used as an input / output device, and the communication interface may include a network interface.
[0147] The embodiments of the present application also provide an electronic device, including: a processor; and a memory, on which executable code is stored, and when the executable code is executed, the processor is caused to execute one or more of the methods as in the embodiments of the present application.
[0148] The embodiments of the present application also provide one or more machine-readable media, on which executable code is stored, and when the executable code is executed, the processor is caused to execute one or more of the methods as in the embodiments of the present application.
[0149] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference may be made to the partial description of the method embodiments.
[0150] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other.
[0151] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks
[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks
[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks
[0154] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application
[0155] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0156] The above has introduced in detail a data processing method, a data processing device, an electronic device and a storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A data processing method, characterized in that, The method includes: Establishing a connection between a first computing device and a second computing device; The second computing device monitors a setup instruction written by the first computing device in a first memory; When the setup instruction is monitored, the second computing device obtains configuration parameters written by the first computing device from a second memory, and configures an accelerator according to the second computing device; The first computing device sends data to be processed to a data cache of an access memory; The first computing device writes a data write message to a third memory; The second computing device obtains the data write message from the third memory, so that the configured accelerator obtains the data to be processed from the data cache in the access memory and processes it; After the accelerator sends the processing result to a result cache of the access memory, the second computing device sends a result write message to the first computing device, so that the first computing device obtains the processing result according to the result write message.
2. The method according to claim 1, wherein The establishing a connection between the first computing device and the second computing device includes: The first computing device determines the idle state of the second computing device; Establishing a connection between the first computing device and the idle second computing device.
3. The method according to claim 1, characterized in that, The sending a result write message to the first computing device, so that the first computing device obtains the processing result according to the result write message, includes: The second computing device writes the result write message in the third memory; The first computing device obtains the result write message from the third memory to obtain the processing result from the result cache of the access memory.
4. The method according to claim 3, characterized in that It further includes at least one of the following steps: The first computing device obtains the cache state of the data cache of the access memory, and determines a target data cache according to the cache state to store the data to be processed; The second computing device obtains the cache state of the result cache of the access memory, and determines a target result cache according to the cache state to store the processing result.
5. The method according to claim 4, wherein It further includes: When the first computing device completes data transmission, the configured accelerator processes the remaining data in the data cache of the access memory.
6. The method according to claim 5, characterized in that It further includes at least one of the following steps: When the target data to be processed in the data to be processed is processed, the first computing device releases the data cache corresponding to the target data to be processed; When the first computing device finishes receiving the target processing result in the processing result, the second computing device releases the result cache corresponding to the target processing result.
7. The method according to claim 3, wherein It further includes: The first computing device determines the cache states of the data cache and the result cache of the target memory, so as to determine the target data cache and the target result cache according to the cache states, where the target data cache is used to store the data to be processed, and the target result cache is used to store the processing result.
8. The method according to claim 7, wherein It further includes: When the first computing device finishes receiving the target processing result, the first computing device releases the result cache corresponding to the target processing result and releases the data cache of the data to be processed corresponding to the target processing result.
9. The method according to claim 3, characterized in that, It further includes: When the first computing device has finished receiving the processing result, the second computing device turns off the accelerator and writes an idle instruction to the first memory to receive a setup instruction.
10. The method according to claim 9, characterized in that, Further included are: The second computing device monitors the result reception completion instruction written by the first computing device in the first memory to determine the processing result reception status of the first computing device.
11. The method according to claim 1, wherein Further included are: When a preset condition is met, the first computing device writes a setup instruction to the first memory, and the preset condition includes at least one of abnormal data writing of the first computing device, abnormal data reception of the first computing device, and abnormal exit of the first computing device.
12. A data processing device, characterized in that, The apparatus includes: a first computing device and a second computing device, wherein the first computing device is connected to the second computing device; The second computing device receives the setup instruction written by the first computing device in the first memory monitored by the second computing device; when the setup instruction is monitored, the second computing device obtains the configuration parameters written by the first computing device from the second memory to configure the accelerator, and obtains a data writing message from the third memory, so that the configured accelerator processes the data to be processed in the data cache of the access memory sent by the first computing device, and sends the processing result to the result cache of the access memory, and sends a result writing message to the first computing device, so that the first computing device obtains the processing result according to the result writing message; The first computing device sends the data to be processed to the data cache of the access memory; writes a data writing message to the third memory.
13. An electronic device, characterized in that, Included are: a processor, and a memory storing executable code thereon, which when executed causes the processor to execute the method according to one or more of claims 1-11.
14. One or more machine-readable media storing executable code thereon, which when executed causes a processor to execute the method according to one or more of claims 1-11.
Citation Information
Patent Citations
Deep belief network acceleration platform based on multi-FPGA and design method thereof
CN108805277A