Data processing system and method

By using multiple FPGAs connected rings in a distributed system for data processing, the problem of low data computing efficiency in a distributed system is solved, efficient data transmission and calculation are achieved, and the dependence on the master node is reduced.

CN114385066BActive Publication Date: 2025-05-30WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011112119.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-16
Publication Date
2025-05-30
Estimated Expiration
2040-10-16

AI Technical Summary

Technical Problem

In distributed systems, data calculation efficiency is low, especially when summing, XOR and other calculations are required for data from multiple nodes. The existing technology relies on the master node for unified calculation, which is low in efficiency.

Method used

Multiple FPGAs connected in a ring form are used to transmit data through optical fibers. Each FPGA has a receiving optical port and a transmit optical port to realize the ring transmission and calculation of data, allowing each FPGA to participate in the computing process and reduce dependence on the master node.

Benefits of technology

Quickly realize data transmission and calculation through FPGA plus optical fiber, improving data processing efficiency in distributed systems, reducing dependence on master nodes, and improving computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114385066B_ABST
    Figure CN114385066B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing system and method. The system includes a plurality of FPGAs connected in a ring. Each FPGA is provided with a receiving optical port and a transmitting optical port. The receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber. Each FPGA is used to perform at least one of the following: receive data sent by the previous FPGA through the receiving optical port; perform calculations based on the received data and the data it has to obtain a calculation result; send data to the next FPGA through the transmitting optical port, where the data sent is the calculation result or the data it has. The above system and method can quickly realize data transmission through the combination of FPGA and optical fiber. Moreover, in a distributed system constructed by a plurality of FPGAs connected in a ring, multiple FPGAs can all participate in the calculation process, effectively improving the calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to a data processing system and method. Background Art

[0002] With the continuous development of computer technology and the continuous improvement of data processing technology, the demand for processing various data is also increasing. In order to improve the processing efficiency, the analysis and processing of massive data can be achieved through a distributed system.

[0003] In a distributed system, data is stored in different nodes. For example, when analyzing the consumption behavior of users, consumption data can be stored on different consumption data servers. Data on different servers may need to be calculated, such as summation, exclusive OR, etc. Currently, the distributed system has the problem of low data processing efficiency. Summary of the Invention

[0004] The main purpose of the present invention is to provide a data processing system and method, aiming to solve the technical problem of low data calculation efficiency in a distributed environment.

[0005] To achieve the above object, an embodiment of the present invention provides a data processing system, which includes: a plurality of FPGAs connected in a ring;

[0006] Each FPGA is provided with a receiving optical port and a transmitting optical port;

[0007] The receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber;

[0008] Each FPGA is used to perform at least one of the following: receive data sent by the previous FPGA through the receiving optical port; calculate according to the received data and the data it has to obtain a calculation result; send data to the next FPGA through the transmitting optical port, where the data sent is the calculation result or the data it has.

[0009] Optionally, the system is used to implement full reduction calculation; the full reduction calculation is used to calculate based on the data of each FPGA and make each FPGA obtain the same final calculation result; the full reduction calculation includes a calculation stage and a synchronization stage;

[0010] In the calculation stage, multiple FPGAs sequentially transfer data and perform calculations to obtain a final calculation result;

[0011] In the synchronization stage, the FPGA that obtains the final calculation result transfers the final calculation result to other FPGAs.

[0012] Optionally, the number of FPGAs is N;

[0013] In the calculation stage, the first FPGA is specifically configured to: send the data it has to the second FPGA;

[0014] Each of the second to the (N - 1)th FPGAs is specifically configured to: perform a calculation based on the data obtained from the previous FPGA and the data it has, obtain a calculation result, and send the calculation result to the next FPGA;

[0015] The Nth FPGA is specifically configured to: perform a calculation based on the data obtained from the previous FPGA and the data it has, to obtain a final calculation result;

[0016] Wherein, N is an integer greater than or equal to 3.

[0017] Optionally, the number of FPGAs is N; each FPGA has K columns of data, and after the full reduction calculation is completed, K columns of final calculation results are obtained;

[0018] In the calculation stage, at least some of the FPGAs are specifically configured to: simultaneously send the data they have to the next FPGA, wherein different FPGAs send different columns of data; if data is obtained from the previous FPGA, then perform a calculation based on the obtained data and the corresponding column of data it has, and if the obtained calculation result is not the final calculation result, then send the calculation result to the next FPGA;

[0019] If there are other FPGAs among the multiple FPGAs other than the at least some of the FPGAs, then each of the other FPGAs is specifically configured to: if data is obtained from the previous FPGA, then perform a calculation based on the obtained data and the corresponding column of data it has, and if the obtained calculation result is not the final calculation result, then send the calculation result to the next FPGA;

[0020] Wherein, N is an integer greater than or equal to 3, and K is an integer greater than or equal to 2.

[0021] Optionally, when K ≤ N, the number of the at least some of the FPGAs is K;

[0022] When K > N, the K columns are divided into M groups, and the number of columns in each group is less than or equal to N, so as to obtain K columns of final calculation results by performing M calculation stages, and each calculation stage is used to calculate a group of data;

[0023] Wherein, when calculating the data of the Mth group, the number of the at least some of the FPGAs is the number of columns corresponding to the Mth group; M is an integer greater than or equal to 1.

[0024] Optionally, if the data sent by each FPGA is the smallest unit for performing full reduction calculation, the FPGA that receives the data is used to perform calculations based on the data after the data reception is completed; if the data is not the smallest unit for performing full reduction calculation, the FPGA that receives the data is used to perform calculations based on the data of the smallest unit that has been received during the data reception process.

[0025] Optionally, the system further includes: a control device; the multiple FPGAs are respectively connected to the control device; the control device is used for:

[0026] Send basic operation instructions to each FPGA, so that each FPGA performs basic operations according to the basic operation instructions to obtain data for performing full reduction calculation;

[0027] Obtain status information sent by each FPGA for indicating whether the basic operation is completed;

[0028] According to the status information sent by each FPGA, select the FPGAs that initiate full reduction calculation from the multiple FPGAs;

[0029] Send a full reduction start instruction to the FPGA that initiates full reduction calculation, so that the FPGA sends the data for performing full reduction calculation to the next FPGA according to the full reduction start instruction;

[0030] Send a wait instruction to other FPGAs, so that the other FPGAs wait to obtain the data for performing full reduction calculation and perform calculations after obtaining it. Optionally, when the control device selects the FPGAs that initiate full reduction calculation from the multiple FPGAs, it is specifically used for: determining the number of FPGAs that initiate full reduction calculation; according to the status information sent by each FPGA, select the corresponding number of FPGAs that finally complete the basic operation as the FPGAs that initiate full reduction calculation;

[0031] Wherein, the corresponding number is the number of FPGAs that initiate full reduction calculation.

[0032] Optionally, the data possessed by the FPGA is stored in the memory, and the FPGA reads data from and / or writes data to the memory through direct memory access technology.

[0033] The present invention also provides a data processing method, and the method includes:

[0034] Each FPGA in at least part of the FPGAs sends the data it possesses to the next FPGA through an optical fiber; wherein, the at least part of the FPGAs are at least part of the multiple FPGAs connected in a ring;

[0035] Repeat the following steps until the final calculation result is obtained: Each FPGA that has obtained data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber.

[0036] Optionally, the method is used to implement full reduction calculation; the full reduction calculation is used to calculate based on the data of each FPGA and make each FPGA obtain the same final calculation result; correspondingly, the method further includes:

[0037] The FPGA that has obtained the final calculation result transmits the final calculation result to other FPGAs.

[0038] Optionally, the number of FPGAs is N; each FPGA in at least part of the FPGAs sends the data it has to the next FPGA through an optical fiber, including:

[0039] The first FPGA sends the data it has to the next FPGA.

[0040] Correspondingly, repeat the following steps until the final calculation result is obtained: Each FPGA that has obtained data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber, including:

[0041] Each of the second to the (N - 1)th FPGAs calculates based on the data obtained from the previous FPGA and the data it has, obtains a calculation result, and sends the calculation result to the next FPGA.

[0042] The Nth FPGA calculates based on the data obtained from the previous FPGA and the data it has, and obtains the final calculation result.

[0043] Wherein, N is an integer greater than or equal to 3.

[0044] Optionally, the number of FPGAs is N; each FPGA has K columns of data. After the full reduction calculation is completed, K columns of final calculation results are obtained.

[0045] Each FPGA in at least part of the FPGAs sends the data it has to the next FPGA through an optical fiber, including:

[0046] The at least part of the FPGAs simultaneously send the data they have to the next FPGA, where different FPGAs send different columns of data.

[0047] Correspondingly, the following steps are repeatedly executed until the final calculation result is obtained: Each FPGA that obtains data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber, including:

[0048] For each of the N FPGAs, if it obtains data from the previous FPGA, it calculates according to the obtained data and the corresponding column data it has. If the obtained calculation result is not the final calculation result, the calculation result is sent to the next FPGA;

[0049] Wherein, N is an integer greater than or equal to 3, and K is an integer greater than or equal to 2.

[0050] In the present invention, the data processing system includes a plurality of FPGAs connected in a ring. Each FPGA is provided with a receiving optical port and a transmitting optical port. The receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber to receive the data sent by the previous FPGA, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber to send data to the next FPGA. Each FPGA is used to perform at least one of the following: receiving the data sent by the previous FPGA; calculating according to the received data and the data it has to obtain a calculation result; sending data to the next FPGA, where the data sent is the calculation result or the data it has. Thus, data transmission can be quickly realized by means of FPGA plus optical fiber, and in a distributed system constructed by a plurality of FPGAs connected in a ring, multiple FPGAs can all participate in the calculation process. Compared with the prior art, when data at different nodes needs to be summed, exclusive-ored, etc., each node needs to send the data it has to the master node uniformly, and the master node completes all data calculations. The present invention does not need to send data to the master node uniformly and rely on the master node for calculation, effectively improving the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the principle of full reduction calculation provided by an embodiment of the present invention;

[0053] Figure 3 It is a schematic diagram of the structure of a data processing system provided by an embodiment of the present invention;

[0054] Figure 4 It is a schematic diagram of the principle of the first calculation stage provided by an embodiment of the present invention;

[0055] Figure 5Schematic diagram of the principle of the second calculation stage provided by the embodiments of the present invention;

[0056] Figure 6 Schematic diagram of a data processing process using DMA technology provided by the embodiments of the present invention;

[0057] Figure 7 Schematic diagram of a data processing method provided by the embodiments of the present invention.

[0058] The implementation, functional features, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0059] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0060] The method provided by the embodiments of the present invention can be applied to any field and scenario that requires data analysis and calculation, for example, it can be applied to the analysis and processing of business data.

[0061] Figure 1 Schematic diagram of an application scenario provided by the embodiments of the present invention. As Figure 1 shown, a distributed system includes multiple devices, each device can be regarded as a Node, each device can store business data, and the business data of multiple devices can be comprehensively calculated to obtain corresponding business results.

[0062] The business data can be any type of business data, for example, it can be user consumption data, user biometric features, etc. In one example, multiple devices can be servers corresponding to different shopping websites, each device can store user consumption data, and after analyzing and processing the consumption data, corresponding consumption preferences can be obtained, so as to push corresponding information to users.

[0063] When analyzing and processing data, methods such as machine learning and deep learning can be used. The method provided by the embodiments of the present invention can be applied to any link in the data analysis and processing process, especially to allreduce calculations.

[0064] Among them, the full reduction calculation described in the embodiments of the present invention may refer to the same operations being performed among the devices participating in the calculation and finally obtaining the same calculation results. In the fields of machine learning and deep learning, full reduction is a commonly used collective communication method. After full reduction communication and calculation, each node in the distributed environment has the final aggregated result.

[0065] For example, during the machine learning process, multiple devices can perform convolution operations based on the data they have. After obtaining the results of the convolution operations, it may be necessary to perform full reduction calculation on the convolution results so that each device obtains the same calculation result and inputs the result into the next layer for processing.

[0066] Figure 2 FIG. is a schematic diagram of the principle of a full reduction calculation provided by an embodiment of the present invention. As Figure 2 shown, taking sum-allreduce as an example, three devices are respectively denoted as Node0, Node1, and Node2. Each device has 3 input elements. After performing full reduction calculation, three results are obtained on each device, which are the results obtained by adding the input elements of the other two devices to the input elements it has.

[0067] In addition to sum, other full reduction calculations can also be implemented, such as max, min, or, etc.

[0068] To implement the above-mentioned full reduction calculation, multiple devices can be set, one of which is used as the master device and the others are used as slave devices. The master device can obtain data from each slave device and perform calculations based on the data obtained from each device. After the calculation is completed, the result is sent to each slave device. In addition, a switch solution can also be used to implement full reduction calculation. However, both of these calculation methods have the problem of low efficiency.

[0069] To solve this problem, an embodiment of the present invention provides a data processing system, which is composed of multiple FPGAs (Field Programmable Gate Array) connected in a ring. For each FPGA participating in data processing, data in each FPGA can be transmitted to the adjacent FPGA through a transmitting optical port, a receiving optical port, and an optical fiber. And each FPGA can perform calculation operations on the received data and the data it has to obtain a calculation result. In this way, through the method of FPGA plus optical fiber, data transmission can be quickly realized, and multiple FPGAs are connected in a ring, and multiple FPGAs can all participate in the calculation process, rather than relying only on the master device for calculation, effectively improving the calculation efficiency.

[0070] The following will describe in detail some embodiments of the present invention in conjunction with the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other.

[0071] Figure 3 FIG. 1 is a schematic structural diagram of a data processing system provided by an embodiment of the present invention. In this embodiment, the system may include: a plurality of FPGAs connected in a ring; each FPGA is provided with a receiving optical port and a transmitting optical port; the receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber; each FPGA is used to perform at least one of the following: receive data sent by the previous FPGA through the receiving optical port; perform calculations based on the received data and the data it has to obtain a calculation result; send data to the next FPGA through the transmitting optical port, where the data sent is the calculation result or the data it has.

[0072] In this embodiment, the system includes a plurality of FPGAs, and the plurality of FPGAs are connected in a ring. The ring connection means that the plurality of FPGAs are connected in sequence, and the last FPGA is connected to the first FPGA to form a ring structure. Each FPGA is provided with a transmitting optical port and a receiving optical port, where the number of transmitting optical ports is at least one, and the number of receiving optical ports is also at least one.

[0073] Take Figure 3 as an example. The system includes 3 FPGAs, and each FPGA includes a transmitting optical port and a receiving optical port. Among them, the transmitting optical port of each FPGA is connected to the receiving optical port of the next FPGA through an optical fiber, and the receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber. For example, the transmitting optical port of the first FPGA is connected to the receiving optical port of the second FPGA through an optical fiber, the transmitting optical port of the second FPGA is connected to the receiving optical port of the third FPGA, and the transmitting optical port of the third FPGA is connected to the receiving optical port of the first FPGA. The small rectangles in the figure represent optical ports, and the lines between the optical ports represent optical fibers, where the arrows indicate the data transmission direction.

[0074] Each FPGA can obtain data from the previous FPGA and can also send data to the next FPGA. Each FPGA has computing capabilities. When any FPGA receives data, it can perform calculations on the received data and all or part of the data it has. After obtaining the calculation result, it can also transmit the calculation result so that the next FPGA can continue to perform calculations based on the calculation result.

[0075] Among them, for any FPGA, the data possessed by the FPGA itself may refer to the data stored within the FPGA or in the memory corresponding to the FPGA. The FPGA performs calculations on this data together with the data obtained from other FPGAs. The data to be calculated is held by different FPGAs respectively, and each FPGA has the ability to perform calculation operations, thereby constructing a distributed system through the multiple FPGAs. Through the solution provided in this embodiment, each FPGA in the distributed system can achieve the calculation of the data it possesses and other data through data interaction.

[0076] The data processing system provided in this embodiment can be applied to full reduction calculation or other calculations. The calculation operations executed in each FPGA can be the same, such as all for summation, or different, such as some for summation and some for multiplication.

[0077] For example, for a system with 3 FPGAs, this system can be used to calculate two types of consumption data. Among them, the first consumption data is stored in the first FPGA, the second consumption data is stored in the second FPGA, and the correction data is stored in the third FPGA. By summing the first consumption data and the second consumption data, the summed consumption data is obtained, and then the summed consumption data is corrected by the correction data to determine the consumption behavior of the user based on the corrected consumption data.

[0078] Based on the above process, the first FPGA can first send the first consumption data it stores to the second FPGA. The second FPGA calculates the sum of the received first consumption data and the second consumption data it possesses and sends the calculation result to the third FPGA. The third FPGA determines the consumption behavior of the user based on the obtained calculation result and the correction data.

[0079] The data processing system provided in this embodiment can be applied to one or more servers. For example, multiple ring-connected FPGAs can be deployed in one server, or multiple ring-connected FPGAs can be deployed in multiple servers. Remote data transmission can be achieved between multiple FPGAs through optical fibers to meet the application requirements in different scenarios.

[0080] In addition, each FPGA can also be provided with multiple transmitting optical ports and multiple receiving optical ports to achieve parallel interaction and calculation of multiple data. For example, each FPGA can be provided with two transmitting optical ports and two receiving optical ports, and each optical port is connected to the corresponding optical port of the previous FPGA or the next FPGA through an optical fiber. In this way, each FPGA can simultaneously send two data to the next FPGA through two transmitting optical ports, further improving the data processing efficiency.

[0081] The data processing system provided in this embodiment includes a plurality of FPGAs connected in a ring. Each FPGA is provided with a receiving optical port and a transmitting optical port. The receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber to receive the data transmitted by the previous FPGA, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber to transmit data to the next FPGA. Each FPGA is used to perform at least one of the following: receive the data transmitted by the previous FPGA, calculate according to the received data and the data it has, obtain a calculation result, and transmit data to the next FPGA, where the transmitted data is the calculation result or the data it has. Thus, data stored at different positions in a distributed system can be processed quickly through FPGAs and optical fibers, effectively improving the calculation efficiency.

[0082] Based on the technical solution provided in the above embodiment, optionally, the system is used to implement full reduction calculation. The full reduction calculation is used to calculate according to the data of each FPGA and make each FPGA obtain the same final calculation result; the full reduction calculation includes a calculation stage and a synchronization stage.

[0083] Among them, full reduction calculation can use different full reduction operators to implement operations such as summation, maximum value, minimum value, and OR operation. In addition, full reduction operators or data structures with other functions can also be implemented through programming. For example, the data structure can be a uniqueness set, and each FPGA can implement a deduplication operation to obtain a data set that does not contain duplicate data.

[0084] Among them, in the calculation stage, multiple FPGAs can sequentially transfer data and perform calculations to obtain the final calculation result; in the synchronization stage, the FPGA that obtains the final calculation result can transfer the final calculation result to other FPGAs.

[0085] Specifically, the calculation stage refers to calculating the data stored in each FPGA and obtaining the final calculation result, and the synchronization stage refers to synchronizing the final calculation result to each FPGA.

[0086] There are many specific implementation methods for multiple FPGAs to perform full reduction calculation. The present invention provides an embodiment to elaborate on the first data processing process in detail. In this embodiment, multiple FPGAs connected in a ring can implement the first data processing process.

[0087] In this embodiment, the number of FPGAs is N. Among them, N is an integer greater than or equal to 3.

[0088] In the calculation stage, the first FPGA is specifically configured to: send the data it has to the second FPGA; each of the second to the (N - 1)th FPGAs is specifically configured to: perform calculations based on the data obtained from the previous FPGA and the data it has, obtain a calculation result, and send the calculation result to the next FPGA; the Nth FPGA is specifically configured to: perform calculations based on the data obtained from the previous FPGA and the data it has, and obtain the final calculation result.

[0089] In the synchronization stage, the Nth FPGA is specifically configured to: send the final calculation result to the next FPGA; the next FPGA of the Nth FPGA is the first FPGA; each of the first to the (N - 2)th FPGAs is specifically configured to: receive and send the final calculation result; the (N - 1)th FPGA is specifically configured to: receive the final calculation result.

[0090] The first FPGA is only used to send the data it has to the second FPGA and receive the final calculation result sent by the Nth FPGA. The second to the (N - 1)th FPGAs can perform calculations on the received data and their own data, and send the calculated data to the next FPGA. The Nth FPGA can obtain the final calculation result.

[0091] Figure 4 This is a schematic diagram of the principle of the first calculation stage provided by the embodiments of the present invention. As Figure 4 shown, taking the execution of addition calculation as an example, the system includes 3 FPGAs, namely FPGA1, FPGA2, and FPGA3 connected in sequence. The transmitting optical port of FPGA3 is connected to the receiving optical port of FPGA1. Among them, FPGA1 is the FPGA that initiates the calculation. Among them, the calculation stage may include the following processes.

[0092] First, FPGA1 sends the first data 1 it has to FPGA2. After receiving the first data 1 from FPGA1, FPGA2 performs calculations on the received first data 1 and its own first data 3, obtains a calculation result 4, and sends the calculation result to FPGA3. At the same time, FPGA1 continues to send subsequent data to FPGA2.

[0093] After receiving the second data 2 from FPGA1, FPGA2 performs calculations on the received second data 2 and its own second data 2, and obtains a calculation result 4. Similarly, after receiving the calculation result of the first data from FPGA2, FPGA3 will perform calculations on the received calculation result 4 of the first data and its own first data 5, and obtain a calculation result 9. And so on, finally, the sum result of the data stored in the 3 FPGAs can be obtained in FPGA3.

[0094] In addition, the data processing process further includes a synchronization phase. After obtaining the calculation result, the calculation result can be synchronized to each FPGA. Specifically, when FPGA3 obtains the final calculation result, it can send the final calculation result to the next FPGA, i.e., FPGA1. After receiving the final result, FPGA1 can send the final calculation result to the next FPGA, i.e., FPGA2, and so on until all FPGAs can obtain the final calculation result.

[0095] The above embodiment illustrates the first data processing process. In this embodiment, each FPGA in multiple ring-connected FPGAs transfers each column of data or each column of calculated data to the adjacent FPGA. Each FPGA can implement data sending and receiving, and at the same time, the FPGA will also perform data calculation, so that the last FPGA can obtain the final calculation result after one calculation. Compared with the method of data calculation through the master-slave device mode, this method can distribute the data calculation process to each FPGA. Since each FPGA can implement data transmission and data calculation, the data processing efficiency can be improved.

[0096] The first data processing process can improve the data processing efficiency. However, at some moments, there will still be idle FPGAs. For example, when FPGA3 does not receive data, it will always be in a waiting state. Therefore, the present invention also provides a second data processing process, which can further improve the data calculation efficiency.

[0097] In this embodiment, it is assumed that the number of FPGAs is N, and each FPGA has K columns of data. After the full reduction calculation is completed, K columns of final calculation results are obtained. Among them, N is an integer greater than or equal to 3, and K is an integer greater than or equal to 2.

[0098] In the calculation phase, at least some FPGAs are specifically configured to: simultaneously send the data they have to the next FPGA, where different FPGAs send different columns of data; if data is obtained from the previous FPGA, calculate according to the obtained data and the corresponding column of data they have, and if the obtained calculation result is not the final calculation result, send the calculation result to the next FPGA.

[0099] If there are other FPGAs among the multiple FPGAs other than the at least some FPGAs, each FPGA in the other FPGAs is specifically configured to: if data is obtained from the previous FPGA, calculate according to the obtained data and the corresponding column of data they have, and if the obtained calculation result is not the final calculation result, send the calculation result to the next FPGA.

[0100] Among them, K can be greater than, less than, or equal to N. Taking Figure 4 as an example, K is equal to N, both being 3, and the calculation of 3 columns of data can be implemented by 3 FPGAs. The data in the first column is 1, 3, 5, the data in the second column is 2, 2, 4, and the data in the third column is 1, 1, 5.

[0101] At least some of the FPGAs can be all 3 FPGAs. Each FPGA simultaneously sends the data of its corresponding column to the next FPGA. When data is obtained from the previous FPGA, calculations can be performed based on the obtained data and the data of the corresponding column it has, and the calculation results are sent to the next FPGA.

[0102] Specifically, as Figure 4 shown in the data distribution, FPGA1 sends its first-column data 1 to FPGA2, while FPGA2 sends its second-column data 2 to FPGA3, and while FPGA3 sends its third-column data 5 to FPGA1. At the next moment, FPGA1, FPGA2, and FPGA3 can all perform calculations based on the received data and the data of the corresponding columns they have.

[0103] For example, if FPGA1 obtains the third-column data 5 from FPGA3, then it adds 5 to its own third-column data 1 to get 6 and sends it to FPGA2. The processing processes of the other two FPGAs are similar, performing the summation operation for the same-column data.

[0104] After newly obtained data, the summation operation can continue. For example, after FPGA2 obtains the current summation result 6 of the third column from FPGA1, it can add this result to its own third-column data 1 to get the final calculation result 7. Similarly, FPGA1 can obtain the final calculation result 8 of the second column, and FPGA3 can obtain the final calculation result 9 of the second column.

[0105] In the synchronization stage, starting from the FPGA that obtains the final calculation result, each FPGA is specifically used for: sending the final calculation result to the next FPGA until the final calculation result is obtained by all FPGAs.

[0106] The above gives the processing process when K = N. In practical applications, the embodiments of the present invention can also be applied to calculations when K ≠ N.

[0107] Figure 5 This is the schematic diagram of the principle of the second calculation stage provided by the embodiments of the present invention. Figure 5 In it, N = 3, K = 2, that is, 3 FPGAs calculate 2 columns of data, and 2 columns of final calculation results can be obtained.

[0108] In this embodiment, the three FPGAs include FPGA1, FPGA2, and FPGA3 connected in sequence, and the transmitting optical port of FPGA3 is connected to the receiving optical port of FPGA1 to form a ring structure. The FPGAs are divided into at least part of the FPGAs and other FPGAs other than at least part of the FPGAs. Among them, the number of at least part of the FPGAs is K, and the number of other FPGAs is N - K.

[0109] As Figure 5 shown, when the number of FPGAs is 3 and the number of columns of data in each FPGA is 2, the number of at least part of the FPGAs that initiate the full reduction calculation is 2, namely FPGA1 and FPGA2, and the other FPGA other than at least part of the FPGAs is 1, namely FPGA3. In the calculation stage, each of the at least part of the FPGAs can transmit data in different columns, and each FPGA can calculate the received data and the corresponding column data it has after receiving the data to obtain a calculation result.

[0110] For example, taking the execution of an addition calculation as an example here, FPGA1 can send the first column of data 1 to FPGA2. At the same time, FPGA2 can send the second column of data 4 to FPGA3. After each FPGA obtains the data, it calculates according to the obtained data and the corresponding column data it has. Specifically, after receiving the first column of data 1 sent by FPGA1, FPGA2 calculates it with its own first column of data 3 to obtain the calculation result 4 corresponding to the first column of data. At the same time, after receiving the second column of data 4 sent by FPGA2, FPGA3 calculates it with its own second column of data 6 to obtain the calculation result 10 corresponding to the second column of data.

[0111] Then, FPGA2 sends the obtained calculation result 4 to FPGA3, and FPGA3 sends the obtained calculation result 10 to FPGA1. Similarly, after obtaining the calculation result 4 corresponding to the first column sent by FPGA2, FPGA3 calculates it with its own first column of data 5 to obtain the final calculation result 9 of the first column. After obtaining the calculation result 10 corresponding to the second column sent by FPGA3, FPGA1 calculates it with its own second column of data 2 to obtain the final calculation result 12 of the second column.

[0112] In the synchronization stage, FPGA1 synchronizes the final calculation result of the second column to the other two FPGAs, and FPGA3 synchronizes the final calculation result of the first column to the other two FPGAs.

[0113] Through the second data processing process in this embodiment, each FPGA in multiple ring-connected FPGAs can send data in different columns to adjacent FPGAs, and each FPGA obtains the calculation result of a certain column. In this process, the calculation process and the data sending process are evenly distributed to each FPGA, and the data is transferred between FPGAs like running water, making the processing process more efficient.

[0114] Optionally, when K is less than or equal to N, the number of at least some of the FPGAs is K; when K is greater than N, the K columns are divided into M groups, and the number of columns in each group is less than or equal to N, so as to obtain the final calculation results of the K columns by performing M calculation stages, and each calculation stage is used to calculate a group of data; wherein, when calculating the data in the Mth group, the number of at least some of the FPGAs is the number of columns corresponding to the Mth group. The M is an integer greater than or equal to 1.

[0115] Specifically, when K is greater than N, the K-column data in each FPGA can be grouped, and the number of columns in each group is less than or equal to N, so the K-column data can be divided into M groups, and the calculation results of the K-column data can be obtained by performing M calculation stages. Among them, when calculating the data in the Mth group, the number of at least some is related to the number of columns of the data in the Mth group.

[0116] For example, when K is 101 and the number of FPGAs N is 3, the 101-column data in each FPGA can be divided into 34 groups of data. Among them, the first 33 groups of data each contain 3 columns of data, and the 34th group of data contains 2 columns of data. Therefore, the 3 FPGAs can obtain the final calculation results after 34 calculation stages. The specific implementation principle and process of each calculation stage can be referred to the previous description and will not be elaborated here.

[0117] After obtaining the final calculation results, data synchronization is performed according to the final calculation results of the corresponding number of columns obtained by each FPGA. Alternatively, synchronization can also be performed every time a calculation stage is passed. The embodiments of the present invention do not limit this.

[0118] Optionally, in any of the foregoing embodiments, the data possessed by the FPGA can be stored in a memory, and the FPGA reads data from and / or writes data to the memory through the Direct Memory Access (DMA) technology.

[0119] Specifically, the memory can be the on-board memory of the FPGA, or it can be other memories such as the CPU memory.

[0120] In any of the foregoing embodiments, when calculating data in the FPGA, it is necessary to first read the data in the FPGA. Among them, the data possessed by the FPGA is stored in the memory, and the access to the FPGA can be achieved through the direct memory access technology. For example, reading data and writing data. Optionally, the FPGA may be integrated with a DMA function to realize reading or writing data from the memory. Alternatively, the FPGA may be connected to a DMA controller to achieve access to the memory through the DMA controller, so as to read or write data from the memory. For example, in the calculation stage, data can be read from the memory through the DMA controller for data transmission and calculation; in the synchronization stage, the received data can be stored in the memory through the DMA controller.

[0121] Through the DMA technology, the access to the data in the FPGA can be realized, enabling the FPGA to obtain data and store data in the FPGA during full reduction calculation, effectively improving the data reading and writing efficiency.

[0122] Optionally, in any of the foregoing embodiments, if the data sent by each FPGA is the smallest unit for full reduction calculation, the FPGA receiving the data is used to calculate according to the data after the data is completely received; if the data is not the smallest unit for full reduction calculation, the FPGA receiving the data is used to calculate according to the data of the smallest unit that has been received during the data reception process.

[0123] For example, if the data to be calculated is a 16-bit number and these 16 bits cannot be separated, then these 16 bits are the smallest unit, and the calculation starts after the 16-bit number is completely received. If the data to be calculated is 16 bits long, but these 16 bits can be split, for example, the first 8 bits represent one number and the last 8 bits represent another number, then the smallest unit is 8 bits, and the calculation can start after receiving 8 bits of data, thus effectively improving the calculation efficiency on the basis of ensuring the calculation accuracy.

[0124] Figure 6 This is a schematic flowchart of data processing through the DMA technology provided by the embodiments of the present invention. As Figure 6 shown, before each FPGA performs data calculation, it first prepares data, that is, completes basic operations to obtain data for full reduction calculation. The data can be stored in the memory, and when full reduction calculation is required, the data can be read from the memory through the DMA technology.

[0125] After the data is read, the data is first judged to determine whether the data is the smallest unit for full reduction calculation. If it is the smallest calculation unit, it means that the data is indivisible, and then the read data is sent in a streaming manner through Remote Direct Memory Access (RDMA) technology. Here, the streaming transmission as a whole can mean that the calculation can only be performed after the data transmission is completed. If it is not the smallest calculation unit, it means that the data is divisible and contains multiple independent data, and then the data is sent in parallel through RDMA technology. The parallel transmission means that the calculation can be performed while sending.

[0126] After the data is obtained, the FPGA can perform calculations on the data and synchronize the final calculation results among the FPGAs through RDMA technology. When each FPGA obtains the final calculation result, it is written into the memory of the FPGA through DMA technology, and the obtained final calculation result can also be applied to other algorithms.

[0127] By judging the sent data and then performing different data sending and calculations according to different types of data, the improvement of the calculation efficiency can be achieved.

[0128] Optionally, in any of the foregoing embodiments, the system further includes: a control device; the multiple FPGAs are respectively connected to the control device; the control device is configured to:

[0129] Send basic operation instructions to each FPGA so that each FPGA performs basic operations according to the basic operation instructions to obtain data for full reduction calculation; obtain status information sent by each FPGA for indicating whether the basic operation is completed; select an FPGA that initiates full reduction calculation from the multiple FPGAs according to the status information sent by each FPGA; send a full reduction start instruction to the FPGA that initiates full reduction calculation so that the FPGA sends the data for full reduction calculation to the next FPGA according to the full reduction start instruction; send a waiting instruction to other FPGAs so that the other FPGAs wait to obtain the data for full reduction calculation and perform calculations after obtaining it.

[0130] Among them, the control device can be any device capable of communicating with the FPGA, such as a central processing unit CPU, etc. Alternatively, the control device can be implemented by a processor plus a memory.

[0131] In any of the foregoing embodiments, the system may further include a control device. By connecting the control device to each FPGA, control over each FPGA can be achieved. First, the control device can send basic operation instructions to each FPGA. The basic operation instructions can be convolution calculations or other calculations during the neural network calculation process to obtain the results of the basic operations. Among them, the result obtained by each FPGA through the basic operation is the data for the full reduction calculation.

[0132] For each FPGA, after obtaining the data for the full reduction calculation, it can send the status information indicating that the basic operation has been completed to the control device. Based on whether the status information sent by each FPGA is obtained, the control device can determine the status of each FPGA. To improve the efficiency of the full reduction calculation, the full reduction calculation can be started after all FPGAs have completed the basic operation.

[0133] For the first data processing process, when performing the full reduction calculation, the control device determines the FPGA that initiates the full reduction calculation. Specifically, the control device can specify an FPGA from all FPGAs as the FPGA to start the full reduction calculation, and the remaining FPGAs are in a waiting state. After determining the FPGA that initiates the full reduction calculation, the control device sends a full reduction start instruction to this FPGA. After receiving the full reduction start instruction, this FPGA can send the data for the full reduction calculation to the next FPGA. At the same time, the control device sends a waiting instruction to other FPGAs, causing the other FPGAs to be in a state of waiting to receive data and perform data calculations after receiving the data.

[0134] In addition, for the second data processing process, the control device first instructs all FPGAs to complete the basic operation. After completing the basic operation, it can send a full reduction start instruction to at least some FPGAs, so that each of the at least some FPGAs simultaneously performs the operations of data sending and data calculation. At the same time, the control device can also send a waiting instruction to the other FPGAs except the at least some FPGAs, causing the other FPGAs except the at least some FPGAs to be in a state of waiting to receive data and perform data calculations after receiving the data.

[0135] Through the method described above, the control device can be used to control multiple FPGAs, achieve the synchronization of multiple FPGAs, effectively improve the accuracy of processing, and ensure the orderliness and stability of data processing.

[0136] Among them, when the control device determines the FPGA that initiates the full reduction calculation, it can be determined based on the status information uploaded from each FPGA. Optionally, when the control device selects the FPGA that initiates the full reduction calculation from the multiple FPGAs, it is specifically used to: determine the number of FPGAs that initiate the full reduction calculation; based on the status information sent by each FPGA, select the corresponding number of FPGAs that finally complete the basic calculation as the FPGA that initiates the full reduction calculation. Specifically, the corresponding number is the number of FPGAs that initiate the full reduction calculation. The FPGA that initiates the full reduction calculation refers to the FPGA that sends its own data to the next FPGA after the calculation phase starts. The number of FPGAs that initiate the full reduction calculation can be set according to actual needs, and can specifically be one or more.

[0137] In practical applications, the FPGA that completes the basic calculation last can be used as the FPGA that initiates the full reduction calculation. Since the FPGA that completes the basic calculation last is often an FPGA with poor computing performance, and the FPGA that initiates the full reduction calculation has one less computing process than other FPGAs, the FPGA that completes the basic calculation last can be used as the FPGA that initiates the full reduction calculation, which can further improve the efficiency of the full reduction calculation.

[0138] Based on the data processing systems provided in the above embodiments, an embodiment of the present invention further provides a data processing method. Figure 7 The flowchart of a data processing method provided by an embodiment of the present invention is shown in FIG. The method in this embodiment can be implemented based on the system described in any of the above embodiments. Figure 7 As shown, the data processing method may include:

[0139] Step S 701 : Each FPGA in at least some of the FPGAs sends its own data to the next FPGA through an optical fiber; wherein the at least some of the FPGAs are at least some of the FPGAs in a plurality of FPGAs connected in a ring shape.

[0140] Step S702, repeat the following steps until the final calculation result is obtained: each FPGA that obtains the data calculates the obtained data with its own data, and if the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through the optical fiber.

[0141] In an optional implementation, the method is used to implement full-reduction calculation; the full-reduction calculation is used to perform calculations based on data held by each FPGA and enable each FPGA to obtain the same final calculation result;

[0142] Correspondingly, the method further includes: the FPGA that obtains the final calculation result transmits the final calculation result to other FPGAs.

[0143] In an alternative implementation, the number of the FPGAs is N;

[0144] Each of at least some of the FPGAs sends the data it has to the next FPGA through an optical fiber, including:

[0145] The first FPGA sends the data it has to the second FPGA;

[0146] Correspondingly, the following steps are repeatedly executed until the final calculation result is obtained: Each FPGA that obtains data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through the optical fiber, including:

[0147] Each of the second to the (N - 1)th FPGAs calculates according to the data obtained from the previous FPGA and the data it has to obtain a calculation result, and sends the calculation result to the next FPGA;

[0148] The Nth FPGA calculates according to the data obtained from the previous FPGA and the data it has to obtain the final calculation result;

[0149] wherein, N is an integer greater than or equal to 3.

[0150] In an alternative implementation, the number of the FPGAs is N; each FPGA has K columns of data, and after the full reduction calculation is completed, K columns of final calculation results are obtained;

[0151] Each of at least some of the FPGAs sends the data it has to the next FPGA through an optical fiber, including:

[0152] The at least some FPGAs simultaneously send the data they have to the next FPGA, wherein different FPGAs send different columns of data;

[0153] Correspondingly, the following steps are repeatedly executed until the final calculation result is obtained: Each FPGA that obtains data calculates the obtained data with the corresponding column of data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA, including:

[0154] For each of the N FPGAs, if it obtains data from the previous FPGA, it calculates according to the obtained data and the corresponding column of data it has. If the obtained calculation result is not the final calculation result, the calculation result is sent to the next FPGA;

[0155] Wherein, N is an integer greater than or equal to 3, and K is an integer greater than or equal to 2.

[0156] In an alternative implementation, when K is less than or equal to N, the number of the at least part of the FPGAs is K;

[0157] When K is greater than N, the K columns are divided into M groups, and the number of columns in each group is less than or equal to N, so as to obtain the final calculation results of the K columns through M calculation phases, and each calculation phase is used to calculate a group of data;

[0158] Wherein, when calculating the data of the Mth group, the number of the at least part of the FPGAs is the number of columns corresponding to the Mth group.

[0159] In an alternative implementation, the data of the FPGA is stored in the memory, and the method further includes: each FPGA reads data from the corresponding memory and / or writes data to the memory through direct memory access technology.

[0160] In an alternative implementation, each FPGA that obtains data calculates the obtained data with the data it has, including:

[0161] If the data sent by each FPGA is the minimum unit for performing full reduction calculation, the FPGA that receives the data is used to perform calculation according to the data after the data is completely received;

[0162] If the data is not the minimum unit for performing full reduction calculation, the FPGA that receives the data is used to perform calculation according to the data of the minimum unit that has been received during the receiving process of the data.

[0163] In an alternative implementation, the method further includes:

[0164] The control device sends basic operation instructions to each FPGA, so that each FPGA performs basic operations according to the basic operation instructions to obtain data for performing full reduction calculation;

[0165] Obtain status information sent by each FPGA for indicating whether the basic operation is completed;

[0166] Select the FPGA that initiates full reduction calculation from the multiple FPGAs according to the status information sent by each FPGA;

[0167] Send a full reduction start instruction to the FPGA that initiates full reduction calculation, so that the FPGA sends the data for performing full reduction calculation to the next FPGA according to the full reduction start instruction;

[0168] Send a wait instruction to other FPGAs so that the other FPGAs wait to obtain data for performing full reduction calculations and perform calculations after obtaining the data.

[0169] In an alternative implementation, selecting the FPGA that initiates the full reduction calculation from the multiple FPGAs includes:

[0170] Determine the number of FPGAs that initiate the full reduction calculation;

[0171] According to the status information sent by each FPGA, select the corresponding number of FPGAs that finally complete the basic operations as the FPGAs that initiate the full reduction calculation;

[0172] Wherein, the corresponding number is the number of FPGAs that initiate the full reduction calculation.

[0173] For the specific implementation principle, process and beneficial effects of the method in this embodiment, reference can be made to the foregoing embodiments, and details are not described herein again.

[0174] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0175] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention.

[0176] It should be understood that the above processor can be a Central Processing Unit (CPU), and can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the present invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0177] The memory may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.

[0178] The above storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium may be any available medium accessible by a general-purpose or special-purpose computer.

[0179] An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium may also exist as discrete components in an electronic device or a master device.

[0180] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0181] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0183] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A data processing system, characterized in that, it includes: a plurality of FPGAs connected in a ring; each FPGA is provided with a receiving optical port and a transmitting optical port; each FPGA is provided with a plurality of transmitting optical ports and a plurality of receiving optical ports; the receiving optical port of each FPGA is connected to the transmitting optical port of the previous FPGA through an optical fiber, and the transmitting optical port is connected to the receiving optical port of the next FPGA through an optical fiber; each FPGA is used to perform at least one of the following: receive data sent by the previous FPGA through the receiving optical port; perform calculations based on the received data and the data it has to obtain a calculation result; send data to the next FPGA through the transmitting optical port, where the data sent is the calculation result or the data it has; wherein, the system is used to implement full reduction calculation; the full reduction calculation is used to perform calculations based on the data of each FPGA and enable each FPGA to obtain the same final calculation result; the full reduction calculation includes a calculation stage and a synchronization stage; in the calculation stage, a plurality of FPGAs sequentially transfer data and perform calculations to obtain a final calculation result; in the synchronization stage, the FPGA that obtains the final calculation result transfers the final calculation result to other FPGAs.

2. The system according to claim 1, characterized in that, the number of the FPGAs is N; in the calculation stage, the 1st FPGA is specifically used to: send the data it has to the 2nd FPGA; each of the 2nd to the N-1th FPGAs is specifically used to: perform calculations based on the data obtained from the previous FPGA and the data it has to obtain a calculation result, and send the calculation result to the next FPGA; the Nth FPGA is specifically used to: perform calculations based on the data obtained from the previous FPGA and the data it has to obtain a final calculation result; wherein, the N is an integer greater than or equal to 3.

3. The system according to claim 1, characterized in that, the number of the FPGAs is N; each FPGA has K columns of data, and after the full reduction calculation is completed, K columns of final calculation results are obtained; in the calculation stage, at least some of the FPGAs are specifically used to: simultaneously send the data they have to the next FPGA, where different FPGAs send different columns of data; if data is obtained from the previous FPGA, then perform calculations based on the obtained data and the corresponding column of data it has, and if the obtained calculation result is not the final calculation result, then send the calculation result to the next FPGA; if there are other FPGAs among the plurality of FPGAs other than the at least some of the FPGAs, then each of the other FPGAs is specifically used to: if data is obtained from the previous FPGA, then perform calculations based on the obtained data and the corresponding column of data it has, and if the obtained calculation result is not the final calculation result, then send the calculation result to the next FPGA; wherein, the N is an integer greater than or equal to 3, and the K is an integer greater than or equal to 2.

4. The system according to claim 3, characterized in that, when K is less than or equal to N, the number of the at least some of the FPGAs is K; When K is greater than N, the K columns are divided into M groups, and the number of columns in each group is less than or equal to N, so as to obtain the final calculation result of the K columns through M calculation phases, and each calculation phase is used to calculate a group of data; wherein, when calculating the data of the Mth group, the number of at least part of the FPGAs is the number of columns corresponding to the Mth group; M is an integer greater than or equal to 1.

5. The system according to any one of claims 1-4, characterized in that, if the data sent by each FPGA is the minimum unit for performing full reduction calculation, the FPGA receiving the data is used to perform calculation according to the data after the data is completely received; if the data is not the minimum unit for performing full reduction calculation, the FPGA receiving the data is used to perform calculation according to the data of the minimum unit that has been received during the receiving process of the data.

6. The system according to any one of claims 1-4, characterized in that, further comprising: a control device; the multiple FPGAs are respectively connected to the control device; the control device is used for: sending a basic operation instruction to each FPGA, so that each FPGA performs a basic operation according to the basic operation instruction to obtain data for performing full reduction calculation; acquiring status information sent by each FPGA for indicating whether the basic operation is completed; selecting an FPGA that initiates full reduction calculation from the multiple FPGAs according to the status information sent by each FPGA; sending a full reduction start instruction to the FPGA that initiates full reduction calculation, so that the FPGA sends data for performing full reduction calculation to the next FPGA according to the full reduction start instruction; sending a waiting instruction to other FPGAs, so that the other FPGAs wait to obtain data for performing full reduction calculation and perform calculation after obtaining it.

7. The system according to claim 6, characterized in that, when the control device selects an FPGA that initiates full reduction calculation from the multiple FPGAs, it is specifically used for: determining the number of FPGAs that initiate full reduction calculation; according to the status information sent by each FPGA, selecting the corresponding number of FPGAs that finally complete the basic operation as the FPGAs that initiate full reduction calculation; wherein, the corresponding number is the number of FPGAs that initiate full reduction calculation.

8. The system according to any one of claims 1-4, characterized in that, the data of the FPGA is stored in the memory, and the FPGA reads data from and / or writes data to the memory through direct memory access technology.

9. A data processing method, characterized in that, the method includes: each FPGA in at least part of the FPGAs sends the data it has to the next FPGA through an optical fiber; wherein, the at least part of the FPGAs is at least part of the multiple FPGAs connected in a ring; each FPGA is provided with a plurality of transmitting optical ports and a plurality of receiving optical ports; Repeat the following steps until the final calculation result is obtained: Each FPGA that has obtained data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber. Among them, the system is used to implement full reduction calculation; the full reduction calculation is used to calculate based on the data of each FPGA and enable each FPGA to obtain the same final calculation result; the full reduction calculation includes a calculation stage and a synchronization stage. In the calculation stage, multiple FPGAs sequentially transfer data and perform calculations to obtain the final calculation result. In the synchronization stage, the FPGA that has obtained the final calculation result transfers the final calculation result to other FPGAs.

10. According to the method described in claim 9, characterized in that the method is used to implement full reduction calculation; the full reduction calculation is used to calculate based on the data of each FPGA and enable each FPGA to obtain the same final calculation result. Correspondingly, the method further includes: The FPGA that has obtained the final calculation result transfers the final calculation result to other FPGAs.

11. According to the method described in claim 9 or 10, characterized in that the number of FPGAs is N; Each of at least some of the FPGAs sends the data it has to the next FPGA through an optical fiber, including: The first FPGA sends the data it has to the next FPGA. Correspondingly, the step of repeating the following steps until the final calculation result is obtained: Each FPGA that has obtained data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber, including: Each of the second to the (N - 1)th FPGAs calculates the data obtained from the previous FPGA with the data it has to obtain a calculation result, and sends the calculation result to the next FPGA. The Nth FPGA calculates the data obtained from the previous FPGA with the data it has to obtain the final calculation result. Among them, N is an integer greater than or equal to 3.

12. According to the method described in claim 10, characterized in that the number of FPGAs is N; each FPGA has K columns of data. After the full reduction calculation is completed, K columns of final calculation results are obtained. Each of at least some of the FPGAs sends the data it has to the next FPGA through an optical fiber, including: The at least some FPGAs simultaneously send the data they have to the next FPGA, where different FPGAs send different columns of data. Correspondingly, the step of repeating the following steps until the final calculation result is obtained: Each FPGA that has obtained data calculates the obtained data with the data it has. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA through an optical fiber, including: For each of the N FPGAs, if data is obtained from the previous FPGA, calculations are performed based on the obtained data and the corresponding column data of its own. If the calculation result is not the final calculation result, the calculation result is sent to the next FPGA; wherein, N is an integer greater than or equal to 3, and K is an integer greater than or equal to 2.

Citation Information

Patent Citations

  • FPGA heterogeneous accelerated computing device and system

    CN106776466A

  • Serial port cascade regulation and control method and serial port equipment

    CN111475368A