Heterogeneous computing clustering acceleration method and system based on Spark and FPGA
By designing a heterogeneous computing clustering acceleration method and system based on Spark and FPGA, the high parallel computing requirements of large-scale data sets and complex algorithms are solved, the computing efficiency and real-time improvement are achieved, and the utilization of hardware resources is optimized.
Patent Information
- Application Number
- CN202510278202.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to meet the high parallel computing needs of large-scale data sets and complex algorithms, especially when computing-intensive tasks on the CPU, there are hardware resource bottlenecks and computing bottlenecks.
A heterogeneous computing clustering acceleration method and system based on Spark and FPGA is designed, and the master node is deployed using the master-slave structure. The master node is responsible for management and task scheduling. The slave node is connected to the FPGA through the PCIe interface, and data exchange is realized using the OpenCL framework, and data transmission is optimized through memory multiplexing and timing analysis to maximize the use of the FPGA storage area.
It significantly reduces the delay of iterative computing, improves the computing efficiency and real-time performance of large-scale data set processing, optimizes the performance bottleneck of CPU processing computing-intensive tasks, and improves the utilization efficiency of hardware resources.
Smart Images

Figure CN120256378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to a heterogeneous computing clustering acceleration method and system based on Spark and FPGA. Background Art
[0002] Today, information technology has developed rapidly. With the continuous popularization of Internet of Things devices and intelligent terminals, information technology has penetrated into people's clothing, food, housing, and transportation. At the same time, this has also made data production faster and more diverse, and the data volume has increased explosively. How to efficiently utilize this data, that is, how to process and analyze this data efficiently and quickly, and mine the value in it has become a hot topic in current research. Among many data analysis methods, clustering analysis has become an important tool in the field of data mining due to its excellent performance in revealing the internal structure of data, pattern recognition, and data compression. Briefly speaking, clustering analysis is to group similar data objects by analyzing the characteristics of a group of data without a pre-given specific classification standard. Through this process, the potential relationships and laws in the data are revealed. Clustering analysis provides a solid foundation for many applications such as decision support, market segmentation, and image recognition, and has also played an important role in many fields such as business and medicine.
[0003] The time complexity of using clustering algorithms to analyze a large amount of data is generally high, resulting in a high time cost for single-machine execution of clustering. To solve this problem, several big data parallel computing frameworks have emerged. Hadoop is a currently widely used big data parallel computing framework. It is based on the MapReduce programming model, which simply divides tasks into two stages: Map and Reduce. When the computing task is relatively complex, developers need to design and decompose the task into multiple jobs, which requires a high coding requirement for developers. In addition, the map mapping stage of Hadoop needs to write to the local disk, which is not suitable for processing a large amount of intermediate data generated by clustering calculations. In contrast, Spark performs in-memory computing, and its scheduling mechanism is better than that of Hadoop. Spark is based on the Resilient Distributed Dataset (RDD), which allows data to be cached and reused in memory, greatly reducing the read and write overhead between different computing stages. On the basis of Spark, Spark Streaming further expands the functions of Spark. Through the Micro-Batching mechanism, real-time data streams are divided into small batches of data for processing, further optimizing the data throughput and processing efficiency.
[0004] To further improve the efficiency of big data analysis using distributed frameworks, especially in the field of real-time data stream processing, edge computing technology has developed rapidly in recent years and has been combined with distributed frameworks. Edge computing can migrate data processing from the centralized cloud to edge devices closer to the data source, thereby reducing latency and improving data processing efficiency. This trend has prompted in-depth exploration and optimization of existing distributed computing frameworks such as Apache Kafka, Apache Flink, Spark Streaming, etc. H. Lee et al. designed a comparison of the performance of three major frameworks, Apache Storm, Apache Flink, and Apache Spark, in resource-constrained devices. May Thet Tun et al. conducted a performance analysis of the integration of Apache Kafka and Spark Streaming and concluded that this framework can perform better in terms of processing time and fault tolerance for massive data, and parallel recovery can provide fault tolerance. Godson Michael D'silva et al. designed a dedicated framework based on Kafka, dashing, and Apache Spark to process real-time data as well as existing or historical data in the Internet of Things system, which is faster and more scalable than existing traditional applications. M. Reza Hoseiny Farahabady et al. discussed the performance degradation caused by resource contention between co-located analysis applications with different priorities and different intrinsic characteristics in the shared Spark platform and proposed an automatic adjustment strategy for computing resources in the distributed Spark platform, which can meet the low-latency requirements of applications that continuously run user-defined operations on large-capacity stream data.
[0005] Although the above research shows that these frameworks have performed well in processing large-scale data, in reality, the storage and processing of big data still face bottlenecks in hardware resources, especially in the processing of computationally intensive tasks on the CPU, and there is still room for improvement for Spark Streaming running on the CPU. There are physical limitations in the number of cores and clock frequency of the CPU, making it difficult to meet the high parallel computing requirements of large-scale data sets and complex algorithms. Moreover, when the CPU faces computationally intensive tasks, it often requires frequent memory access and data transfer, further exacerbating the computing bottleneck. Summary of the Invention
[0006] The present invention provides a heterogeneous computing clustering acceleration method and system based on Spark and FPGA to solve the technical problem that the existing technology is difficult to meet the high parallel computing requirements of large-scale data sets and complex algorithms.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] On the one hand, the present invention provides a heterogeneous computing clustering acceleration method based on Spark and FPGA. The heterogeneous computing clustering acceleration method based on Spark and FPGA includes:
[0009] Design a parallel heterogeneous clustering acceleration framework based on Spark and FPGA; wherein, the parallel heterogeneous clustering acceleration framework is deployed in a master-slave structure, including a master node and multiple slave nodes; wherein, the master node undertakes the management tasks of the entire framework, including communication with each slave node, generation, allocation and scheduling of jobs; each slave node is respectively connected to at least one FPGA, and after receiving a task, hands over the calculation process of the corresponding task to the FPGA for operation, and each slave node is equipped with Spark Streaming;
[0010] Use the parallel heterogeneous clustering acceleration framework to run a preset clustering algorithm to achieve computing acceleration.
[0011] Furthermore, the slave node uses a PCIe interface to connect to the FPGA, and realizes data exchange between the slave node and the FPGA through the PCIe communication mechanism.
[0012] Furthermore, the method further includes:
[0013] Write the content of the OpenCL framework in C++ language and compile it into a.so dynamic link library file dedicated to the Linux system, and use it as the bridge between Spark and FPGA.
[0014] Furthermore, the parallel heterogeneous clustering acceleration framework reserves a common storage area, aiming to directly write the externally input data and make it available for common reading by different levels of execution environments; and with the help of the underlying address mapping mechanism, reduce the need for redundant copying on the local side; the local side only needs to keep an address pointer for identifying the storage location, and can find the corresponding storage start position according to the address pointer, so as to complete data access or processing.
[0015] Furthermore, when using the parallel heterogeneous clustering acceleration framework to run a preset clustering algorithm to achieve computing acceleration, the parallel heterogeneous clustering acceleration framework performs a timing analysis on the application of the external DRAM mapped to the FPGA, so as to try to use the same piece of memory for several different calculation stages.
[0016] On the other hand, the present invention also provides a heterogeneous computing clustering acceleration system based on Spark and FPGA, which is used to run a preset clustering algorithm to achieve computing acceleration. The heterogeneous computing clustering acceleration system based on Spark and FPGA is deployed in a master-slave structure, including a master node and multiple slave nodes; wherein, the master node undertakes the management tasks of the entire system, including communication with each slave node, job generation, allocation, and scheduling; each slave node is respectively connected to at least one FPGA, and after receiving a task, the computing process of the corresponding task is handed over to the FPGA for operation, and each slave node is equipped with Spark Streaming.
[0017] Further, the slave node is connected to the FPGA through a PCIe interface, and data exchange between the slave node and the FPGA is achieved through a PCIe communication mechanism.
[0018] Further, the system uses the C++ language to write the content of the OpenCL framework and compiles it into a.so dynamic link library file dedicated to the Linux system, which is used as the bridge between Spark and FPGA.
[0019] Further, the system reserves a common storage area, aiming to directly write the externally input data and make it available for common reading by different levels of execution environments; and with the help of the underlying address mapping mechanism, the need for redundant replication on the local side is reduced; the local side only needs to retain an address pointer for identifying the storage location, and it can find the corresponding storage start position according to the address pointer, so as to complete data access or processing.
[0020] Further, when using the heterogeneous computing clustering acceleration system to run a preset clustering algorithm to achieve computing acceleration, the heterogeneous computing clustering acceleration system performs timing analysis on the application of the external DRAM mapped to the FPGA, so as to try to use the same piece of memory for several different computing stages.
[0021] On yet another aspect, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0022] On still another aspect, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.
[0023] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0024] The present invention uses FPGA to accelerate the clustering calculation tasks in Spark Streaming, integrates a variety of optimization techniques and connection methods, designs a parallel heterogeneous clustering acceleration framework based on Spark and FPGA, optimizes the data transmission of the framework for the big data environment, and also conducts timing analysis on the data of various clustering methods to reuse the memory, making the most of the precious storage area of FPGA and providing the possibility to accommodate more data calculations. The experimental results show that the framework can significantly reduce the latency of iterative calculations for a large amount of data, and is applicable not only to various optimized K-means algorithms, but also to other clustering algorithms such as DBSCAN, ensuring the scalability of the clustering method. It optimizes the performance bottleneck when the CPU processes computationally intensive tasks in the existing computing framework, improves the computing efficiency and real-time performance when processing large-scale data sets, provides an optimization strategy for the FPGA kernel design, and enhances the utilization efficiency of hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 is a schematic execution flowchart of the heterogeneous computing clustering acceleration method based on Spark and FPGA provided by the embodiment of the present invention;
[0027] Figure 2 is a system block diagram of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0029] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of the word "exemplarily" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0030] The First Embodiment
[0031] This embodiment provides a heterogeneous computing clustering acceleration method based on Spark and FPGA. It should be noted that through parallel task execution and pipeline processing, FPGA can achieve efficient data flow control and computing optimization at the hardware level, significantly improving computing throughput. Different from the CPU, FPGA does not need to rely on complex control logic, can execute specific algorithms more directly, and reduce the latency of instruction execution. FPGA has a large number of programmable logic resources and dedicated hardware acceleration units, which has great advantages for the high parallelization of data and computing. Secondly, the programmability and customization ability of FPGA enable developers to design and optimize the hardware architecture according to specific algorithm requirements, realize the fine scheduling and efficient utilization of computing resources, and can adapt to diverse clustering tasks. In addition, through customized hardware acceleration design, FPGA can significantly reduce power consumption and be more environmentally friendly. Therefore, this embodiment uses FPGA to accelerate the clustering computing task in Spark Streaming. The integration and improvement of data transmission optimization mechanism, memory reuse technology and the customized optimization of FPGA kernel in the implementation of hardware-accelerated clustering algorithm.
[0032] This method can be implemented by an electronic device, and the execution process of this method is as Figure 1 shown, including the following steps:
[0033] S1, design a parallel heterogeneous clustering acceleration framework based on Spark and FPGA;
[0034] S2, use the parallel heterogeneous clustering acceleration framework to run a preset clustering algorithm to achieve computing acceleration.
[0035] Among them, it should be noted that in this embodiment, the entire system architecture is deployed in a master-slave structure, with one server as the Master and two servers as Slavers. The master node undertakes the management responsibilities of the entire architecture, including functions such as communication with each slave node, job generation, allocation, and scheduling. Each slave node is connected to an FPGA board through a PCIe interface, which can not only enable the computing to be handed over to the FPGA for operation after receiving the task, but also realize the expansion of multiple FPGA boards. Each sub-node is equipped with a Spark Streaming and an optimized clustering algorithm kernel program group, waiting for the master node to allocate jobs.
[0036] To achieve the efficient integration of FPGA in the system, a PCIe interface is adopted in the slave node to connect the FPGA to the server, and a high-bandwidth and low-latency data exchange between the two is realized through the PCIe communication mechanism. At the same time, to enable the communication between Spark framework tasks and FPGA, the content of the OpenCL framework is written in C++ language and compiled into a.so dynamic link library file dedicated to the Linux system, which serves as a bridge between Spark and FPGA. Specifically, the dynamic link library realizes functions such as checking the currently available hardware platform, creating an OpenCL process and context environment, and achieving efficient data transfer between the host and FPGA.
[0037] The FPGA development process based on OpenCL organically integrates hardware design and high-level language programming. By writing the core logic of various clustering algorithms in the form of OpenCL C in.cl and performing appropriate algorithm optimization and parallel design according to the hardware characteristics of FPGA (such as parallelism, storage hierarchy, etc.), and then using the aoc compilation tool chain provided by Intel FPGA to automatically convert the OpenCL C code into a high-level hardware description, and performing synthesis, place and route, and routing on the hardware description to generate the final configuration bitstream that can be deployed to the FPGA and encapsulate it into an.aocx binary image file, which greatly simplifies the workload of writing HDL (Verilog / VHDL) code in traditional FPGA development, reduces the hardware programming threshold, and also realizes the deep combination of the software layer and the hardware layer, enabling flexible change of the clustering calculation core logic under the same architecture, ensuring portability and scalability.
[0038] Before optimization, after the system framework was completed, it was found that although FPGA could improve the computing efficiency, it correspondingly increased the number of data reception times. To enable the FPGA to execute computing tasks, the entire framework needed to first read the data from the disk or HDFS into the Spark's resilient distributed dataset in local memory, and then distribute it to the appropriate nodes. When the data is passed into the child node, it also needs to be converted into a Java array and then into a C++ array through JNI, and after the calculation, it needs to be returned along the original path, and the data transfer time cost is too high, offsetting the computing acceleration of FPGA.
[0039] To reduce excessive data transmission and memory occupation, first, a general storage area is reserved in the system, aiming to directly write the externally input data and make it available for different levels of execution environments to read jointly. Through this shared storage method, there is no need to frequently perform data transfer operations between the application layer and the local side, thus significantly reducing the overall data transmission overhead. Subsequently, with the help of the underlying address mapping mechanism, the need for redundant replication on the local side is further reduced. The local execution end only needs to retain an address pointer for identifying the storage location, and it can find the corresponding storage start position according to this mapping, thus efficiently completing data access or processing. With this strategy, not only can the data flow efficiency between different computing units be maximized, but also the system communication and management processes can be significantly simplified, thereby enhancing the scalability and maintainability of the overall architecture in high-concurrency and large-scale data processing scenarios.
[0040] In addition, in the kernel program, timing analysis is performed on the application for the external DRAM mapped to the FPGA board, so as to try to use the same piece of memory for several different computing stages.
[0041] Through this framework, the operation of the clustering algorithm can be accelerated, providing a solid foundation for many application scenarios such as decision support, market segmentation, image recognition, and image segmentation. For example, the clustering algorithm can be run through this framework. By analyzing data such as customers' purchase history and browsing behavior, customers can be divided into different consumer groups, and then products and services suitable for each group can be launched to achieve market segmentation. Another example is that in social media and e-commerce platforms, the clustering algorithm can be run through this framework to group users according to their behaviors, interests, and preferences. This can help the platform better understand users' needs and provide personalized recommendations and services. For another example, in image processing, the clustering algorithm can be run through this framework to segment pixels or regions in the image according to features such as color and texture. This method has wide applications in fields such as medical image analysis and satellite image processing. By dividing different parts of the image into different clusters, image analysis and processing can be better carried out. It can also be used in information retrieval and natural language processing systems. By running the clustering algorithm through this framework, documents can be classified according to themes or contents. By analyzing information such as keywords and sentence structures in the documents, the documents can be divided into different theme categories. This method can be used for content management of news websites, classification of academic papers, etc.
[0042] Next, take the selection of the initial points and subsequent iterative steps of Kmeans++ as an example to systematically introduce the memory reuse mechanism used. The goal of K-Means++ is to select more high-quality initial centroids for the traditional K-Means clustering, thereby accelerating iterative convergence and improving the result quality. The core process of its selection of initial points is as follows:
[0043] (1) Randomly select the first centroid: Randomly select a data point from the dataset with equal probability as the first centroid.
[0044] (2) Calculate the distances and select a new centroid based on weighted probability: Let D(x) denote the distance from data point x to its nearest selected centroid. The total distance sum is:
[0045]
[0046] Select a new centroid with probability .
[0047] (3) Repeat until K centroids are selected.
[0048] Subsequent iterations use the iterative steps of traditional K-Means. In a hardware-accelerated environment with limited resources, allocating large chunks of storage for each array separately will waste precious external DRAM resources; if the data structures required in the initialization and iteration phases can coexist at different times, the same buffer can be reused.
[0049] Reusing the global memory buffer can avoid allocating large memory for different process data separately in an FPGA accelerator environment where the on-board DRAM is very limited, thus effectively reducing the overall occupancy of the global memory. Moreover, it can also reduce the initiation of additional initialization, copy, or memory management operations, and reduce the movement and management of useless data. Using the same buffer can enable related data to be updated and accessed sequentially or continuously in the same area, thereby improving data locality to a certain extent and reducing the random access overhead of external memory. For FPGAs, synthesis and placement and routing will allocate resources and schedule according to the number and size of global buffers declared in the kernel. Reusing buffers can reduce the competition for storage resources in high-level synthesis, help the compiler better perform pipelining optimization, and reduce the logical complexity, providing the possibility of deploying more functions on the same FPGA board.
[0050] Suppose there are currently N two-dimensional data, each data type occupies 8 bytes, and we want to cluster them into K clusters. Using the Kmeans++ method to determine the initial points, and the number of iterations of the clustering algorithm is T. Without using the memory reuse mechanism, the memory requirements are analyzed as follows:
[0051] 1. Memory of the original dataset:
[0052] Memory0 = 2 × 8 × N bytes
[0053] 2. Calculate the distance from each point to the current centroid and the total distance when selecting the initial centroids. Since the distance between all points and the last initial center point does not need to be calculated after it is determined, only the distances between N points and K - 1 points need to be calculated:
[0054] Memory1 = 8(K - 1)(N - 1)+(K - 1) bytes
[0055] 3. Store K centroids, each centroid occupying 8 bytes.
[0056] Memory KC = 16K bytes
[0057] 4. In each iteration, the distances from each data point to all K centroids are calculated, and a cluster label is assigned to each data point. Calculating the distances requires storing N×K distance values, each distance value occupying 8 bytes. Also, an array of size N is needed to store the cluster labels of each data point, and each cluster label can be stored in 4 bytes (assuming integer storage). So the memory usage for each iteration is:
[0058] Memory C = 8NK + 4N bytes
[0059] So the total memory usage is:
[0060] Memory = 16N + 8(K - 1)(N - 1)+(K - 1)+16K+(8NK + 4N) bytes
[0061] = 12N + 16NK + 9K + 7 bytes
[0062] If Memory1 and Memory share the same space when applying for kernel space, the total memory usage drops to: C The total memory usage drops to:
[0063] Memory = 16N + 16K+(8NK + 4N) bytes = 20N + 16K + 8NK bytes
[0064] And so on, timing analysis and optimization have been carried out for various clustering algorithms (such as K - means, DBSCAN, etc.) for the FPGA environment.
[0065] In summary, this embodiment provides a heterogeneous computing clustering acceleration method based on Spark and FPGA, which uses FPGA to accelerate the clustering calculation tasks in Spark Streaming, integrates a variety of optimization techniques and connection methods, designs a parallel heterogeneous clustering acceleration framework based on Spark and FPGA, optimizes the data transmission of the framework for the big data environment, and also conducts timing analysis on the data of various clustering methods to reuse the memory, making the most of the precious storage area of FPGA and providing the possibility to accommodate more data calculations. The experimental results show that this framework can significantly reduce the latency of iterative calculations for a large amount of data, and is not only applicable to various optimized K-means algorithms, but also applicable to other clustering algorithms such as DBSCAN, ensuring the scalability of clustering methods. It optimizes the performance bottleneck when the CPU processes computationally intensive tasks in the existing computing framework, improves the computing efficiency and real-time performance when processing large-scale data sets, provides an optimization strategy for FPGA kernel design, and enhances the utilization efficiency of hardware resources.
[0066] Second Embodiment
[0067] This embodiment provides a heterogeneous computing clustering acceleration system based on Spark and FPGA, which is used to run a preset clustering algorithm to achieve computing acceleration. The heterogeneous computing clustering acceleration system based on Spark and FPGA is deployed in a master-slave structure, including a master node and multiple slave nodes; among them, the master node undertakes the management tasks of the entire system, including communication with each slave node, generation, allocation, and scheduling of jobs; each slave node is respectively connected to at least one FPGA, and after receiving a task, hands over the calculation process of the corresponding task to the FPGA for operation, and each slave node is equipped with Spark Streaming.
[0068] It should be noted that the heterogeneous computing clustering acceleration system based on Spark and FPGA in this embodiment corresponds to the heterogeneous computing clustering acceleration method based on Spark and FPGA in the above first embodiment; among them, the functions implemented by each functional module in the heterogeneous computing clustering acceleration system based on Spark and FPGA in this embodiment correspond one by one to the respective process steps in the heterogeneous computing clustering acceleration method based on Spark and FPGA in the above first embodiment; therefore, it will not be elaborated here.
[0069] Third Embodiment
[0070] This embodiment provides an electronic device, such as Figure 2As shown, the electronic device includes: a processor and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the first embodiment above. In addition, the electronic device may further include a transceiver, the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.
[0071] Next, a specific introduction to the various components of the electronic device will be given in conjunction with Figure 2 :
[0072] Among them, the processor is the control center of the electronic device. The electronic device may include multiple processors, and each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor may be a single processor or a collective term for multiple processing elements. For example, the processor is one or more central processing units (central processing unit, CPU), or may be other general-purpose processors, application specific integrated circuits (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0073] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 2 CPU0 and CPU1 shown in, of course, this is only an exemplary illustration.
[0074] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner may refer to the above method embodiments and will not be elaborated here.
[0075] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 2 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0076] The transceiver may include a receiver and a transmitter ( Figure 2 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 2 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0077] In addition, it should be noted that Figure 2 the structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine certain components, or have a different component layout. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above may refer to the technical effects described in the first embodiment above, so they will not be elaborated here.
[0078] Fourth Embodiment
[0079] This embodiment provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be ROM, random access memory, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.
[0080] In addition, it should be noted that the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented using software, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center containing one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0081] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide for implementing in the process Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.
[0083] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element. In addition, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally means that the contextually related objects are in an "or" relationship, but it may also mean an "and / or" relationship, which can be understood specifically with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0084] In addition, it can be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0085] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0086] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0087] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0088] Finally, it should be noted that the above are only the preferred embodiments of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle of the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A heterogeneous computing clustering acceleration method based on Spark and FPGA, characterized in that, Including: Design a parallel heterogeneous clustering acceleration framework based on Spark and FPGA. Among them, the parallel heterogeneous clustering acceleration framework is deployed in a master-slave structure, including a master node and multiple slave nodes. Among them, the master node undertakes the management tasks of the entire framework, including communication with each slave node, job generation, allocation, and scheduling. Each slave node is respectively connected to at least one FPGA, and after receiving a task, it hands over the calculation process of the corresponding task to the FPGA for operation, and each slave node is equipped with Spark Streaming. Use the parallel heterogeneous clustering acceleration framework to run a preset clustering algorithm to achieve computing acceleration.
2. The heterogeneous computing clustering acceleration method based on Spark and FPGA according to claim 1, characterized in that, The slave node uses a PCIe interface to connect to the FPGA and realizes data exchange between the slave node and the FPGA through the PCIe communication mechanism.
3. The heterogeneous computing clustering acceleration method based on Spark and FPGA according to claim 1, wherein The method further includes: Write the content of the OpenCL framework in C++ language and compile it into a.so dynamic link library file dedicated to the Linux system, which is used as the bridge between Spark and FPGA.
4. The heterogeneous computing clustering acceleration method based on Spark and FPGA according to claim 1, characterized in that, The parallel heterogeneous clustering acceleration framework reserves a common storage area, aiming to directly write external input data and make it available for different levels of execution environments to read together; and with the help of the underlying address mapping mechanism, reduce the need for redundant copying on the local side; The local side only needs to keep an address pointer for identifying the storage location, and it can find the corresponding storage start position according to the address pointer, so as to complete data access or processing.
5. The heterogeneous computing clustering acceleration method based on Spark and FPGA according to claim 1, characterized in that, When using the parallel heterogeneous clustering acceleration framework to run a preset clustering algorithm to achieve computing acceleration, the parallel heterogeneous clustering acceleration framework performs timing analysis on the application of the external DRAM mapped to the FPGA, so as to try to use the same piece of memory for several different calculation stages.
6. A heterogeneous computing clustering acceleration system based on Spark and FPGA is used to run a preset clustering algorithm to achieve computing acceleration, and is characterized in that The heterogeneous computing clustering acceleration system based on Spark and FPGA is deployed in a master-slave structure, including a master node and multiple slave nodes. Among them, the master node undertakes the management tasks of the entire system, including communication with each slave node, job generation, allocation, and scheduling. Each slave node is respectively connected to at least one FPGA, and after receiving a task, it hands over the calculation process of the corresponding task to the FPGA for operation, and each slave node is equipped with Spark Streaming.
7. The heterogeneous computing clustering acceleration system based on Spark and FPGA according to claim 6, characterized in that, The slave node uses a PCIe interface to connect to the FPGA and realizes data exchange between the slave node and the FPGA through the PCIe communication mechanism.
8. The heterogeneous computing clustering acceleration system based on Spark and FPGA according to claim 6, characterized in that, The system writes the content of the OpenCL framework in C++ language and compiles it into a.so dynamic link library file dedicated to the Linux system, which is used as the bridge between Spark and FPGA.
9. The heterogeneous computing clustering acceleration system based on Spark and FPGA according to claim 6, wherein, The system reserves a common storage area, aiming to directly write external input data and make it available for different levels of execution environments to read together; and with the help of the underlying address mapping mechanism, reduce the need for redundant copying on the local side; The local side only needs to keep an address pointer for identifying the storage location, and it can find the corresponding storage start position according to the address pointer, so as to complete data access or processing.
10. The heterogeneous computing clustering acceleration system based on Spark and FPGA according to claim 6, characterized in that, When using the heterogeneous computing cluster acceleration system to run a preset clustering algorithm to achieve computing acceleration, the heterogeneous computing cluster acceleration system performs timing analysis on the application for external DRAM mapped to the FPGA, so as to try to use the same piece of memory for several different computing stages.