A Heterogeneous Acceleration Board Computing Method, Device, Equipment and Medium
By generating management pools in management nodes and optimizing execution trees, segmenting query tasks and using intelligent network cards for data preprocessing and aggregation operations, the problem of reducing the total data transmission amount under the condition of fixed network communication bandwidth is solved, and efficient calculation and resource release of heterogeneous acceleration boards is achieved.
Patent Information
- Application Number
- CN202310148335.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-21
AI Technical Summary
Under the condition of fixed network communication bandwidth, how to reduce the total effective data transmission of distributed databases.
By sending device resource query instructions to each acceleration node in the management node, generating a management pool and optimizing the execution tree, segmenting query tasks and using intelligent network cards for data preprocessing and aggregation operations, the calculation of heterogeneous acceleration board cards is realized.
It effectively reduces data communication delay, reduces the communication process between heterogeneous boards and cards, and frees up the computing resources of the host nodes.
Smart Images

Figure CN116192849B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interconnection architectures of distributed databases, and particularly to a heterogeneous acceleration board computing method, device, equipment and medium. Background Art
[0002] From the perspective of downstream application fields, current distributed databases are mainly applied to industries that generate massive data such as finance, telecommunications, and the Internet, as well as industries with a large number of customers such as express logistics, catering services, and tourism services. The common feature of these industries is that they have massive data and the foreseeable continuous growth of data in the future. To meet the data transmission between each distributed node, various 40G / 100G network communications and various network optimization technologies have emerged.
[0003] As can be seen from the above, how to reduce the total amount of effective data transmission under the condition of a fixed network communication bandwidth is a problem to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a heterogeneous acceleration board computing method, device, equipment and medium, which can reduce the total amount of effective data transmission under the condition of a fixed network communication bandwidth. The specific solutions are as follows:
[0005] In a first aspect, the present application discloses a heterogeneous acceleration board computing method, which is applied to a management node and includes:
[0006] Sending device resource query instructions to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information;
[0007] Obtaining all the resource pool information, generating a management pool based on all the resource pool information, obtaining a query operation instruction sent by a user, parsing the query operation instruction to obtain an execution tree, optimizing the execution tree based on the management pool to obtain a new execution tree, and generating a query task according to the new execution tree and the query operation instruction;
[0008] Sending the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sends each query subtask to its corresponding heterogeneous acceleration board;
[0009] Obtain data from local storage resources, and send the data to the local smart network card. Use the smart network card to preprocess the data to obtain preprocessed data, and send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results.
[0010] Optionally, optimizing the execution tree based on the management pool to obtain a new execution tree includes:
[0011] Determine the resource occupancy information from the query operation instruction, and determine the usage status of the acceleration node from the management pool; wherein, the usage status includes an idle state and a non-idle state;
[0012] Optimize the execution tree based on the resource occupancy information and the usage status to obtain a new execution tree.
[0013] Optionally, sending the query task to the acceleration node so that the acceleration node divides the query task based on its own resource pool information and current actual resource information includes:
[0014] Send the query task to the acceleration node, so that the acceleration node determines the current actual resource information in the idle state and the current actual resource information in the non-idle state of itself, then estimates the duration of the current actual resource information in the non-idle state to obtain an estimated processing duration, and divides the query task based on the current actual resource information in the idle state and the estimated processing duration.
[0015] Optionally, using the smart network card to preprocess the data to obtain preprocessed data and sending the preprocessed data to the acceleration node includes:
[0016] Determine the data volume of the data, call the local data filtering array and row-column conversion array, use the smart network card and based on the data volume and the query operation instruction to preprocess the data to obtain preprocessed data, and send the preprocessed data to the acceleration node through the local network port module.
[0017] Optionally, before sending the device resource query instruction to each acceleration node, it further includes:
[0018] Screen out the management node from all host nodes, and use the remaining all host nodes except the management node as acceleration nodes;
[0019] Establish a connection relationship between the management node and each acceleration node.
[0020] Optionally, the device resource query instructions are respectively sent to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information, including:
[0021] The device resource query instructions are respectively sent to each acceleration node by using the connection relationship, so that each acceleration node uses the device management module and memory management module in its own intelligent network card, and determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information.
[0022] Optionally, the step of sending the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result, includes:
[0023] The preprocessed data is sent to the acceleration node by using the connection relationship, so that the acceleration node uses its own intelligent network card to send the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result, and then determines whether the calculation result contains an aggregation operation. If the calculation result contains an aggregation operation, the aggregation operation acceleration module is called to process the calculation result to obtain the final calculation result.
[0024] In a second aspect, the present application discloses a heterogeneous acceleration board computing device, including:
[0025] An instruction sending module, configured to send device resource query instructions to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information;
[0026] A management pool generation module, configured to obtain all the resource pool information, generate a management pool based on all the resource pool information, obtain a query operation instruction sent by a user, parse the query operation instruction to obtain an execution tree, optimize the execution tree based on the management pool to obtain a new execution tree, and generate a query task according to the new execution tree and the query operation instruction;
[0027] A query task sending module, configured to send the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sends each query subtask to its corresponding heterogeneous acceleration board;
[0028] The heterogeneous acceleration board computing module is used to obtain data from local storage resources, send the data to the local smart network card, preprocess the data by using the smart network card to obtain preprocessed data, and send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results.
[0029] In a third aspect, the present application discloses an electronic device, including:
[0030] A memory for storing a computer program;
[0031] A processor for executing the computer program to implement the foregoing heterogeneous acceleration board computing method.
[0032] In a fourth aspect, the present application discloses a computer storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed heterogeneous acceleration board computing method are implemented.
[0033] As can be seen, the present application provides a heterogeneous acceleration board computing method, which includes sending device resource query instructions to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; obtaining all the resource pool information, generating a management pool based on all the resource pool information, obtaining a query operation instruction sent by a user, parsing the query operation instruction to obtain an execution tree, optimizing the execution tree based on the management pool to obtain a new execution tree, generating a query task according to the new execution tree and the query operation instruction; sending the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sending each query subtask to its corresponding heterogeneous acceleration board; obtaining data from the local storage resource, sending the data to the local smart network card, preprocessing the data by using the smart network card to obtain preprocessed data, and sending the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result. By attaching data preprocessing such as data filtering and row-column data conversion to the smart network card, the present application can effectively reduce the network communication pressure with the lower-level computing nodes; by attaching post-processing of partial aggregation operations to the smart network card, the calculation results of multiple heterogeneous acceleration boards can be summarized, reducing the communication process between different heterogeneous boards; by performing resource pool processing on each host node and acceleration node, the present application can achieve proximity processing of device management and task scheduling, reduce data communication latency, and effectively release the computing resources of the host node. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0035] Figure 1 Flowchart of a heterogeneous acceleration board computing method disclosed in the present application;
[0036] Figure 2 Flowchart of a heterogeneous acceleration board computing method disclosed in the present application;
[0037] Figure 3 Schematic diagram of the structure of a host node with a heterogeneous acceleration board disclosed in the application;
[0038] Figure 4 Schematic diagram of an acceleration node structure with multiple heterogeneous acceleration boards disclosed in the application;
[0039] Figure 5 Schematic diagram of a host node structure without heterogeneous acceleration boards disclosed in the application;
[0040] Figure 6 Schematic diagram of an intelligent network card on the host node side disclosed in the application;
[0041] Figure 7 Schematic diagram of an intelligent network card on the acceleration node side disclosed in the application;
[0042] Figure 8 Specific flowchart of a heterogeneous acceleration board calculation method disclosed in the application;
[0043] Figure 9 Schematic diagram of a heterogeneous acceleration board calculation device structure disclosed in the present application;
[0044] Figure 10 Structure diagram of an electronic device provided by the application. Specific implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] From the perspective of downstream application fields, current distributed databases are mainly applied to industries that generate massive data, such as finance, telecommunications, and the Internet, as well as industries with a large number of customers, such as express logistics, catering services, and tourism services. The common feature of these industries is that they have massive data and the foreseeable continuous growth of data in the future. To meet the data transmission between each distributed node, various 40G / 100G network communications and various network optimization technologies have emerged. As can be seen from the above, how to reduce the total amount of effective data transmission under the condition of a fixed network communication bandwidth is a problem to be solved in this field.
[0047] Refer to Figure 1 As shown, an embodiment of the present invention discloses a heterogeneous acceleration board calculation method, which is applied to a management node and specifically may include:
[0048] Step S11: Send device resource query instructions to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information.
[0049] In this embodiment, before sending the device resource query instructions to each acceleration node respectively, it further includes: screening out a management node from all host nodes, and using the remaining host nodes except the management node as acceleration nodes; establishing a connection relationship between the management node and each acceleration node. Then, use the connection relationship to send the device resource query instructions to each acceleration node respectively, so that each acceleration node uses the device management module and memory management module in its own intelligent network card, and determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information. That is to say, first select a management node from all host nodes, and then use all other hosts except this management node as acceleration nodes. The management node monitors the effective heterogeneous acceleration boards mounted on the acceleration nodes in real time, detects the available resource status of the heterogeneous acceleration boards of the computing nodes, and sends the device resource query instructions to each acceleration node. When an acceleration node receives the device resource query instructions, it determines the status of the heterogeneous acceleration boards in the acceleration node, such as the idle state and the non-idle state, records the computing tasks executed by the currently online heterogeneous acceleration boards, and when a heterogeneous acceleration board is abnormal, reallocates the computing tasks, records the idle computing resources of the heterogeneous acceleration boards, and finally generates resource pool information including heterogeneous acceleration board device information and resource information. Then each node sends its own resource pool information to the management node respectively.
[0050] Step S12: Obtain all the resource pool information, generate a management pool based on all the resource pool information, obtain a query operation instruction sent by a user, parse the query operation instruction to obtain an execution tree, optimize the execution tree based on the management pool to obtain a new execution tree, and generate a query task according to the new execution tree and the query operation instruction.
[0051] Step S13: Send the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sends each query subtask to its corresponding heterogeneous acceleration board.
[0052] Step S14: Obtain data from local storage resources, send the data to the local smart network card, use the smart network card to preprocess the data to obtain preprocessed data, and send the preprocessed data to the acceleration node so that the acceleration node can send the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results.
[0053] In this embodiment, the smart network card of the host node has a preprocessing function, which is used to realize the communication between the host and other computing node hosts and the heterogeneous acceleration board array of the computing node, and realize the data preprocessing of query operations such as data filtering and row-column data conversion of query operations. The smart network card of the acceleration node has a postprocessing function, which is used to realize the communication between the heterogeneous acceleration board array and the hosts of other nodes and the heterogeneous acceleration board array of the computing node, and realize the data postprocessing of query operations such as partial aggregation operations of query operations.
[0054] In this embodiment, after sending the data to the local smart network card, determine the data volume of the data, call the local data filtering array and row-column conversion array, use the smart network card and based on the data volume and the query operation instruction to preprocess the data to obtain preprocessed data, and send the preprocessed data to the acceleration node through the local network interface module. After obtaining the preprocessed data, use the connection relationship between the management node and each acceleration node to send the preprocessed data to the acceleration node so that the acceleration node can use its own smart network card to send the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results, and then determine whether the calculation results contain aggregation operations. If the calculation results contain aggregation operations, call the aggregation operation acceleration module to process the calculation results to obtain the final calculation results.
[0055] In this embodiment, the device resource query instructions are respectively sent to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; all the resource pool information is obtained, a management pool is generated based on all the resource pool information, a query operation instruction sent by a user is obtained, the query operation instruction is parsed to obtain an execution tree, the execution tree is optimized based on the management pool to obtain a new execution tree, and a query task is generated according to the new execution tree and the query operation instruction; the query task is sent to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and the current actual resource information to obtain each query subtask, and each query subtask is sent to its corresponding heterogeneous acceleration board; data is obtained from the local storage resource, and the data is sent to the local smart network card, and the smart network card is used to preprocess the data to obtain preprocessed data, and the preprocessed data is sent to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results. By attaching data preprocessing such as data filtering and row-column data conversion to the smart network card, the network communication pressure with the lower-level computing nodes can be effectively reduced; by attaching data postprocessing of partial aggregation operations to the smart network card, the calculation results of multiple heterogeneous acceleration boards can be summarized, reducing the communication process between different heterogeneous boards; by performing resource pool processing on each host node and acceleration node, proximity processing of device management and task scheduling can be realized, reducing data communication latency and effectively releasing the computing resources of the host node.
[0056] See Figure 2 As shown, an embodiment of the present invention discloses a heterogeneous acceleration board calculation method, which is applied to a management node and specifically may include:
[0057] Step S21: The device resource query instructions are respectively sent to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information.
[0058] Step S22: All the resource pool information is obtained, a management pool is generated based on all the resource pool information, a query operation instruction sent by a user is obtained, the query operation instruction is parsed to obtain an execution tree, the resource occupancy information is determined from the query operation instruction, and the usage status of the acceleration node is determined from the management pool; wherein, the usage status includes an idle status and a non-idle status.
[0059] Step S23: Optimize the execution tree based on the resource occupancy information and the usage status to obtain a new execution tree, and then generate a query task according to the new execution tree and the query operation instruction.
[0060] Step S24: Send the query task to the acceleration node, so that the acceleration node determines the current actual resource information in the idle state and the current actual resource information in the non-idle state of itself, then estimates the duration of the current actual resource information in the non-idle state to obtain an estimated processing duration, and divides the query task based on the current actual resource information in the idle state and the estimated processing duration to obtain each query subtask, and sends each query subtask to the corresponding different heterogeneous acceleration boards of itself.
[0061] In this embodiment, after the query task is sent to the acceleration node, the task scheduling module in the acceleration node will evaluate the query operations in the first generated execution tree that can be accelerated by the heterogeneous acceleration board, combine the current actual resource information in the idle state and the current actual resource information in the non-idle state of the current heterogeneous acceleration board, and the duration estimation of the computing tasks in the task queue of the non-idle computing resources, and divide the query task into multiple subtasks and allocate them to different heterogeneous acceleration boards.
[0062] Step S25: Obtain data from the local storage resource, send the data to the local smart network card, preprocess the data by using the smart network card to obtain preprocessed data, and send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result.
[0063] Among them, the smart network card is divided into a smart network card with a preprocessing function and a smart network card with a post-processing function. There are two composition methods for the host node. One is the composition method of the management node composed of the host and the smart network card (i.e., Figure 5 ) and the other is the composition method of the computing node of the host equipped with heterogeneous acceleration boards and the smart network card (i.e., Figure 3 ); the heterogeneous acceleration board array of the acceleration node is composed of FPGA acceleration boards and is used to realize heterogeneous acceleration in the form of a query operation pipeline (i.e., Figure 4 ); the heterogeneous computing board communicates with other nodes through the smart network card, and different heterogeneous computing boards in a single computing node perform routing communication through the smart network card; the smart network card with a post-processing function (i.e., Figure 7, the smart network card in the acceleration node), including a data post - processing calculation module, a network communication module, a data routing communication module, a device management and memory management module; the device management and memory management module is used to detect the status of heterogeneous acceleration boards in the current node in real time, allocate and recycle the memory of each heterogeneous acceleration board, generate a resource pool for all heterogeneous acceleration boards of this node, and respond to the queries of the host - side node. According to the task instructions sent by the host - side node, resource scheduling of heterogeneous computing boards is performed. The smart network card is composed of a network card with an FPGA, and flexible configuration of the smart network card can be realized; the smart network card in the host node (i.e., Figure 6 ) includes a data pre - processing module, a data sending module, and a data receiving module.
[0064] The specific process of this application is as Figure 8 shown, and taking host node A (i.e., Figure 5 ) as the management node, host node B (i.e., Figure 3 ) and host node C (i.e., Figure 4)Taking the acceleration node as an example, (1) The management node A periodically sends device resource query instructions to the acceleration node B and the acceleration node C; (2) The acceleration node B and the acceleration node C will periodically query whether they are loaded with heterogeneous acceleration boards, detect the status of the heterogeneous acceleration boards, count the resource usage of the current heterogeneous acceleration boards, and estimate the current task queue status. When the acceleration node B and the acceleration node C receive the device resource query instructions, they generate resource pool information containing the device information and resource information of the heterogeneous acceleration boards, and send the resource pool information to the management node A; (3) After A receives the resource pool information replied by other nodes, it generates a management pool, and the management pool contains a device list and a resource pool table; (4) After A receives the query operation instruction sent by the user, through instruction parsing and optimization, it generates an execution tree. The task scheduling module in A queries the management pool, further optimizes the execution tree, generates a new execution tree, and starts the scheduling and distribution of task instructions and data; For example, assume that in this query operation, the query operation that can be accelerated needs to call 1.5 heterogeneous acceleration boards, and in the management pool, the heterogeneous acceleration board resource pool shows that there are three heterogeneous acceleration boards in the C idle state; A sends a query task to C, and at the same time sends the data filtering operation or row-column data conversion task of this query operation to the smart network card of A. After receiving the query task, C checks the current actual resource usage of the heterogeneous acceleration boards and the memory / cache usage (i.e., the current actual resource information). Since resource pooling has been performed on all heterogeneous acceleration boards of this node, the query task needs to be disassembled by the task scheduling module in C. Through the mapping relationship between the idle heterogeneous acceleration board resource table (in the resource pool) and the actual idle heterogeneous acceleration boards, the query subtasks are sent to different heterogeneous acceleration boards; The memory management module in C applies for sufficient memory space and cache space for the query operation and records it; (5) A calls data from the local storage resource and sends it to the smart network card of A. The smart network card calls the data filtering array and row-column conversion array according to the device resource query instruction and the data volume received from the host node A, performs preprocessing on the data, caches the preprocessed data, and sends it to the smart network card of C through the network interface module; (6) After the smart network card of C receives the data, it determines the specified heterogeneous acceleration board for data distribution through the task scheduling module in C, and distributes the data to the specified heterogeneous acceleration board through the routing communication module. According to different requirements of the acceleration operation, on C, through the task scheduling module, such as after the sorting operation in the query operation, the connection operation is sequentially executed; Control different heterogeneous acceleration boards, and use the routing communication module of the smart network card to achieve point-to-point communication between different heterogeneous acceleration boards; (7) After the heterogeneous acceleration board of C completes the calculation, the smart network card of C notifies A and waits for the data reading instruction from A; (8) When C receives the data reading instruction from A, the acceleration node C reads the calculation result from the heterogeneous acceleration board through the task scheduling module, and the memory management module reclaims the memory and cache resources used in this sub-operation; (9) The smart network card of C judges whether the calculation result contains an aggregation operation. If the calculation result contains some aggregation operations, such as sum, count, minimum, maximum, average, etc., the smart network card calls the aggregation operation acceleration module in the data post-processing after passing the data, performs data post-processing on the calculation result to obtain the final calculation result, and then transmits the processed data to the specified host node through network communication according to the query operation instruction of A, that is, according to the task scheduling algorithm of A, C needs to transmit the final calculation result to other nodes for further acceleration processing. When subsequent acceleration processing involves preprocessing operations such as data filtering or row-column conversion, a preprocessing module can be deployed on the smart network card of the acceleration node C. After the data passes through the preprocessing module, it is transmitted to other nodes through network communication. In addition, Figure 3 The host node in can also be used as an acceleration node, and the smart network card of this node can also deploy a data post-processing module to accelerate the aggregation operation in data post-processing,Figure 3 The host node in
[0065] In this embodiment, the device resource query instructions are respectively sent to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; all the resource pool information is obtained, a management pool is generated based on all the resource pool information, a query operation instruction sent by a user is obtained, the query operation instruction is parsed to obtain an execution tree, the execution tree is optimized based on the management pool to obtain a new execution tree, a query task is generated according to the new execution tree and the query operation instruction; the query task is sent to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and the current actual resource information to obtain each query subtask, and each query subtask is sent to its corresponding heterogeneous acceleration board; data is obtained from the local storage resource, and the data is sent to the local smart network card, and the data is preprocessed by using the smart network card to obtain preprocessed data, and the preprocessed data is sent to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results. By attaching data preprocessing such as data filtering and row-column data conversion to the smart network card, the network communication pressure with the lower-level computing nodes can be effectively reduced; by attaching post-processing of partial aggregation operations to the smart network card, the calculation results of multiple heterogeneous acceleration boards can be summarized, reducing the communication process between different heterogeneous boards; by performing resource pool processing on each host node and acceleration node, the proximity processing of device management and task scheduling can be realized, reducing data communication latency and effectively releasing the computing resources of the host node.
[0066] See Figure 9 As shown in
[0067] An instruction sending module 11, configured to respectively send device resource query instructions to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information;
[0068] A management pool generation module 12, configured to obtain all the resource pool information, generate a management pool based on all the resource pool information, obtain a query operation instruction sent by a user, parse the query operation instruction to obtain an execution tree, optimize the execution tree based on the management pool to obtain a new execution tree, and generate a query task according to the new execution tree and the query operation instruction;
[0069] A query task sending module 13, configured to send the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sends each query subtask to its corresponding different heterogeneous acceleration boards;
[0070] A heterogeneous acceleration board computing module 14, configured to obtain data from local storage resources, send the data to a local smart network card, preprocess the data by using the smart network card to obtain preprocessed data, and send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for computing to obtain a computing result.
[0071] In this embodiment, the device resource query instructions are respectively sent to each acceleration node, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; all the resource pool information is obtained, a management pool is generated based on all the resource pool information, a query operation instruction sent by a user is obtained, the query operation instruction is parsed to obtain an execution tree, the execution tree is optimized based on the management pool to obtain a new execution tree, and a query task is generated according to the new execution tree and the query operation instruction; the query task is sent to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and each query subtask is sent to its corresponding heterogeneous acceleration board; data is obtained from the local storage resource and the data is sent to the local smart network card, the smart network card is used to preprocess the data to obtain preprocessed data, and the preprocessed data is sent to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain calculation results. By attaching data preprocessing such as data filtering and row-column data conversion to the smart network card in this application, the network communication pressure with the lower-level computing nodes can be effectively reduced; by attaching data postprocessing of partial aggregation operations to the smart network card, the calculation results of multiple heterogeneous acceleration boards can be summarized, reducing the communication process between different heterogeneous boards; by performing resource pool processing on each host node and acceleration node, proximity processing of device management and task scheduling can be achieved, reducing data communication latency and effectively releasing the computing resources of the host node.
[0072] In some specific embodiments, the management pool generation module 12 may specifically include:
[0073] A resource occupancy information determination module, configured to determine resource occupancy information from the query operation instruction and determine the usage status of the acceleration node from the management pool; wherein, the usage status includes an idle status and a non-idle status;
[0074] An execution tree optimization module, configured to optimize the execution tree based on the resource occupancy information and the usage status to obtain a new execution tree.
[0075] In some specific embodiments, the query task sending module 13 may specifically include:
[0076] The query task sending module is used to send the query task to the acceleration node, so that the acceleration node determines the current actual resource information in the idle state and the current actual resource information in the non-idle state of itself, then estimates the duration of the current actual resource information in the non-idle state to obtain the estimated processing duration, and divides the query task based on the current actual resource information and the estimated processing duration in the idle state.
[0077] In some specific embodiments, the heterogeneous acceleration board computing module 14 may specifically include:
[0078] The preprocessed data sending module is used to determine the data volume of the data, call the local data filtering array and row-column conversion array, use the intelligent network card and based on the data volume and the query operation instruction to preprocess the data to obtain preprocessed data, and send the preprocessed data to the acceleration node through the local network interface module.
[0079] In some specific embodiments, the instruction sending module 11 may specifically include:
[0080] The management node screening module is used to screen out the management node from all host nodes, and use the remaining all host nodes except the management node as acceleration nodes;
[0081] The connection relationship establishing module is used to establish the connection relationship between the management node and each acceleration node.
[0082] In some specific embodiments, the instruction sending module 11 may specifically include:
[0083] The resource pool information determining module is used to use the connection relationship to send the device resource query instruction to each acceleration node respectively, so that each acceleration node uses the device management module and memory management module in its own intelligent network card, and determines its own resource pool information based on the device resource query instruction; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information.
[0084] In some specific embodiments, the heterogeneous acceleration board computing module 14 may specifically include:
[0085] The judgment module is used to use the connection relationship to send the preprocessed data to the acceleration node, so that the acceleration node uses its own intelligent network card to send the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result, and then judges whether the calculation result contains an aggregation operation. If the calculation result contains an aggregation operation, the aggregation operation acceleration module is called to perform data processing on the calculation result to obtain the final calculation result.
[0086] Figure 10 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the heterogeneous acceleration board computing method executed by the electronic device disclosed in any of the foregoing embodiments.
[0087] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and specific limitations are not imposed herein.
[0088] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223, etc., and the storage method can be temporary storage or permanent storage.
[0089] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222 to implement the operation and processing of the data 223 in the memory 22 by the processor 21, and it can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the heterogeneous acceleration board computing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks. The data 223 may include not only the data transmitted by external devices received by the heterogeneous acceleration board computing device, but also the data collected by its own input / output interface 25, etc.
[0090] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0091] Further, an embodiment of the present application also discloses a computer-readable storage medium. A computer program is stored in the storage medium. When the computer program is loaded and executed by a processor, the steps of the heterogeneous acceleration board computing method disclosed in any of the foregoing embodiments are implemented.
[0092] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0093] The above has introduced in detail a heterogeneous acceleration board computing method, device, equipment and storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A heterogeneous acceleration board computing method, characterized in that, Applied to the management node, including: Sending device resource query instructions to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instructions; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; Obtaining all the resource pool information, generating a management pool based on all the resource pool information, obtaining a query operation instruction sent by a user, parsing the query operation instruction to obtain an execution tree, optimizing the execution tree based on the management pool to obtain a new execution tree, and generating a query task according to the new execution tree and the query operation instruction; Sending the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information to obtain each query subtask, and sends each query subtask to its corresponding heterogeneous acceleration board; Obtaining data from the local storage resource, sending the data to the local smart network card, preprocessing the data by using the smart network card to obtain preprocessed data, and sending the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result; Sending device resource query instructions to each acceleration node respectively, so that each acceleration node determines the state of its own heterogeneous acceleration board, records the calculation tasks executed by the currently online heterogeneous acceleration boards, and the idle computing resources of the heterogeneous acceleration boards, and generates resource pool information.
2. The heterogeneous acceleration board computing method according to claim 1, wherein The optimizing the execution tree based on the management pool to obtain a new execution tree includes: Determining resource occupancy information from the query operation instruction, and determining the usage status of the acceleration node from the management pool; wherein, the usage status includes an idle status and a non-idle status; Optimizing the execution tree based on the resource occupancy information and the usage status to obtain a new execution tree.
3. The heterogeneous acceleration board computing method according to claim 1, wherein The sending the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and current actual resource information includes: Sending the query task to the acceleration node, so that the acceleration node determines its current actual resource information in the idle state and its current actual resource information in the non-idle state, then estimates the duration of the current actual resource information in the non-idle state to obtain an estimated processing duration, and divides the query task based on the current actual resource information in the idle state and the estimated processing duration.
4. The heterogeneous acceleration board calculation method according to any one of claims 1 to 3, characterized in that, The preprocessing the data by using the smart network card to obtain preprocessed data, and sending the preprocessed data to the acceleration node includes: Determine the data volume of the said data, call the local data filtering array and row-column conversion array, and use the intelligent network card to preprocess the data based on the data volume and the query operation instruction to obtain preprocessed data, and send the preprocessed data to the acceleration node through the local network port module.
5. The heterogeneous acceleration board computing method according to claim 1, characterized in that Before sending the device resource query instruction to each acceleration node respectively, it further includes: Screen out the management node from all host nodes, and use the remaining all host nodes except the management node as acceleration nodes; Establish the connection relationship between the management node and each acceleration node.
6. The heterogeneous acceleration board computing method according to claim 5, wherein Send the device resource query instruction to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instruction; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information, including: Use the connection relationship to send the device resource query instruction to each acceleration node respectively, so that each acceleration node uses the device management module and memory management module in its own intelligent network card, and determines its own resource pool information based on the device resource query instruction; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information.
7. The heterogeneous acceleration board computing method according to claim 5, wherein Send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result, including: Use the connection relationship to send the preprocessed data to the acceleration node, so that the acceleration node uses its own intelligent network card to send the preprocessed data to the corresponding heterogeneous acceleration boards for calculation to obtain a calculation result, and then determine whether the calculation result contains an aggregation operation. If the calculation result contains an aggregation operation, call the aggregation operation acceleration module to process the calculation result to obtain the final calculation result.
8. An heterogeneous acceleration board computing device, characterized in that, It includes: An instruction sending module, configured to send the device resource query instruction to each acceleration node respectively, so that each acceleration node determines its own resource pool information based on the device resource query instruction; wherein, the resource pool information includes heterogeneous acceleration board device information and resource information; A management pool generation module, configured to obtain all the resource pool information, generate a management pool based on all the resource pool information, obtain the query operation instruction sent by the user, parse the query operation instruction to obtain an execution tree, optimize the execution tree based on the management pool to obtain a new execution tree, and generate a query task according to the new execution tree and the query operation instruction; A query task sending module, configured to send the query task to the acceleration node, so that the acceleration node divides the query task based on its own resource pool information and the current actual resource information to obtain each query subtask, and send each query subtask to its corresponding heterogeneous acceleration board; The heterogeneous acceleration board computing module is used to obtain data from local storage resources, send the data to the local smart network card, preprocess the data by using the smart network card to obtain preprocessed data, and send the preprocessed data to the acceleration node, so that the acceleration node sends the preprocessed data to each corresponding heterogeneous acceleration board for calculation to obtain a calculation result; Send device resource query instructions to each acceleration node respectively, so that each acceleration node determines the status of its own heterogeneous acceleration board, records the computing tasks executed by the currently online heterogeneous acceleration boards, and the idle computing resources of the heterogeneous acceleration boards, and generates resource pool information.
9. An electronic device, characterized in that, Comprising: A memory for storing computer programs; A processor for executing the computer program to implement the heterogeneous acceleration board computing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing computer programs; wherein, when the computer program is executed by the processor, it implements the heterogeneous acceleration board computing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and device, distributed data flow programming framework and related components
CN111324558A
Method and device for accelerating database operation
CN113448967A
Method, system and device for managing computing resources
CN115220902A