A training and inference method, product, electronic device and medium for a decision-making model

By dividing the training inference process of the decision model into multiple computing tasks and assigning them to different FPGA cluster nodes, the problem of unreasonable node load allocation in the FPGA cluster is solved, and a more efficient training inference process is achieved.

CN119647604BActive Publication Date: 2025-05-27SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510169233.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-27
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The node load allocation in the FPGA cluster is unreasonable, resulting in data transmission and synchronization delays, affecting the training and inference efficiency of the decision model.

Method used

By obtaining the network characteristics of the decision model, the training inference process is divided into multiple computing tasks, and different computing tasks are assigned to different cluster nodes. At the same time, the input data is split and allocated to different computing nodes to achieve parallel processing.

Benefits of technology

It reduces the load pressure on the node, improves the utilization rate of node resources, shortens training time, and improves model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119647604B_ABST
    Figure CN119647604B_ABST
Patent Text Reader

Abstract

The present invention discloses a training and inference method, product, electronic device and medium for a decision-making model, which relates to the field of artificial intelligence technology. In this method, when the training and inference process of the decision-making model is divided into multiple computing tasks and assigned to different cluster nodes, the load pressure on the nodes is reduced; and multiple data subsets obtained by splitting the input data corresponding to the computing tasks are assigned to different computing nodes, which reduces the load pressure on the computing nodes, improves the utilization rate of node resources, and realizes parallel processing of the data subsets corresponding to the input data, thereby improving the efficiency of training and inference. When the training and inference process of the decision-making model is divided into one computing task, by splitting the input data corresponding to the computing task into multiple data subsets and assigning them to different computing nodes, the load balance on the computing nodes is ensured, the load is reasonably allocated to the nodes in the cluster, and the utilization rate of node resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a training and inference method, product, electronic device and medium for a decision-making model. Background Art

[0002] Intelligent autonomous decision-making technology is an important branch in the field of artificial intelligence, aiming to enable a system to automatically make reasonable decisions in a complex environment. It covers a variety of methods and technologies, including various decision-making models, such as rule-based systems, machine learning algorithms, and reinforcement learning-based methods, etc. In order to implement the training and inference of a decision-making model, in related technologies, the training and inference of the model are carried out on a platform based on a central processing unit (CPU) + graphics processing unit (GPU). However, the number of cores and the structure of traditional CPUs and GPUs are fixed, and the hardware is not reconfigurable, so the structure cannot be dynamically adjusted according to application requirements.

[0003] Considering that a field-programmable gate array (FPGA) has unique reconfigurable characteristics, and a single FPGA board is difficult to support the operation of multiple components in training and inference, therefore, during the training and inference of a decision-making model, an FPGA cluster is constructed to achieve parallel distributed computing, so as to accelerate the model training and inference process, shorten the training time, and improve the model performance. In a decision-making model training and inference system, if the load assigned to the nodes in the FPGA cluster is unreasonable, it may cause data transmission and synchronization delays, resulting in some tasks not being completed in time, or the model training effect being poor.

[0004] Therefore, it can be seen that how to reasonably assign loads to the nodes in the FPGA cluster is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a training and inference method, product, electronic device and medium for a decision-making model, so as to solve the technical problem of unreasonable load distribution of nodes in the FPGA cluster.

[0006] To solve the above technical problem, the present invention provides a training and inference method for a decision-making model, which is applied to a processor of a host computer in a training and inference acceleration system; the accelerator cluster of the training and inference acceleration system is in a cluster structure; the method includes:

[0007] Obtain a decision-making model and obtain the network characteristics of the decision-making model;

[0008] Divide the training and inference process of the decision model into one or more computing tasks according to the network characteristics of the decision model, and allocate different computing tasks to different cluster nodes;

[0009] Allocate multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes to different computing nodes, and transfer the computing tasks on the cluster nodes to the computing nodes to which the input data is allocated;

[0010] Determine the training and inference results of the decision model according to the output characteristics of the network model on the computing nodes to which the input data is allocated.

[0011] On the one hand, dividing the training and inference process of the decision model into multiple computing tasks according to the network characteristics of the decision model, and allocating different computing tasks to different cluster nodes includes:

[0012] Divide the training and inference process of the decision model into one or more tasks according to the training and inference phases included in the decision model, and allocate the tasks to different cluster nodes of the first type;

[0013] Divide the tasks into one or more subtasks according to the network structure used when processing the tasks, and allocate different subtasks to different cluster nodes of the second type;

[0014] According to the complexity of processing the subtasks, regard the subtasks as one computing task or divide them into multiple computing tasks, and allocate different computing tasks to different cluster nodes of the third type.

[0015] On the other hand, regarding the subtasks as one computing task or dividing them into multiple computing tasks according to the complexity of processing the subtasks, and allocating different computing tasks to different cluster nodes of the third type includes:

[0016] When it is detected that the amount of computation for processing the subtasks is less than the pre-designed amount of computation, regard the subtasks as one computing task and allocate the computing task to the cluster nodes of the third type;

[0017] When it is detected that the amount of computation for processing the subtasks is greater than or equal to the pre-designed amount of computation, divide the subtasks into multiple computing tasks; classify the network structures of the network models used when processing the computing tasks according to the complexity of the network structures of the network models used when processing the computing tasks; allocate the computing tasks corresponding to different categories of network structures to different cluster nodes of the third type.

[0018] On the other hand, obtaining the input data corresponding to the computing tasks on the cluster nodes includes:

[0019] If the computing task is the initial computing task among all computing tasks, determine that the input data corresponding to the initial computing task is the data sent by the processor of the host computer received by the cluster nodes of the first type;

[0020] If the computing task is not the initial computing task, determine that the input data corresponding to the computing task is the output features obtained by passing the input data corresponding to the previous computing task of the computing task through the network model on the computing node.

[0021] On the other hand, splitting the input data corresponding to the computing task on the cluster node to obtain multiple data subsets includes:

[0022] In the case where it is detected that the input data corresponding to the computing task on the cluster node is image data, if the data batch size is greater than or equal to the number of input channels, split the input data corresponding to the computing task on the cluster node along the data batch dimension to obtain multiple data subsets; if the data batch size is less than the number of input channels, split the input data corresponding to the computing task on the cluster node along the input channel dimension to obtain multiple data subsets;

[0023] In the case where it is detected that the input data corresponding to the computing task on the cluster node is not image data, obtain the dimension with the largest number of dimensions from the data batch dimension and multiple feature dimensions; split the input data corresponding to the computing task on the cluster node along the dimension with the largest number of dimensions to obtain multiple data subsets.

[0024] On the other hand, in the case where it is detected that the dimension with the largest number of dimensions is a feature dimension, the step of allocating the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different computing nodes includes:

[0025] Obtain the number of computing nodes to which the input data is to be allocated as preset;

[0026] Obtain the result of taking the remainder of the dimension with the largest number of dimensions and the number of computing nodes to which the data is to be allocated to determine the number of the first group of computing nodes;

[0027] Determine the number of the second group of computing nodes according to the difference between the number of computing nodes to which the data is to be allocated and the number of the first group of computing nodes;

[0028] The input data corresponding to the computing tasks on the cluster nodes is split according to the dimension with the largest number of dimensions to determine a data subset with the first data size and a data subset with the second data size; wherein, the first data size is determined by the product of the result obtained by performing a floor operation on the dimension with the largest number of dimensions and the number of computing nodes to be allocated plus 1 and the data size on the remaining dimensions; the second data size is determined by the product of the result obtained by the floor operation and the data size on the remaining dimensions; the remaining dimensions are the dimensions other than the dimension with the largest number of dimensions among all dimensions.

[0029] The data subset with the first data size is respectively allocated to the first number of computing nodes, and the data subset with the second data size is respectively allocated to the second number of computing nodes.

[0030] On the other hand, determining the training inference result of the decision model according to the output features of the network model on the computing nodes to which the input data is allocated includes:

[0031] Obtain the current output features of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node; wherein, the current subtask starts from the initial subtask of the training inference process of the decision model.

[0032] Obtain the next subtask of the current subtask in the first order of precedence of the subtasks in the training inference process of the decision model, and use the next subtask of the current subtask as the new current subtask.

[0033] Transmit the current output features from the third type of cluster node where the computing task corresponding to the current subtask is located to the second type of cluster node where the current subtask is located.

[0034] On the second type of cluster node where the new current subtask is located, transmit the current output features from the second type of cluster node where the current subtask is located to the second type of cluster node where the new current subtask is located, so as to obtain the input data corresponding to the new current subtask, return the current output features of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node, until the current subtask is the last subtask in the training inference process of the decision model, stop returning, and determine the training inference result of the decision model according to the current output features.

[0035] On the other hand, transmitting the current output features from the second type of cluster node where the current subtask is located to the second type of cluster node where the new current subtask is located includes:

[0036] Control the communication between the cluster nodes of the second type where the current subtask is located and the cluster nodes of the second type where the new current subtask is located through the first signal transmission method and obtain communication information; wherein, at least the status monitoring information and the status feedback situation are included in the communication information;

[0037] When it is detected that the communication information is normal, transmit the current output feature to the cluster nodes of the second type where the new current subtask is located through the cluster nodes of the second type where the current subtask is located by the second signal transmission method.

[0038] On the other hand, obtaining the current output feature of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node includes:

[0039] When it is detected that the current subtask has not been divided into the current computing task, obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, and use the sum of the current output features as the current output feature of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node;

[0040] When it is detected that the current subtask is divided into multiple computing tasks, obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node; obtain the next computing task of the current sub-computing task according to the second order in the current subtask of the computing task, and use the next computing task of the current computing task as the new current computing task; transmit the sum of the current output features to the cluster nodes of the third type where the new current computing task is located through the cluster nodes of the third type where the current computing task is located, so as to obtain the input data corresponding to the new current computing task, return to the step of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, until the current computing task is the last computing task in the current subtask, stop returning, and output the sum of the current output features, and use the sum of the current output features as the current output feature of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node.

[0041] On the other hand, obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node includes:

[0042] Obtain the first type of computing node and the second type of computing node used when processing the data subset corresponding to the current computing task; wherein, the network model is set in the first type of computing node, and the network model is not set in the second type of computing node; the second type of computing node is one.

[0043] Obtain the output features of each data subset corresponding to the current computing task on the corresponding first type of computing nodes;

[0044] Send the output features on all the first type of computing nodes to the second type of computing nodes;

[0045] Perform an aggregation operation on all the output features through the second type of computing nodes to obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing nodes.

[0046] On the other hand, obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing nodes includes:

[0047] Obtain the communication order among all the computing nodes processing the current computing task; among all the computing nodes arranged in the communication order, the last computing node is communicatively connected to the first computing node;

[0048] Obtain the data segments of each data subset corresponding to the current computing task after passing through the network model on the corresponding computing nodes and store them in the sending buffer of the computing node for sending data;

[0049] Determine whether there are data segments of all other nodes in the current computing nodes processing the current computing task; where other nodes are the computing nodes other than the current computing node among all the computing nodes processing the current computing task;

[0050] If not, send the data segments on the current computing node from the sending buffer to the receiving buffer of the next computing node of the current computing node in accordance with the communication order; take the next computing node of the current computing node as the new current computing node; perform an aggregation operation on the data in the receiving buffer and the data in the sending buffer of the new current computing node and use it as the data segment in the sending buffer of the new current computing node, and return to the step of determining whether there are data segments of all other nodes in the current computing nodes processing the current computing task;

[0051] If so, the data segments on the current computing node will be sent from the sending buffer to the receiving buffer for receiving data of the next computing node of the current computing node in the communication order; the next computing node of the current computing node will be used as the new current computing node; the data in the receiving buffer of the new current computing node will be overwritten with the data in the sending buffer of the new current computing node and used as the data segments in the sending buffer of the new current computing node, and return to the step of determining whether there are data segments of all other nodes in the current computing node processing the current computing task until it is detected that all computing nodes processing the current computing task have data segments of all other nodes, and then stop returning.

[0052] On the other hand, transmitting the sum of the current output features to the cluster node of the third type where the new current computing task is located through the cluster node of the third type where the current computing task is located includes:

[0053] Controlling the communication between the cluster node of the third type where the current computing task is located and the cluster node of the third type where the new current computing task is located through the first signal transmission method and obtaining communication information; wherein, the communication information at least includes status monitoring information and status feedback situation;

[0054] When it is detected that the communication information is normal, transmitting the sum of the current output features to the cluster node of the third type where the new current computing task is located through the cluster node of the third type where the current computing task is located by the second signal transmission method.

[0055] On the other hand, before obtaining the output features of each data subset corresponding to the current computing task on the corresponding first type of computing node, it further includes:

[0056] If it is detected that the load conditions on the first type of computing node and the second type of computing node used when processing the data subset corresponding to the current computing task are less than the preset load amount, then enter the step of obtaining the output features of each data subset corresponding to the current computing task on the corresponding first type of computing node.

[0057] On the other hand, before obtaining the communication order between all computing nodes processing the current computing task, it further includes:

[0058] Obtaining the load conditions on the first type of computing node and the second type of computing node used when processing the data subset corresponding to the current computing task;

[0059] If it is detected that the load conditions on the first type of computing nodes and the second type of computing nodes used for processing the data subset corresponding to the current computing task are greater than or equal to the preset load amount, then enter the step of obtaining the communication order among all the computing nodes for processing the current computing task.

[0060] On the other hand, the computing nodes include multiple types, and different types of computing nodes are used to process different types of network layer data;

[0061] The step of allocating the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different computing nodes includes:

[0062] Obtain the network layer characteristics of the computing task;

[0063] Determine the target computing nodes according to the network layer characteristics;

[0064] Allocate the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different target computing nodes.

[0065] On the other hand, the first type of cluster nodes are located on the host computer; the second type of cluster nodes to which the subtasks corresponding to the same task are allocated and the third type of cluster nodes to which the same computing task is allocated are located on the same host; the third type of cluster nodes to which different computing tasks are allocated are located on different hosts.

[0066] On the other hand, the step of allocating the task to different first type of cluster nodes includes:

[0067] Allocate the task to different first type of cluster nodes through the third signal transmission method;

[0068] The step of allocating different subtasks to different second type of cluster nodes includes:

[0069] Use the first type of cluster nodes and allocate different subtasks to different second type of cluster nodes through the third signal transmission method;

[0070] The step of allocating different computing tasks to different third type of cluster nodes includes:

[0071] Use the second type of cluster nodes and allocate different computing tasks to different third type of cluster nodes through the third signal transmission method.

[0072] On the other hand, the training and inference of the decision model include an interaction stage and a model update stage; before sending data to the first type of cluster nodes, it further includes:

[0073] Obtain the data sent to the first type of cluster node from the process of the graphics processing unit for training and inference;

[0074] The determining the training and inference result of the decision model according to the current output feature includes:

[0075] Receive, in the interaction stage, the first current output feature that sequentially passes through the third type of cluster node, the second type of cluster node, and the first type of cluster node;

[0076] Store the first current output feature in the data storage module located in the host computer;

[0077] In the case where it is detected that the amount of data in the data storage module is greater than the preset amount of data, receive, in the model update stage, the second current output feature that sequentially passes through the third type of cluster node, the second type of cluster node, and the first type of cluster node;

[0078] Calculate the loss function based on the second current output feature to determine the parameters of the decision model, so as to obtain the training and inference result of the decision model;

[0079] After obtaining the training and inference result of the decision model, it further includes:

[0080] Transmit the parameters of the decision model to the first type of cluster node to update the parameters of the decision model located in the first type of cluster node.

[0081] On the other hand, establishing the decision model located on the first type of cluster node includes:

[0082] Obtain the complexity of the decision task;

[0083] Establish decision models with different structures on the first type of cluster node according to the complexity of the decision task; wherein, the complexity of the decision task is positively correlated with the complexity of the decision model structure.

[0084] On the other hand, determining the structure of the decision model includes:

[0085] Obtain the operating state of the hardware platform; wherein, the operating state of the hardware platform at least includes the hardware resource utilization rate, latency, and power consumption;

[0086] Obtain the current expected cumulative reward value under the current state and the current decision model structure; wherein, the current state at least includes the current input data feature and the current operating state of the hardware platform;

[0087] Determine the expected value according to the current expected cumulative reward and the current reward function; wherein, the current reward function is determined by the current operating state of the hardware platform;

[0088] Obtain the current decision model structure whose expected value meets the requirements;

[0089] Use the current decision model structure whose expected value meets the requirements as the structure of the decision model.

[0090] On the other hand, determining the structure of the decision model includes:

[0091] Structures of multiple pre-set decision models;

[0092] Obtain the hardware performance indicators under decision models with different structures, and establish hardware constraint conditions according to the hardware performance indicators; wherein, the hardware performance indicators at least include delay, power consumption, and resource occupancy;

[0093] Obtain the fitness of decision models with different structures;

[0094] Under the condition of meeting the hardware constraint conditions, obtain the decision model structure with the maximum fitness among all decision models with different structures;

[0095] Use the decision model structure with the maximum fitness as the structure of the decision model.

[0096] On the other hand, determining the structure of the decision model includes:

[0097] Structures of multiple pre-set decision models and extracting features of input data;

[0098] Predict the selection probabilities of decision models with different structures according to the features of the input data;

[0099] Obtain the structure of the decision model with the maximum probability among all decision models with different structures;

[0100] Use the structure of the decision model with the maximum probability as the structure of the decision model.

[0101] To solve the above technical problems, the present invention also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above decision model training and inference method are implemented.

[0102] To solve the above technical problems, the present invention also provides an electronic device, including:

[0103] A memory for storing computer programs;

[0104] A processor for implementing the steps of the above decision model training and inference method when executing the computer program.

[0105] To solve the above technical problems, the present invention also provides a non-volatile storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned training and inference method of the decision model are implemented.

[0106] The beneficial effects of the present invention are as follows. During the training and inference process of the decision model, the training and inference process of the decision model is divided into one or more computing tasks according to the network characteristics of the decision model. Among them, for the case where the training and inference process of the decision model is divided into multiple computing tasks according to the network characteristics of the decision model, in the present invention, different computing tasks are assigned to different cluster nodes. Compared with the method of directly assigning the entire training and inference process of the decision model to nodes without division, the method provided by the present invention reduces the load pressure on the nodes; and multiple data subsets obtained by splitting the input data corresponding to the computing tasks are assigned to different computing nodes, which reduces the load pressure on the computing nodes and improves the utilization rate of node resources compared with the method of not splitting the input data; when assigning computing nodes to the computing tasks, the computing tasks are assigned to the multiple computing nodes to which the input data is assigned, realizing parallel processing of the data subsets corresponding to the input data and improving the training and inference efficiency. For the case where the training and inference process of the decision model is divided into one computing task according to the network characteristics of the decision model, by splitting the input data corresponding to the computing task into multiple data subsets and assigning them to different computing nodes, the load balance on the computing nodes is ensured, the load is reasonably assigned to the nodes in the cluster, and the utilization rate of node resources is improved.

[0107] In addition, when the training and inference process of the decision model is divided into multiple computing tasks according to the network characteristics of the decision model and different computing tasks are assigned to different cluster nodes, the training and inference process of the decision model is divided into one or more tasks, and then the tasks are divided into one or more subtasks according to the network structure used when processing the tasks; and further, according to the complexity of processing the subtasks, the subtasks are used as one computing task or divided into multiple computing tasks, that is, the training and inference process of the decision model is divided into smaller granularities and handed over to the cluster nodes for processing. And the tasks, subtasks, and computing tasks obtained by dividing the training and inference of the decision model are assigned to different types of cluster nodes, so that the data traffic of each cluster node is about the same in scale, ensuring the load balance of each node, and making full use of node resources and improving the utilization rate of node resources.

[0108] When the amount of computation in detecting a processing subtask is greater than or equal to a pre-designed amount of computation, the subtask is divided into multiple computing tasks; the network structure of the network model used in processing the computing tasks is classified according to the complexity of the network structure of the network model used in processing the computing tasks; the computing tasks corresponding to different types of network structures are allocated to different cluster nodes of the third type, which as much as possible ensures the balance of the network parameters allocated on the cluster nodes of the third type.

[0109] In the process of splitting the input data corresponding to the computing task on the cluster node into multiple data subsets, different data splitting methods are adopted according to the type of the input data corresponding to the computing task (such as image data or non-image data), realizing the reasonable splitting of the input data; and when it is detected that the input data corresponding to the computing task on the cluster node is not image data, the input data corresponding to the computing task on the cluster node is split along the dimension with the largest number of dimensions to obtain multiple data subsets, improving the parallelism of data processing and the resource utilization rate.

[0110] Obtain the result of taking the remainder of the dimension with the largest number of dimensions and the number of computing nodes to which the data is to be allocated to determine the number of the first computing nodes; determine the number of the second computing nodes according to the difference between the number of computing nodes to which the data is to be allocated and the number of the first computing nodes; split the input data corresponding to the computing task on the cluster node according to the dimension with the largest number of dimensions to determine the data subsets with the first data size and the second data size; allocate the data subsets with the first data size to the number of the first computing nodes respectively, and allocate the data subsets with the second data size to the number of the second computing nodes respectively. Through this method, a way of allocating data subsets to computing nodes is provided in the scenario where the number of computing nodes set cannot be evenly divided by the feature dimension, which as much as possible ensures the reasonable allocation of input data to the computing nodes.

[0111] In the process of determining the training and inference result of the decision model according to the output features of the network model on the computing nodes to which the input data is allocated, the output features of the previous subtask are transmitted to the next subtask in the first order of precedence of the subtask in the training and inference process of the decision model, realizing the training and inference of the decision model. And in the process of transmitting the current output feature from the cluster node of the second type where the current subtask is located to the cluster node of the second type where the new current subtask is located, the communication information between the cluster node of the second type where the current subtask is located and the cluster node of the second type where the new current subtask is located is detected by the first signal transmission method, and only when the communication information is detected to be normal, the current output feature is transmitted from the cluster node of the second type where the current subtask is located to the cluster node of the second type where the new current subtask is located by the second signal transmission method, which as much as possible ensures the stable operation of the whole system and the accurate transmission of data.

[0112] In the process of obtaining the current output features after the input data of the computing task corresponding to the current subtask passes through the network model on the computing node, in the case where the current subtask is not divided to obtain the current computing task, the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node is directly used as the current output features after the input data of the computing task corresponding to the current subtask passes through the network model on the computing node. In the case where the current subtask is divided into multiple computing tasks, the sum of the current output features is transmitted to the cluster node of the third type where the new current computing task is located through the cluster node of the third type where the current computing task is located, so as to obtain the input data corresponding to the new current computing task. Further, the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node is obtained, realizing the aggregation of the results of all computing tasks of the same subtask on the computing node, that is, the current output features after the input data of the computing task corresponding to the current subtask passes through the network model on the computing node are obtained.

[0113] In the process of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, different types of computing nodes are set, and the output features on all the first-type computing nodes are sent to the second-type computing node; the second-type computing node performs an aggregation operation on all the output features to obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node. That is, centralized data parallelization is realized based on the structure of the central node - multiple sub-nodes.

[0114] In the process of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, the communication order between all the computing nodes processing the current computing task is obtained, and the last computing node among all the computing nodes arranged in the communication order is communicatively connected to the first computing node. Each computing node sends its own data to the next computing node and receives data from the previous computing node. As the communication progresses, each computing node gradually integrates the data segments from other nodes, and finally all nodes have the completely synchronized data, that is, the data integration of each computing node is completed based on the decentralized node ring communication method.

[0115] In the process of transmitting the sum of the current output features to the cluster node of the third type where the new current computing task is located through the cluster node of the third type where the current computing task is located, first monitor the communication status between the cluster node of the third type where the current computing task is located and the cluster node of the third type where the new current computing task is located. Only when the communication information is normal, the data transmission is carried out, ensuring the stable operation of the entire system and the accurate transmission of data as much as possible.

[0116] Set different computing nodes for processing different types of network layer data. Obtain the network layer characteristics of the computing task; determine the target computing node according to the network layer characteristics; allocate multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different target computing nodes, which improves the efficiency of training and inference.

[0117] Locate the first type of cluster nodes on the host computer; locate the second type of cluster nodes assigned to subtasks corresponding to the same task and the third type of cluster nodes assigned to the same computing task on the same host; locate the third type of cluster nodes assigned to different computing tasks on different hosts. Through different hosts, parallel processing can be achieved, which improves the efficiency of training and inference of the decision model.

[0118] Use different decision models for decision tasks with different complexities, so that the selected decision model is more appropriate; when determining the structure of the decision model, different methods can be used to select the structure of the decision model, which realizes the reasonable setting of the structure of the decision model and enables the structure of the decision model to be adaptively selected, improving the flexibility of the structure of the decision model.

[0119] In addition, the present invention also provides a computer product, an electronic device, and a non-volatile storage medium, which have the same or corresponding technical features as the training and inference method of the above-mentioned decision model, and the effects are the same. BRIEF DESCRIPTION OF THE DRAWINGS

[0120] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0121] Figure 1 Schematic diagram of a training and inference acceleration system provided by an embodiment of the present invention;

[0122] Figure 2 Flowchart of a training and inference method of a decision model provided by an embodiment of the present invention;

[0123] Figure 3 Schematic diagram of a decision model training and inference acceleration system based on an FPGA cluster provided by an embodiment of the present invention;

[0124] Figure 4 Schematic diagram of a heterogeneous computing framework based on an FPGA cluster provided by an embodiment of the present invention;

[0125] Figure 5Schematic diagram of a decision model structure for complex decision-making tasks provided by an embodiment of the present invention;

[0126] Figure 6 Schematic diagram of a multi-child node - central node data parallelization structure provided by an embodiment of the present invention;

[0127] Figure 7 Schematic diagram of a model update process based on multi-child node - central node data parallelization provided by an embodiment of the present invention;

[0128] Figure 8 Schematic diagram of a ring communication topology structure for decentralized node data parallelization provided by an embodiment of the present invention;

[0129] Figure 9 Schematic diagram of a multi - Actor parallel interaction structure based on FPGA provided by an embodiment of the present invention;

[0130] Figure 10 Structural diagram of an FPGA single computing node for convolutional computing provided by an embodiment of the present invention;

[0131] Figure 11 Structural diagram of an FPGA single computing node for attention computing provided by an embodiment of the present invention;

[0132] Figure 12 Structural diagram of an FPGA single computing node for convolutional computing provided by an embodiment of the present invention;

[0133] Figure 13 Schematic diagram of the structure of a multi - level cluster of accelerators provided by an embodiment of the present invention;

[0134] Figure 14 Logical schematic diagram of an action cluster node provided by an embodiment of the present invention;

[0135] Figure 15 Logical schematic diagram of a trainer cluster node provided by an embodiment of the present invention;

[0136] Figure 16 Logical structure diagram of a state encoding task cluster for the action part provided by an embodiment of the present invention;

[0137] Figure 17 Logical structure diagram of a state encoding task cluster for the trainer part provided by an embodiment of the present invention;

[0138] Figure 18 Logical schematic diagram of a module cluster for convolutional computing in the action part provided by an embodiment of the present invention;

[0139] Figure 19Logical schematic diagram of the module cluster for convolutional computing in the trainer part provided by the embodiment of the present invention;

[0140] Figure 20 Logical schematic diagram of the module cluster for attention computing in the action part provided by the embodiment of the present invention;

[0141] Figure 21 Logical schematic diagram of the module cluster for attention computing in the trainer part provided by the embodiment of the present invention;

[0142] Figure 22 Logical schematic diagram of the module cluster for point cloud feature encoding in the action part provided by the embodiment of the present invention;

[0143] Figure 23 Logical schematic diagram of the module cluster for point cloud feature encoding in the trainer part provided by the embodiment of the present invention;

[0144] Figure 24 Logical schematic diagram of a variational autoencoder decoder generation model provided by the embodiment of the present invention;

[0145] Figure 25 Logical schematic diagram of a generative adversarial network generation model provided by the embodiment of the present invention;

[0146] Figure 26 Logical schematic diagram of a diffusion generation model provided by the embodiment of the present invention;

[0147] Figure 27 Structural diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0148] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0149] The core of the present invention is to provide a training and inference method, product, electronic device and medium for a decision model to solve the technical problem of unreasonable node load distribution in the FPGA cluster.

[0150] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. Figure 1 Schematic diagram of a training and inference acceleration system provided by the embodiment of the present invention, as Figure 1As shown in the figure, the training and inference acceleration system includes: a host computer and an accelerator cluster. The accelerator cluster of the training and inference acceleration system is a cluster structure. Figure 2 The flowchart of a method for training and inferring a decision model provided by an embodiment of the present invention is shown. The method includes:

[0151] S10: Obtain a decision model and obtain the network features of the decision model;

[0152] S11: Divide the training and inference process of the decision model into one or more computing tasks according to the network features of the decision model, and allocate different computing tasks to different cluster nodes;

[0153] S12: Allocate multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes to different computing nodes, and transfer the computing tasks on the cluster nodes to the computing nodes to which the input data is allocated;

[0154] S13: Determine the training and inference result of the decision model according to the output features of the network model on the computing nodes to which the input data is allocated.

[0155] There is no limitation on the set decision model, which is determined according to the actual situation. In order to make the set decision model more reasonable, the structure of the decision model can be established by considering the complexity of the decision task. As the task complexity and hardware heterogeneity increase, the design method of the coding module based on predefined rules will lack flexibility in a dynamic environment. In order to improve the flexibility of the decision model structure, in implementation, the ways to determine the structure of the decision model include at least one of the following ways:

[0156] Way 1: Obtain the operating state of the hardware platform; wherein, the operating state of the hardware platform at least includes the hardware resource utilization rate, delay, and power consumption;

[0157] Obtain the current expected cumulative reward value under the current state and the current decision model structure; wherein, the current state at least includes the current input data features and the current operating state of the hardware platform;

[0158] Determine the expected value according to the current expected cumulative reward and the current reward function; wherein, the current reward function is determined by the current operating state of the hardware platform;

[0159] Obtain the current decision model structure whose expected value meets the requirements;

[0160] Take the current decision model structure whose expected value meets the requirements as the structure of the decision model.

[0161] Way 2: Structures of multiple pre-set decision models;

[0162] Obtain the hardware performance metrics under decision models with different structures, and establish hardware constraint conditions based on the hardware performance metrics; wherein, the hardware performance metrics at least include latency, power consumption, and resource occupancy;

[0163] Obtain the fitness of decision models with different structures;

[0164] Under the condition of meeting the hardware constraint conditions, obtain the decision model structure with the maximum fitness among all decision models with different structures;

[0165] Take the decision model structure with the maximum fitness as the structure of the decision model.

[0166] Method 3: Structures of multiple pre-set decision models and extract features of input data;

[0167] Predict the selection probabilities of decision models with different structures according to the features of the input data;

[0168] Obtain the structure of the decision model with the maximum probability among all decision models with different structures;

[0169] Take the structure of the decision model with the maximum probability as the structure of the decision model.

[0170] That is, the above method for determining the structure of the decision model realizes the adaptive selection of the model structure.

[0171] After determining the decision model, obtain the network features of the decision model. The network features of the decision model include the training and inference stages included in the decision model, the network structures used when processing each training and inference stage, and the complexity when implementing each layer of the network structure. In order to implement the training and inference of the decision model, an accelerator cluster is established in the embodiments of the present invention. The accelerator cluster includes a first type of cluster node, a second type of cluster node, and a third type of cluster node that undertake task allocation, control, and status feedback tasks, as well as computing nodes that undertake the implementation of computing tasks. These nodes are all accelerators. In order to improve the communication efficiency between the processor of the host computer and the first type of cluster node, in implementation, the first type of cluster node can be set in the host computer; in order to implement parallel processing, the second type of cluster node, the third type of cluster node, and the computing nodes that undertake the implementation of computing tasks can be set in the host connected to the host computer.

[0172] In implementation, dividing the training and inference process of the decision model into multiple computing tasks according to the network features of the decision model, and allocating different computing tasks to different cluster nodes includes:

[0173] Divide the training and inference process of the decision model into one or more tasks according to the training and inference stages included in the decision model, and allocate the tasks to different first type of cluster nodes;

[0174] Divide the task into one or more subtasks according to the network structure used when processing the task, and allocate different subtasks to different cluster nodes of the second type;

[0175] According to the complexity of processing the subtask, regard the subtask as a computing task or divide it into multiple computing tasks, and allocate different computing tasks to different cluster nodes of the third type.

[0176] Specifically, regarding the subtask as a computing task or dividing it into multiple computing tasks according to the complexity of processing the subtask, and allocating different computing tasks to different cluster nodes of the third type includes:

[0177] When it is detected that the amount of computation for processing the subtask is less than the pre-designed amount of computation, regard the subtask as a computing task and allocate the computing task to a cluster node of the third type;

[0178] When it is detected that the amount of computation for processing the subtask is greater than or equal to the pre-designed amount of computation, divide the subtask into multiple computing tasks; classify the network structure of the network model used when processing the computing task according to the complexity of the network structure of the network model used when processing the computing task; allocate the computing tasks corresponding to different types of network structures to different cluster nodes of the third type.

[0179] There is no limit to the pre-designed amount of computation, which is determined according to the actual situation. Allocating the task to different cluster nodes of the first type includes:

[0180] Allocate the task to different cluster nodes of the first type through the third signal transmission method;

[0181] Allocating different subtasks to different cluster nodes of the second type includes:

[0182] Use the cluster nodes of the first type and allocate different subtasks to different cluster nodes of the second type through the third signal transmission method;

[0183] Allocating different computing tasks to different cluster nodes of the third type includes:

[0184] Use the cluster nodes of the second type and allocate different computing tasks to different cluster nodes of the third type through the third signal transmission method.

[0185] The third signal transmission method is, for example, the Ethernet transmission method. Taking the reinforcement learning model as an example, the training and inference stages of the reinforcement learning model include an environment interaction stage and two processes of model update. According to the environment interaction stage, the training and inference process of the decision-making model is divided into one or more action (Actor) tasks; according to the model update, the training and inference process of the decision-making model is divided into one or more trainer (Learner) tasks. The Actor tasks are assigned to the cluster nodes of the first type, which can be called Actor cluster nodes, and the Learner tasks are also assigned to the cluster nodes of the first type, which can be called Learner cluster nodes. It should be noted that when there are multiple tasks, different tasks can be assigned to the cluster nodes of the first type (including Actor cluster nodes or Learner cluster nodes) located on different hosts. A decision-making model is set on the cluster nodes of the first type. Establishing the decision-making model on the cluster nodes of the first type includes: obtaining the complexity of the decision-making task; establishing decision-making models with different structures on the cluster nodes of the first type according to the complexity of the decision-making task; where the complexity of the decision-making task is positively correlated with the complexity of the decision-making model structure.

[0186] After the tasks are assigned to the cluster nodes of the first type, the complexity of some tasks is still relatively high. Therefore, further, the tasks are divided into one or more subtasks according to the network structure used when processing the tasks, and different subtasks are assigned to different cluster nodes of the second type. For example, if the network structure included in the Actor task is: a state encoding layer, a context encoding layer, a decision encoding layer, and an action encoding layer, the Actor task can be divided into a state encoding subtask, a context encoding subtask, a decision encoding subtask, and an action encoding subtask. Each subtask is assigned to the cluster nodes of the second type (also called task cluster nodes). The subtasks corresponding to the same task can be set on the same host. After the subtasks are assigned to the cluster nodes of the second type, there may still be subtasks with relatively high complexity. Therefore, further, according to the complexity of processing the subtasks, the subtasks are regarded as a computing task or divided into multiple computing tasks, and different computing tasks are assigned to different cluster nodes of the third type. For example, if the loads of the state encoding subtask and the context encoding subtask are relatively large, the state encoding subtask and the context encoding subtask can be further divided into multiple computing tasks, and then the divided computing subtasks are assigned to different cluster nodes of the third type (which can be called module cluster nodes). For example, the load of the action encoding task is relatively small. Therefore, the action encoding task can be directly regarded as a computing task without division, and then the computing task is assigned to the cluster nodes of the third type. The method of assigning tasks, subtasks, and computing tasks to cluster nodes described above can be called the process of model parallelization.

[0187] After allocating computing nodes to cluster nodes, in order to process computing tasks and improve the efficiency of training and inference, further, parallelize the data. First, it is necessary to obtain the input data corresponding to the computing nodes. In implementation, obtaining the input data corresponding to the computing tasks on the cluster nodes includes:

[0188] If the computing task is the initial computing task among all computing tasks, determine that the input data corresponding to the initial computing task is the data sent by the processor of the host computer received by the first type of cluster node;

[0189] If the computing task is not the initial computing task, determine that the input data corresponding to the computing task is the output feature obtained by passing the input data corresponding to the previous computing task of the computing task through the network model on the computing node.

[0190] When the training and inference of the decision model include an interaction phase and a model update phase; before sending data to the first type of cluster node, the data sent to the first type of cluster node can be obtained from the process for training and inference of the graphics processing unit.

[0191] After obtaining the input data corresponding to the computing task, splitting the input data corresponding to the computing task on the cluster node to obtain multiple data subsets includes:

[0192] In the case where it is detected that the input data corresponding to the computing task on the cluster node is image data, if the data batch size is greater than or equal to the number of input channels, split the input data corresponding to the computing task on the cluster node along the data batch dimension to obtain multiple data subsets; if the data batch size is less than the number of input channels, split the input data corresponding to the computing task on the cluster node along the input channel dimension to obtain multiple data subsets;

[0193] In the case where it is detected that the input data corresponding to the computing task on the cluster node is not image data, obtain the dimension with the largest number of dimensions from the data batch dimension and multiple feature dimensions; split the input data corresponding to the computing task on the cluster node along the dimension with the largest number of dimensions to obtain multiple data subsets.

[0194] In practice, there may be a scenario where the number of feature dimensions cannot be evenly divided by the number of set computing nodes. In this scenario, in order to reasonably allocate data to the computing nodes, allocating the multiple data subsets obtained after splitting the input data corresponding to the computing task on the cluster node to different computing nodes includes:

[0195] Obtain the number of computing nodes to which the input data is to be allocated as preset;

[0196] Obtain the result of taking the remainder of the dimension with the largest number of dimensions and the number of computing nodes to which it is to be allocated to determine the first number of computing nodes;

[0197] Determine the number of computing nodes of the second quantity according to the difference between the number of computing nodes to be allocated and the first quantity;

[0198] Split the input data corresponding to the computing task on the cluster node according to the dimension with the largest number of dimensions, and then determine the data subset with the first data size and the data subset with the second data size; wherein, the first data size is determined by the product of the result obtained by rounding down the dimension with the largest number of dimensions and the number of computing nodes to be allocated plus 1 and the data size on the remaining dimensions; the second data size is determined by the product of the result obtained by the rounding-down operation and the data size on the remaining dimensions; the remaining dimensions are the dimensions other than the dimension with the largest number of dimensions among all dimensions;

[0199] Allocate the data subset with the first data size to the first quantity of computing nodes respectively, and allocate the data subset with the second data size to the second quantity of computing nodes respectively.

[0200] To improve the efficiency of input data processing, allocating the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different computing nodes includes:

[0201] Obtain the network layer features of the computing task;

[0202] Determine the target computing node according to the network layer features;

[0203] Allocate the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different target computing nodes.

[0204] After dividing the input data into multiple data subsets, in order to implement the processing of the data subsets, further transfer the computing tasks on the cluster node to the computing nodes to which the input data is allocated. Finally, determine the training and inference results of the decision model according to the output features of the network model on the computing nodes to which the input data is allocated.

[0205] In implementation, determining the training and inference results of the decision model according to the output features of the network model on the computing nodes to which the input data is allocated includes:

[0206] Obtain the current output features of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node; wherein, the current subtask starts from the initial subtask of the training and inference process of the decision model;

[0207] Obtain the next subtask of the current subtask according to the first sequence in the training and inference process of the decision model for the subtasks, and use the next subtask of the current subtask as the new current subtask;

[0208] Transfer the current output feature to the cluster node of the second type where the current subtask is located through the cluster node of the third type where the computing task corresponding to the current subtask is located;

[0209] Obtain on the cluster node of the second type where the new current subtask is located, transfer the current output feature to the cluster node of the second type where the new current subtask is located through the cluster node of the second type where the current subtask is located, so as to obtain the input data corresponding to the new current subtask, return the current output feature after the input data corresponding to the current subtask passes through the network model on the computing node, until the current subtask is the last subtask in the training and inference process of the decision model, stop returning, and determine the training and inference result of the decision model according to the current output feature.

[0210] Transferring the current output feature to the cluster node of the second type where the new current subtask is located through the cluster node of the second type where the current subtask is located includes:

[0211] Control the communication between the cluster node of the second type where the current subtask is located and the cluster node of the second type where the new current subtask is located through the first signal transmission method and obtain the communication information; wherein, the communication information at least includes status monitoring information and status feedback situation;

[0212] When it is detected that the communication information is normal, transfer the current output feature to the cluster node of the second type where the new current subtask is located through the cluster node of the second type where the current subtask is located by the second signal transmission method.

[0213] The first signal transmission method is, for example, the Low-Voltage Differential Signaling (LVDS) interface transmission method, and the second signal transmission method is, for example, the Peripheral Component Interconnect Express (PCIe) transmission method.

[0214] Obtaining the current output feature after the input data corresponding to the current subtask passes through the network model on the computing node includes:

[0215] When it is detected that the current subtask has not been divided into the current computing task, obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, and use the sum of the current output features as the current output feature after the input data corresponding to the current subtask passes through the network model on the computing node;

[0216] In the case where multiple computing tasks are obtained after the current subtask is partitioned, obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node; obtain the next computing task of the current sub-computing task according to the second sequence of the computing tasks in the current subtask, and use the next computing task of the current computing task as the new current computing task; transmit the sum of the current output features to the third-type cluster node where the new current computing task is located through the third-type cluster node where the current computing task is located, so as to obtain the input data corresponding to the new current computing task, and return to the step of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, until the current computing task is the last computing task in the current subtask, stop returning, and output the sum of the current output features, and use the sum of the current output features as the current output feature after passing through the network model on the computing node of the input data corresponding to the computing task of the current subtask.

[0217] Specifically, obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node includes:

[0218] Obtain the first-type computing node and the second-type computing node used when processing the data subset corresponding to the current computing task; among them, a network model is set in the first-type computing node, and a network model is not set in the second-type computing node; the second-type computing node is one.

[0219] Obtain the output features of each data subset corresponding to the current computing task on the corresponding first-type computing node.

[0220] Send the output features on all the first-type computing nodes to the second-type computing node.

[0221] Perform an aggregation operation on all the output features through the second-type computing node to obtain the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node.

[0222] The above method of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node can be called a model update method based on multi-subnode - central node data parallelization. Before using this method, the load condition on the computing node can be judged. Specifically, before obtaining the output features of each data subset corresponding to the current computing task on the corresponding first-type computing node, it further includes:

[0223] If it is detected that the load conditions on the first type of computing nodes and the second type of computing nodes used for processing the data subset corresponding to the current computing task are less than the preset load amount, then enter the step of obtaining the output features of each data subset corresponding to the current computing task on the corresponding first type of computing nodes.

[0224] In this method, by judging the load conditions on the first type of computing nodes and the second type of computing nodes used for processing the data subset corresponding to the current computing task, and only when the load conditions are less than the preset load amount, the model update method based on multi-subnode - central node data parallelization is used, avoiding excessive load pressure on the computing nodes.

[0225] In addition, the sum of the current output features of the data subsets corresponding to the current computing task on the corresponding computing nodes can also be achieved in the following way:

[0226] Obtain the communication order among all the computing nodes for processing the current computing task; among them, the last computing node in all the computing nodes arranged according to the communication order is communicatively connected to the first computing node;

[0227] Obtain the data segments of each data subset corresponding to the current computing task after passing through the network model on the corresponding computing nodes and store them in the sending buffer of the computing node for sending data;

[0228] Judge whether there are data segments of all other nodes in the current computing node for processing the current computing task; among them, other nodes are the computing nodes except the current computing node among all the computing nodes for processing the current computing task;

[0229] If not, then send the data segment on the current computing node from the sending buffer to the receiving buffer of the next computing node of the current computing node according to the communication order; take the next computing node of the current computing node as the new current computing node; perform an aggregation operation on the data in the receiving buffer and the data in the sending buffer of the new current computing node and use it as the data segment in the sending buffer of the new current computing node, and return to the step of judging whether there are data segments of all other nodes in the current computing node for processing the current computing task;

[0230] If so, the data segments on the current computing node will be sent from the sending buffer to the receiving buffer for receiving data of the next computing node of the current computing node in the communication order; the next computing node of the current computing node will be used as the new current computing node; the data in the receiving buffer of the new current computing node will overwrite the data in the sending buffer of the new current computing node and be used as the data segments in the sending buffer of the new current computing node, and return to the step of determining whether there are data segments of all other nodes in the current computing node processing the current computing task until it is detected that all computing nodes processing the current computing task have data segments of all other nodes, and then stop and return.

[0231] Transmitting the sum of the current output features through the cluster nodes of the third type where the current computing task is located to the cluster nodes of the third type where the new current computing task is located includes:

[0232] Controlling the communication between the cluster nodes of the third type where the current computing task is located and the cluster nodes of the third type where the new current computing task is located through the first signal transmission method and obtaining communication information; wherein, the communication information at least includes status monitoring information and status feedback situation;

[0233] When it is detected that the communication information is normal, the sum of the current output features will be transmitted through the cluster nodes of the third type where the current computing task is located to the cluster nodes of the third type where the new current computing task is located through the second signal transmission method.

[0234] The method of the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node can be called the ring communication method of decentralized node data parallelization.

[0235] Before obtaining the communication order among all computing nodes processing the current computing task, it further includes:

[0236] Obtaining the load conditions on the first type of computing nodes and the second type of computing nodes used when processing the data subset corresponding to the current computing task;

[0237] If it is detected that the load conditions on the first type of computing nodes and the second type of computing nodes used when processing the data subset corresponding to the current computing task are greater than or equal to the preset load amount, then enter the step of obtaining the communication order among all computing nodes processing the current computing task.

[0238] The preset load amount is not limited and is determined according to the actual situation. Before using the ring communication method for decentralized node data parallelization, first judge the load conditions on the first type of computing node and the second type of computing node. If there is a computing node with a load greater than the preset load amount, use the ring communication method for decentralized node data parallelization, ensuring load balancing on the computing nodes.

[0239] When the decision-making model is a reinforcement learning model, determining the training and inference results of the decision-making model according to the current output features includes:

[0240] In the interaction stage, receive the first current output features sequentially output by the third type of cluster node, the second type of cluster node, and the first type of cluster node;

[0241] Store the first current output features in the data storage module located in the host computer;

[0242] In the case where it is detected that the data volume in the data storage module is greater than the preset data volume, in the model update stage, receive the second current output features sequentially output by the third type of cluster node, the second type of cluster node, and the first type of cluster node;

[0243] Calculate the loss function based on the second current output features to determine the parameters of the decision-making model, so as to obtain the training and inference results of the decision-making model;

[0244] After obtaining the training and inference results of the decision-making model, it further includes:

[0245] Transmit the parameters of the decision-making model to the first type of cluster node to update the parameters of the decision-making model located in the first type of cluster node.

[0246] To enable those skilled in the art to better understand the training and inference method of the decision-making model described above, the above method will be described in detail below by taking the decision-making model as a reinforcement learning model as an example.

[0247] Figure 3 The following is a schematic diagram of an acceleration system for training and inference of a decision-making model based on an FPGA cluster provided by an embodiment of the present invention. As Figure 3 shown, it mainly includes the following parts:

[0248] 1. Design the network structure of the decision-making model. According to the task and space complexity, divide the decision-making model into a state encoding layer, a decision encoding layer, and an action encoding layer, select the network structure according to the task requirements, and determine whether to adopt an adaptive selection strategy for model structure design to optimize the performance.

[0249] Second, design a distributed parallel training framework for the decision-making model. Using the data parallelization method, the input data is divided into multiple subsets, and a parallel computing structure is composed of one (or more) central computing nodes and multiple sub-computing nodes. At the same time, the model parallelization method is used to set up multiple data collection and model update processes, and the decision-making model calculation tasks are gradually split to different nodes to improve device utilization and throughput.

[0250] Third, design the hardware structure of a single FPGA computing node, including different node architectures for convolutional computing, attention computing, and fully connected computing to meet the data computing requirements of different network layers; configure multiple FPGA unit node interconnection and communication methods to meet the computing requirements of heterogeneous platforms.

[0251] Fourth, design an FPGA cluster with a multi-level cluster structure. Use multi-level FPGA cluster nodes to build a computing framework. Each level of cluster node includes a master node and subordinate sub-cluster nodes. Different communication methods are adopted between nodes according to the data type. The master node of the bottom-level cluster node is connected to the FPGA unit node to ensure load balancing of each node.

[0252] Fifth, design a heterogeneous computing framework based on the FPGA cluster to achieve stable interconnection between the GPU environment process module, CPU main controller, data storage module, and FPGA cluster, and support efficient training and inference of the decision-making model.

[0253] Next, each part of the decision-making model training and inference acceleration system based on the FPGA cluster will be described in detail.

[0254] First, the part of the heterogeneous computing framework based on the FPGA cluster will be described.

[0255] The decision-making model based on reinforcement learning adopts an interaction-update iterative training mechanism, and two modules are set up to implement the computing tasks in different stages. The Actor module is responsible for the interaction between the decision-making model and the environment, only involving forward inference computing. The Learner module is responsible for updating the decision-making model parameters, including forward inference and backward gradient update. Due to the differences in the operations of the two for the decision-making model, each will maintain its own local decision-making model.

[0256] Figure 4 It is a schematic diagram of a heterogeneous computing framework based on the FPGA cluster provided by the embodiment of the present invention, as Figure 4As shown in the figure, the decision model training and inference computing framework based on the FPGA cluster includes an environment process module created based on the GPU, a CPU main controller for controlling the FPGA cluster and the environment process, a data storage module for storing interaction experience samples and model parameters, and an FPGA cluster for implementing the Actor and Learner tasks. Among them, in order to ensure the load balance of each FPGA computing node, the present invention adopts a multi-level cluster structure to design the FPGA cluster interconnection scheme, which can effectively avoid computing latency and enhance data consistency. The CPU main controller creates an environment process through the Application Programming Interface (API) of the environment simulator, selects and calls the GPU according to the running requirements of the simulator, and at the same time realizes the feedback of data such as actions, states, and rewards through API communication. Here, in order to improve the interaction efficiency, an environment process pipelining parallelization design is adopted, that is, multiple simulation environments are created in each environment process, and the action execution and action inference are alternately performed in two groups.

[0257] During the interaction, the CPU main controller sends the interaction experience, including data such as actions, states, and rewards, to the data storage module through the data bus for updating the model parameters of the Learner module. In terms of action inference, each environment process will correspond to an Actor (action) cluster node in the corresponding Actor module to achieve task parallelization. The CPU main controller sends program instructions to the FPGA cluster through PCIe to implement computing scheduling and start-stop control, and after receiving the feedback information of each environment process, it will be transmitted to the Actor cluster master node through PCIe. The Actor master node will further send it down to each level of cluster nodes through Ethernet. The master nodes of each level of task cluster or module cluster form a tightly coupled FPGA embedded system. The local memory of each master node is mapped to a shared memory space through the PCIe interface by using the DMA controller, so that any master node can directly access the memory area of other master nodes to achieve high-speed data exchange. After completing the model inference calculation, the Actor cluster master node sends the output result back to the CPU main controller through PCIe and issues it to each environment process through the API to complete the action execution.

[0258] When the interaction experience collected in the data storage module reaches a certain amount, the CPU main controller controls the Learner cluster nodes to start model update calculation through PCIe, and sends the experience samples to the main nodes of each task cluster and module cluster step by step through Ethernet until they are assigned to each unit computing node. The main nodes of each level of task cluster and module cluster also form a tightly coupled FPGA embedded system, and realize high-speed data interconnection through PCIe and DMA. After the model update is completed, the Learner cluster transmits the updated model parameters of each node back to the data storage module for backup. At the same time, the model parameters are copied to each Actor cluster node to update the local model.

[0259] Second, the network structure part of the design decision model will be described.

[0260] (1) The setting scheme of the decision model is as follows:

[0261] According to the task requirements, the complexity of the state and action spaces, the decision model structure can have various design methods. According to the functional modules, the decision model can be divided into three parts: the state encoding layer, the decision encoding layer, and the action encoding layer:

[0262] State encoding layer: Maps the input original state space to the latent feature space;

[0263] Decision encoding layer: Approximates the mapping distribution between the input state and the output action;

[0264] Action encoding layer: Maps the latent feature space after decision encoding to the action space.

[0265] In order to handle different types of decision tasks, the complexity of the model structure is also different.

[0266] (1) Simple decision tasks:

[0267] For decision tasks with a low-dimensional state space, the state encoding layer, the decision encoding layer, and the action encoding layer can be constructed using a simple fully connected network to form an end-to-end decision model. The core calculations of this model mainly involve matrix multiplication and activation function calculations.

[0268] (2) Complex decision tasks:

[0269] For tasks with a high-dimensional state space (such as images or multi-modal perception data), the structure of the decision model will be more complex. Figure 5 This is a schematic diagram of a decision model structure provided by an embodiment of the present invention for complex decision tasks, as Figure 5Shown as follows: State encoding layer: For tasks with low-pixel intermediate feature images as input, the state encoding layer is usually constructed using a Convolutional Neural Network (CNN). To reduce the computational load, the decision encoding layer and the action encoding layer still use fully connected networks. The state encoding layer can choose to be pre-trained or trained end-to-end together with the decision encoding layer and the action encoding layer.

[0270] High-pixel image input: For the input of high-pixel original camera images, the number and structure of convolutional layers in the state encoding layer will be more complex, and pre-trained deep networks (such as Residual Network (ResNet), Visual Geometry Group (VGG), and object detection algorithms, etc.) are used to implement it.

[0271] Multi-modal perception data: For tasks of multi-view images or multi-modal perception data, a multi-branch perception encoding network is usually used to preprocess different data, and then a network structure with a cross-attention mechanism (such as Transformer) is used for feature fusion and alignment. In such tasks, in addition to matrix multiplication, convolution, and activation function calculations, attention calculation operations are also required.

[0272] (3) Temporal dependence and experience replay:

[0273] To make full use of the perceptual state information and improve the decision-making model's understanding ability of system dynamics and environmental context, a bypass network can be added to the state encoding layer to process the serialized information of historical states and capture temporal dependence relationships. This part can be implemented through a Transformer network structure, involving recurrent calculations or attention calculation operations. In addition, to improve the utilization efficiency of training data and the training stability of the model, a generative experience replay module can be added. This module is constructed based on generative models (such as Variational Auto-Encoder (VAE), Generative Adversarial Network (GAN), or diffusion models, etc.), and can enhance the model's learning of online interaction experiences. The update of the generative experience replay module can be used as a bypass branch and alternated with the training process of the decision-making model to jointly promote the optimization of the model.

[0274] (2). Model structure adaptive selection strategy.

[0275] The above three optional adaptive design methods are described in detail below:

[0276] (1) Adaptive design method based on reinforcement learning:

[0277] In a decision-making model, the invocation of each encoding module can be modeled as an action in reinforcement learning, and the module selection strategy can be optimized through trial and error and feedback. This method is applicable to:

[0278] Dynamic tasks: The distribution or characteristics of the input data change dynamically, and the module invocation strategy needs to be adjusted according to the real-time environment.

[0279] No clear rules: When the module selection cannot be judged by fixed rules, the best choice needs to be found through exploration and optimization.

[0280] Online learning: The system needs to continuously improve the selection strategy during operation.

[0281] Its specific implementation steps include:

[0282] a) Construct a reinforcement learning framework:

[0283] In the adaptive design based on reinforcement learning, the optimization goal is to dynamically adjust the invocation strategy of each encoding module in the decision-making model so that the system performance (such as accuracy, execution speed, and resource utilization) reaches the optimal. To achieve this goal, it is first necessary to build a reinforcement learning framework and define the state, action, reward, and environment.

[0284] State : Represents the input data characteristics of the current system and the operating state of the hardware platform. The data characteristics include the distribution and complexity (such as dimension and sparsity) of the input data, and the system state includes the hardware resource utilization rate, latency, and power consumption, etc.

[0285] Action : Represents the invocation strategy of the encoding module. For example: at the state encoding layer, choose to use a convolutional neural network (CNN) or Transformer; at the decision encoding layer, choose a model based on the attention mechanism or a lightweight network.

[0286] Reward : Evaluates the performance of the module selection strategy, such as the weighted sum of accuracy, execution speed, and resource occupancy rate.

[0287] Environment: The environment is determined by the characteristics of the task and the hardware platform, including the dynamic changes of the input data and the resource allocation of the FPGA hardware.

[0288] b) Design a reward function:

[0289] The reward function is the core of optimization. Its role is to guide the agent to learn the optimal module invocation strategy through feedback. The set reward function needs to balance the weights of performance and resource consumption. Let the performance after model selection be , the execution latency be , and the resource utilization rate be , the reward function can be expressed as:

[0290] ;

[0291] where, , , are weight parameters, which are adjusted according to the application scenario. When the real-time requirement is high, increase the weight of to reduce the latency. When the hardware resources are limited, increase the weight of to reduce the resource occupancy.

[0292] c) Optimization objective:

[0293] The goal of reinforcement learning is to find an optimal policy to maximize the cumulative reward. Its optimization objective function is the optimal state-action value function :

[0294] ;

[0295] where, is the discount factor, which controls the importance of future rewards; represents the expected cumulative reward that can be obtained after taking action in state , and E represents the expected value.

[0296] d) Optimization process:

[0297] Use a neural network to approximate . The network structure is a fully connected network with 2 hidden layers. The input layer receives the state and the action . The number of nodes in each hidden layer is adjusted according to the task complexity. The output layer outputs the predicted value of . The following mean square error is used as the loss function :

[0298] ;

[0299] where, the expression of the target network is: ; is the current network parameter, is the target network parameter, which is updated regularly to stabilize the training. The training steps are as follows:

[0300] 1) Initialize the network and the target network;

[0301] 2) Obtain a randomly initialized state from the environment ;

[0302] 3) Select an action according to the -greedy policy: Select an action randomly with probability : or select the action corresponding to the maximum value with probability ; value;

[0303] 4) Execute the action , observe the reward and the new state ;

[0304] 5) Update the network parameters: , where is the learning rate and is the gradient;

[0305] 6) Repeat the above process, and synchronize the network parameters to the target network through operations such as copying and accumulation at regular intervals until the network converges.

[0306] During the execution of the decision-making model, the reinforcement learning agent dynamically selects the optimal network design for different coding modules according to the current state , and adjusts the weight of the reward function according to the real-time feedback to adapt to the changes of tasks and hardware. After the agent completes training, it can use the trained network to quickly infer the optimal module call strategy.

[0307] (2) Adaptive neural architecture search:

[0308] Based on adaptive neural architecture search for the optimal neural network structure, by incorporating the call strategy of coding modules into the search space, this method can automatically optimize the module design according to task characteristics and hardware constraints. This method is applicable to:

[0309] Module structures that require customization: Specific modules such as state encoding and context encoding require different architectures such as convolutional layers and Transformers.

[0310] Multi-objective optimization: The performance and hardware limitations (such as latency, power consumption) need to be balanced.

[0311] Its specific implementation steps include:

[0312] a) Define the search space:

[0313] ​Design candidate modules: Design modules such as the state encoding module and context encoding module in the decision-making model as searchable candidate modules:

[0314] ;

[0315] Parameterized search space : It includes network depth (number of layers, such as 3, 6, or 12 layers), network width (number of neurons or channels per layer), layer type (convolutional layer, fully connected layer, attention layer, etc.), activation function (such as Rectified Linear Unit (ReLU), etc.). For example, a candidate module design can be represented as: = (number of layers = 6, number of channels = 128, activation function = ReLU).

[0316] b) Design a multi-objective optimization problem:

[0317] In practical applications, the optimization of neural networks usually requires a trade-off between multiple objectives. For example, seeking a balance between performance (such as accuracy) and hardware constraints (such as latency, power consumption). The optimization problem can be formulated as:

[0318] ;

[0319] where, is the performance objective (such as classification accuracy, object detection precision) of candidate design , is the hardware constraint (such as execution latency, memory occupancy, power consumption) of candidate design . In this solution, the performance objective can be determined by the evaluation metrics of the decision-making model, and the hardware constraint is represented by hardware performance metrics, including latency (time required for the model to execute once), power consumption (energy consumed by the model during operation), and resource occupancy (usage rate of FPGA's logic units or memory).

[0320] c) Search process:

[0321] The search process uses an algorithm to explore the search space , and find the optimal solution between performance and hardware constraints. First, randomly generate a set of initial designs in the search space as a population. Then calculate the fitness of each architecture, which is defined here as a combined function of performance and hardware constraints :

[0322] ;

[0323] where, is a penalty factor used to restrict designs that exceed hardware constraints. In each iteration, selection, crossover, and mutation operations are performed to gradually generate a better architecture until the performance or hardware requirements are met. Among them, the selection operation is used to select excellent individuals in the population according to fitness, the crossover operation is used to combine the characteristics of two designs to generate a new design, and the mutation operation is used to randomly change some parameters of the architecture (such as the number of layers, width). Then, the generated new architecture replaces the individuals with poor performance to form the next generation population. Repeat the above process until the performance reaches the expected goal under the condition that the hardware constraints are met.

[0324] (3) Data-driven dynamic model selection:

[0325] Data-driven dynamic model selection is a flexible adaptive design method that dynamically selects the optimal coding module by analyzing the characteristics of input data in real time, thereby improving the adaptability and performance of the model. This method is applicable to:

[0326] Variable data characteristics: such as large differences in image resolution and modal characteristics;

[0327] Real-time applications: Require quick decisions on call strategies.

[0328] Its specific implementation steps include:

[0329] a) Extract data characteristics:

[0330] In dynamic model selection, the characteristics of the input data are the core basis for module selection. By extracting the statistical characteristics of the input data, it can provide a basis for subsequent module selection. The following two operations can be taken:

[0331] Feature engineering: Use simple statistical methods or feature engineering to extract key characteristics. For example, for image data, the main data characteristics include resolution (width and height of the input image), color distribution (average pixel value, standard deviation), and sparsity (proportion of non-zero pixels); for multimodal data, the main data characteristics include data type (audio, image, or text, etc.), temporal characteristics (periodicity or trend of time series data), and real-time applications (require quick decisions on call strategies).

[0332] Use lightweight networks to extract characteristics: To efficiently process complex data, lightweight neural networks can be used to extract high-dimensional features. Define the input as the original data (such as images, audio), and define the output as a low-dimensional feature vector representation , which contains the key characteristics of the data. The feature vector extraction operation can be expressed as:

[0333] ;

[0334] b) Set up a dynamic control network:

[0335] The dynamic control network is responsible for predicting the selection probability of each candidate module based on the feature vector of the input data ; the control network is a lightweight classifier, and its output is the selection probability of each module: ;

[0336] ;

[0337] Among them, the weight matrix is used for the learnable parameters of the control network, and the output probability distribution represents the selection probability of each candidate module. For example, assume there are three candidate modules: Convolutional Network (CNN), Transformer, and Fully Convolutional Network (FCN). After the input feature vector is calculated by the control network, the following are obtained respectively:

[0338] ;

[0339] .

[0340] c) Adaptive selection:

[0341] According to the probability output by the control network , dynamically select the optimal module design and generate the final output. First, determine the selection strategy, and there are two methods:

[0342] Hard selection: directly select the module with the highest probability Design:

[0343] ;

[0344] This method has a relatively small computational cost and is suitable for scenarios with high real-time requirements.

[0345] Soft selection: According to perform weighted fusion on the outputs of all modules to obtain:

[0346] ;

[0347] Among them is the output of the module design, is the final output. This method can retain the output information of multiple module designs and improve the robustness of the model.

[0348] Furthermore, perform weighted summation on the output selections of each encoding module design:

[0349] ;

[0350] Final output Can be used for subsequent tasks.

[0351] d) Implement optimization and deployment:

[0352] The goal of the dynamic control network is to predict the selection probability of the optimal module design , and the following cross-entropy loss is used to measure the difference between the predicted probability and the true selection :

[0353] ;

[0354] where represents the cross-entropy loss, is the distribution of the true selection. During training, the parameters of the control network are updated using stochastic gradient descent or the Adam optimizer. With the trained dynamic control network, the model design can adapt to different input data characteristics in real time, thereby improving system performance and flexibility.

[0355] (III). Hardware adaptation of the decision model:

[0356] When performing hardware adaptation of the decision model, the final calculation and deployment of the decision model must take into account resource allocation and task characteristics in the hardware acceleration environment, mainly considering the following aspects:

[0357] Hardware resource utilization: The hardware resources of FPGA are limited. By accelerating computationally intensive tasks (such as convolution, matrix multiplication, activation functions, etc.) using a large-scale parallel processing element (PE) array, the performance can be significantly improved. When designing the computing logic of FPGA nodes, consider how to reasonably allocate computing tasks to FPGA resources, and improve the overall performance by optimizing the computing graph, reducing redundant calculations, and reducing memory bandwidth consumption.

[0358] Modular design: To better adapt to FPGA hardware, by decomposing different layers in the decision-making process (such as the state encoding layer, decision encoding layer, action encoding layer) into independent modules, according to the hardware resource situation of FPGA, computing nodes with different logical designs can be allocated to each module. For example, the state encoding layer (mainly involving convolution operations) is more suitable for deployment on FPGA nodes dedicated to convolution computing, while the decision encoding layer and action encoding layer (more involving fully connected layers and activation functions) can be deployed on FPGA nodes dedicated to matrix multiplication computing.

[0359] Computing parallelization: Through appropriate data partitioning and task scheduling, multiple computing tasks can be executed simultaneously, giving full play to the advantages of multi-device parallel computing. For example, in a Convolutional Neural Network (CNN), the convolution operation is a computationally intensive and parallelizable task that can be directly mapped to multiple FPGA nodes for efficient collaborative computing.

[0360] III. Describe the distributed parallel training framework part.

[0361] Since the decision-making model conducts trial-and-error exploration based on environmental feedback signals during training, the model needs to interact with the environment to collect online experience as training data before each parameter update. After collecting the data, the model updates the parameters through a forward calculation + backpropagation, and the above process is iteratively advanced. Existing model training platforms based on this interactive exploration training method mainly adopt a single-environment configuration of the CPU – GPU framework, where the CPU is responsible for environment establishment and data control, and the GPU is used for high-throughput parallel computing. Due to the single-environment configuration, only one state can be predicted for actions at a time, and the overhead of scheduling the GPU for operation is sometimes longer than the parallel computing time, resulting in low computing efficiency. To address this problem, the present invention designs a distributed parallel computing architecture for decision-making models, decouples tasks such as environment interaction, model inference, and experience storage, and sets multiple computing nodes to execute in parallel for different tasks, thereby making full use of computing resources and improving computing efficiency.

[0362] The main idea of decision-making model parallel computing is to decompose tasks, data, and computing processes and execute them synchronously on multiple processing units to improve computing efficiency. Here, two design schemes of data parallelization and model parallelization are mainly adopted.

[0363] (1) Data parallelization.

[0364] The data parallelization scheme mainly divides the input data into multiple subsets and processes them in parallel on multiple processing units to improve the processing efficiency of large batches of computing data. The specific structure consists of one (or more) central computing nodes and multiple sub-computing nodes. Each computing node is constructed and completes its respective computing tasks through 1 FPGA board. Figure 6 FIG. is a schematic diagram of a data parallelization structure of multiple sub-nodes - central node provided for an embodiment of the present invention. As Figure 6 shown, the data is sliced into multiple subsets and distributed in parallel to multiple sub-nodes for parallel processing. Each sub-computing node maintains 1 local model. By partitioning the data onto different processing units, locally calculating the results of each small batch of data in parallel, and finally summarizing and updating, the computing efficiency can be significantly improved.

[0365] During the training process of the decision-making model, data parallelization design can be adopted in both the model forward inference and parameter update phases to improve processing efficiency. Assume the input feature is image data with a size of , where is the data batch size, is the number of input channels, is the image size. If , it is recommended to divide along the dimension when partitioning the data subsets. If , then it can be divided according to the input channel dimension. For non-image data input, its shape can be represented in a general form , where corresponds to the size of the feature dimension of the th dimension. Generally, does not exceed 3. For this input case, if the feature dimensions are independent of each other, that is, there is no correlation during data calculation, it can be considered that feature data can be regarded as a feature individual during calculation, and the size of each individual is . If the data in each dimension is mutually correlated, is the dimension that can be parallelized for calculation. When partitioning the data subsets, select the dimension with the largest corresponding dimension number in the middle according to the principle of maximizing parallelism to improve resource utilization.

[0366] Assume the data is divided along . The size of the data subset allocated to each sub-computation node is determined by the number of sub-computation nodes set by the system, where corresponds to the data parallelism of the system. The size of the data that can be allocated to each node should be . Since in many cases cannot be divided evenly by , there will be an inconsistent situation in the amount of data allocated to the sub-computation nodes. Among them, sub-computation nodes have data of size , while the remaining sub-computation nodes have data of size , represents the remainder operation, represents the floor operation.

[0367] During the training process of the decision-making model, the two stages of environment interaction and model update are iteratively advanced. Among them, the environment interaction stage mainly involves the model forward inference calculation. Refer to Figure 6, the input data is divided into multiple subsets according to the system data parallelism and sent to each sub-computation node in parallel. Each computation node has the same model parameters and independently completes the layer-by-layer inference calculation of the model to obtain the corresponding output features. Then, each node sends its own output to the central computation node, and the overall output features are obtained through a single data aggregation operation.

[0368] In the model update stage, in addition to including the forward inference calculation of the model, it is also necessary to calculate the parameter gradients based on the feature results and complete the parameter update through backpropagation. Figure 7 This is a schematic diagram of a model update process based on multi-subnode-central node data parallelization provided by an embodiment of the present invention. As Figure 7 shown, its data parallelization processing method is similar to that in the inference stage. Here, a reordering (Resort) operation is added in the data subset division link. This operation shuffles and rearranges the order of each feature individual along the data division dimension to achieve the function of random sampling. Since the model update stage will be triggered only when enough empirical samples are collected, the number of samples in the dataset is often much larger than other feature dimensions at this time. Therefore, the data division here is equivalent to batch data sampling from the dataset. Each sub-computation node also maintains 1 local model. Each node independently completes the forward propagation and backpropagation to obtain the corresponding gradients. At this time, the gradients on each node are different from each other. Then, each node sends its own gradient to the central computation node, and the merged gradient is obtained through a single data aggregation operation. The merged gradient is returned to each node. At this time, the gradients on each node are exactly the same. Finally, each node independently completes the update of the local model parameters.

[0369] The above solution integrates the information of sub-computation nodes by setting a data aggregation operation and a central computation node. However, load imbalance is likely to occur during the communication process of each node. When it is found that a certain node has a heavy load, the present invention proposes a data parallelization scheme without a central node to achieve this. Figure 8 This is a schematic diagram of a ring communication topology structure of data parallelization without a central node provided by an embodiment of the present invention. As Figure 8As shown in the figure, this solution is based on a ring communication topology structure and there is no central computing node. All sub-computing nodes form a logically ring. First, the data to be calculated is evenly divided into multiple segments, so that each node only needs to process and transmit a part of the data during communication, thereby reducing the communication volume and computing complexity. Here, the same division strategy as the previous solution is adopted. Similarly, each node initially has a part of the complete data and independently completes the local model inference calculation. The difference is that this solution does not require a central computing node to assist in implementing the data aggregation operation, but completes the data integration of each sub-node through ring communication. During the ring communication process, each node sends its own data to the next node and receives data from the previous node. As the communication progresses, each node gradually integrates the data segments from other nodes, and finally all nodes have the completely synchronized data. This ring communication process specifically includes the following two stages:

[0370] Decentralized processing stage:

[0371] a. Initialization: For each node the data segment obtained after local calculation , set the initial send buffer , and store it into ;

[0372] b. Ring communication: Node sends the content in the send buffer to node . If , then send it to node 0; at the same time, receive data from node and store it into the receive buffer . If , then receive data from node ;

[0373] c. Data update: Node performs a specific operation (such as splicing, accumulation, etc.) on the content of the receive buffer and the local data segment to obtain a new data segment and update the data in the send buffer to the new ;

[0374] d. Repeat the above steps times so that each node gradually accumulates the partial sum of the corresponding data segments of all nodes.

[0375] Central synchronization stage:

[0376] a. Initialization: After the end of the decentralized processing stage, each node has partial aggregated data fragments , sets an initial send buffer , and stores into ;

[0377] b. Ring communication: Similar to the decentralized processing stage, node sends the content of the send buffer to node , and at the same time receives data from node into the receive buffer ;

[0378] c. Data update: Node performs specific operations (such as combination, overwrite, etc.) on the content of the receive buffer and the local data fragment to obtain a new data fragment , and updates the send buffer to the new ;

[0379] d. Repeat the above steps times, and finally each node has the fully synchronized data.

[0380] In the data parallelization scheme of the decentralized node, each node makes full use of the uplink and downlink bandwidth, with high bandwidth utilization and good scalability. However, as the number of nodes increases, the communication delay will gradually increase, and this ring topology requires that the bandwidth and delay between nodes be as consistent as possible. Therefore, when adopting this scheme, it is necessary to control the number of nodes and ensure that the performance of each node is similar. In addition, this scheme has a risk of error accumulation during multi-node data transmission and aggregation, and each node needs to have enough memory space to cache intermediate results. Therefore, for modules that are greatly affected by cumulative errors and have large memory occupancy, it is more recommended to adopt the centralized data parallel scheme of multiple sub-nodes - central node.

[0381] (2) Model parallelization.

[0382] To achieve a higher throughput, various computational tasks in the decision model training process are fully parallelized. In implementation, an Actor module is set up to collect data, and a Learner module is responsible for updating the model. Multiple Actors are parallel to each other, and one Learner can correspond to one Actor or multiple Actors. Considering the requirements of the simulator operation, the multi-environment process simulation will be implemented on multiple GPU cores. Multiple environment processes are arranged on each GPU card. During this period, the CPU controls the environment interaction through the API interface and realizes the interconnection and communication with the Actor through PCIe to control data transmission and caching. Each Actor can be responsible for the interaction of one or more simulation environments, and the model inference process is implemented based on FPGA. Figure 9 It is a schematic diagram of a multi-Actor parallel interaction structure based on FPGA provided by an embodiment of the present invention. Figure 9 In it, each Actor and Learner will maintain a local decision model for action inference and model iterative update in environment interaction. Considering that when dealing with complex decision-making tasks with visual images or multi-modal perception data as the original state input, the network scale of the state encoding part is relatively large. Under the data parallelization setting, each FPGA node corresponding to an Actor and a Learner is required to retain the complete model parameters, which will cause redundancy on the one hand. On the other hand, when functional modules such as context encoding modules and experience replay modules are introduced, the number of decision model parameters will be further expanded, especially for functional modules based on complex network structures such as Transformers and diffusion models. Due to the limited memory of a single FPGA board, the decision model will be difficult to load. To solve the above problems, the present invention divides the decision model into multiple parts and uses multiple cluster nodes to cooperate to complete the computational tasks of the entire model. Each cluster node contains multiple unit computational nodes based on a single FPGA board, and different data parallelization design schemes are applied according to the characteristics and requirements of the corresponding task modules. In the decision model division stage, the following three division methods are specifically adopted:

[0383] ① Task division:

[0384] The computational process of the decision model is regarded as a pipeline operation, and different computational tasks during model inference are divided into different task cluster nodes. Each task cluster node is responsible for the work of one task, and the intermediate features obtained by the calculation will be sequentially transmitted between the task cluster nodes. Take Figure 5Taking the decision model structure for complex decision-making tasks as an example, according to the computational tasks, it can be divided into a state encoding layer, a decision encoding layer, and an action encoding layer. Among them, the state encoding layer contains a branch network for processing historical data, and this part can be further split out as a context encoding layer. In addition, during the training process of the decision model, an experience replay module is used for data augmentation. Since the training and inference of the generative model in this module are relatively independent, it can also be split as an independent model. Among them, the computational tasks divided by the generative model according to different network structures are slightly different. The generative model based on the VAE structure can be divided into an encoder and a decoder, the generative model based on the GAN structure can be divided into a generator and a discriminator, and the generative model based on the diffusion structure can be divided into forward diffusion and reverse denoising.

[0385] After the decision model is split, each computational subtask will be assigned to a corresponding task cluster node. Each task cluster node can work independently in the stage it is responsible for, achieving a high device utilization rate and improving the overall processing efficiency.

[0386] ② Module division:

[0387] After taking the decision model task division and task cluster node settings, if the computational amount of a certain computational subtask is too large or the performance of the task cluster node is poor, it will lead to a decrease in the computational efficiency of the entire model. To address this problem, the present invention further divides tasks with heavy computational burdens such as state encoding and context encoding, and sets corresponding module cluster nodes inside the corresponding task cluster nodes to cooperate in completing each part of the computation. Specifically, in the state encoding task, it is divided into image processing, point cloud processing, and multi-scale feature fusion according to the data processing stage. If the input of image processing is multi-view images, after image encoding, a multi-view image feature fusion operation is also required, which can be implemented using a network structure with cross-attention calculation operations such as Transformer; the network for the point cloud processing part can be divided into point cloud feature encoding, pseudo-image encoding, single-shot multi-object detection, and feature aggregation. In the above tasks, such as Figure 5Among them, image encoding, pseudo-image encoding, and single-shot multi-object detection are all implemented based on convolutional networks. Point cloud feature encoding is mainly achieved through rotation transformation operations and multilayer perceptrons (MLPs). Multi-view image feature fusion and multi-scale feature fusion are both based on Transformers. Each part is separately assigned to different module clusters for processing. For multi-view image feature fusion, multi-scale feature fusion, and context encoding, considering the complex structure of the Transformer model, the multi-head attention module and the feed-forward neural network module in it are further divided into different module clusters. The multi-head attention module is responsible for capturing the dependencies at different positions in the input sequence, with a relatively large computational load and relatively independent computational logic. The feed-forward neural network module further performs non-linear transformation on the output of the attention module.

[0388] By partitioning the task clusters with heavy computational burdens, the computing capabilities of different devices can be fully utilized, improving the training and inference efficiency of the model. At the same time, it helps to improve the parallel efficiency of the model. Especially when the computational amounts and parameter scales of different module clusters are relatively balanced, the resources of multiple computing devices can be better utilized.

[0389] ③ Network layer partitioning:

[0390] For the computational tasks of image encoding, pseudo-image encoding, and single-shot multi-object detection, since the computational amount of feature data and the number of model parameters in the convolutional network part are relatively large, in order to ensure relatively balanced computational loads for all cluster nodes, a layer partitioning method is adopted to assign the convolutional network with a large number of layers to multiple module clusters for collaborative processing. The partitioning basis can be network layer block groups, scales, or computational functions, which need to be determined according to the model structure. Taking three convolutional models commonly used in image processing, namely ResNet, VGG, and You Only Look Once (YOLO) (unified real-time object detection), as examples, the specific partitioning methods can refer to the following schemes:

[0391] ① ResNet:

[0392] ResNet is mainly composed of multiple residual blocks, which contain convolutional layers, batch normalization layers, and activation function layers. The network depth can reach dozens or even hundreds of layers. The following two partitioning schemes can be adopted:

[0393] a. Division by residual block groups: Several consecutive residual blocks can be divided and placed on one computing device. For example, in ResNet-50, there are residual blocks in 4 stages, and the number of residual blocks in each stage is different. The residual blocks in the first stage (for example, containing 3 residual blocks) can be placed on a module cluster node. This stage is mainly used to extract low-level features of the image, such as edges and simple textures. The residual blocks in the second stage (for example, containing 4 residual blocks) are placed on another module cluster node to further extract more complex features. The residual blocks in the third and fourth stages follow the same pattern. Such division helps to balance the computing load between different devices and enables resource allocation according to the functional characteristics of the residual blocks in each stage.

[0394] b. Division of convolutional layers and shortcut connection layers: Inside each residual block, there are convolutional layers and shortcut connections (used to skip one or more layers and add directly). The convolutional layers can be placed on one module cluster node, and the shortcut connection part can be placed on another device. However, this division method may increase the communication overhead between devices because intermediate results need to be transmitted for addition operations. But it may be useful in certain specific hardware architectures or optimization scenarios. For example, when testing the contribution of different layer types to performance, this division method can be used to evaluate the computing efficiency of convolutional layers and shortcut connection layers separately.

[0395] ②VGG:

[0396] The VGG network is mainly composed of multiple convolutional layers and pooling layers. The network structure is relatively regular, with multiple convolutional layers stacked, followed by fully connected layers. Taking VGG-16 as an example, it has 13 convolutional layers and 3 fully connected layers. The following two division schemes can be adopted:

[0397] a. Division by convolutional layer groups and fully connected layers: The previous convolutional layers (for example, the first 6 convolutional layers) can be placed on one module cluster node. This part is mainly used to extract features of the image, such as edges, textures, and simple shapes. Then, the subsequent convolutional layers and pooling layers (for example, the remaining 7 convolutional layers and all pooling layers) are placed on another module cluster node to further extract high-level semantic features. Finally, the 3 fully connected layers are placed on a third module cluster node to perform further feature fusion based on the features extracted previously.

[0398] b. Division by the functions of convolutional layers and pooling layers: Divide according to the functions of convolutional layers and pooling layers. Place the layers that only perform convolutional operations on one device, and place the layers that include pooling operations on another device. In the VGG network, the pooling layer is mainly used to reduce the data dimension and extract the main features, while the convolutional layer is used to mine various features in the image. This division allows each device to focus on one type of computational operation, but may increase the number of communications between devices because the pooling layer needs to receive the output of the convolutional layer for operation.

[0399] ③ YOLO:

[0400] YOLO is a real-time object detection model that can be used for single-shot multi-object detection in point cloud encoding. It mainly includes convolutional layers and fully connected layers. Taking YOLOv3 as an example, it has a Darknet-53 backbone network (mainly composed of convolutional layers) for feature extraction, followed by a detection head part (including convolutional layers and fully connected layers) for predicting the category, location, and confidence of the target. The following two division schemes can be adopted:

[0401] a. Division of the backbone network and the detection head: The Darknet-53 backbone network can be placed on a module cluster node, which is responsible for extracting rich features from the input image. For example, it can first extract low-level features of the image, such as edges and textures, and then, as the network depth increases, extract more high-level semantic features, such as the shape of the object and category-related features. Place the detection head part on another module cluster node, and the detection head uses the features extracted by the backbone network to perform calculations related to object detection, such as generating prediction boxes of different scales and predicting the object category and location.

[0402] b. Division of prediction layers of different scales: In the detection head of YOLOv3, there are layers for predicting objects of different scales. These layers for predicting objects of different scales can be divided onto different devices. For example, place the layers for predicting small-scale objects on a module cluster node because the detection of small-scale objects usually requires more refined features and more computational resources to distinguish different small objects; place the layers for predicting medium-scale and large-scale objects on another module cluster node, so that objects of different scales can be detected in parallel, improving the detection efficiency.

[0403] IV. Explanation of a single FPGA computing node:

[0404] (1) Node architecture:

[0405] To meet the data calculation requirements of different types of network layers in the decision-making model, the present invention abstracts the calculation operations in the forward inference and backpropagation processes of the model into multiple core calculation units, and designs a single calculation node hardware structure for different calculation tasks based on FPGA, aiming to enable the unit nodes in each cluster node to allocate resources more efficiently to complete a certain type of network calculation.

[0406] ① Node architecture for convolutional calculation:

[0407] Figure 10 FIG. is a structural diagram of an FPGA single calculation node for convolutional calculation provided by an embodiment of the present invention. As Figure 10 shown, it includes a communication interface, a cache unit, a scheduling kernel, a forward inference module, a backpropagation module, and a multiplexer. The communication interface includes an Ethernet interface, a Peripheral Component Interconnect Express (PCIe) interface, a Low-Voltage Differential Signaling (LVDS) interface, a Double Data Rate Synchronous Dynamic Random Access Memory Interface (DDR) memory interface, a host Random Access Memory (RAM) memory interface, etc., and is used to connect to the upper computer software, other calculation nodes, and the internal cache unit of the node for data transmission. The cache unit includes two parts, DDR and RAM, and is used to store data. Among them, DDRA is used to cache feature data, and DDRB is used to store model parameter data. The RAM connected to DDRB is used to store data results, and DDRB will transmit the data back to the upper computer or other nodes through the communication interface after receiving the data. The communication interface is connected to the scheduling kernel, and the upper computer or the main node can modify the internal logic and data flow of the FPGA in a programmed manner through PIO (Programmed Input / Output) to execute data transmission or specific control commands. The scheduling kernel is used to issue memory access, calculation control, and start signals for each layer to each calculation unit. After the calculation operation is completed, each calculation unit will return the running state and completion signal. In addition, the scheduling kernel is also connected to a multiplexer for result write-back and debugging control.

[0408] The input data of the data operation unit is directly obtained from the feature buffer, and the calculation result is then passed back to the feature buffer. In the forward inference module, the general matrix multiplication calculation unit and the convolution calculation unit are respectively used to complete the basic matrix multiplication and convolution calculation operations. Among them, the feature data is read from the feature buffer, the parameter data is read from the parameter buffer, and the bias data is directly read from the bias buffer integrated inside the calculation unit. After the calculation is completed, the feature result will be passed to the parameter buffer and the multiplexer. In the backpropagation module, the loss gradient calculation unit and the parameter gradient calculation unit are respectively used to complete the error loss and gradient calculation of each network layer. Among them, the loss gradient calculation unit needs to read the feature and parameter data from the feature buffer and the parameter buffer respectively for gradient calculation, and its result will be passed back to the parameter buffer and sent to the multiplexer at the same time. The parameter gradient calculation unit directly reads the network parameters and the corresponding error gradients from the parameter buffer, and sends the result to the parameter buffer and the multiplexer after completing the parameter update calculation.

[0409] The FPGA single node adopting this calculation architecture is mainly used for constructing the module clusters of the three parts of image encoding, pseudo-image encoding, and single-shot multi-object detection in state encoding. It should be noted that if the node is assigned to the task cluster node and module cluster node of the decision-making model in the Actor, since the Actor only needs to perform the forward inference calculation of the model, the backpropagation module in the node can be omitted. In addition, if the decision-making model directly calls the pre-trained convolutional model to implement state encoding, since the model parameters in this part will be frozen during the model update stage, there is no backpropagation calculation either, and the backpropagation module in the corresponding node of the Learner can also be omitted.

[0410] ② Node architecture for attention calculation:

[0411] Figure 11A structural diagram of an FPGA single computing node for attention calculation provided by an embodiment of the present invention. Except for the forward inference module, the rest is set the same as the FPGA single-node architecture for convolutional calculation. The FPGA node adopting this type of computing architecture is mainly used for constructing module clusters in three parts: multi-view image feature fusion, multi-scale feature fusion, and context encoding in state encoding, focusing on processing complex networks based on Transformer. Therefore, in the forward inference module, it mainly includes three key computing units: general matrix multiplication, layer normalization (LN), and Softmax. Among them, the input data of Softmax and LN are both obtained from the feature buffer, and the calculation results will be passed back to the feature buffer. Softmax and LN units are also connected to a multiplexer for result feedback. When performing LN calculation, the weight data and bias data are read from the weight buffer and bias buffer respectively. However, different from the general matrix multiplication unit, these two parts are integrated in the LN unit.

[0412] ③ Node architecture for fully connected calculation:

[0413] Figure 12 A structural diagram of an FPGA single computing node for convolutional calculation provided by an embodiment of the present invention. The node architecture for fully connected calculation can be regarded as a simplified version of the node architecture for convolutional calculation. Since there is only parallel multiply-accumulate operation in the fully connected layer, the forward inference module in this node architecture only includes a general matrix multiplication computing unit, which is equivalent to all the computing resources corresponding to the convolutional computing unit being transferred to the general matrix multiplication computing unit, so that it has a larger-scale PE parallel array and higher computing density, and can achieve more efficient and dense parallel multiply-accumulate operations.

[0414] (2) Communication method:

[0415] There are various interconnection and communication methods between FPGA unit nodes and other devices. In the present invention, FPGA unit nodes are used to construct module cluster nodes for different computing tasks. Since the main nodes in each module cluster and task cluster need to undertake tasks such as data distribution, control signal / instruction transmission, and status monitoring, an FPGA board with special management functions is used as the main node to achieve coordination between module clusters and task clusters. Compared with unit nodes, in addition to having basic FPGA logic processing capabilities, the main node also integrates an additional control module, which realizes functions such as task distribution and status monitoring of the entire cluster node through an Ethernet controller or a high-speed serial interface, and uses its internal embedded processor to run a management program to send instructions to other FPGA nodes, such as starting a specific hardware acceleration function and reading the status information of the node.

[0416] In terms of the interconnection method, the interconnection communication between FPGAs mainly occurs between the master node and the unit node. If a certain module cluster node adopts a decentralized node ring communication topology to achieve data parallel processing, there is also an interconnection requirement between adjacent unit nodes. In addition, the master node is involved in the interconnection communication with the host computer, where the host computer includes two computing devices, CPU and GPU. Therefore, there is communication transmission between FPGA and CPU as well as between FPGA and GPU. For the above three interconnection communication requirements, there are the following specific implementation methods:

[0417] ① FPGA-FPGA interconnection communication.

[0418] The interconnection communication between FPGAs is mainly achieved through three channels: PCIe, LVDS, and Ethernet. In this solution, the large-scale data transmission of features, parameters, etc. between single computing nodes in each task cluster and module cluster is realized by PCIe, and the control signal / command transmission between nodes and the feedback of status flag information are realized by LVDS. Considering the limitations of FPGA interface quantity and the bandwidth and communication delay problems existing in the expansion interface, the cross-level data transmission between the master node and the subordinate sub-cluster nodes in each task cluster and module cluster is realized by Ethernet.

[0419] ② FPGA-CPU interconnection communication.

[0420] The interconnection communication requirements between FPGA and CPU are mainly reflected in the mutual transmission of computing data such as features and parameters, as well as control, debugging information, and status feedback signals between the master node of each task cluster and the host computer. It can be realized by PCIe or Compute Express Link (CXL). CXL provides higher bandwidth and lower latency and supports cache coherence. Although it has advantages in performance and function, considering that it is a relatively new technical standard and device compatibility, this solution chooses to use PCIe to realize the data transmission between FPGA and CPU. Through the PCIe interface, data transmission is carried out in the form of Direct Memory Access (DMA), which can effectively reduce the burden on the CPU and improve the efficiency of data transmission. Since PCIe has the characteristic of backward compatibility, the new version of the PCIe interface can be compatible with the old version of the device, which provides convenience for subsequent system upgrade and expansion.

[0421] ③ FPGA-GPU interconnection communication.

[0422] In this solution, the GPU is only used to create an interactive environment. The main communication requirements involved mainly include sending the state and reward signals from the simulation environment in the GPU to the Actor module, and the Actor module sending the decision actions obtained through inference back to the simulation environment in the GPU for execution. Since the simulation environment needs to call the API interface from the CPU side of the host computer to achieve creation and interaction on the GPU side, the direct communication method between FPGA and GPU is not adopted in this solution. The FPGA master nodes and GPUs in each task cluster and module cluster are used as PCIe devices and are connected to the CPU through the PCIe bus. Among them, the FPGA first transfers the data to the RAM through DMA, and then the GPU reads the data from the host RAM through DMA.

[0423] V. The FPGA cluster with a multi-level cluster structure will be described below.

[0424] Considering the differences in the computing task requirements of each module in the parallel training architecture of the decision-making model, this solution will build a cluster through multiple FPGA task clusters and module clusters to implement model inference and update tasks. At the same time, it will combine the CPU and GPU to build an environment interaction framework to achieve efficient online inference training of the model. Figure 13 It is a schematic diagram of the structure of an accelerator multi-level cluster provided by an embodiment of the present invention. Each level of cluster node includes a master node and subordinate sub-cluster nodes. Among them, the task assignment, monitoring, and data transmission between the master node and each cluster node are realized through Ethernet. Each subordinate sub-cluster node also includes a master node and subordinate sub-cluster nodes. The master nodes at all levels are responsible for receiving and parsing the information management packets, algorithm model packets, configuration parameter packets, etc. from the upper-level master node, deploying different model computing tasks and modules to each sub-cluster node through Ethernet, and controlling the start and stop of the computing inside each cluster node. At the same time, the model parameters and feature data are distributed to each cluster node according to the computing task settings through Ethernet, thus forming a multi-level FPGA cluster structure, which is used as the Actor or Learner module respectively to realize model inference and parameter update. In addition, there is also interconnected communication between the master nodes and unit nodes of each cluster node according to the task requirements. Among them, large-scale streaming data such as model parameters and training features is realized through the PCIe interface, and the data related to signal management and status monitoring is transmitted between nodes through LVDS. For the bottom-level cluster nodes, its master node will be directly connected to each FPGA unit node, and each unit node corresponds to a single FPGA board. Similarly, the data transmission and task assignment between the master node and the unit nodes inside the cluster node are realized through Ethernet, and the status monitoring and data transmission between the unit nodes are also realized through PCIe and LVDS respectively.

[0425] (1) Logic structure of the Actor / Learner cluster node:

[0426] Designed according to the model structure and computational task requirements, the Actor and Learner cluster nodes have different logical structures respectively. Figure 14 It is a logical schematic diagram of an action cluster node provided by an embodiment of the present invention, as Figure 14 shown. For the Actor cluster node, its subordinate sub-cluster nodes include four parts: the state encoding task cluster, the context encoding task cluster, the decision encoding task cluster, and the action encoding task cluster. In terms of data caching logic, the Actor master node communicates with the host computer through PCIe, and distributes tasks and transmits status back through Ethernet with each task cluster. For the network parameters of each module of the decision model, they are also transmitted from the Actor master node to each task cluster through Ethernet. Among them, the state encoding task cluster will also read the original state data obtained in real time from the environment interaction from the master node. The encoded features obtained after processing will be further sent to the context encoding task cluster through PCIe for caching and subsequent encoding processing, and at the same time sent to the decision encoding task cluster as input data. Similarly, the encoded features obtained after processing by the context encoding task cluster will also be sent to the decision encoding task cluster through PCIe as input data. Similarly, the encoded features obtained after processing by the decision encoding task cluster will be sent to the action encoding task cluster through PCIe as input data. The action features obtained after processing will finally be sent back to the master node through Ethernet and sent to the host computer to execute the action to complete the environment transfer. Since there is a certain degree of independence and timing in the data processing process of each task cluster, and the data processing speeds and load conditions of different task clusters may be different, in order to ensure the stable operation of the entire system and the accurate transmission of data, adjacent task clusters need to monitor and transmit status through LVDS. Figure 15 It is a logical schematic diagram of a trainer cluster node provided by an embodiment of the present invention, as Figure 15 shown. For the Learner cluster node, its subordinate sub-cluster nodes add a generation replay task cluster on the basis of the Actor cluster node and are also interconnected through Ethernet. This task cluster needs to read the experience samples collected online from the master node for the training of the diffusion model. Each experience sample contains three parts: the original state data, the action, and the reward. Then, the virtual experience samples generated by the diffusion model are sent back to the master node to expand the training data of the decision model. Since it involves the update of the decision model parameters, in the Learner cluster node, in order to implement the backpropagation operation, the action encoding task cluster calculates the loss according to the action and reward data in the experience sample and performs gradient backpropagation. Each task cluster needs to send the updated local parameters back to the master node. Therefore, the data transmission method here is two-way.

[0427] (2) Logical structure of task cluster nodes.

[0428] Considering that the computing load of some task cluster nodes is too heavy, in order to balance the resource allocation of each cluster node, this solution divides multiple module cluster nodes to construct each task cluster node. For the specific module division, refer to the model parallelization part described above. Figure 16 It is the logical structure diagram of a state encoding task cluster for the action part provided by an embodiment of the present invention. Figure 17 It is the logical structure diagram of a state encoding task cluster for the trainer part provided by an embodiment of the present invention. The state encoding task cluster in the Learner adds module parameters and loss gradient backpropagation operations on the basis of the Actor part.

[0429] a) For the multi-view image encoding, pseudo-image encoding, and single-shot multi-object detection module clusters that focus on convolutional computing, they are further divided into multiple CNN module clusters according to the network layer. Each CNN module cluster contains multiple unit computing nodes, and data parallelization of decentralized nodes is realized based on the ring topology structure.

[0430] Figure 18 It is the logical schematic diagram of the module cluster for convolutional computing in the action part provided by an embodiment of the present invention. As Figure 18 shown, each CNN module cluster transmits intermediate features through PCIe. Considering that there are residual connections in some convolutional computing operations, in addition to data interaction between adjacent CNN module clusters, there may also be cross-module cluster data transmission to ensure the accuracy of feature transfer and calculation involved in the residual connection. Figure 19 It is the logical schematic diagram of the module cluster for convolutional computing in the trainer part provided by an embodiment of the present invention. As Figure 19 shown, compared with the Actor part, the backpropagation of the loss gradient is increased.

[0431] b) For the multi-view feature fusion and multi-scale feature fusion module clusters that focus on Transformer attention computing, they are divided into multi-head attention and feed-forward neural network module clusters according to functions. Each module cluster realizes centralized data parallelization based on the structure of a central node - multiple child nodes. Figure 20 It is the logical schematic diagram of the module cluster for attention computing in the action part provided by an embodiment of the present invention. As Figure 20 shown, the input features are first transmitted from the master node to the multi-head attention module cluster through Ethernet for processing, and then, according to the network block group setting, the feed-forward neural network and multi-head attention calculations are alternately performed. After the last feed-forward neural network calculation is completed, the output features are transmitted back to the master node through Ethernet. Therefore, there is feature mutual transmission between the multi-head attention and feed-forward neural network module clusters. Figure 21 It is the logical schematic diagram of the module cluster for attention computing in the trainer part provided by an embodiment of the present invention. As Figure 21 shown, compared with the Actor part, the backpropagation of the loss gradient is increased.

[0432] c) The point cloud feature encoding module clusters are divided into rotation transformation module clusters and MLP module clusters according to their functions. Each module cluster also realizes centralized data parallelization based on the structure of a central node - multiple sub - nodes. Figure 22 This is the logical schematic diagram of the module cluster for point cloud feature encoding in the action part provided by the embodiment of the present invention. That is, the module cluster for point cloud feature encoding in the Actor part. As Figure 22 shown, the input features are first transmitted from the master node to the rotation transformation module cluster through Ethernet for processing. Then, according to the network block group setting, MLP and rotation transformation calculations are alternately performed. After the last MLP calculation is completed, the output features are transmitted back to the master node through Ethernet. Therefore, similar to the design of the module cluster for attention calculation, there is also feature mutual transmission between rotation transformation and MLP in the point cloud feature encoding module cluster. Figure 23 This is the logical schematic diagram of the module cluster for point cloud feature encoding in the trainer part provided by the embodiment of the present invention. That is, the module cluster for point cloud feature encoding in the Learner part. Compared with the Actor part, the back - propagation of loss gradients is added.

[0433] d) The context encoding task cluster is similar to the multi - view feature fusion and multi - scale feature fusion module clusters in that they all focus on attention calculation. Therefore, it can also be divided into multi - head attention and feed - forward neural network module clusters according to their functions, and realizes centralized data parallelization based on the structure of a central node - multiple sub - nodes. Since the decision encoding and action encoding task clusters have a relatively simple module structure and fewer network layers, there is no need to further divide them. The master node of the task cluster will be directly interconnected with multiple unit nodes for task allocation and data transmission, and each unit node is connected in a central node - multiple sub - nodes structure to realize centralized data parallelization.

[0434] e) The generation replay task cluster has different model division methods. Figure 24 This is the logical schematic diagram of a variational auto - encoder decoder generation model provided by the embodiment of the present invention. That is, the VAE generation model. As Figure 24 shown, the VAE generation model is divided into an encoder and a decoder. Figure 25 This is the logical schematic diagram of a generative adversarial network generation model provided by the embodiment of the present invention. That is, the GAN generation model. The GAN generation model is divided into a discriminator and a generator, and the diffusion generation model is divided into a forward diffusion and a reverse denoising module.

[0435] During the training phase, for the VAE generative model, the encoder module cluster reads real experience samples from the master node via Ethernet, and transmits the processed feature data to the decoder part through PCIe. After the decoder completes the loss calculation, it then transmits the loss gradient back to the encoder via PCIe to complete the parameter update; for the GAN generative model, the generator module cluster takes random noise as input to generate virtual experience samples, and sends them to the discriminator through PCIe. The discriminator simultaneously reads real experience samples from the master node via Ethernet to calculate the sample error, and then updates the parameters of the discriminator through the backpropagation algorithm. Further, the parameters of the discriminator are fixed. The generator regenerates virtual samples and sends them to the discriminator, and then the discriminator transmits the discrimination probability back to the generator to calculate the generation loss, and then updates the parameters of the generator through the backpropagation algorithm. Figure 26 This is a logical schematic diagram of a diffusion generative model provided by an embodiment of the present invention. For the diffusion generative model, the forward diffusion module cluster reads real experience samples from the master node via Ethernet, gradually adds a small amount of Gaussian noise to the data at a certain step size, and then transmits the noisy sample data to the reverse denoising part through PCIe, and adjusts the denoising model parameters continuously to enable it to have the denoising ability.

[0436] During the generation phase, the VAE decoder, the GAN generator, and the diffusion model reverse denoising module cluster all take randomly sampled Gaussian noise as input to generate virtual experience samples, and finally send the sample data to the master node via Ethernet. Generally, the experience samples should include three parts: original state data, actions, and rewards. However, the original state data includes multi-view images and complex point cloud data, and the generation difficulty is too high. To ensure the sample generation quality and save computing resources at the same time, this solution constructs implicit experience samples based on the feature space. This sample uses the latent state features after state encoding to replace the original state data, which can greatly reduce the dimension of the sample data. Only a shallow fully connected network can be used to construct the generative model to achieve data augmentation. Therefore, each module cluster can adopt a central node - multi-subnode structure to achieve centralized data parallelization. It should be noted that when using implicit experience samples to update the parameters of the decision model, the state encoding module part should be frozen and no backpropagation calculation is performed.

[0437] It can be seen that in view of the problems existing in the existing intelligent autonomous decision-making technology based on reinforcement learning when combined with FPGA applications, such as unreasonable resource allocation, difficult data synchronization, and hardware-limited computing efficiency, the decision model training and inference acceleration system based on FPGA clusters proposed by the present invention adopts a distributed parallel computing architecture design, decouples tasks such as environment interaction, model inference, and experience storage, sets multi-node parallel execution, combines data parallelization and model parallelization schemes to improve computing efficiency. Data parallelization includes various subset partitioning methods and solutions for dealing with load imbalance. Model parallelization improves device utilization through module settings. In terms of the decision model structure design, multi-layer functional modules are divided according to task and space complexity, different network structures are adopted for different state spaces and flexibly trained, and at the same time, three model structure adaptive selection strategies are proposed to optimize the model design according to the task scenario. Further, in view of the FPGA single computing node architecture and communication method, a node architecture for different computing tasks is designed to improve resource allocation efficiency, and at the same time, multiple interconnection communication methods are set to ensure efficient and accurate data transmission between different devices. Finally, for the decision model training and inference tasks, an FPGA cluster computing framework and a heterogeneous computing platform are built. A multi-level FPGA cluster node is used to build the computing framework, which improves hardware utilization and reduces system latency while ensuring load balance of each computing node, combines CPU and GPU to implement model inference and update, and improves training and inference efficiency through the cooperation of each module, enhancing the adaptability and flexibility of the system.

[0438] In terms of computing efficiency, a distributed parallel computing architecture is adopted to improve efficiency through task decoupling and multi-node parallel execution, combined with data and model parallelization. Data parallelization reasonably partitions data subsets and solves the problem of load imbalance, while model parallelization improves device utilization and throughput. In terms of resource allocation, the decision model structure design reasonably constructs each layer according to task and data characteristics, and distributes computing tasks to cluster nodes through various partitioning methods, improving FPGA resource utilization and overall computing resource utilization. In terms of data synchronization, reasonable parallelization schemes are used to ensure timely data synchronization, solve the problem of difficult data synchronization in distributed clusters, and improve model accuracy and stability. In terms of model performance, the model structure design and its adaptive selection strategy will effectively improve the decision model's understanding and adaptation ability to system dynamics and the environment, enhance accuracy, and then improve system stability and generalization ability in combination with reasonable resource allocation and training methods.

[0439] In the above embodiments, the training and inference method of the decision model is described in detail. The present invention also provides corresponding embodiments of the training and inference device of the decision model. It should be noted that the present invention describes the embodiments of the device part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware.

[0440] The training and inference device for the decision-making model provided by the embodiment of the present invention, from the perspective of functional modules, includes:

[0441] An acquisition module, configured to acquire a decision-making model and the network features of the decision-making model;

[0442] A first allocation module, configured to divide the training and inference process of the decision-making model into one or more computing tasks according to the network features of the decision-making model, and allocate different computing tasks to different cluster nodes;

[0443] A second allocation module, configured to allocate multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes to different computing nodes, and transfer the computing tasks on the cluster nodes to the computing nodes where the input data is allocated;

[0444] A determination module, configured to determine the training and inference result of the decision-making model according to the output features of the network model on the computing nodes where the input data is allocated.

[0445] Since the embodiments in the device part correspond to the embodiments in the method part, please refer to the description of the embodiments in the method part for the embodiments in the device part, and details are not described here for the time being.

[0446] Figure 27 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Based on the hardware perspective, as Figure 27 shown, the electronic device includes:

[0447] A memory 20, configured to store a computer program;

[0448] A processor 21, configured to implement the steps of the method for training and inferring the decision-making model as mentioned in the above embodiments when executing the computer program.

[0449] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array. The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0450] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the training and inference method of the decision model disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the training and inference method of the decision model mentioned above.

[0451] In some embodiments, the electronic device may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0452] Those skilled in the art can understand that Figure 27 the structure shown in

[0453] does not constitute a limitation on the electronic device, and it may include more or fewer components than those shown in the figure.

[0454] An embodiment of the present invention also provides a computer program product, including computer programs / instructions, which, when executed by a processor, implement the steps of the above-mentioned training and inference method of the decision model.

[0455] Finally, the present invention also provides a corresponding embodiment of a non-volatile storage medium. A computer program is stored on the non-volatile storage medium, and when the computer program is executed by a processor, it implements the steps recorded in the above method embodiments.

[0456] It can be understood that if the methods in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0457] The non-volatile storage medium provided by the present invention includes the above-mentioned training and inference method of the decision model, and the effect is the same.

[0458] The above has introduced in detail a training and inference method, product, electronic device, and medium of a decision model provided by the present invention. The various embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

[0459] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. A training and reasoning method for a decision model, characterized in that: Processor used for host computer in training and reasoning acceleration system; The accelerator cluster of the training and reasoning acceleration system is a cluster structure; the method comprises: Obtaining a decision model and obtaining network characteristics of the decision model; Dividing the training and reasoning process of the decision model into one or more computing tasks according to the network characteristics of the decision model, and assigning different computing tasks to different cluster nodes; Distribute the multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes to different computing nodes, and transfer the computing tasks on the cluster nodes to the computing nodes to which the input data are distributed; The training and reasoning results of the decision model are determined according to the output characteristics of the network model on the computing node to which the input data is distributed.

2. The decision model training and reasoning method according to claim 1, characterized in that: The training and reasoning process of the decision model is divided into multiple computing tasks according to the network characteristics of the decision model, and different computing tasks are assigned to different cluster nodes, including: Dividing the training and reasoning process of the decision model into one or more tasks according to the training and reasoning stages included in the decision model, and allocating the tasks to different cluster nodes of the first type; Dividing the task into one or more subtasks according to the network structure used when processing the task, and allocating different subtasks to different cluster nodes of the second type; The subtask is treated as one computing task or divided into a plurality of computing tasks according to the complexity of processing the subtask, and different computing tasks are allocated to different cluster nodes of the third type.

3. The training and reasoning method of the decision model according to claim 2, characterized in that: According to the complexity of processing the subtask, the subtask is treated as one computing task or divided into multiple computing tasks, and different computing tasks are allocated to different cluster nodes of the third type, including: When it is detected that the amount of computation when processing the subtask is less than the preset amount of computation, the subtask is regarded as a computation task, and the computation task is allocated to a cluster node of the third type; When it is detected that the amount of computation when processing the subtask is greater than or equal to the preset amount of computation, the subtask is divided into multiple computing tasks; the network structure of the network model used when processing the computing task is classified according to the complexity of the network structure of the network model used when processing the computing task; and the computing tasks corresponding to different categories of network structures are allocated to different cluster nodes of the third type.

4. The training and reasoning method of the decision model according to claim 2 is characterized in that: Obtaining the input data corresponding to the computing task on the cluster node includes: If the computing task is an initial computing task among all computing tasks, determining that the input data corresponding to the initial computing task is data sent by the processor of the host computer and received by the cluster node of the first type; If the computing task is not the initial computing task, it is determined that the input data corresponding to the computing task is the output feature obtained by computing the input data corresponding to the previous computing task of the computing task through the network model on the computing node.

5. The training and reasoning method of the decision model according to claim 4 is characterized in that: The input data corresponding to the computing tasks on the cluster nodes are split into multiple data subsets including: When it is detected that the input data corresponding to the computing task on the cluster node is image data, if the data batch size is greater than or equal to the number of input channels, the input data corresponding to the computing task on the cluster node is split along the data batch dimension to obtain multiple data subsets; if the data batch size is less than the number of input channels, the input data corresponding to the computing task on the cluster node is split along the input channel dimension to obtain multiple data subsets; When it is detected that the input data corresponding to the computing task on the cluster node is not image data, the dimension with the largest number of dimensions is obtained from the data batch dimension and multiple feature dimensions; the input data corresponding to the computing task on the cluster node is split along the dimension with the largest number of dimensions to obtain multiple data subsets.

6. The decision model training and reasoning method according to claim 5, characterized in that: In the case where it is detected that the dimension with the largest number of dimensions is the feature dimension, the step of allocating the multiple data subsets obtained by splitting the input data corresponding to the computing task on the cluster node to different computing nodes includes: Obtaining the number of computing nodes to which the preset input data is to be distributed; Obtain the result of performing a modulo operation on the dimension with the largest number of dimensions and the number of computing nodes to be allocated to determine a first number of computing nodes; Determining a second number of computing nodes according to a difference between the number of computing nodes to be allocated and the first number; Splitting the input data corresponding to the computing task on the cluster node according to the dimension with the largest number of dimensions determines a data subset of a first data size and a data subset of a second data size; wherein the first data size is determined by multiplying the result obtained by rounding down the dimension with the largest number of dimensions and the number of computing nodes to be allocated plus 1 and the data size on the remaining dimensions; the second data size is determined by multiplying the result obtained after the rounding down operation and the data size on the remaining dimensions; the remaining dimensions are dimensions other than the dimension with the largest number of dimensions among all dimensions; The data subsets of the first data size are respectively distributed to a first number of computing nodes, and the data subsets of the second data size are respectively distributed to a second number of computing nodes.

7. The decision model training and reasoning method according to claim 6, characterized in that: Determining the training and reasoning results of the decision model according to the output characteristics of the network model on the computing node to which the input data is distributed includes: Obtaining the current output features of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node; wherein the current subtask starts from the initial subtask of the training and reasoning process of the decision model; Obtain the next subtask of the current subtask according to the first order of the subtasks in the training and reasoning process of the decision model, and use the next subtask of the current subtask as the new current subtask; Transmitting the current output feature to the second type of cluster node where the current subtask is located via the third type of cluster node where the computing task corresponding to the current subtask is located; Obtain the second type of cluster node where the new current subtask is located, transmit the current output feature to the second type of cluster node where the new current subtask is located through the second type of cluster node where the current subtask is located, so as to obtain the input data corresponding to the new current subtask, and return the current output feature after the input data of the computing task corresponding to the current subtask passes through the network model on the computing node, until the current subtask is the last subtask in the training and reasoning process of the decision model, stop returning, and determine the training and reasoning result of the decision model according to the current output feature.

8. The decision model training and reasoning method according to claim 7, characterized in that: Transmitting the current output feature to the new second type cluster node where the current subtask is located via the second type cluster node where the current subtask is located comprises: Controlling the communication between the second type of cluster node where the current subtask is located and the second type of cluster node where the new current subtask is located through the first signal transmission mode and obtaining communication information; wherein the communication information at least includes status monitoring information and status feedback information; When it is detected that the communication information is normal, the current output feature is transmitted to the second type of cluster node where the new current subtask is located through the second signal transmission mode via the second type of cluster node where the current subtask is located.

9. The decision model training and reasoning method according to claim 7, characterized in that: The current output features of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node include: In the case where it is detected that the current computing task is not obtained by dividing the current subtask, the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node is obtained, and the sum of the current output features is used as the current output feature of the computing task corresponding to the current subtask after the input data passes through the network model on the computing node; When it is detected that the current subtask is divided into multiple computing tasks, the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node is obtained; the next computing task of the current subcomputing task is obtained according to the second priority order of the computing tasks in the current subtask, and the next computing task of the current computing task is used as the new current computing task; the sum of the current output features is transmitted to the third type of cluster node where the new current computing task is located via the third type of cluster node where the current computing task is located, so as to obtain the input data corresponding to the new current computing task, and return to the step of obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node, until the current computing task is the last computing task in the current subtask, stop returning, output the sum of the current output features, and use the sum of the current output features as the current output feature of the input data of the computing task corresponding to the current subtask after passing through the network model on the computing node.

10. The decision model training and reasoning method according to claim 9, characterized in that: Obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node includes: Obtaining a computing node of a first type and a computing node of a second type used when processing a data subset corresponding to a current computing task; wherein a network model is set in the computing node of the first type, and a network model is not set in the computing node of the second type; and there is one computing node of the second type; Obtain output features of each data subset corresponding to the current computing task on the corresponding computing node of the first type; Sending output features on all computing nodes of the first type to computing nodes of the second type; An aggregation operation is performed on all output features through the computing nodes of the second type to obtain the sum of current output features of the data subset corresponding to the current computing task on the corresponding computing nodes.

11. The decision model training and reasoning method according to claim 9, characterized in that: Obtaining the sum of the current output features of the data subset corresponding to the current computing task on the corresponding computing node includes: Obtaining the communication order between all computing nodes processing the current computing task; wherein the last computing node among all computing nodes arranged in the communication order is connected to the first computing node in communication; Obtain data segments of each data subset corresponding to the current computing task after being passed through the network model on the corresponding computing node and store them in a sending buffer of the computing node for sending data; Determine whether the current computing node processing the current computing task has data fragments of all other nodes; wherein the other nodes are computing nodes other than the current computing node among all computing nodes processing the current computing task; If not, then the data fragment on the current computing node is sent from the sending buffer to the receiving buffer for receiving data of the next computing node of the current computing node in accordance with the communication order; the next computing node of the current computing node is used as the new current computing node; the data of the receiving buffer of the new current computing node is aggregated with the data of the sending buffer of the new current computing node and used as the data fragment of the sending buffer of the new current computing node, and the step of determining whether there are data fragments of all other nodes in the current computing node processing the current computing task is returned; If so, the data fragments on the current computing node will be sent from the sending buffer to the receiving buffer for receiving data of the next computing node of the current computing node in accordance with the communication order; the next computing node of the current computing node will be used as the new current computing node; the data in the receiving buffer of the new current computing node will be overwritten with the data in the sending buffer of the new current computing node and used as the data fragments of the sending buffer of the new current computing node, and the step of determining whether there are data fragments of all other nodes in the current computing node processing the current computing task will be returned until it is detected that there are data fragments of all other nodes in all computing nodes processing the current computing task and the process stops returning.

12. The decision model training and reasoning method according to claim 10, characterized in that: Transmitting the sum of the current output features to the cluster node of the third type where the new current computing task is located via the cluster node of the third type where the current computing task is located comprises: Controlling the communication between the third type of cluster node where the current computing task is located and the third type of cluster node where the new current computing task is located through the first signal transmission mode and obtaining communication information; wherein the communication information at least includes status monitoring information and status feedback information; When it is detected that the communication information is normal, the sum of the current output features is transmitted via the third type of cluster node where the current computing task is located to the third type of cluster node where a new current computing task is located through the second signal transmission mode.

13. The decision model training and reasoning method according to claim 10, characterized in that: Before obtaining the output features of each data subset corresponding to the current computing task on the corresponding computing node of the first type, the method further includes: If it is detected that the load on the first type of computing nodes and the second type of computing nodes used to process the data subset corresponding to the current computing task is less than the preset load, the step of obtaining the output characteristics of each data subset corresponding to the current computing task on the corresponding first type of computing node is entered.

14. The decision model training and reasoning method according to claim 11, characterized in that: Before obtaining the communication order between all computing nodes processing the current computing task, it also includes: Obtaining load conditions on the first type of computing nodes and the second type of computing nodes used when processing the data subset corresponding to the current computing task; If it is detected that the load on the first type of computing nodes and the second type of computing nodes used to process the data subset corresponding to the current computing task is greater than or equal to the preset load, the step of obtaining the communication order between all computing nodes processing the current computing task is entered.

15. The training and reasoning method of the decision model according to any one of claims 1 to 14, characterized in that: The computing nodes include multiple types, and different types of computing nodes are used to process different types of network layer data; The step of allocating multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes to different computing nodes includes: Obtaining network layer characteristics of the computing task; Determine the target computing node based on the network layer characteristics; The multiple data subsets obtained by splitting the input data corresponding to the computing tasks on the cluster nodes are distributed to different target computing nodes.

16. The decision model training and reasoning method according to any one of claims 7 to 14, characterized in that: The first type of cluster nodes are located on the host computer; the second type of cluster nodes assigned to the subtasks corresponding to the same task and the third type of cluster nodes assigned to the same computing task are located on the same host; the third type of cluster nodes assigned to different computing tasks are located on different hosts.

17. The decision model training and reasoning method according to claim 16, characterized in that: Allocating the tasks to different cluster nodes of the first type includes: Allocating the task to different cluster nodes of the first type by a third signal transmission method; The allocating different subtasks to different cluster nodes of the second type comprises: Using the first type of cluster nodes and allocating different subtasks to different second type of cluster nodes through the third signal transmission mode; The allocating different computing tasks to different third-type cluster nodes includes: Different computing tasks are allocated to different cluster nodes of the third type by utilizing the cluster nodes of the second type and through the third signal transmission manner.

18. The decision model training and reasoning method according to claim 17, characterized in that: The training and reasoning of the decision model includes an interaction phase and a model update phase; before sending data to the first type of cluster nodes, it also includes: Obtaining data sent to a cluster node of the first type from a process of the graphics processing unit for training inference; The training inference result of determining the decision model according to the current output feature includes: receiving, in an interactive phase, a first current output feature outputted sequentially by the cluster nodes of the third type, the cluster nodes of the second type, and the cluster nodes of the first type; Storing the first current output feature in a data storage module located in a host computer; In the case where it is detected that the amount of data in the data storage module is greater than a preset amount of data, receiving, in the model updating stage, a second current output feature outputted sequentially by the cluster nodes of the third type, the cluster nodes of the second type, and the cluster nodes of the first type; Calculating a loss function based on the second current output feature to determine parameters of a decision model, so as to obtain a training inference result of the decision model; After obtaining the training and reasoning results of the decision model, the following steps are also included: The parameters of the decision model are transmitted to the cluster nodes of the first type to update the parameters of the decision model in the cluster nodes of the first type.

19. The decision model training and reasoning method according to claim 18, characterized in that: Establishing a decision model located on the first type of cluster node includes: Capturing the complexity of the decision-making task; Decision models with different structures are established on the first type of cluster nodes according to the complexity of the decision task; wherein the complexity of the decision task is positively correlated with the complexity of the decision model structure.

20. The decision model training and reasoning method according to claim 18, characterized in that: Determining the structure of the decision model includes: Acquire the operating status of the hardware platform; wherein the operating status of the hardware platform at least includes hardware resource utilization, latency and power consumption; Obtaining the current expected cumulative reward value under the current state and the current decision model structure; wherein the current state at least includes the current input data characteristics and the current operating state of the hardware platform; Determining an expected value according to a current expected cumulative reward and a current reward function; wherein the current reward function is determined by a current operating state of the hardware platform; Obtain the current decision model structure whose expected value meets the requirements; The current decision model structure whose expected value meets the requirements is used as the structure of the decision model.

21. The decision model training and reasoning method according to claim 18, characterized in that: Determining the structure of the decision model includes: The structure of multiple pre-defined decision models; Obtaining hardware performance indicators under decision models of different structures, and establishing hardware constraints based on the hardware performance indicators; wherein the hardware performance indicators at least include latency, power consumption, and resource occupancy; Obtain the fitness of decision models with different structures; Under the condition of satisfying hardware constraints, obtain the decision model structure with the greatest fitness among all decision models of the structure; The decision model structure with the greatest fitness is taken as the structure of the decision model.

22. The decision model training and reasoning method according to claim 18, characterized in that: Determining the structure of the decision model includes: Pre-set structures of various decision models and features of extracted input data; Predicting the probability of selecting decision models of different structures according to the characteristics of the input data; Obtain the structure of the decision model with the highest probability among all decision models of the structure; The structure of the decision model with the highest probability is taken as the structure of the decision model.

23. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the training reasoning method of the decision model described in any one of claims 1 to 22 are implemented.

24. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the training and reasoning method for the decision model as described in any one of claims 1 to 22 when executing the computer program.

25. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the training and reasoning method of the decision model as described in any one of claims 1 to 22 are implemented.

Citation Information

Patent Citations

  • Enterprise group distributed decision-making method, device and equipment and storage medium

    CN117852745A

  • Distributed training method, device and equipment based on heterogeneous equipment and medium

    CN118733282A