Heterogeneous computing acceleration method and system based on deep learning framework network
By collecting equipment characteristic data in heterogeneous computing environments and optimizing task decomposition of DNA encoding genetic recombination algorithms, the problems of unbalanced resource utilization and inefficient equipment synchronization are solved, and efficient computing acceleration and resource utilization in heterogeneous computing environments are achieved.
Patent Information
- Application Number
- CN202510920161.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The prior art has problems such as uneven resource utilization, unoptimized task decomposition, inefficient equipment synchronization and poor adaptability when dealing with highly heterogeneous computing environments, resulting in waste of computing resources and inefficient execution, especially in edge computing scenarios.
By collecting and analyzing performance parameters of multiple types of computing devices in heterogeneous computing environments, an accurate device characteristic data set is formed, and a task decomposition is optimized by using DNA encoding and genetic recombination algorithms, combining real-time monitoring and dynamic adjustment mechanisms to achieve intelligent matching and load balancing between tasks and devices.
It improves the calculation acceleration ratio in heterogeneous environments, improves resource utilization and execution efficiency, can dynamically adapt to equipment performance fluctuations and load changes, and achieves efficient computing task allocation.
Smart Images

Figure CN120429123B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a heterogeneous computing acceleration method and system based on a deep learning framework network. Background Art
[0002] With the widespread adoption of deep learning technologies, computing demands are growing exponentially, driving the increasing use of heterogeneous computing environments. Current heterogeneous computing environments typically include a variety of computing devices, including central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), and field-programmable gate arrays (FPGAs), each with distinct computing characteristics and advantages. Traditional deep learning frameworks such as TensorFlow and PyTorch have implemented basic heterogeneous computing support, including distributed training methods such as data parallelism and model parallelism, as well as hardware-specific optimization operators. Parameter server architectures and decentralized training methods are also widely used in large-scale distributed training to improve training efficiency and resource utilization.
[0003] Existing technologies have obvious shortcomings when dealing with highly heterogeneous computing environments. First, most frameworks adopt static resource allocation strategies and cannot dynamically adapt to changes in device performance during training, resulting in uneven resource utilization. Second, existing task division methods are mainly based on empirical rules or simple heuristic algorithms, lacking global optimization search capabilities, making it difficult to find the optimal solution in the vast task decomposition solution space, and failing to fully tap the computing potential of heterogeneous environments. Third, traditional methods often use greedy or local search strategies for task decomposition, which can easily fall into local optimality and fail to achieve the goal of global minimization of execution time. In addition, existing task decomposition algorithms lack adaptive learning capabilities and cannot continuously optimize decomposition strategies based on historical execution experience, resulting in poor adaptability when facing new computing modes. Finally, the synchronization mechanism between devices with different performance is inefficient, and high-performance devices often need to wait for low-performance devices, resulting in a waste of computing resources, especially in resource-constrained edge computing scenarios. Summary of the Invention
[0004] This application provides a heterogeneous computing acceleration method and system based on a deep learning framework network, which is used to achieve adaptive task decomposition and resource allocation, as well as a load balancing mechanism in a highly heterogeneous computing environment based on the computing characteristics of deep learning tasks and the physical constraints of hardware devices, thereby maximizing the utilization of heterogeneous computing resources and improving the execution efficiency of deep learning tasks.
[0005] In the first aspect, the present application provides a heterogeneous computing acceleration method based on a deep learning framework network, and the heterogeneous computing acceleration method based on the deep learning framework network includes: collecting and analyzing performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a heterogeneous computing device characteristic data set; receiving computing task input, parsing the operation sequence of the computing task, and obtaining a task operation sequence; based on the heterogeneous computing device characteristic data set, encoding the task operation sequence into a gene sequence, optimizing the task decomposition through a genetic recombination algorithm to minimize the execution time, and obtaining a subtask set marked with acceleration characteristics; matching and analyzing the subtask set marked with acceleration characteristics with the heterogeneous computing device characteristic data set to obtain a task allocation plan; deploying the subtask set to the corresponding heterogeneous computing device according to the task allocation plan to obtain a distributed execution computing framework; monitoring the running status of the distributed execution computing framework in real time to obtain device load balancing data, and dynamically adjusting the heterogeneous computing resource allocation based on the device load balancing data to obtain a heterogeneous environment computing result with improved acceleration ratio.
[0006] In a second aspect, the present application provides a heterogeneous computing acceleration system based on a deep learning framework network, wherein the heterogeneous computing acceleration system based on a deep learning framework network includes:
[0007] The acquisition module is used to collect and analyze the performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a characteristic data set of heterogeneous computing devices;
[0008] An analysis module is used to receive a computing task input, perform operation sequence analysis on the computing task, and obtain a task operation sequence;
[0009] a mapping module for encoding the task operation sequence into a gene sequence based on the heterogeneous computing device characteristic data set, optimizing task decomposition by a genetic recombination algorithm to minimize execution time, and obtaining a subtask set marked with acceleration characteristics;
[0010] A matching module, configured to perform matching analysis on the subtask set marked with acceleration characteristics and a heterogeneous computing device characteristic data set to obtain a task allocation solution;
[0011] A deployment module, configured to deploy the subtask set onto corresponding heterogeneous computing devices according to the task allocation scheme to obtain a distributed execution computing framework;
[0012] The monitoring module is used to monitor the running status of the distributed execution computing framework in real time, obtain device load balancing data, and dynamically adjust the allocation of heterogeneous computing resources based on the device load balancing data to obtain heterogeneous environment computing results with improved acceleration ratio.
[0013] In a third aspect, a heterogeneous computing acceleration device based on a deep learning framework network is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the heterogeneous computing acceleration device based on the deep learning framework network executes the above-mentioned heterogeneous computing acceleration method based on the deep learning framework network.
[0014] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned heterogeneous computing acceleration method based on a deep learning framework network.
[0015] The technical solution provided in this application meticulously collects and analyzes the performance parameters of multiple types of computing devices in a heterogeneous computing environment to generate a precise dataset of heterogeneous computing device characteristics, thus avoiding the irrational resource allocation issues caused by insufficient understanding of device characteristics in traditional approaches. By analyzing the operation sequence of computing tasks, the system accurately captures the internal operation flow relationships and execution dependencies within the task, providing a structured foundation for task decomposition and addressing the existing issues of inadequate task decomposition granularity or unclear handling of operation relationships. Specifically, the system introduces a DNA encoding and genetic recombination optimization mechanism that combines the global search capabilities of biological evolutionary algorithms with the optimization goal of minimizing execution time. This mechanism transforms abstract task operation sequences into concrete genetic encoding representations and intelligently optimizes task decomposition solutions by simulating biological evolutionary processes, demonstrating the innovative application value of biologically inspired algorithms in the field of heterogeneous computing. The subtask-device matching analysis based on this mechanism not only considers computational efficiency but also incorporates multiple factors such as physical resource constraints and communication overhead. The resulting task allocation solution is more comprehensive and balanced, surpassing traditional allocation strategies based on single metrics. The distributed execution computing framework established during the deployment phase enables efficient collaboration among heterogeneous devices, reduces communication overhead, and improves overall execution efficiency. Crucially, this solution implements real-time monitoring and dynamic adjustment of the distributed execution framework, enabling timely adjustments to resource allocation strategies based on operational status, overcoming the drawback of static allocation solutions that struggle to cope with performance fluctuations. This dynamic adaptability enables the entire system to maintain efficient operation despite complex conditions such as load changes and fluctuating device performance, effectively improving computational speedup in heterogeneous environments. Overall, by applying bio-inspired algorithms (DNA encoding and genetic recombination optimization) to heterogeneous computing resource scheduling, the solution leverages the unique contributions of evolutionary algorithms to resource matching, achieving intelligent matching between computing tasks and hardware resources, and significantly improving the execution efficiency and resource utilization of deep learning frameworks in heterogeneous environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 This is a schematic diagram of an embodiment of a heterogeneous computing acceleration method based on a deep learning framework network in an embodiment of the present application;
[0018] Figure 2 This is a schematic diagram of an embodiment of a heterogeneous computing acceleration system based on a deep learning framework network in an embodiment of the present application;
[0019] Figure 3 It is a schematic block diagram of the structure of a heterogeneous computing acceleration device based on a deep learning framework network in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The embodiments of the present application provide a heterogeneous computing acceleration method and system based on a deep learning framework network. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.
[0021] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a heterogeneous computing acceleration method based on a deep learning framework network includes:
[0022] Step S101: collecting and analyzing performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a heterogeneous computing device characteristic data set;
[0023] Step S102: receiving a computing task input, performing an operation sequence analysis on the computing task, and obtaining a task operation sequence;
[0024] Step S103: Based on the heterogeneous computing device characteristic data set, the task operation sequence is encoded into a gene sequence, and the task decomposition is optimized by a genetic recombination algorithm to minimize the execution time, thereby obtaining a subtask set marked with acceleration characteristics;
[0025] Step S104: performing matching analysis on the subtask set marked with acceleration characteristics and the heterogeneous computing device characteristic data set to obtain a task allocation plan;
[0026] Step S105: deploy the subtask set to the corresponding heterogeneous computing devices according to the task allocation plan to obtain a distributed execution computing framework;
[0027] Step S106: Monitor the running status of the distributed execution computing framework in real time to obtain device load balancing data, and dynamically adjust the heterogeneous computing resource allocation based on the device load balancing data to obtain a heterogeneous environment computing result with improved acceleration ratio.
[0028] It is understandable that the execution subject of this application can be a heterogeneous computing acceleration system based on a deep learning framework network, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0029] Specifically, the performance parameters of multiple types of computing devices in a heterogeneous computing environment are collected and analyzed. Device scanning identifies all computing devices in the environment, including CPUs, GPUs, tensor processors, and field-programmable gate arrays (FPGAs). Key parameters such as the specific model, number of compute cores, clock frequency, memory capacity, memory bandwidth, and cache size are recorded for each device. Performance benchmarks are then conducted on each device using standardized compute-intensive, memory-intensive, and communication-intensive tasks. For a specific deep learning framework platform, performance data recorded for a typical GPU device when executing a 1024×1024 matrix multiplication includes completion time, throughput, and memory usage. The same operation performed significantly differently on a CPU. This data is organized into device performance feature vectors. A topological map of device nodes and communication links is then constructed, mapping device parameters to operation execution time. This yields a dataset of heterogeneous computing device characteristics encompassing both computational power and communication costs.
[0030] When receiving computing task input and parsing the operation sequence, the system receives user-submitted computing tasks in standardized framework formats, such as deep learning models in TensorFlow or PyTorch. The system then parses the computing task to extract its initial computational graph structure, which contains all computing nodes and their connections. For each computing node, the system extracts its operation type, input and output tensor dimensions, computational complexity, and memory requirements. For a deep neural network model, for example, the parsing process identifies convolutional layers requiring a large number of matrix multiplications, while fully connected layers require significant memory and network bandwidth. The system then analyzes data dependencies between computing nodes to form a data flow graph. A comprehensive analysis of the computational graph is performed to identify computational bottlenecks and opportunities for parallel execution. Optimization strategies are then applied, including operation fusion, memory optimization, computational graph rewriting, precision adjustment, and operator replacement. Finally, the operation sequence is extracted according to execution order, organizing the computing node's operation type, parameter configuration, and execution order into a linear sequence structure to obtain the task operation sequence.
[0031] Based on a dataset of heterogeneous computing device characteristics, task operation sequences are genetically encoded and optimized. First, for each operation node in the task operation sequence, features such as operation type, data dimension, computational density, memory access pattern, and degree of parallelism are extracted to form an operation feature vector. Gene encoding mapping rules are established to map different operation types, such as convolution, pooling, and fully connected, into quaternary gene segments. For example, convolution operations are encoded as "ATCG" and pooling operations are encoded as "TACG." Operation parameters, such as kernel size and step size, are encoded as gene expression sequences. A genetic algorithm framework is constructed, with minimizing total execution time as the fitness function. A population containing multiple task decomposition schemes is initialized, with each individual representing a strategy for decomposing the original task into subtasks. A fitness score is calculated for each individual in the initial evolutionary population based on the heterogeneous computing device characteristic dataset. Simulations are performed to evaluate the execution efficiency of different decomposition schemes on devices such as GPUs, CPUs, and TPUs. A genetic recombination operator is applied to crossover high-fitness individuals, such as by exchanging some gene segments from two excellent decomposition schemes to generate new task decomposition combinations. Mutation operations are also performed on some individuals to explore new decomposition possibilities. The fitness evaluation and genetic recombination process is repeated until the algorithm converges, and the individual with the highest fitness is selected as the optimal decomposition scheme. Its gene sequence is decoded into subtask divisions to obtain a set of subtasks marked with acceleration characteristics.
[0032] A matching analysis is performed between a set of subtasks marked with acceleration characteristics and a dataset of heterogeneous computing device characteristics. A subtask-device affinity matrix is constructed to quantify the execution affinity between subtasks and devices. The matrix elements are normalized. When the execution time of a subtask on a device exceeds a preset threshold, the corresponding element value is set to zero; otherwise, it is assigned a value based on the degree of matching. For example, convolution-intensive subtasks have a high affinity for GPUs and a low affinity for CPUs, which is reflected in the corresponding values in the affinity matrix. An integer linear programming model is then constructed, with the objective function being to minimize total execution time. Constraints include that each subtask must be assigned to exactly one device and that device resource capacity is limited. A branch-and-bound algorithm is used to solve the model. If the computational scale is too large, a Lagrangian relaxation method is used. After obtaining a preliminary task allocation solution, the data transmission cost between subtasks is calculated, a communication overhead graph is constructed, and the amount of data transferred is plotted. Finally, task allocation is optimized by minimizing the total communication delay, taking into account the physical connection topology between devices. When communication delay exceeds the computational benefit, the subtask allocation position is adjusted to obtain the optimal task allocation solution.
[0033] According to the task allocation plan, a collection of subtasks is deployed to the corresponding heterogeneous computing devices. A task execution engine, comprising a task queue manager, memory manager, execution scheduler, and performance monitor, is deployed on each heterogeneous computing device. Parameters for the subtask collection are initialized according to the task allocation plan, and the computation parameters and initial weights are converted to a device-compatible format. For example, when a convolution subtask is assigned to a GPU, its weights must be converted to a GPU-compatible storage format and precision type. Execution memory is allocated to each heterogeneous computing device, the memory space required by the subtasks is calculated, and data buffers are established to generate a memory allocation table. Based on the data transfer requirements between the memory allocation table and the subtask collection, inter-subtask communication channels are established, and data transfer interfaces are established for subtasks with data transfer relationships. A distributed computing schedule is created based on the communication topology, including the subtask execution order, synchronization point locations, and data transfer timing. Finally, the global execution sequence is converted into device-level instructions, deployed to the corresponding heterogeneous computing devices, and the execution engine is started, resulting in a distributed execution computing framework.
[0034] The distributed computing framework's operational status is monitored in real time. Performance monitors on each heterogeneous computing device collect parameters such as computing load, memory usage, execution progress, and communication latency to form a heterogeneous device state vector. These state vectors are analyzed using a sliding time window to calculate the average execution time and time standard deviation for each device, generating device performance statistics. Based on these statistics, a performance prediction model is constructed to predict the future performance trends of each device and identify performance bottlenecks. A four-layer adjustment strategy is implemented for load-unbalanced computing devices: internal device parameter optimization, load redistribution, subtask reconfiguration, and global reoptimization. When a GPU device is detected to be consistently highly loaded while a CPU device is less loaded, load redistribution is triggered, migrating subtasks suitable for CPU processing from the GPU to the CPU. Time-window-based parameter aggregation is performed on the relocated subtasks, and a weighted average is calculated using the number of samples processed by the device as a weight coefficient. Finally, knowledge from high-performance devices is transferred to low-performance devices, with transfer priority determined by adjusting the ratio of knowledge importance to communication cost. This results in improved speedup for heterogeneous computing environments.
[0035] In this application, by meticulously collecting and analyzing the performance parameters of multiple types of computing devices in a heterogeneous computing environment, an accurate dataset of heterogeneous computing device characteristics is generated, avoiding the problem of irrational resource allocation caused by insufficient understanding of device characteristics in traditional methods. By parsing the operation sequence of computing tasks, the internal operation flow relationships and execution dependencies of the tasks are accurately captured, providing a structured foundation for task decomposition, and addressing the problems of inadequate task decomposition granularity or unclear operation relationship handling in existing technologies. Specifically, a DNA encoding and genetic recombination optimization mechanism is introduced, which combines the global search capabilities of biological evolution algorithms with the optimization goal of minimizing execution time. This mechanism transforms abstract task operation sequences into concrete genetic encoding representations and intelligently optimizes task decomposition solutions by simulating biological evolutionary processes, demonstrating the innovative application value of biologically inspired algorithms in the field of heterogeneous computing. The subtask and device matching analysis based on this mechanism not only considers computational efficiency but also incorporates multiple factors such as physical resource constraints and communication overhead. The resulting task allocation solution is more comprehensive and balanced, surpassing traditional allocation strategies based on single indicators. The distributed execution computing framework established during the deployment phase enables efficient collaboration between heterogeneous devices, reduces communication overhead, and improves overall execution efficiency. Crucially, this solution implements real-time monitoring and dynamic adjustment of the distributed execution framework, enabling timely adjustments to resource allocation strategies based on operational status, overcoming the drawback of static allocation solutions that struggle to cope with performance fluctuations. This dynamic adaptability enables the entire system to maintain efficient operation despite complex conditions such as load changes and fluctuating device performance, effectively improving computational speedup in heterogeneous environments. Overall, by applying bio-inspired algorithms (DNA encoding and genetic recombination optimization) to heterogeneous computing resource scheduling, the solution leverages the unique contributions of evolutionary algorithms to resource matching, achieving intelligent matching between computing tasks and hardware resources, and significantly improving the execution efficiency and resource utilization of deep learning frameworks in heterogeneous environments.
[0036] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0037] Automatically identify CPUs, GPUs, tensor processors, and FPGAs in heterogeneous computing environments and obtain a list of device types.
[0038] Performing standardized computing-intensive tasks, memory-intensive tasks, and communication-intensive tasks on each computing device in the device type list to obtain performance benchmark test results;
[0039] Based on the performance benchmark test results, the number of computing cores, clock frequency, memory capacity, memory bandwidth, and cache size parameters of each device are recorded to obtain the device performance feature vector;
[0040] Based on the device performance feature vector, a heterogeneous network topology diagram including device nodes and communication links is constructed to obtain the network relationship data between the device nodes and communication links;
[0041] Based on the network relationship data, a polynomial regression method is used to establish a mapping relationship between device parameters and deep learning operation execution time, and a mathematical model of device performance is obtained;
[0042] A communication cost model is established by combining the mathematical model of device performance with the data transmission overhead between devices, and a characteristic dataset of heterogeneous computing devices including computing power and communication cost is obtained.
[0043] Specifically, the device discovery protocol is used to scan computing devices on the network. The device discovery protocol uses multicast DNS and service discovery mechanisms to broadcast query requests to the network, receive responses from each device, and record its IP address and device type. For directly connected devices, automatic detection is performed through the hardware interface. In a Linux environment, the lspci command is used to identify PCI devices. For NVIDIA GPUs, the nvidia-smi tool is used to obtain detailed information. TPUs are detected through a dedicated API, and FPGAs are identified using the toolchain provided by the manufacturer. Each identified device is assigned a unique ID and recorded in a device type list, which contains basic information such as device type, manufacturer, model, and connection method.
[0044] A standardized set of performance test tasks was executed for each computing device in the device type list. Compute-intensive tasks included generalized matrix multiplication (GEMM) operations, setting up matrices of varying dimensions and recording completion times and floating-point operations per second. Memory-intensive tasks performed large-scale data read and write operations, measuring memory throughput and access latency, including both sequential and random access modes. Communication-intensive tasks measured inter-device data transfer capabilities, testing both point-to-point transmission and collective communication performance. Each test task was executed multiple times and averaged to eliminate the influence of random factors. For example, in the GEMM test, each device performed matrix multiplication of the same dimensions and recorded completion times, directly reflecting differences in computing power. After testing, all test results were compiled into a performance benchmark result dataset. Based on the performance benchmark results, key hardware parameters for each device were further extracted and recorded. The number of compute cores was obtained through device query APIs or operating system commands, such as the number of CUDA cores on a GPU. Clock frequency recorded the processor's main frequency in GHz. Memory capacity recorded the total available memory on the device in GB. Memory bandwidth was calculated using the memory-intensive tasks in the benchmark test, in GB / s. Cache size recorded the capacity of each cache level. In addition, special features such as the number of dedicated function units and supported precision types are recorded. All of these parameters are combined to form a performance characteristic vector for each device. For a particular GPU, an example performance characteristic vector includes key metrics such as 5120 CUDA cores, a base frequency of 1.4 GHz, 32 GB of video memory, 900 GB / s of video memory bandwidth, and 6 MB of L2 cache.
[0045] A heterogeneous network topology is constructed based on device performance feature vectors. The physical connection relationships between devices are identified, and the physical connection methods and versions between devices are obtained using the operating system's network interface query tool. Bandwidth and latency tests are performed on each link, and the actual transmission performance is recorded. A directed graph is then constructed to represent the inter-device communication topology. Nodes in the graph represent computing devices, and edges represent communication links. Each node stores the performance feature vector for the corresponding device, and each edge is annotated with bandwidth and latency information. For multi-node environments, an adjacency matrix is used to store the topology. Matrix elements represent the communication performance between the corresponding devices, and a value of 0 indicates no direct connection. This method generates network relationship data between device nodes and communication links. Based on this network relationship data, a polynomial regression method is used to establish a mapping between device parameters and deep learning operation execution time. For common operations such as convolution, pooling, and fully connected operations, a series of different parameter configurations are selected as training samples. These operations are executed on each device and the execution time is recorded. For convolution operations, for example, the input features include the input tensor size, filter size, and number of channels, and the output is the execution time. A polynomial regression model is then trained for each type of operation. The model input is a combination of the operation parameters and the device feature vector, and the output is the predicted execution time. The model is trained using the least squares method, with polynomials of different orders selected for operations of varying complexity. Typically, quadratic or cubic polynomials are sufficient to capture the performance characteristics of most operations. After training, a mathematical model of device performance is obtained, which can predict the execution time of any operation on a specific device.
[0046] A communication cost model is established by combining a mathematical model of device performance with the overhead of inter-device data transmission. This model considers data size, communication link characteristics, and communication mode to calculate the time required for data transmission. For two computing tasks that require communication, if they are assigned to the same device, the communication cost is zero; otherwise, the transmission time is calculated based on the connection information in the topology graph. Considering the data transmission characteristics of deep learning training, the communication cost model specifically focuses on the communication overhead of gradient synchronization and parameter updates. The mathematical model of device performance and the communication cost model are then integrated to form a heterogeneous computing device characteristic dataset that includes the performance characteristics of all devices and the communication characteristics between them.
[0047] For example, training a ResNet model requires executing in a heterogeneous environment consisting of CPUs, GPUs, and TPUs. The heterogeneous computing device characterization dataset obtained through the above steps shows that: CPUs excel at processing small batches of serial operations, with feature vectors indicating multi-core architectures and large memory capacities; GPUs excel at massively parallel floating-point operations, with feature vectors indicating thousands of computing cores and high memory bandwidth; and TPUs are optimized for matrix operations, with feature vectors reflecting their matrix processing unit architecture. A mathematical model of device performance predicts that convolutional layers execute 10 times faster on GPUs than on CPUs, while TPUs are even faster than GPUs. Communication cost models indicate that the CPU and GPU are connected via PCIe with a bandwidth of 16 GB / s, while the TPU has a higher bandwidth via a dedicated interconnect. The task allocation process optimizes the allocation of convolutional layers to the TPU and certain specialized operations to the CPU, achieving optimal overall execution efficiency.
[0048] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0049] Receive computing tasks in a standardized framework format submitted by users, parse the computing tasks, and obtain the initial computing graph structure;
[0050] Extract the operation type, input and output tensor dimensions, computational complexity, and memory requirement characteristics of all computing nodes in the initial computational graph structure to obtain a node feature set;
[0051] Analyze and calculate the data dependency between nodes based on the node feature set to obtain a data flow graph;
[0052] Analyze the data flow graph for computational bottlenecks, data dependencies, and parallel execution opportunities to identify computational graph optimization opportunities.
[0053] Based on the computational graph optimization opportunity, operation fusion, memory optimization, computational graph rewriting, precision adjustment, and operator replacement are applied to the initial computational graph structure to obtain an optimized computational graph;
[0054] The operation sequence of the optimized calculation graph is extracted according to the execution order, and the operation type, parameter configuration, and execution order of the calculation node are arranged into a linear sequence structure to obtain the task operation sequence.
[0055] Specifically, standardized framework formats refer to the common intermediate representation formats used by mainstream deep learning frameworks such as TensorFlow, PyTorch, MXNet, or ONNX. When a user submits a model such as ResNet-50 or BERT, the receiving module verifies the file integrity and then identifies the framework type based on the file extension or header information. The corresponding parser is invoked for each framework type, such as the TensorFlow parser for TensorFlow models and the ONNX parser for ONNX models. The parser reads the model file contents and extracts the computational graph definition, including all operation nodes and their connections. The parsed result is represented as a graph data structure, where nodes represent computational operations and edges indicate data flow. This representation is called the initial computational graph structure.
[0056] Each node in the computational graph is traversed one by one, extracting the operation type—that is, the specific operation performed by the node. The input and output tensor dimensions are then extracted. For example, for a convolutional layer, the input feature map dimensions (batch size, height, width, number of channels) and the output feature map dimensions are recorded. The computational complexity of the operation is then calculated, typically expressed in floating-point operations (FLOPs). For example, the FLOPs for a standard convolutional layer are calculated as: output height × output width × number of output channels × kernel height × kernel width × number of input channels. The memory requirements of the operation are also estimated, including the total storage space required for the input tensors, output tensors, and parameters. Other characteristics of the operation are also extracted and recorded, such as whether it supports parallelization and its computational intensity, which can be calculated as the ratio of computational effort to data volume. All of this feature information is combined to form a feature set for each node, forming the node feature set.
[0057] The processing begins by identifying the producer-consumer relationships between nodes—that is, which node's output serves as the input for another node. By traversing all nodes, the producer node and all consuming nodes of each tensor are recorded, establishing a mapping relationship between tensors and nodes. Based on this mapping relationship, a directed graph is constructed to represent the data flow. Nodes represent computational operations, and directed edges indicate the direction of data flow. The edges are annotated with information about the transferred tensors, including their dimensions and data types. Taking into account the branching and merging patterns in deep learning models, special attention is paid to nodes with multiple inputs and multiple outputs. For nodes with multiple inputs, their data flow relationships are analyzed to determine the execution preconditions. The resulting data flow graph clearly illustrates the data flow relationships within the computational task.
[0058] A comprehensive data flow graph analysis is performed to identify optimization opportunities across multiple dimensions. Computational bottleneck analysis is performed to identify compute-intensive nodes. This is done by calculating the theoretical execution time of each node and ranking them. The top-ranked nodes are identified as potential bottlenecks. Subsequently, data flow relationship analysis is performed to identify the critical path—the longest execution path from input to output, which determines the lower bound on the overall execution time. Parallel execution opportunity analysis constructs a parallelism model for the task graph to identify the set of nodes that can execute simultaneously. This approach uses the concept of "ready time" for a node. A node is considered ready when all its inputs are available. Multiple nodes that are simultaneously ready can execute in parallel. Memory access pattern analysis is also performed to identify memory-intensive operations and memory bottlenecks. Through these multi-dimensional analyses, various optimization directions for the computation graph are comprehensively evaluated to form a set of computation graph optimization opportunities.
[0059] Based on the identified graph optimization opportunities, a series of optimizations are applied to transform the initial graph. Operation fusion combines multiple consecutive small operations into a single larger operation, reducing intermediate result storage and data transfer. For example, three consecutive Conv+BatchNorm+ReLU operations are fused into a CombinedConvBNReLU operation. Memory optimization uses in-place operations to reduce memory allocations and reuse already allocated memory space. For example, the intermediate results of two consecutive Reshape operations are shared. Graph rewriting replaces the original computational pattern with an equivalent but more efficient one. For example, multiple small convolutions are replaced with a single large convolution, or a fully connected layer is rewritten as a 1x1 convolutional layer. Precision adjustment selectively reduces the computational precision of certain operations based on their importance and device characteristics, for example, reducing the precision of non-critical layers from FP32 to FP16 or INT8. Operator replacement replaces common operators with hardware-specific optimized implementations. For example, replacing a standard GEMM implementation with a GPU-optimized cuBLAS implementation. By applying these optimizations, we obtain an optimized computation graph that has higher execution efficiency while maintaining the functional equivalence of the original.
[0060] Extracting operation sequences from the optimized computation graph to form a linear execution structure is the foundation of task serialization. Topological sorting is performed based on data flow relationships to ensure that all dependent input nodes have completed execution before a node is executed. The topological sorting algorithm selects nodes with zero in-degree (nodes without dependencies) from the graph and adds them to the execution sequence. This node and its associated edges are then removed, and this process is repeated until all nodes have been added to the sequence. For multiple nodes with zero in-degree, a heuristic algorithm is used to prioritize them, taking into account factors such as the node's computational complexity, its location on the critical path, and data locality with other nodes. Based on parallel execution opportunities, groups of nodes eligible for parallel execution are identified, and each group is assigned a parallelism level. Finally, the computational node's operation type, parameter configuration, and execution order are organized into a linear sequence structure to form a task operation sequence. This sequence contains not only the original computational operation information but also execution priority, parallel group information, and estimated execution time, providing structured input data for subsequent DNA encoding.
[0061] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0062] Extract the operation type, data dimension, computation density, memory access mode, and parallelism of each operation node in the task operation sequence to obtain the operation feature vector;
[0063] Establish gene coding mapping rules, map different operation types into quaternary gene segments, encode operation parameters into gene expression sequences, convert operation feature vectors into standardized gene coding, and obtain task gene sequences;
[0064] Construct a genetic algorithm framework, set the minimization of total execution time as the fitness function, initialize a population containing multiple task decomposition schemes, each individual represents a task decomposition strategy, and obtain the initial evolutionary population;
[0065] The fitness score of each individual in the initial evolutionary population is calculated based on the characteristic data set of heterogeneous computing devices, and the execution efficiency of different decomposition schemes in heterogeneous environments is evaluated to obtain the population fitness distribution;
[0066] Apply genetic recombination operators to perform crossover operations on high-fitness individuals to generate new task decomposition combinations, perform mutation operations on some individuals to explore new decomposition possibilities, and obtain a new generation of evolutionary populations;
[0067] The fitness evaluation and genetic recombination process is repeated until the algorithm converges, and the individual with the highest fitness is selected as the optimal decomposition scheme. Its gene sequence is decoded into subtask divisions to obtain a set of subtasks marked with acceleration characteristics.
[0068] Specifically, each operation node in the task operation sequence is traversed to extract the operation type, such as "Conv" for convolution, "MatMul" for matrix multiplication, and "Pool" for pooling. This information is directly obtained from the node definition of the operation sequence. Next, data dimensionality information is extracted, including the shapes of the input and output tensors, such as the input dimensions (batch size, number of channels, height, width) and output dimensions of the convolutional layer. Computational density refers to the ratio of the amount of computation to the amount of data in an operation. It is calculated by dividing the number of floating-point operations performed by the total number of bytes of input and output data. A high computational density indicates that the operation is suitable for devices with high computing power. Memory access pattern analyzes how the operation accesses memory, including characteristics such as continuous or random access, read-write ratio, and reuse. This is obtained by analyzing the implementation details of the operation. Parallelism indicates the degree to which the operation can be parallelized. It is evaluated based on the operation type and algorithmic characteristics. For example, matrix multiplication has high parallelism, while certain order-dependent operations have low parallelism. All these features are combined into an operation feature vector, which comprehensively characterizes the computational characteristics of the operation node.
[0069] Establishing gene encoding mapping rules is the core step in converting operation sequences to gene sequences. This mapping rule comprises three levels of encoding: the operation type encoding layer maps different computational operations into quaternary gene fragments, such as "ATCG" for convolution, "TACG" for matrix multiplication, "CGAT" for pooling, and "GCAT" for activation functions, ensuring that each operation type has a unique quaternary identifier. The operation parameter encoding layer encodes key operation parameters, such as kernel size, stride, and padding, into gene expression sequences by converting numerical parameters into quaternary representations. For example, a kernel size of 3×3 is encoded as "AT-AT" and a stride of 2 is encoded as "TC." The feature vector normalization layer normalizes the numerical features extracted from the operation feature vector to the range [0, 1] and then quantizes them into quaternary values to form a standardized gene encoding. Through the combined application of these three encoding rules, the entire task operation sequence is converted into a structured task gene sequence. This sequence retains all key information about the original operation while maintaining the data format for genetic operations.
[0070] Constructing a genetic algorithm framework is the core mechanism for achieving task decomposition optimization. This framework uses minimizing total execution time as its fitness function and searches for the optimal task decomposition strategy by simulating biological evolution. During the initialization phase, a population containing multiple task decomposition schemes is randomly generated. Each individual represents a strategy for splitting the original task gene sequence into multiple subtasks. The choice of split points and the combination of subtasks determine the individual's genotype. The population size is typically set between 50 and 200 individuals to ensure sufficient genetic diversity. Each individual contains two components: a split point gene and a subtask combination gene. The split point gene determines the position in the gene sequence where the split occurs, while the subtask combination gene determines whether adjacent operations should be merged into a single subtask. This encoding method forms an initial evolutionary population rich in genetic information.
[0071] Evaluating the fitness of each individual in the initial evolutionary population is a key step in driving the evolutionary process. The fitness evaluation process first decodes the individual's genotype into a specific task decomposition scheme, determining the number of subtasks, the operations contained in each subtask, and the data transfer relationships between subtasks. The expected execution time of each subtask on different devices is then calculated based on a dataset of heterogeneous computing device characteristics, taking into account factors such as the device's computing power, memory capacity, and communication bandwidth. The data transfer overhead between subtasks is then evaluated, calculating the communication delay between different devices in the heterogeneous environment. Combining the computational time and communication overhead, the total execution time of the decomposition scheme is calculated as the fitness score, with shorter execution times indicating higher fitness. By evaluating the fitness of all individuals, a population fitness distribution is formed, providing a basis for subsequent selection and genetic operations.
[0072] Applying genetic recombination operators is a core step in generating a new generation of optimization solutions. The crossover operation selects two individuals with high fitness as parents and exchanges gene segments through single-point crossover, two-point crossover, or uniform crossover, producing offspring individuals that inherit the optimal characteristics of both parents. For example, single-point crossover randomly selects a crossover point in the gene sequences of the two parents and exchanges the gene segments after the crossover point, generating two new task decomposition combinations. The mutation operation randomly modifies some individuals, including changing the location of the decomposition point, adjusting the subtask combination method, and modifying local gene segments. This allows the algorithm to explore new areas in the search space and prevent premature convergence to local optimal solutions. The mutation probability is typically set to 0.01-0.1 to ensure sufficient genetic diversity while maintaining population stability. The selection operation retains outstanding individuals based on the fitness distribution. Common selection strategies include roulette wheel selection, tournament selection, or elite retention strategies to ensure that individuals with high fitness have a greater probability of being passed on to the next generation.
[0073] Repeating the fitness evaluation and genetic recombination process until the algorithm converges is an iterative process to obtain the optimal solution. Convergence criteria include reaching a preset maximum number of iterations, no significant improvement in the optimal fitness over several consecutive generations, or a decrease in population diversity below a threshold. The maximum number of iterations is typically set between 100 and 500 generations, with a fitness improvement threshold of 1% and a diversity threshold of 10% of the total number of individuals. When convergence conditions are met, the individual with the highest fitness is selected from the final population as the optimal decomposition solution. The genetic sequence of this individual is decoded into specific subtasks, and each subtask is determined to contain a list of operations, a subtask's computational feature tags, and expected execution time and resource requirements. Each subtask is tagged with acceleration characteristics, including suitable hardware type, expected speedup ratio, memory requirements, and communication mode. This forms a collection of subtasks tagged with acceleration characteristics, providing optimized input for subsequent device matching and task allocation.
[0074] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0075] Based on the set of subtasks marked with acceleration characteristics and the heterogeneous computing device characteristic dataset, a subtask-device affinity matrix is constructed, where the matrix element value represents the execution affinity between the subtask and the device, and a matching quantization matrix is obtained;
[0076] Each element in the matching quantization matrix is normalized. If the execution time of a subtask on a device exceeds a preset threshold, the corresponding element value is set to zero. If it does not exceed the threshold, the element value is assigned according to the degree of match between the subtask characteristics and the device characteristics to obtain the optimized matching matrix.
[0077] An integer linear programming model is constructed based on the optimized matching matrix. The objective function is set to minimize the total execution time. The constraints include that each subtask must be assigned to only one device and the device resource capacity is limited. This results in a resource allocation mathematical model.
[0078] Apply the branch and bound algorithm to the resource allocation mathematical model. If the calculation scale exceeds the preset limit, the Lagrangian relaxation method is used to solve it. If it does not exceed the limit, the solution is directly obtained to obtain a preliminary task allocation plan.
[0079] Based on the preliminary task allocation plan, the data transmission cost between subtasks is calculated and a data transmission plan graph is constructed. If two subtasks with a data transmission relationship are assigned to different devices, the corresponding data transmission edges are added to the graph and the amount of data transmitted is marked to obtain a communication cost graph.
[0080] According to the communication cost graph and the physical connection topology between devices, task allocation is optimized by minimizing the total communication delay. If the communication delay exceeds the computational benefit, the subtask allocation position is adjusted to obtain the task allocation solution.
[0081] Specifically, a subtask-device affinity matrix is constructed based on a set of subtasks marked with acceleration features and a dataset of heterogeneous computing device characteristics. During the construction process, each subtask marked with acceleration features is iterated over, and execution affinity is calculated for each heterogeneous computing device. Execution affinity is a measure of the subtask's efficiency on a specific device. The calculation comprehensively considers the degree of match between the subtask's compute characteristics and the device's characteristics. For each subtask-device pair, the compatibility of the subtask's primary compute characteristics with the device's characteristics is examined, and an affinity score is generated through a weighted summation. For example, compute-intensive subtasks have a high affinity for GPUs with high floating-point performance, while memory-intensive subtasks have a high affinity for devices with large memory bandwidth. Subtasks derived through genetic optimization inherit the optimal characteristics of their parent tasks, making affinity calculation more accurate. All affinity scores are organized into a matrix, with rows corresponding to subtasks and columns corresponding to devices. Each element represents the execution affinity between the corresponding subtask and the device, resulting in a matching quantization matrix.
[0082] Optimizing the matching metric matrix is a key step in ensuring feasible task allocation. Each element in the matrix is normalized, mapping the affinity scores of different subtasks and devices to a uniform range for easier comparison and decision-making. Next, execution time threshold filtering is performed. This process is based on the estimated execution time of the subtasks on the device, calculated using the device performance model. Because the subtasks are optimized using a genetic algorithm, their execution time estimates are more accurate, enabling better identification of inappropriate device assignments. If the estimated execution time exceeds a preset threshold, the corresponding matrix element is set to zero, indicating that the subtask cannot be assigned to that device. If the threshold is within the threshold, the affinity value is retained or adjusted based on the ratio of the execution time to the threshold. This threshold filtering mechanism ensures that task allocation avoids severe execution bottlenecks. Furthermore, the current load of the device is taken into consideration; if a device is nearing full capacity, the corresponding element value is reduced. After these steps, an optimized matching matrix is obtained.
[0083] Constructing an integer linear programming model based on the optimized matching matrix is a scientific method to formalize the task allocation problem. Define a decision variable to indicate whether the subtask is assigned to a specific device, and take the value of zero or one. Set the objective function to minimize the total execution time, which is essentially to minimize the execution time of the device with the longest execution time among all devices, also known as optimizing the bottleneck device. Constraints include two types: the first type of constraint is that each subtask must and can only be assigned to one device to ensure the uniqueness of task allocation; the second type of constraint is the device resource capacity limit, which means that the total demand of the subtask for device resources cannot exceed the device capacity, including restrictions on computing resources, memory resources and communication resources. In this way, a mathematical model of resource allocation is constructed, and the task allocation problem is converted into a standard integer linear programming problem. Among them, the objective function is as follows:
[0084]
[0085] in, is the execution time of subtask p on device q (unit: seconds); is a binary decision variable, indicating whether subtask p is assigned to device q; is the data transmission time between subtasks i and j (unit: seconds); is a binary variable indicating whether subtasks i and j are assigned to different devices and need to communicate; is the set of all subtasks, is the set of all heterogeneous computing devices; E is the set of subtask pairs with data transmission relationships, Represents subtask pairs Belongs to set E, that is, there is a data transmission dependency relationship between subtask i and subtask j, and the output data of subtask i needs to be used as the input data of subtask j.
[0086] The objective function is to minimize the total execution time, which includes computation time and communication time.
[0087] Solving the mathematical model of resource allocation involves two algorithms: branch-and-bound and Lagrangian relaxation. The branch-and-bound algorithm is a classic approach for integer linear programming. Its core concept is to find the optimal solution through branching and bounding. In practice, the linear programming relaxation problem is solved to obtain a lower bound. A non-integer variable is then selected for branching, and ceiling and floor constraints are added, forming two subproblems. The subproblems are then solved recursively, with branches pruned based on the bounds of the solutions to find an integer solution. If the problem size exceeds a preset limit, typically determined by the number of variables and constraints, the Lagrangian relaxation method is used. Lagrangian relaxation introduces difficult constraints into the objective function using Lagrangian multipliers, transforming the problem into a more manageable one. For task allocation problems, device resource constraints are typically relaxed using Lagrangian multipliers. The Lagrangian dual problem is then iteratively solved using a subgradient method to find a near-optimal solution. This process yields a preliminary task allocation solution, clearly defining which device each subtask should be assigned to.
[0088] Calculating the data transfer cost between subtasks based on the preliminary task assignment plan is the first step in optimizing communication overhead. The analysis begins by identifying data transfer relationships between subtasks, which are inherited from the data flow of the original task operation sequence. For each pair of subtasks with a data transfer relationship, their target devices in the preliminary assignment plan are examined. If two subtasks are assigned to the same device, the data transfer cost is zero; if they are assigned to different devices, a data transfer cost is calculated. The data transfer cost is determined by the amount of data transferred and the communication bandwidth between the devices. The amount of data transferred is derived from the data flow relationships between subtasks and represents the amount of data output from one subtask to another. Based on these calculation results, a data transfer plan graph is constructed. This graph uses a directed graph structure, with nodes representing subtasks and edges representing data transfer relationships. The edge weights are the amount of data transferred. If two subtasks with a data transfer relationship are assigned to different devices, an edge is added from the source subtask to the target subtask in the graph, labeled with the amount of data transferred.
[0089] The final step in the task allocation process is optimizing the task allocation based on the communication cost graph and the inter-device physical connection topology. This optimization process considers the impact of data transmission latency on total execution time and seeks the optimal balance between computational latency and communication latency by adjusting the subtask allocation locations. The communication latency parameters for each pair of devices, including bandwidth and base latency, are extracted from the inter-device physical connection topology. The latency of each data transmission is then calculated based on the communication cost graph and inter-device communication parameters. Next, the relationship between communication latency and computational benefit is evaluated, defined as the reduction in computational time after moving a subtask from its current device to its target device. If the communication latency of a subtask exceeds the computational benefit, the subtask's allocation location is considered for adjustment. Adjustment strategies include attempting to assign closely related subtasks to the same device, even if this results in a slight decrease in computational performance, or rerouting data transmission paths to devices with higher communication bandwidth. Through iterative local adjustments, the task allocation solution is continuously optimized until a balance between communication latency and computational performance is found, resulting in the optimized task allocation solution.
[0090] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0091] Deploy a task execution engine on heterogeneous computing devices. The task execution engine includes a task queue manager, a memory manager, an execution scheduler, and a performance monitor module to obtain a subtask execution environment.
[0092] Initialize the parameters of the subtask set according to the task allocation plan, convert the calculation parameters and initial weights of each subtask into a device-adaptable format, and obtain device-specific subtask parameters;
[0093] Allocate execution memory to each heterogeneous computing device based on device-specific subtask parameters, calculate the memory space required by the subtask and establish a data buffer to obtain a memory allocation table;
[0094] According to the data transmission requirements between the memory allocation table and the subtask set, a communication channel between subtasks is constructed, a data transmission interface is established for subtasks with data transmission relationships, and a subtask communication topology is obtained;
[0095] Create a distributed computing schedule based on the subtask communication topology, including the subtask execution order, synchronization point location, and data transmission timing to obtain the global execution sequence;
[0096] The global execution sequence is converted into device-level instructions, deployed to the corresponding heterogeneous computing devices and the execution engine is started to obtain a distributed execution computing framework.
[0097] Specifically, task execution engines are deployed on heterogeneous computing devices. A task execution engine is a software component running on each heterogeneous device, responsible for actually executing the computing tasks assigned to that device. For each heterogeneous device, a corresponding version of the execution engine is deployed based on its hardware type and operating system. For example, a CUDA version of the engine is deployed for GPU devices, an OpenMP version of the engine is deployed for CPU devices, and a dedicated HDL-based engine is deployed for FPGA devices. The task execution engine consists of four core functional modules: a task queue manager is responsible for receiving subtasks optimized using genetic algorithms, maintaining a queue of pending tasks, prioritizing them, and managing the subtask lifecycle; a memory manager is responsible for allocating and reclaiming memory resources required during the computation process, implementing memory pool reuse and garbage collection; an execution scheduler is responsible for scheduling the specific execution timing of subtasks, handling data transfer relationships between tasks, and ensuring execution in the correct order; and a performance monitor collects performance metrics during execution in real time, such as execution time, memory usage, and computing resource utilization, providing feedback data. These four modules work together to form the subtask execution environment.
[0098] Initializing the parameters of a collection of subtasks according to the task allocation scheme is a key step in adapting to heterogeneous devices. For each subtask generated through DNA encoding and genetic optimization, the target execution device is determined from the task allocation scheme, and then the task parameters are converted based on the device type. Specifically, the computational and weight parameters of the subtask are extracted, which inherit the optimization characteristics of the most fit individuals in the genetic algorithm, and these parameters are converted into a device-specific format. For example, for convolution subtasks assigned to the GPU, the weight matrix is converted to a CUDA-compatible storage format, such as NCHW (batch size-number of channels-height-width) layout; for subtasks assigned to the CPU, it may be converted to NHWC (batch size-height-width-number of channels) layout to optimize cache usage. Furthermore, precision adjustment is performed based on the data precision supported by the device, such as converting the original FP32 (32-bit floating point) format to FP16 (16-bit floating point) on the GPU or INT8 (8-bit integer) format for specific accelerators. For specialized hardware such as FPGAs, the computational parameters must also be compiled into a hardware configuration bitstream.
[0099] Allocating execution memory to each heterogeneous computing device based on device-specific subtask parameters is essential for ensuring smooth subtask execution. The allocation process calculates the memory space required for each subtask, which consists of three components: input data memory, which stores the subtask's input tensors; weight parameter memory, which stores model parameters; and working memory, which stores intermediate results and output data. Because the subtasks are optimized using a genetic algorithm, their memory requirements are more accurately estimated, avoiding issues such as wasted or insufficient memory. For example, for a convolution subtask, the input tensor size is determined by the input feature map dimensions, the weight parameter size is determined by the number and size of the convolution kernels, and the working memory includes the space required for the output feature map and intermediate results. After the computation is complete, memory blocks of the appropriate size are allocated on the target device, and a memory mapping table is established to record the starting address, size, and purpose of each memory block. For adjacent subtasks that frequently exchange data, contiguous memory space is allocated to reduce data movement overhead. Data buffers are then set up for each memory block, including input buffers, output buffers, and intermediate result buffers. These buffers serve as the actual storage medium for the subtask data. A memory allocation table is then created to record the memory allocation information for all subtasks.
[0100] Building inter-subtask communication channels based on the data transfer requirements between the memory allocation table and the subtask collection is key to achieving distributed execution. Data transfer requirements between subtasks are identified from the data flow relationships within the original task operation sequence, determining which subtask outputs serve as inputs for other subtasks. For adjacent subtasks assigned to the same device, data transfer is achieved through a memory sharing mechanism, directly connecting the output buffer of the previous subtask with the input buffer of the next subtask using memory address pointers without actual data copying. For adjacent subtasks assigned to different devices, cross-device communication channels must be established. The appropriate communication mechanism is selected based on the device type: PCIe channels and unified memory mapping are used between CPUs and GPUs; network sockets or MPI (Message Passing Interface) protocols are used between distributed nodes; and DMA controllers are used for specialized hardware such as FPGAs. A buffer is allocated for each communication channel, and the format, size, and layout conversion method of the transferred data are determined. A subtask communication topology graph is constructed, which illustrates the data flow paths and communication methods between all subtasks, providing a basis for coordinating subtask execution.
[0101] Creating a distributed computing schedule based on the subtask communication topology is crucial for ensuring efficient task execution. The schedule consists of three core elements: subtask execution order, synchronization point locations, and data transfer timing. A topological sort is performed based on the data transfer relationships between subtasks to generate an execution sequence that satisfies the transmission constraints. This sequence then identifies subtasks on the critical path and assigns them higher execution priority. Synchronization points are then determined—the points where multiple parallel execution paths converge—and synchronization instructions are inserted at these locations to ensure data consistency. For each data transfer request, the optimal transmission timing is determined, employing a prefetch strategy that initiates data transfers before the consumer subtask needs the data, thereby masking communication delays. The schedule also includes error handling and recovery mechanisms, defining retry strategies and alternative paths for subtask execution failures. This results in a global execution sequence, a directed acyclic graph (DAG). Nodes represent subtask execution or communication operations, and edges represent execution order constraints. Each node is annotated with information such as the target device, estimated execution time, and resource requirements.
[0102] The final step in achieving distributed execution is converting the global execution sequence into device-level instructions and deploying them to the corresponding devices. This conversion process generates a specific instruction sequence for each target device, the format of which depends on the device type and execution engine implementation. For GPU devices, these generate CUDA or OpenCL kernel launch instructions; for CPU devices, these generate multithreaded execution instructions; and for FPGA devices, these generate configuration and control instructions. The instruction sequence consists of four types of instructions: compute instructions, which perform the actual computation; memory instructions, which allocate and release memory and move data; synchronization instructions, which ensure execution order and data consistency; and communication instructions, which control inter-device data transfer. The generated instruction sequence is encapsulated in a device executable format and sent to the task execution engine on each heterogeneous device via a control channel. After the execution engine receives the instruction sequence, the task queue manager adds the instructions to the execution queue. The execution scheduler retrieves the instructions for execution in sequence, the memory manager handles memory-related operations, and the performance monitor records the execution status. When the execution engines on all devices are activated and working together, a distributed execution computing framework is formed, enabling efficient execution of deep learning tasks in heterogeneous environments.
[0103] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0104] The performance monitors on each heterogeneous computing device collect computing load, memory usage, execution progress, and communication delay parameters to obtain the heterogeneous device state vector;
[0105] Apply sliding time window analysis to the state vectors of heterogeneous devices, calculate the average execution time and time standard deviation of each device, and obtain device performance statistics;
[0106] Build a performance prediction model based on device performance statistics. The performance prediction model predicts the future performance trend of each device and obtains performance bottleneck identification results.
[0107] Based on the performance bottleneck identification results, a four-layer adjustment strategy is implemented for the unbalanced computing devices, including device internal parameter optimization, load redistribution, subtask reconstruction, and global reoptimization, to obtain the adjusted resource allocation plan;
[0108] Perform parameter aggregation based on time windows on the subtasks redeployed based on the adjusted resource allocation scheme, and use the number of device processed samples as the weight coefficient for weighted averaging to obtain a synchronized and coordinated training strategy;
[0109] A synchronous and coordinated training strategy transfers knowledge from high-performance devices to low-performance devices. The transmission priority is determined by adjusting the ratio of knowledge importance to communication cost, resulting in heterogeneous environment computing results with improved acceleration ratio.
[0110] Specifically, real-time data is collected through performance monitors on each heterogeneous computing device. A performance monitor is a lightweight agent deployed on each computing device, responsible for regularly collecting device operational status data. For CPU devices, information such as processor utilization, memory usage, and thread status is obtained through operating system APIs. For GPU devices, core utilization, video memory usage, and temperature parameters are collected through driver interfaces (such as NVIDIA SMI). For specialized accelerators such as FPGAs, operating status is read through hardware monitoring registers. Monitoring includes four key parameters: Computational load metrics reflect the device's computing resource utilization, such as processor utilization and the number of active threads; Memory usage metrics record memory resource usage, including total memory usage and peak memory usage; Execution progress metrics track subtask completion, recording the current data batch and iteration number; and Communication latency metrics monitor inter-device data transmission performance, such as transmission delay and bandwidth utilization. These parameters are sampled at a fixed frequency (typically 10-100 milliseconds) and, after preliminary processing (such as outlier removal and unit normalization), form a heterogeneous device state vector, which comprehensively reflects the device's current operational status. Applying sliding time window analysis to heterogeneous device state vectors is a key step in identifying performance trends. Sliding time window analysis is a time series data processing technique that analyzes the temporal characteristics of data by sliding a fixed-size window across the time dimension. In implementation, a state history queue is maintained for each device, storing state vectors from a past period. The window size is typically set to 10-30 seconds, meaning that the queue stores all sampled data from that period. When a new state vector arrives, it is added to the end of the queue, and the oldest data at the head of the queue is removed, maintaining a fixed window size. Statistical analysis is performed on the data within the window to calculate the average execution time for each device. This is done by dividing the number of subtasks completed within the window by the window duration to obtain the average processing rate. The standard deviation of the execution time is then calculated to reflect the stability of device performance; a smaller standard deviation indicates a more stable execution speed. Trends in other performance indicators, such as the growth rate of memory usage and fluctuation patterns in communication latency, are also analyzed. These statistical results together constitute device performance statistics.
[0111] Building a performance prediction model based on device performance statistics is a core component of dynamic resource adjustment. A performance prediction model is a time series prediction model used to predict future device performance changes based on historical performance data. During model building, an appropriate prediction algorithm is selected. For short-term predictions of a few seconds to minutes, a moving average model is used; for medium-term predictions of a few minutes to hours, an autoregressive moving average model is used; and for performance patterns with significant periodicity, a seasonal decomposition method is used. For example, a moving average model predicts the performance value at the next time point by taking a weighted average of performance data from N past time points. The weight coefficients can be either uniformly distributed or decreasingly distributed. During model training, the model is fitted using earlier performance data. The prediction accuracy is then evaluated on a validation set, and the prediction results are optimized by adjusting model parameters. The trained model is applied to current performance statistics to predict performance changes for each device over the next time period, including trends in metrics such as computing power, memory usage, and communication latency. By comparing the predicted performance of different devices, potential performance bottlenecks—those compute nodes whose performance is expected to significantly degrade or fall far below that of other devices—are identified, generating performance bottleneck identification results.
[0112] Optimizing heterogeneous resource allocation involves implementing a four-tiered adjustment strategy for load-unbalanced computing devices based on performance bottleneck identification. The first tier involves optimizing internal device parameters, adjusting the internal execution parameters of identified performance bottleneck devices. For example, for devices with high memory pressure, batch sizes are reduced to reduce memory usage; for devices with high computational loads, the number of threads or workgroup size is adjusted to optimize resource utilization; and for GPUs with excessively high temperatures, core frequency is reduced to prevent thermal throttling. The second tier involves load redistribution. When internal device optimization fails to resolve performance bottlenecks, some subtasks are migrated from the bottleneck device to less-loaded devices. The migration process involves identifying migratable subtasks, prioritizing tasks with minimal inter-device dependencies and moderate computational load, calculating migration benefits, and then executing the task migration, including state preservation, target device preparation, and execution resumption. The third tier involves subtask refactoring, redesigning the subtask structure to address performance mismatches. Specific methods include task segmentation, breaking down large subtasks into smaller units; task merging, combining multiple small, related tasks to reduce scheduling overhead; and computational precision adjustment, adjusting the computational precision of different subtasks based on their importance. The fourth layer is global re-optimization. When local adjustments fail to meet performance requirements, the entire task allocation process is re-executed, re-optimizing the task allocation plan based on the latest performance data and device status. Through these four layers of adjustment strategies, an optimized resource allocation plan is formed for the current device status.
[0113] Performing time-windowed parameter aggregation on subtasks redeployed based on the adjusted resource allocation plan is key to ensuring distributed training consistency. Traditional synchronous parameter aggregation strategies require all devices to complete training for the current batch before performing parameter updates. This causes fast devices to wait for slower devices, resulting in wasted resources. Time-windowed parameter aggregation sets a fixed time window and aggregates the current model parameters of each device at the end of each window, regardless of whether the device has completed training for the current batch. In implementation, the aggregation window size is determined based on inter-device performance differences, network communication latency, and model convergence characteristics. Within each time window, each device independently trains and records the number of samples processed. At the end of the window, the current model parameters and number of samples processed are collected from each device. Parameter aggregation uses a weighted average method, with weights proportional to the number of samples processed by the device. The calculation formula is: parameter = sum of each device's parameter × (number of device samples / total number of samples). This ensures that devices that contribute more computation have a greater impact on the model while allowing devices of different speeds to participate in training without waiting for each other. The aggregated parameters are broadcast back to each device and serve as the training starting point for the next time window, forming a synchronized and coordinated training strategy.
[0114] Knowledge distillation in heterogeneous environments is achieved by transferring knowledge from high-performance devices to low-performance devices based on a synchronized and coordinated training strategy. Knowledge distillation is a model compression technique that trains a small model to mimic the behavior of a larger model. In a heterogeneous environment, a functional teacher model runs on a high-performance device, while a simplified student model runs on a low-performance device. The knowledge transfer process determines the content to be transferred, including soft-label knowledge, feature knowledge, and relationship knowledge. Knowledge importance is then calculated to measure the contribution of different pieces of knowledge to the student model's performance. Communication cost is also calculated to assess the network resources required to transmit that knowledge. Knowledge transfer priority is determined by the ratio of knowledge importance to communication cost, prioritizing knowledge with high importance and low communication cost. In actual transmission, an incremental update strategy is employed, transferring only the knowledge with significant changes to reduce communication overhead. After receiving the knowledge, the low-performance device incorporates it into the student model's training process and guides the student model to learn the teacher model's behavior by adjusting the loss function. This knowledge transfer mechanism enables low-performance devices to benefit from the computing power of high-performance devices, achieving good performance even under resource constraints, and improving the speedup of computational results in heterogeneous environments.
[0115] The above describes the heterogeneous computing acceleration method based on the deep learning framework network in the embodiment of the present application. The following describes the heterogeneous computing acceleration system based on the deep learning framework network in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a heterogeneous computing acceleration system based on a deep learning framework network includes:
[0116] The acquisition module 201 is used to collect and analyze performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a characteristic data set of the heterogeneous computing devices;
[0117] The analysis module 202 is used to receive a computing task input, perform an operation sequence analysis on the computing task, and obtain a task operation sequence;
[0118] A mapping module 203 is configured to encode a task operation sequence into a gene sequence based on a heterogeneous computing device characteristic data set, optimize task decomposition by a genetic recombination algorithm to minimize execution time, and obtain a subtask set marked with acceleration characteristics;
[0119] A matching module 204 is configured to perform matching analysis on the subtask set marked with acceleration characteristics and the heterogeneous computing device characteristic data set to obtain a task allocation solution;
[0120] A deployment module 205 is configured to deploy the subtask set to corresponding heterogeneous computing devices according to the task allocation scheme to obtain a distributed execution computing framework;
[0121] The monitoring module 206 is used to monitor the running status of the distributed execution computing framework in real time, obtain device load balancing data, and dynamically adjust the allocation of heterogeneous computing resources based on the device load balancing data to obtain heterogeneous environment computing results with improved acceleration ratio.
[0122] above Figure 2 The heterogeneous computing acceleration system based on the deep learning framework network in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The heterogeneous computing acceleration device based on the deep learning framework network in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0123] Figure 3This is a schematic diagram of the structure of a heterogeneous computing acceleration device based on a deep learning framework network, provided by an embodiment of the present invention. This heterogeneous computing acceleration device 300 based on a deep learning framework network can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing application programs 333 or data 332. The memory 320 and storage medium 330 may be either transient or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instruction operations within the heterogeneous computing acceleration device 300 based on a deep learning framework network. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instruction operations stored in the storage medium 330 on the heterogeneous computing acceleration device 300 based on a deep learning framework network, thereby implementing the steps of the aforementioned heterogeneous computing acceleration method based on a deep learning framework network.
[0124] The heterogeneous computing acceleration device 300 based on the deep learning framework network may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the heterogeneous computing acceleration device based on the deep learning framework network shown does not constitute a limitation on the heterogeneous computing acceleration device based on the deep learning framework network provided by the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0125] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the heterogeneous computing acceleration method based on a deep learning framework network.
[0126] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a heterogeneous computing acceleration device based on a deep learning framework network (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program code.
Claims
1. A heterogeneous computing acceleration method based on a deep learning framework network, characterized in that: The method comprises: Collect and analyze performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a characteristic data set of heterogeneous computing devices; Receiving a computing task input, performing an operation sequence analysis on the computing task, and obtaining a task operation sequence; Based on the heterogeneous computing device characteristic data set, the task operation sequence is encoded into a gene sequence, and the task decomposition is optimized by a genetic recombination algorithm to minimize the execution time, thereby obtaining a subtask set marked with acceleration characteristics, including: Extracting the operation type, data dimension, computation density, memory access mode, and parallelism of each operation node in the task operation sequence to obtain an operation feature vector; Establish gene coding mapping rules, map different operation types into quaternary gene segments, encode operation parameters into gene expression sequences, convert operation feature vectors into standardized gene coding, and obtain task gene sequences; Construct a genetic algorithm framework, set the minimization of total execution time as the fitness function, initialize a population containing multiple task decomposition schemes, each individual represents a task decomposition strategy, and obtain the initial evolutionary population; Calculating a fitness score for each individual in the initial evolutionary population based on the heterogeneous computing device characteristic data set, evaluating the execution efficiency of different decomposition schemes in the heterogeneous environment, and obtaining a population fitness distribution; Apply genetic recombination operators to perform crossover operations on high-fitness individuals to generate new task decomposition combinations, perform mutation operations on some individuals to explore new decomposition possibilities, and obtain a new generation of evolutionary populations; Repeat the fitness evaluation and genetic recombination process until the algorithm converges, select the individual with the highest fitness as the optimal decomposition solution, decode its gene sequence into subtask divisions, and obtain a set of subtasks marked with acceleration characteristics; Performing matching analysis on the subtask set marked with acceleration characteristics and the heterogeneous computing device characteristic data set to obtain a task allocation plan; Deploying the subtask set onto corresponding heterogeneous computing devices according to the task allocation scheme to obtain a distributed execution computing framework; The running status of the distributed execution computing framework is monitored in real time to obtain device load balancing data, and the heterogeneous computing resource allocation is dynamically adjusted based on the device load balancing data to obtain a heterogeneous environment computing result with improved acceleration ratio.
2. The heterogeneous computing acceleration method based on a deep learning framework network according to claim 1 is characterized in that: The performance parameter collection and analysis of multiple types of computing devices in a heterogeneous computing environment to obtain a heterogeneous computing device characteristic data set includes: Automatically identify CPUs, GPUs, tensor processors, and FPGAs in heterogeneous computing environments and obtain a list of device types. executing standardized computation-intensive tasks, memory-intensive tasks, and communication-intensive tasks on each computing device in the device type list to obtain performance benchmark test results; Record the number of computing cores, clock frequency, memory capacity, memory bandwidth, and cache size parameters of each device based on the performance benchmark test results to obtain a device performance feature vector; Constructing a heterogeneous network topology graph including device nodes and communication links based on the device performance characteristic vector, and obtaining network relationship data between the device nodes and the communication links; Establishing a mapping relationship between device parameters and deep learning operation execution time using a polynomial regression method based on the network relationship data to obtain a device performance mathematical model; A communication cost model is established by combining the device performance mathematical model with the data transmission overhead between devices to obtain a heterogeneous computing device characteristic data set including computing power and communication cost.
3. The heterogeneous computing acceleration method based on a deep learning framework network according to claim 1 is characterized in that: The receiving of a computing task input and performing an operation sequence analysis on the computing task to obtain a task operation sequence includes: Receive computing tasks in a standardized framework format submitted by users, parse the computing tasks, and obtain an initial computing graph structure; Extracting operation type, input and output tensor dimension, computational complexity, and memory requirement characteristics of all computing nodes in the initial computational graph structure to obtain a node feature set; Analyze and calculate the data dependency relationship between nodes according to the node feature set to obtain a data flow graph; Analyze the data flow graph for computational bottlenecks, data dependencies, and parallel execution opportunities to obtain computational graph optimization opportunities; Based on the computation graph optimization opportunity, operation fusion, memory optimization, computation graph rewriting, precision adjustment, and operator replacement are applied to the initial computation graph structure to obtain an optimized computation graph; The operation sequence of the optimized calculation graph is extracted according to the execution order, and the operation type, parameter configuration and execution order of the calculation node are arranged into a linear sequence structure to obtain a task operation sequence.
4. The heterogeneous computing acceleration method based on a deep learning framework network according to claim 1 is characterized in that: The matching analysis of the subtask set marked with the acceleration characteristic and the heterogeneous computing device characteristic data set to obtain a task allocation solution includes: Constructing a subtask-device affinity matrix based on the subtask set marked with acceleration characteristics and the heterogeneous computing device characteristic data set, wherein the matrix element value represents the execution affinity between the subtask and the device, and obtaining a matching quantization matrix; Normalizing each element in the matching quantization matrix. If the execution time of a subtask on a device exceeds a preset threshold, the corresponding element value is set to zero. If it does not exceed the threshold, the element value is assigned according to the degree of matching between the subtask characteristics and the device characteristics to obtain an optimized matching matrix. An integer linear programming model is constructed based on the optimized matching matrix, and the objective function is set to minimize the total execution time. The constraints include that each subtask must be assigned to only one device and the device resource capacity is limited, thereby obtaining a resource allocation mathematical model. Applying a branch-and-bound algorithm to the resource allocation mathematical model, if the computational scale exceeds a preset limit, switching to a Lagrangian relaxation method for solution, if not, directly solving the problem to obtain a preliminary task allocation plan; Based on the preliminary task allocation plan, the data transmission cost between subtasks is calculated and a data transmission plan graph is constructed. If two subtasks with a data transmission relationship are assigned to different devices, the corresponding data transmission edges are added to the graph and the amount of data transmitted is marked to obtain a communication cost graph. According to the communication cost graph and the physical connection topology between devices, task allocation is optimized by minimizing the total communication delay. If the communication delay exceeds the computing benefit, the subtask allocation position is adjusted to obtain a task allocation solution.
5. The heterogeneous computing acceleration method based on a deep learning framework network according to claim 1 is characterized in that: The step of deploying the subtask set onto corresponding heterogeneous computing devices according to the task allocation scheme to obtain a distributed execution computing framework includes: Deploying a task execution engine on the heterogeneous computing device, the task execution engine comprising a task queue manager, a memory manager, an execution scheduler and a performance monitor module, to obtain a subtask execution environment; Initializing the parameters of the subtask set according to the task allocation scheme, converting the calculation parameters and initial weights of each subtask into a device-adaptive format to obtain device-specific subtask parameters; Allocate execution memory to each heterogeneous computing device based on the device-specific subtask parameters, calculate the memory space required by the subtask and establish a data buffer to obtain a memory allocation table; According to the data transmission requirements between the memory allocation table and the subtask set, a communication channel between subtasks is constructed, a data transmission interface is established for subtasks with data transmission relationships, and a subtask communication topology is obtained; Creating a distributed computing scheduling plan based on the subtask communication topology, including the subtask execution order, synchronization point location, and data transmission timing, to obtain a global execution sequence; The global execution sequence is converted into device-level instructions, deployed on corresponding heterogeneous computing devices and the execution engine is started to obtain a distributed execution computing framework.
6. The heterogeneous computing acceleration method based on a deep learning framework network according to claim 1 is characterized in that: The real-time monitoring of the running status of the distributed execution computing framework to obtain device load balancing data, and dynamically adjusting the heterogeneous computing resource allocation based on the device load balancing data to obtain a heterogeneous environment computing result with improved acceleration ratio, including: The performance monitors on each heterogeneous computing device collect computing load, memory usage, execution progress, and communication delay parameters to obtain the heterogeneous device state vector; Applying a sliding time window analysis to the heterogeneous device state vectors, calculating the average execution time and time standard deviation of each device, and obtaining device performance statistics; Building a performance prediction model based on the device performance statistical data, the performance prediction model predicts the future performance trend of each device and obtains a performance bottleneck identification result; Based on the performance bottleneck identification results, a four-layer adjustment strategy is implemented on the computing device with unbalanced load, including device internal parameter optimization, load redistribution, subtask reconstruction and global reoptimization, to obtain an adjusted resource allocation plan; Performing parameter aggregation based on a time window on the subtasks redeployed based on the adjusted resource allocation scheme, performing weighted averaging using the number of samples processed by the device as a weight coefficient, and obtaining a synchronous and coordinated training strategy; Based on the synchronous and coordinated training strategy, the knowledge of high-performance devices is transferred to low-performance devices. The transmission priority is determined by adjusting the ratio of knowledge importance and communication cost, thereby obtaining a heterogeneous environment computing result with improved acceleration ratio.
7. A heterogeneous computing acceleration system based on a deep learning framework network, characterized in that: Used to implement the heterogeneous computing acceleration method based on a deep learning framework network according to any one of claims 1 to 6, the heterogeneous computing acceleration system based on a deep learning framework network includes: The acquisition module is used to collect and analyze the performance parameters of multiple types of computing devices in a heterogeneous computing environment to obtain a characteristic data set of heterogeneous computing devices; An analysis module is used to receive a computing task input, perform operation sequence analysis on the computing task, and obtain a task operation sequence; A mapping module is configured to encode the task operation sequence into a gene sequence based on the heterogeneous computing device characteristic data set, optimize task decomposition by a genetic recombination algorithm to minimize execution time, and obtain a subtask set marked with acceleration characteristics, including: Extracting the operation type, data dimension, computation density, memory access mode, and parallelism of each operation node in the task operation sequence to obtain an operation feature vector; Establish gene coding mapping rules, map different operation types into quaternary gene segments, encode operation parameters into gene expression sequences, convert operation feature vectors into standardized gene coding, and obtain task gene sequences; Construct a genetic algorithm framework, set the minimization of total execution time as the fitness function, initialize a population containing multiple task decomposition schemes, each individual represents a task decomposition strategy, and obtain the initial evolutionary population; Calculating a fitness score for each individual in the initial evolutionary population based on the heterogeneous computing device characteristic data set, evaluating the execution efficiency of different decomposition schemes in the heterogeneous environment, and obtaining a population fitness distribution; Apply genetic recombination operators to perform crossover operations on high-fitness individuals to generate new task decomposition combinations, perform mutation operations on some individuals to explore new decomposition possibilities, and obtain a new generation of evolutionary populations; Repeat the fitness evaluation and genetic recombination process until the algorithm converges, select the individual with the highest fitness as the optimal decomposition solution, decode its gene sequence into subtask divisions, and obtain a set of subtasks marked with acceleration characteristics; A matching module, configured to perform matching analysis on the subtask set marked with acceleration characteristics and a heterogeneous computing device characteristic data set to obtain a task allocation solution; A deployment module, configured to deploy the subtask set onto corresponding heterogeneous computing devices according to the task allocation scheme to obtain a distributed execution computing framework; The monitoring module is used to monitor the running status of the distributed execution computing framework in real time, obtain device load balancing data, and dynamically adjust the allocation of heterogeneous computing resources based on the device load balancing data to obtain heterogeneous environment computing results with improved acceleration ratio.
8. A heterogeneous computing acceleration device based on a deep learning framework network, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the heterogeneous computing acceleration method based on the deep learning framework network described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor executes the heterogeneous computing acceleration method based on a deep learning framework network according to any one of claims 1 to 6.
Citation Information
Patent Citations
Task scheduling method and device based on heterogeneous computing
CN112328380A
GPU program optimization method based on CUDA parallel environment
CN119690684A