Data stream parallel acceleration method and system in heterogeneous computing environment

CN122412144BActive Publication Date: 2026-09-08BEIJING YUANFENGCHENG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610551534.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-09-08
Estimated Expiration
2046-04-24

AI Technical Summary

Technical Problem

[0003]但是,现有技术难以动态识别数据单元之间细粒度的前驱后继约束,以及不同计算单元在处理特定计算模式时的真实效率差异,导致任务分配可能违背数据依赖,或未能将计算特征与硬件特长有效匹配,造成部分计算资源闲置而其他资源过载,整体并行效率低下,此外,现有技术对由任务分配引发的数据通信开销优化不足,容易在关键路径上引入巨大的通信瓶颈,使得数据传输时间成为整体性能的主要制约因素,无法实现计算与访存开销的全局均衡优化

Benefits of technology

[0015]本发明中,通过对数据单元序列进行依赖关系分析和特征提取,实现了计算任务与异构资源的精准匹配,单元特征向量的提取与聚类划分,将原始任务流智能地分离为计算密集型和访存密集型两类子集,为针对性的优化策略提供了清晰的输入,针对计算密集型任务,算子融合分析能够识别出可合并的算子组,有效减少了内核启动开销和中间结果的存储与传输,提升了计算核心的利用效率,针对访存密集型任务,内存访问模式识别与冲突区域标记,提前揭示了潜在的访存瓶颈,为优化数据布局和传输路径创造了条件,基于初始映射方案分析由前驱后继关系引出的跨资源通信边及其开销,能够精确计算出全局任务的延迟分布,实现了计算负载与通信开销的均衡。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412144B_ABST
    Figure CN122412144B_ABST
Patent Text Reader

Abstract

The application provides a data stream parallel acceleration method and system in a heterogeneous computing environment, and relates to the technical field of heterogeneous computing, which comprises the following steps: analyzing data unit dependency and characteristics, dividing and optimizing a computation-intensive subset and a memory-intensive subset; combining heterogeneous resource parameters, generating an initial mapping scheme through computation and memory affinity analysis, and determining an optimized mapping scheme based on communication overhead and global delay optimization of a critical path and executing the optimized mapping scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heterogeneous computing technology, and in particular to a method and system for accelerating parallel data flow in a heterogeneous computing environment. Background Technology

[0002] In heterogeneous computing environments, efficiently processing data stream tasks is a key challenge. Existing technologies typically pre-allocate computational tasks within the data stream to different computing units on a heterogeneous platform, such as CPUs, GPUs, or dedicated accelerators, based on fundamental attributes of the tasks, such as estimated computational or data volume. Allocation decisions often rely on simple modeling of hardware performance parameters, such as peak computing power or memory bandwidth. After task allocation, the system executes tasks according to a pre-defined scheduling order, performing necessary data transfers between different computing units to complete the overall processing flow.

[0003] However, existing technologies struggle to dynamically identify fine-grained predecessor-successor constraints between data units, as well as the actual efficiency differences between different computing units when processing specific computing modes. This can lead to task allocation violating data dependencies or failing to effectively match computational characteristics with hardware strengths, resulting in some computing resources being idle while others are overloaded, leading to overall low parallel efficiency. Furthermore, existing technologies are insufficient in optimizing data communication overhead caused by task allocation, which can easily introduce huge communication bottlenecks on the critical path, making data transmission time the main constraint on overall performance and failing to achieve a global balance between computation and memory access overhead. Summary of the Invention

[0004] This invention provides a method and system for accelerating parallel data flow in a heterogeneous computing environment, which can at least solve some of the problems existing in the prior art.

[0005] A first aspect of this invention provides a method for accelerating parallel data flow in a heterogeneous computing environment, comprising: The process involves acquiring and parsing the data stream to be processed to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traversing the data unit sequence and extracting unit feature vectors, clustering the unit feature vectors to obtain computationally intensive subsets and memory-intensive subsets, performing operator fusion analysis on the computationally intensive subsets to obtain a fusion operator group, and performing memory access pattern recognition on the memory-intensive subsets and marking access conflict regions. The processing capability parameters and bandwidth parameters of computing resources in a heterogeneous computing environment are obtained. The feature vectors corresponding to the fusion operator group are matched with the processing capability parameters to obtain the computing affinity matrix. Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is obtained by solving the computing affinity matrix and the memory access affinity matrix. Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. The global delay distribution is obtained based on the communication overhead vector. The critical path is identified according to the global delay distribution to determine the optimized mapping scheme. Based on the optimized mapping scheme, the data unit sequence is allocated to the corresponding computing resources and parallel execution is started to obtain the output data stream.

[0006] In one alternative implementation, The process involves acquiring and parsing the data stream to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, and traversing the data unit sequence to extract unit feature vectors, including: The process involves acquiring a data stream to be processed, parsing the data format of the data stream and identifying data boundaries to obtain a set of data segments, performing semantic parsing on each data segment in the set of data segments and marking the operation type to obtain an operation mark sequence, and encapsulating each data segment into a data unit based on the operation mark sequence and arranging them in chronological order to obtain a data unit sequence. Traverse the data unit sequence, extract the input variable identifier and output variable identifier of each data unit and perform intersection matching to identify variable dependencies. Based on the variable dependencies, construct a directed dependency graph and perform topological traversal on the directed dependency graph. Mark the predecessor data unit set and successor data unit set corresponding to each data unit to obtain the predecessor-successor relationship set. The operation type and variable access pattern of each data unit are extracted from the data unit sequence. The operation type is one-hot encoded to obtain the operation type vector. The access frequency and access span are calculated based on the variable access pattern to obtain the access pattern vector. An adjacency matrix is ​​constructed based on the predecessor and successor relationship set, and the adjacency matrix is ​​spectral decomposed to obtain the graph topology feature vector. The operation type vector and the access pattern vector are combined and linearly projected to obtain the unit feature vector.

[0007] In one alternative implementation, Clustering the unit feature vectors yields computationally intensive and memory-intensive subsets. Operator fusion analysis is performed on the computationally intensive subsets to obtain a fusion operator group. Memory access pattern recognition and access conflict region marking are performed on the memory-intensive subsets, including: Calculate the variance contribution rate of different feature dimensions in the feature vector of the unit and determine the dominant dimension. Construct a feature subspace based on the dominant dimension. Calculate the distance metric between data units in the feature subspace and determine the cluster center. Iteratively cluster the data units based on the distance metric until the cluster center converges. Calculate the memory access ratio based on the converged clusters and divide the clusters into a computationally intensive subset and a memory-intensive subset based on the calculated memory access ratio. Traverse the data units in the computationally intensive subset and extract the data flow relationship to identify matching data unit pairs. Determine the fusion feasibility of the matching data unit pairs and calculate the number of variable eliminations. Based on the number of variable eliminations, perform filtering and fusion operations on the matching data unit pairs to obtain a fusion operator group. Traverse the data units in the memory-intensive subset and extract the variable access sequence. Calculate the step size feature of the access address and the interval feature of the access time in the variable access sequence. Determine the access pattern based on the step size feature. Identify the access window based on the interval feature. Align the variable access sequence with the time axis based on the access pattern and the access window and detect the overlapping interval of the access address. Determine and record the access conflict area based on the overlapping interval.

[0008] In one alternative implementation, Obtaining the processing capacity and bandwidth parameters of computing resources in a heterogeneous computing environment, and matching the feature vectors corresponding to the fusion operator group with the processing capacity parameters to obtain a computational affinity matrix includes: The microarchitecture pipeline depth and functional unit configuration of each computing resource in the heterogeneous computing environment are obtained and the processing capacity parameters are obtained by performance modeling. The interconnection topology between different computing resources is obtained and a topology adjacency graph is constructed. The topology adjacency graph is traversed and the effective bandwidth and congestion window size are obtained as bandwidth parameters. The fusion operators in the fusion operator group are traversed and aggregated to obtain the fusion operator feature vector. The operation type feature and data scale feature are separated from the fusion operator feature vector. The scheduling feature vector is generated based on the operation type feature and matched with the preset operation mode prototype to determine the dominant operation type. The operation type label is generated based on the dominant operation type. The data throughput requirement is calculated based on the data scale feature and the throughput label is generated. The processing capability parameters are arranged into capability vectors according to computing resources, and a type sensitivity matrix is ​​constructed. A type-weighted capability vector is calculated based on the operation type label and the type sensitivity matrix. A type matching degree is calculated based on the operation type label and the type-weighted capability vector. A capacity matching degree is calculated based on the throughput label and the capability vector, and a weighted sum of the capacity matching degree and the type matching degree is obtained to obtain a comprehensive matching degree. The comprehensive matching degree is filled into a computing affinity matrix based on the fusion operator and the computing resources.

[0009] In one alternative implementation, Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset, and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is then obtained by solving the calculated affinity matrix and the memory access affinity matrix, including: Extract the physical address range and access port configuration of storage resources from the access conflict region, construct an address space mapping table based on the physical address range, identify mutually exclusive access paths based on the access port configuration, traverse the memory-intensive subset and extract the read operation set and the write operation set, match the read operation set and the write operation set with the address space mapping table, determine potential conflicting memory access operations and group them according to a preset time window, calculate the concurrent access conflict probability based on the grouping results and generate a conflict penalty coefficient; A bandwidth constraint graph is constructed based on the bandwidth parameters and the mutually exclusive access paths. In the bandwidth constraint graph, a multi-path search is performed with the memory access operation of the fusion operator as the starting point and the storage resource as the ending point to obtain candidate paths and calculate the corresponding path capacity. The transmission delay is obtained based on the path capacity and the amount of memory accessed, and the transmission cost is determined by combining the conflict penalty coefficient. The transmission cost is filled into a memory access affinity matrix based on the fusion operator and the storage resource. Align the calculated affinity matrix and the memory access affinity matrix to construct a joint affinity tensor. Perform slice decomposition on the joint affinity tensor to obtain a preference vector. Solve the mapping relationship based on the preference vector and encode it as an initial mapping scheme.

[0010] In one alternative implementation, Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. Based on the communication overhead vector, the global delay distribution is obtained, including: Data dependencies between fusion operators are extracted from the predecessor-successor relationship set. Fusion operator pairs mapped to different computing resources are identified and cross-resource data transmission edges are marked in conjunction with the initial mapping scheme. Data transmission volume is extracted from the cross-resource data transmission edges and data transmission time is calculated in conjunction with the bandwidth parameter. The topology level crossed by the cross-resource data transmission edge is identified based on the data transmission time and a level penalty factor is generated according to the topology level. The communication overhead of each cross-resource data transmission edge is calculated based on the data transmission time and the level penalty factor. Identify cross-resource data transmission edges that converge to the same fusion operator from the set of predecessor and successor relationships, sort them according to communication overhead, identify the maximum communication overhead, calculate the difference between the maximum communication overhead and other communication overheads to obtain the synchronization overhead, update the communication overhead of each cross-resource data transmission edge based on the synchronization overhead, and arrange the updated communication overhead into a communication overhead vector according to the order of the cross-resource data transmission edges. The computation time of each fusion operator is extracted from the initial mapping scheme. The execution delay of each fusion operator is calculated based on the predecessor-successor relationship set and the communication overhead vector, and statistical analysis is performed to obtain the global delay distribution.

[0011] In one alternative implementation, Based on the global delay distribution, critical paths are identified, and an optimized mapping scheme is determined. Based on this optimized mapping scheme, the data unit sequence is allocated to corresponding computing resources, and parallel execution is initiated to obtain the output data stream, including: The sequence of fusion operators with the largest execution latency is identified from the global latency distribution as the critical path and the key fusion operators are extracted. The computing resources mapped by the key fusion operators in the initial mapping scheme are subjected to load analysis and the load saturation is calculated. Based on the load saturation, overloaded computing resources are identified and idle computing resources are searched in the heterogeneous computing environment. The key fusion operators are mapped from the overloaded computing resources to the idle computing resources to obtain the corrected mapping relationship. The communication overhead vector is updated based on the corrected mapping relationship, and the corrected global delay distribution is solved based on the updated communication overhead vector. It is determined whether the maximum execution delay in the corrected global delay distribution is less than the maximum execution delay in the global delay distribution. If it is less, the corrected mapping relationship is retained; otherwise, the corrected mapping relationship is rolled back. The mapping adjustment and delay evaluation are repeated until the maximum execution delay converges to obtain the optimized mapping scheme. The computational resources of each fusion operator mapping are extracted from the optimized mapping scheme, the data units in the data unit sequence are allocated to the corresponding computational resources, each computational resource is started to execute the data units in parallel and perform cross-resource data transmission based on the predecessor and successor relationship set, the data output by the computational resources is collected and spliced ​​to obtain the output data stream.

[0012] A second aspect of this invention provides a data stream parallel acceleration system for heterogeneous computing environments, comprising: The parsing unit is used to acquire the data stream to be processed and parse it to obtain a data unit sequence, perform dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traverse the data unit sequence and extract unit feature vectors, perform clustering on the unit feature vectors to obtain a computationally intensive subset and a memory-intensive subset, perform operator fusion analysis on the computationally intensive subset to obtain a fusion operator group, and perform memory access pattern recognition on the memory-intensive subset and mark access conflict regions. The analysis unit is used to obtain the processing capability parameters and bandwidth parameters of computing resources in the heterogeneous computing environment, match the feature vectors corresponding to the fusion operator group with the processing capability parameters to obtain the computing affinity matrix, perform path search on the memory-intensive subset based on the bandwidth parameters and the access conflict region and calculate the transmission cost to obtain the memory access affinity matrix, and solve the initial mapping scheme based on the computing affinity matrix and the memory access affinity matrix. The extraction unit is used to identify cross-resource data transmission edges in the initial mapping scheme based on the predecessor-successor relationship set and solve for the communication overhead vector, solve for the global delay distribution based on the communication overhead vector, identify the critical path according to the global delay distribution and determine the optimized mapping scheme, allocate the data unit sequence to the corresponding computing resources based on the optimized mapping scheme and start parallel execution to obtain the output data stream.

[0013] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0014] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0015] In this invention, by performing dependency analysis and feature extraction on data unit sequences, precise matching of computing tasks and heterogeneous resources is achieved. The extraction and clustering of unit feature vectors intelligently separates the original task flow into two subsets: computationally intensive and memory-intensive. This provides clear input for targeted optimization strategies. For computationally intensive tasks, operator fusion analysis can identify mergeable operator groups, effectively reducing kernel startup overhead and the storage and transmission of intermediate results, thus improving the utilization efficiency of the computing core. For memory-intensive tasks, memory access pattern identification and conflict region marking reveal potential memory bottlenecks in advance, creating conditions for optimizing data layout and transmission paths. Based on the analysis of the cross-resource communication edges and their overhead derived from the predecessor-successor relationship using the initial mapping scheme, the latency distribution of the global task can be accurately calculated, achieving a balance between computing load and communication overhead. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the data stream parallel acceleration method in a heterogeneous computing environment according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the communication overhead analysis of the fusion operator in the data stream parallel acceleration method under a heterogeneous computing environment, as described in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0019] Figure 1 This is a flowchart illustrating the data stream parallel acceleration method in a heterogeneous computing environment according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: The process involves acquiring and parsing the data stream to be processed to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traversing the data unit sequence and extracting unit feature vectors, clustering the unit feature vectors to obtain computationally intensive subsets and memory-intensive subsets, performing operator fusion analysis on the computationally intensive subsets to obtain a fusion operator group, and performing memory access pattern recognition on the memory-intensive subsets and marking access conflict regions. The processing capability parameters and bandwidth parameters of computing resources in a heterogeneous computing environment are obtained. The feature vectors corresponding to the fusion operator group are matched with the processing capability parameters to obtain the computing affinity matrix. Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is obtained by solving the computing affinity matrix and the memory access affinity matrix. Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. The global delay distribution is obtained based on the communication overhead vector. The critical path is identified according to the global delay distribution to determine the optimized mapping scheme. Based on the optimized mapping scheme, the data unit sequence is allocated to the corresponding computing resources and parallel execution is started to obtain the output data stream.

[0020] In one alternative implementation, The process involves acquiring and parsing the data stream to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, and traversing the data unit sequence to extract unit feature vectors, including: The process involves acquiring a data stream to be processed, parsing the data format of the data stream and identifying data boundaries to obtain a set of data segments, performing semantic parsing on each data segment in the set of data segments and marking the operation type to obtain an operation mark sequence, and encapsulating each data segment into a data unit based on the operation mark sequence and arranging them in chronological order to obtain a data unit sequence. Traverse the data unit sequence, extract the input variable identifier and output variable identifier of each data unit and perform intersection matching to identify variable dependencies. Based on the variable dependencies, construct a directed dependency graph and perform topological traversal on the directed dependency graph. Mark the predecessor data unit set and successor data unit set corresponding to each data unit to obtain the predecessor-successor relationship set. The operation type and variable access pattern of each data unit are extracted from the data unit sequence. The operation type is one-hot encoded to obtain the operation type vector. The access frequency and access span are calculated based on the variable access pattern to obtain the access pattern vector. An adjacency matrix is ​​constructed based on the predecessor and successor relationship set, and the adjacency matrix is ​​spectral decomposed to obtain the graph topology feature vector. The operation type vector and the access pattern vector are combined and linearly projected to obtain the unit feature vector.

[0021] After acquiring the data stream to be processed, it is parsed and segmented. During data stream parsing, specific data boundary markers, such as newline characters, delimiters, or special character sequences, need to be identified. For example, for a data stream containing multiple computational tasks, segmentation can be performed by identifying task boundary identifiers such as "TASK_BEGIN" and "TASK_END". Assuming the original data stream contains computational tasks such as "TASK_BEGIN matrix_multiply ABC TASK_END" and "TASK_BEGIN vector_add DEF TASK_END", the parsed data segment set will contain two segments, corresponding to matrix multiplication and vector addition operations, respectively.

[0022] Semantic parsing is performed on each data segment to identify the operation type and parameter information according to predefined syntax rules. For example, for the segment "matrix_multiply ABC", semantic analysis identifies the operation as matrix multiplication, with input variables matrix A and matrix B, and output variable matrix C, tagged as "COMPUTE_MATRIX_MULTIPLY". Similarly, for the segment "vector_add DEF", it is identified as vector addition, with input variables vectors D and vector E, and output variable vector F, tagged as "COMPUTE_VECTOR_ADD". In this way, semantic parsing of all data segments is completed, resulting in a sequence of operation tags.

[0023] Based on the operation tag sequence, each data segment is encapsulated into a data unit. Each data unit contains information such as operation type, input variable identifier, output variable identifier, and timestamp. The timestamp can be the position index of the data segment in the original data stream or the actual time information. All data units are arranged in chronological order to form a data unit sequence.

[0024] Traverse the sequence of data units, extract the input and output variable identifiers for each data unit, and determine the dependencies between variables by matching the intersection of these identifiers. For example, if the output variable X of data unit A is an input variable of data unit B, then B is considered dependent on A. By identifying all variable dependencies, construct a directed dependency graph, where nodes represent data units and edges represent dependencies.

[0025] A topological traversal is performed on the constructed directed dependency graph, marking the predecessor and successor sets of each data unit. For example, suppose there are three units in the data unit sequence: unit 1 executes "matrix_multiply AB C", unit 2 executes "vector_add CDE", and unit 3 executes "scalar_multiply EFG". Through variable dependency analysis, it can be seen that unit 2 depends on the output C of unit 1, and unit 3 depends on the output E of unit 2. Therefore, the successor set of unit 1 is unit 2, the predecessor set of unit 2 is unit 1, the successor set of unit 2 is unit 3, and the predecessor set of unit 3 is unit 2, resulting in the complete set of predecessor and successor relationships.

[0026] Extract the operation type feature for each data unit from the data unit sequence. Perform one-hot encoding on the operation type, mapping each operation type to a vector with 1 in one position and 0 in the others. For example, assuming the system supports 5 operation types: matrix multiplication, vector addition, scalar multiplication, matrix transpose, and vector dot product, then the operation type vector for matrix multiplication is (1, 0, 0, 0, 0), for vector addition it is (0, 1, 0, 0, 0), and so on.

[0027] Analyze the variable access patterns of each data unit, calculating the variable access frequency and access span. Variable access frequency refers to the number of times the variable is accessed throughout the entire data stream, and access span refers to the number of data units between the first and last access of the variable. For each data unit, construct an access pattern vector based on the access frequency and access span of its involved variables. For example, for data unit 2 "vector_add CDE", variable C has an access frequency of 2 and an access span of 1, variable D has an access frequency of 1 and an access span of 0, and variable E has an access frequency of 2 and an access span of 1. The access pattern vector constructed based on this information is (2, 1, 2, 0, 1, 0).

[0028] An adjacency matrix is ​​constructed based on the set of predecessor and successor relationships, where the matrix elements represent the dependencies between data units. For example, for a sequence containing three data units, the adjacency matrix can be represented as follows: the first row is (0, 0, 1), the second row is (0, 0, 0), and the third row is (0, 1, 0). Spectral decomposition is performed on the adjacency matrix to extract eigenvalues ​​and eigenvectors, generating graph topological eigenvectors.

[0029] The operation type vector, access mode vector, and graph topology feature vector are concatenated to form a comprehensive feature vector. A linear projection operation is then applied to this comprehensive feature vector to map the high-dimensional features to a low-dimensional space suitable for parallel scheduling decisions, resulting in the final unit feature vector. For example, for data unit 2, its operation type vector is (0, 1, 0, 0, 0), its access mode vector is (2, 1, 2, 0, 1, 0), and its graph topology feature vector is (0.5, 0.7, 0.1). The concatenated comprehensive feature vector is (0, 1, 0, 0, 0, 2, 1, 2, 0, 1, 0, 0.5, 0.7, 0.1), which, when mapped using a linear projection matrix, yields the unit feature vector (0.72, 0.45, 0.89, 0.36, 0.51).

[0030] In this embodiment, by performing structured parsing and semantic reconstruction on the original data stream, the originally continuous data stream without explicit structural identifiers is decomposed into data units with clear semantic boundaries and operation types. This can accurately depict the real causal relationship between data units from the perspective of variable dependency, avoiding analysis bias caused by the failure to identify implicit dependencies, and significantly improving the completeness and accuracy of data stream structure modeling. By encoding operation types and quantifying variable access patterns, and combining them with the topological features of the dependency graph for fusion expression, the discriminative and generalization capabilities are improved. By constructing a directed dependency graph and performing topological traversal, the set of predecessor and successor relationships is clarified, which can maintain the consistency and traceability of dependency order in complex data stream scenarios, enhance the transparency and interpretability of the data processing flow, and help improve system debugging efficiency and operational reliability.

[0031] In one alternative implementation, Clustering the unit feature vectors yields computationally intensive and memory-intensive subsets. Operator fusion analysis is performed on the computationally intensive subsets to obtain a fusion operator group. Memory access pattern recognition and access conflict region marking are performed on the memory-intensive subsets, including: Calculate the variance contribution rate of different feature dimensions in the feature vector of the unit and determine the dominant dimension. Construct a feature subspace based on the dominant dimension. Calculate the distance metric between data units in the feature subspace and determine the cluster center. Iteratively cluster the data units based on the distance metric until the cluster center converges. Calculate the memory access ratio based on the converged clusters and divide the clusters into a computationally intensive subset and a memory-intensive subset based on the calculated memory access ratio. Traverse the data units in the computationally intensive subset and extract the data flow relationship to identify matching data unit pairs. Determine the fusion feasibility of the matching data unit pairs and calculate the number of variable eliminations. Based on the number of variable eliminations, perform filtering and fusion operations on the matching data unit pairs to obtain a fusion operator group. Traverse the data units in the memory-intensive subset and extract the variable access sequence. Calculate the step size feature of the access address and the interval feature of the access time in the variable access sequence. Determine the access pattern based on the step size feature. Identify the access window based on the interval feature. Align the variable access sequence with the time axis based on the access pattern and the access window and detect the overlapping interval of the access address. Determine and record the access conflict area based on the overlapping interval.

[0032] A variance contribution analysis is performed on the aforementioned unit feature vectors to calculate the variance contribution rate of each feature dimension. The variance contribution rate refers to the proportion of the variance of a particular feature dimension to the total variance, reflecting the degree of influence of that dimension on the data distribution. Taking the aforementioned unit feature vectors as an example, the variance of each dimension is calculated, assuming the obtained variances are 0.21, 0.15, 0.38, 0.08, and 0.13, with a total variance of 0.95. Then, the variance contribution rates of each dimension are 0.22, 0.16, 0.40, 0.08, and 0.14, respectively. The dominant dimension, i.e., the dimension with the highest contribution rate, is determined based on the variance contribution rate. In this embodiment, the contribution rate of the third dimension is 0.40, which is much higher than that of the other dimensions; therefore, it is determined as the dominant dimension.

[0033] Based on the determined dominant dimension, a feature subspace is constructed. The selection of the dominant dimension can be based on a threshold, such as selecting a dimension with a contribution rate greater than 0.15 as the dominant dimension. In this example, the contribution rates of the first, third, and fifth dimensions are 0.22, 0.40, and 0.14, respectively. Considering that the fifth dimension is close to the threshold and has a certain influence on the data distribution, the first, third, and fifth dimensions can be selected to construct a three-dimensional feature subspace.

[0034] Calculate the distance metric between data units within the constructed feature subspace. The distance metric can be Euclidean distance, Manhattan distance, or Mahalanobis distance, etc. Taking Euclidean distance as an example, for two data points in the feature subspace with coordinates (0.72, 0.89, 0.14) and (0.65, 0.92, 0.17), the calculated Euclidean distance between the two points is 0.09.

[0035] Cluster centers are determined based on distance metrics. Initial cluster centers can be determined by randomly selecting data points or density peak points. Suppose two data units are randomly selected from the data unit set as initial cluster centers, with feature subspace coordinates of (0.72, 0.89, 0.14) and (0.35, 0.42, 0.68), respectively.

[0036] Iterative clustering of data units is performed based on a distance metric. In each iteration, the data unit is assigned to the nearest cluster center, and the center point of each cluster is recalculated. This process is repeated until the cluster centers converge, i.e., the change is less than a preset threshold of 0.001. Assume that after 5 iterations, two clusters are finally formed, with center points of (0.68, 0.87, 0.16) and (0.31, 0.38, 0.72), respectively.

[0037] The memory access ratio is calculated based on the converged clusters. The memory access ratio refers to the ratio of computational operations to memory access operations in a data unit, reflecting the computationally or memory-intensive characteristics of that data unit. Operation type information is extracted from the data units of each cluster, and the number of computational operations and memory access operations are statistically analyzed to calculate the ratio. For example, in the first cluster, there are 156 computational operations and 52 memory access operations, resulting in a memory access ratio of 3.0; in the second cluster, there are 37 computational operations and 98 memory access operations, resulting in a memory access ratio of 0.38.

[0038] Clusters are divided into compute-intensive and memory-intensive subsets based on their computed memory access ratio (CMR). A threshold of 1.0 is set; clusters with a CMR greater than this threshold are classified as compute-intensive subsets, and those with a CMR less than this threshold are classified as memory-intensive subsets. In this embodiment, the first cluster is classified as a compute-intensive subset, and the second cluster is classified as a memory-intensive subset.

[0039] The algorithm iterates through the data units in a computationally intensive subset, extracting data flow relationships to identify matching data unit pairs. A data flow relationship refers to the input-output dependency between data units. For example, if the output variable of data unit A serves as the input variable of data unit B, then A and B constitute a matching data unit pair, and there is a direct data flow relationship between them. Multiple such matching data unit pairs may exist within a computationally intensive subset.

[0040] The identified matching data unit pairs are fused to determine feasibility, and the number of variables eliminated is calculated. The fusion feasibility determination includes checking whether the two data units have an adjacent execution order and whether there are other dependencies hindering fusion. The number of variables eliminated refers to the number of intermediate variables that can be eliminated after fusion. For example, for matching data unit pairs A and B, if the output variable X of A is only used as an input variable of B and is not used by other data units, then variable X can be eliminated after fusion, and the number of variables eliminated is 1.

[0041] The matching data unit pairs are filtered and fused based on the number of variables eliminated, resulting in a fusion operator group. The more variables eliminated, the higher the fusion efficiency. A threshold of 2 is set, meaning only matching data unit pairs with 2 or more variables eliminated are fused. For example, if a matching data unit pair can eliminate 3 variables, exceeding the threshold, it is fused into a new data unit, forming part of the fusion operator group.

[0042] Traverse the data units in the memory-intensive subset and extract the variable access sequence. The variable access sequence records the address and time information of each variable being accessed. For example, variable M is accessed at time points 1, 3, and 5, with address offsets of 0, 64, and 128, respectively.

[0043] Calculate the step size characteristic and the interval characteristic of the access time in the variable access sequence. The step size characteristic refers to the address offset difference between two consecutive accesses, and the interval characteristic refers to the time difference between two consecutive accesses. For the access sequence of variable M, the step size is 64, 64, exhibiting an arithmetic progression characteristic; the time interval is 2, 2, also exhibiting an arithmetic progression characteristic. Based on the step size characteristic, the access pattern is determined, such as continuous access, equal step size access, random access, etc. For example, if the step size of variable M is always 64, it is determined to be an equal step size access pattern, which is suitable for a prefetch optimization strategy.

[0044] Access windows are identified based on interval features. An access window refers to a concentrated area of ​​access to a variable within a specific time period. Accesses with similar time intervals can be grouped into the same access window. For example, accesses to variable M from time points 1 to 5 can be classified into the same access window.

[0045] The system performs time-axis alignment of variable access sequences based on access patterns and access windows, and detects overlapping regions of access addresses. Time-axis alignment refers to aligning the access sequences of different variables along the time dimension to detect potential access conflicts. Overlapping regions refer to multiple variables accessing the same or adjacent memory regions within the same time period.

[0046] Identify and record conflicting regions based on overlapping intervals. Conflicting regions are areas that may lead to memory access contention and require special attention and optimization. Record information such as variable identifiers, address ranges, and time windows for each conflicting region to provide a basis for subsequent memory access optimization. For example, if variables M and N are found to access address ranges 64 to 256 within time windows 3 to 7, this is recorded as a conflicting region.

[0047] In this embodiment, by performing variance contribution rate analysis on the unit feature vectors and constructing a dominant feature subspace, effective dimensionality reduction and information reconstruction of high-dimensional features are achieved. This reduces the interference of redundant features on distance calculation, improves the stability and separability of clustering results, and automatically divides data units into computationally intensive subsets and memory-intensive subsets by performing distance measurement and iterative clustering in the feature subspace and calculating the memory access ratio based on the convergence results. This adapts to the dynamic changes of different data flow scenarios and significantly improves the adaptive and generalization capabilities of classification. By identifying and matching data units and performing fusion filtering based on the number of variable eliminations, the number of intermediate variables and redundant operations can be reduced while ensuring semantic correctness, thereby reducing the computation path length and intermediate data storage overhead.

[0048] In one alternative implementation, Obtaining the processing capacity and bandwidth parameters of computing resources in a heterogeneous computing environment, and matching the feature vectors corresponding to the fusion operator group with the processing capacity parameters to obtain a computational affinity matrix includes: The microarchitecture pipeline depth and functional unit configuration of each computing resource in the heterogeneous computing environment are obtained and the processing capacity parameters are obtained by performance modeling. The interconnection topology between different computing resources is obtained and a topology adjacency graph is constructed. The topology adjacency graph is traversed and the effective bandwidth and congestion window size are obtained as bandwidth parameters. The fusion operators in the fusion operator group are traversed and aggregated to obtain the fusion operator feature vector. The operation type feature and data scale feature are separated from the fusion operator feature vector. The scheduling feature vector is generated based on the operation type feature and matched with the preset operation mode prototype to determine the dominant operation type. The operation type label is generated based on the dominant operation type. The data throughput requirement is calculated based on the data scale feature and the throughput label is generated. The processing capability parameters are arranged into capability vectors according to computing resources, and a type sensitivity matrix is ​​constructed. A type-weighted capability vector is calculated based on the operation type label and the type sensitivity matrix. A type matching degree is calculated based on the operation type label and the type-weighted capability vector. A capacity matching degree is calculated based on the throughput label and the capability vector, and a weighted sum of the capacity matching degree and the type matching degree is obtained to obtain a comprehensive matching degree. The comprehensive matching degree is filled into a computing affinity matrix based on the fusion operator and the computing resources.

[0049] Obtaining the microarchitecture pipeline depth and functional unit configuration information of each computing resource in a heterogeneous computing environment is fundamental to performance modeling. Microarchitecture pipeline depth refers to the number of pipeline stages required from instruction fetch to execution completion. Functional unit configuration includes the number and type of arithmetic logic units (ALUs), floating-point units (Floating-point units), and vector processing units (Vector units). For example, a general-purpose processor might have a pipeline depth of 14, with 4 integer ALUs, 2 floating-point units, and 1 vector processing unit; a graphics processing unit (GPU) might have a pipeline depth of 24, with 128 single-precision floating-point units, 64 double-precision floating-point units, and 32 tensor cores. Based on these hardware parameters, performance modeling yields the processing power parameters of each computing resource. For instance, a general-purpose processor might have an integer processing power of 12.8 billion operations per second (12.8 billion) and a floating-point processing power of 6.4 billion operations per second (6.4 billion). A GPU might have a single-precision floating-point processing power of 12 trillion operations per second (12 trillion) and a double-precision floating-point processing power of 6 trillion operations per second (6 trillion).

[0050] Obtaining the interconnection topology between different computing resources is a prerequisite for bandwidth modeling. The interconnection topology describes the physical connection methods between computing resources, such as bus, star, ring, and mesh topologies. A topological adjacency graph is constructed based on the interconnection structure, where nodes represent computing resources and edges represent direct connections. For example, in a heterogeneous system containing one main processor, two coprocessors, and one dedicated accelerator, the topological adjacency graph can be represented as a four-node star topology, with the main processor at the center directly connected to the other three nodes. Each edge in the topological adjacency graph is traversed, and the effective bandwidth and congestion window size are measured or obtained as bandwidth parameters. Effective bandwidth refers to the actual available data transfer rate, and the congestion window size refers to the maximum amount of data allowed to be transmitted simultaneously under network congestion control mechanisms. For example, the effective bandwidth between the main processor and coprocessor 1 is 16 GB / s, and the congestion window size is 2 MB; the effective bandwidth between the main processor and the dedicated accelerator is 32 GB / s, and the congestion window size is 4 MB.

[0051] The fusion operators in the fusion operator group are traversed, and the fusion operator feature vectors are aggregated. A fusion operator is a computational unit formed by fusing data unit pairs matched in the aforementioned computationally intensive subset. Each fusion operator's feature vector contains operation type features and data size features. For example, the feature vector of a certain fusion operator might be (0.8, 0.1, 0.0, 0.1, 0.0, 1024, 1024, 1024), where the first 5 elements represent operation type features, and the last 3 elements represent data size features. The operation type features and data size features are separated from the fusion operator feature vectors; the operation type features reflect the computational characteristics of the operator, and the data size features reflect the data volume.

[0052] A scheduling feature vector is generated based on the operation type characteristics. This vector reflects the operator's adaptability to different types of computing resources. The scheduling feature vector is then matched with predefined operation mode prototypes to determine the dominant operation type. Operation mode prototypes are predefined typical computing mode characteristics, such as dense matrix operation mode, sparse matrix operation mode, and reduction operation mode. For example, the scheduling feature vector (0.8, 0.1, 0.0, 0.1, 0.0) has a similarity of 0.96 with the dense matrix operation mode prototype (0.75, 0.15, 0.0, 0.1, 0.0) and a similarity of 0.65 with the sparse matrix operation mode prototype (0.3, 0.3, 0.2, 0.1, 0.1), thus determining the dominant operation type as dense matrix operation. An operation type label, such as "DENSE_MATRIX_COMPUTE", is generated based on the dominant operation type.

[0053] Data throughput requirements are calculated based on data size characteristics. Data throughput requirements refer to the bandwidth needed by the operator to process data, and are related to data size and access patterns. For example, for dense matrix multiplication operations with data size characteristics (1024, 1024, 1024), assuming each element is 4 bytes, the calculated data throughput requirement is 12 GB / s. Throughput labels are generated based on these requirements, such as "HIGH_THROUGHPUT" indicating high throughput requirements and "MEDIUM_THROUGHPUT" indicating medium throughput requirements.

[0054] The processing capacity parameters are arranged into a capacity vector according to the computing resources. For example, in the aforementioned heterogeneous system containing four computing resources, the capacity vector can be represented as a quadruple, corresponding to the processing capacity of each computing resource, such as (128, 64, 12800, 6400) billion operations / second. A type sensitivity matrix is ​​constructed, which describes the processing efficiency of different computing resources for different types of operations. For example, the sensitivity of the four computing resources to dense matrix operations can be represented as (0.4, 0.8, 1.0, 0.6), and the sensitivity to sparse matrix operations can be represented as (0.6, 0.5, 0.4, 1.0).

[0055] A type-weighted capability vector is calculated based on the operation type label and the type sensitivity matrix. This vector considers the processing efficiency of computing resources for a specific operation type. For example, for dense matrix operations, the type-weighted capability vector is (51.2, 51.2, 12800, 3840), which is the original capability vector multiplied by the corresponding sensitivity. The type matching degree is then calculated based on the operation type label and the type-weighted capability vector. This degree reflects the adaptability of computing resources to a specific operation type. For example, for dense matrix operations, the type matching degree for each computing resource might be (0.004, 0.004, 0.76, 0.23).

[0056] Capacity matching degree is calculated based on throughput labels and capacity vectors. The capacity matching degree reflects the degree of matching between the processing capacity of computing resources and the data throughput requirements of the operator. For example, for an operator labeled "HIGH_THROUGHPUT", the capacity matching degree of each computing resource might be (0.1, 0.15, 0.95, 0.8). The overall matching degree is obtained by weighting and summing the type matching degree and the capacity matching degree. Assuming the type matching degree has a weight of 0.7 and the capacity matching degree has a weight of 0.3, the overall matching degree is (0.033, 0.048, 0.817, 0.4).

[0057] The overall matching degree is filled into a computational affinity matrix based on the fusion operator and computational resources. The rows of the computational affinity matrix correspond to the fusion operator, the columns correspond to the computational resources, and the matrix element values ​​represent the affinity between the operator and the resource. For the fusion operator and four computational resources in the previous example, the corresponding rows in the computational affinity matrix are (0.033, 0.048, 0.817, 0.4).

[0058] In this embodiment, by modeling the microarchitecture pipeline depth and functional unit configuration of various computing resources in a heterogeneous computing environment, and combining parameters such as interconnect topology, effective bandwidth, and congestion window for unified characterization, a synergistic quantitative expression of computing and communication capabilities is achieved. This more realistically reflects the actual processing capabilities and transmission bottlenecks of different resources in specific operating scenarios, thereby improving the accuracy and reliability of resource performance evaluation. By performing feature aggregation and decomposition on fusion operators, operation type features and data scale features are extracted respectively, and similarity matching with preset operation mode prototypes is performed to determine the dominant operation type. This achieves refined identification of task computing attributes and can automatically infer the core operation mode based on operator-level semantic features, improving the objectivity and adaptability of task characteristic identification. By constructing a type sensitivity matrix, the performance sensitivity of different computing resources to different operation types is incorporated into the modeling, and the capacity matching degree is calculated in conjunction with data throughput requirements. This achieves a two-dimensional synergistic evaluation of type matching and capacity matching, which can simultaneously take into account the adaptability of operation type and resource carrying capacity, significantly improving the matching accuracy between tasks and resources, and avoiding performance waste or congestion problems caused by type mismatch or insufficient capacity.

[0059] In one alternative implementation, Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset, and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is then obtained by solving the calculated affinity matrix and the memory access affinity matrix, including: Extract the physical address range and access port configuration of storage resources from the access conflict region, construct an address space mapping table based on the physical address range, identify mutually exclusive access paths based on the access port configuration, traverse the memory-intensive subset and extract the read operation set and the write operation set, match the read operation set and the write operation set with the address space mapping table, determine potential conflicting memory access operations and group them according to a preset time window, calculate the concurrent access conflict probability based on the grouping results and generate a conflict penalty coefficient; A bandwidth constraint graph is constructed based on the bandwidth parameters and the mutually exclusive access paths. In the bandwidth constraint graph, a multi-path search is performed with the memory access operation of the fusion operator as the starting point and the storage resource as the ending point to obtain candidate paths and calculate the corresponding path capacity. The transmission delay is obtained based on the path capacity and the amount of memory accessed, and the transmission cost is determined by combining the conflict penalty coefficient. The transmission cost is filled into a memory access affinity matrix based on the fusion operator and the storage resource. Align the calculated affinity matrix and the memory access affinity matrix to construct a joint affinity tensor. Perform slice decomposition on the joint affinity tensor to obtain a preference vector. Solve the mapping relationship based on the preference vector and encode it as an initial mapping scheme.

[0060] Extract the physical address range and access port configuration of storage resources from the access conflict region. The physical address range refers to the actual address distribution of the storage resource in the memory system, and the access port configuration refers to the number and type of channels that can perform read and write operations simultaneously on the storage resource. For example, a cache may have a physical address range of 0x10000000 to 0x1FFFFFFF, configured with 4 read ports and 2 write ports; a shared memory may have a physical address range of 0x20000000 to 0x2FFFFFFF, configured with 8 read ports and 4 write ports. An address space mapping table is constructed based on the physical address range. This table records the mapping relationship from logical addresses to physical addresses and the corresponding storage resource types. For example, logical addresses 0x00100000 to 0x001FFFFF are mapped to physical addresses 0x10100000 to 0x101FFFFF, corresponding to cache; logical addresses 0x00200000 to 0x002FFFFF are mapped to physical addresses 0x20200000 to 0x202FFFFF, corresponding to shared memory.

[0061] Identifying mutually exclusive access paths based on access port configuration. A mutually exclusive access path refers to an access channel that cannot be used simultaneously due to port limitations. For example, for a cache configured with 4 read ports, when there are 5 concurrent read requests, at least one request must wait, forming a mutually exclusive access relationship; for a cache configured with 2 write ports, when there are 3 concurrent write requests, at least one request must wait, also forming a mutually exclusive access relationship. By analyzing the port configuration of each storage resource, potential mutually exclusive access paths can be identified, such as mutual exclusion between cache write operations, where a maximum of 2 write operations can be executed concurrently.

[0062] Traverse the memory-intensive subset and extract the read and write operation sets. Data units within the memory-intensive subset typically contain a large number of memory access operations. These memory access operations are categorized according to read and write types. For example, extract 20 read operations and 15 write operations from a memory-intensive data unit to form read and write operation sets. Match the read and write operation sets with the address space mapping table to determine which physical address region of which storage resource each operation accesses. For example, in the read operation set, 12 operations access the cache, and 8 operations access shared memory; in the write operation set, 5 operations access the cache, and 10 operations access shared memory.

[0063] Identify potentially conflicting memory access operations and group them according to a preset time window. Potentially conflicting memory access operations refer to operations that access the same storage resource and may execute at the same time. Assuming a preset time window of 100 nanoseconds, analyze the execution time of each operation and group operations that may execute within the same time window into the same group. For example, within a 100-nanosecond time window, 6 read operations and 3 write operations simultaneously access the cache, forming one potential conflict group; within another time window, 5 read operations and 7 write operations simultaneously access shared memory, forming another potential conflict group. Calculate the concurrent access conflict probability based on the grouping results. The concurrent access conflict probability reflects the likelihood of a memory access conflict occurring within a given time window and is related to the number of concurrent operations and the number of available ports. For example, the probability of concurrent access conflicts for cache read operations is (6-4) / 6 = 0.33; for cache write operations, the probability is (3-2) / 3 = 0.33; for shared memory read operations, the probability is 0; and for shared memory write operations, the probability is (7-4) / 7 = 0.43. A conflict penalty coefficient is generated based on the concurrent access conflict probability, reflecting the degree of impact of memory access conflicts on performance. For example, the conflict penalty coefficient for cache read operations is 1.33, and the conflict penalty coefficient for cache write operations is 1.33; the conflict penalty coefficient for shared memory read operations is 1.0, and the conflict penalty coefficient for shared memory write operations is 1.43.

[0064] A bandwidth constraint graph is constructed based on bandwidth parameters and mutually exclusive access paths. The bandwidth constraint graph is a directed graph where nodes represent fusion operators and storage resources, edges represent possible data transfer paths, and edge weights represent bandwidth constraints. For example, there are two paths from the fusion operator to the cache, passing through different memory controllers, with bandwidths of 24GB / s and 16GB / s respectively; there is one path from the fusion operator to shared memory with a bandwidth of 32GB / s. In the bandwidth constraint graph, a multi-path search is performed starting from the memory access operation of the fusion operator and ending at the storage resource to obtain candidate paths and calculate the corresponding path capacity. Path capacity refers to the minimum bandwidth on the path, reflecting the maximum data transfer rate that the path can support. For example, the path capacities of the two paths from the fusion operator to the cache are 24GB / s and 16GB / s respectively; the path capacity from the fusion operator to shared memory is 32GB / s.

[0065] The transmission latency is calculated based on path capacity and the amount of data accessed. Transmission latency refers to the time required for data to be transferred from the fusion operator to the storage resource, and it is related to the amount of data and the path capacity. For example, if a fusion operator needs to write 192MB of data to the cache and chooses a path with a capacity of 24GB / s, the calculated transmission latency is 192 MB / 24 GB / s = 8 milliseconds; if the same fusion operator needs to read 320MB of data from shared memory and chooses a path with a capacity of 32GB / s, the calculated transmission latency is 320 MB / 32 GB / s = 10 milliseconds. The transmission cost is determined by combining the transmission latency with a conflict penalty coefficient. The transmission cost considers the pure transmission time and the additional latency caused by memory access conflicts. For example, the transmission cost of writing data to the cache is the transmission latency multiplied by the write operation conflict penalty coefficient, i.e., 8 milliseconds × 1.33 = 10.64; the transmission cost of reading data from shared memory is the transmission latency multiplied by the read operation conflict penalty coefficient, i.e., 10ms × 1.0 = 10. The memory access affinity matrix is ​​filled with the transmission cost based on the fusion operator and storage resources. The rows of the memory access affinity matrix correspond to the fusion operator, and the columns correspond to the storage resources. The matrix element values ​​represent the memory access efficiency between the operator and the resource, typically expressed as the reciprocal or negative value of the transmission cost; a larger value indicates a higher affinity. For example, for the aforementioned fusion operator and two types of storage resources, the corresponding row in the memory access affinity matrix might be (-10.64, -10).

[0066] The computation affinity matrix and memory access affinity matrix are aligned and a joint affinity tensor is constructed. The joint affinity tensor is a three-dimensional data structure, with the three dimensions corresponding to the fusion operator, computational resources, and storage resources, respectively. The tensor element values ​​represent the overall affinity among the three. For example, for the aforementioned fusion operator, four types of computational resources, and two types of storage resources, the corresponding elements in the joint affinity tensor might be (0.033, -10.64), (0.048, -10.64), (0.817, -10.64), (0.4, -10.64), (0.033, -10), (0.048, -10), (0.817, -10), and (0.4, -10), representing the affinity between the fusion operator and different combinations of computational and storage resources, respectively. The joint affinity tensor is then sliced ​​to obtain a preference vector. Slicing decomposition is a dimensionality reduction operation that transforms a three-dimensional tensor into a one-dimensional vector, reflecting the preference of the fusion operator for different resource combinations. For example, by weighted summing of the aforementioned joint affinity tensor, a preference vector of (-10.607, -10.622, -9.853, -10.27, -9.967, -9.982, -9.183, -9.6) can be obtained, representing the preference of the fusion operator for eight different resource combinations; a larger value indicates a higher preference. Based on the preference vector, a mapping relationship is obtained and encoded as an initial mapping scheme. The mapping relationship refers to the allocation scheme of the fusion operator to computing and storage resources. Typically, the resource combination corresponding to the element with the largest value in the preference vector is selected as the mapping scheme. For example, the maximum value in the aforementioned preference vector is -9.183, corresponding to the 7th element, which maps the fusion operator to the combination of the 3rd computing resource and the 2nd storage resource. This mapping relationship is encoded as the initial mapping scheme, serving as the basis for subsequent resource allocation.

[0067] In this embodiment, by performing fine-grained analysis of the physical address range and access port configuration of the access conflict region, constructing an address space mapping table and identifying mutually exclusive access paths, potential conflict sources can be identified at the physical address level and port level. This significantly improves the accuracy and foresight of conflict detection, reduces the performance uncertainty caused by hidden resource contention, and generates a conflict penalty coefficient by performing address matching and time window grouping on the read and write operation set and calculating the probability of concurrent access conflicts. This allows for the prediction of conflict risks during the mapping decision stage, enabling the scheduling strategy to have preventive optimization capabilities. This reduces the overhead of dynamic rollback and rescheduling during the runtime phase, improves the stability of system operation, and performs multi-path search in the bandwidth constraint graph. By combining path capacity, memory access data volume, and conflict penalty coefficient to calculate the transmission cost, the entire link model of the data transmission process is realized. This enables a comprehensive evaluation of path capacity differences and potential congestion impacts, making memory access path selection more reasonable, significantly reducing transmission delays caused by link bottlenecks or port mutual exclusion, and improving overall data flow efficiency.

[0068] In one alternative implementation, Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. Based on the communication overhead vector, the global delay distribution is obtained, including: Data dependencies between fusion operators are extracted from the predecessor-successor relationship set. Fusion operator pairs mapped to different computing resources are identified and cross-resource data transmission edges are marked in conjunction with the initial mapping scheme. Data transmission volume is extracted from the cross-resource data transmission edges and data transmission time is calculated in conjunction with the bandwidth parameter. The topology level crossed by the cross-resource data transmission edge is identified based on the data transmission time and a level penalty factor is generated according to the topology level. The communication overhead of each cross-resource data transmission edge is calculated based on the data transmission time and the level penalty factor. Identify cross-resource data transmission edges that converge to the same fusion operator from the set of predecessor and successor relationships, sort them according to communication overhead, identify the maximum communication overhead, calculate the difference between the maximum communication overhead and other communication overheads to obtain the synchronization overhead, update the communication overhead of each cross-resource data transmission edge based on the synchronization overhead, and arrange the updated communication overhead into a communication overhead vector according to the order of the cross-resource data transmission edges. The computation time of each fusion operator is extracted from the initial mapping scheme. The execution delay of each fusion operator is calculated based on the predecessor-successor relationship set and the communication overhead vector, and statistical analysis is performed to obtain the global delay distribution.

[0069] The data dependencies between fusion operators are extracted from the predecessor-successor relationship set. This set records the execution order and data transfer between fusion operators and is represented as a directed graph, where nodes represent fusion operators and edges represent data dependencies. For example, in a computation graph containing five fusion operators, the output of fusion operator 1 serves as the input to fusion operators 2 and 3, the outputs of fusion operators 2 and 3 serve as the input to fusion operator 4, and the output of fusion operator 4 serves as the input to fusion operator 5. This forms a data flow graph with multiple dependency paths, such as 1→2→4→5 and 1→3→4→5.

[0070] The initial mapping scheme identifies fusion operator pairs mapped to different computing resources and marks cross-resource data transfer edges. The initial mapping scheme determines which computing and storage resources each fusion operator is assigned to. For example, fusion operators 1 and 2 are mapped to computing resource A, fusion operator 3 is mapped to computing resource B, and fusion operators 4 and 5 are mapped to computing resource C. Based on the initial mapping scheme and the previously extracted data dependencies, cross-resource data transfer edges can be identified, including 1→3 (transfer from resource A to resource B), 2→4 (transfer from resource A to resource C), and 3→4 (transfer from resource B to resource C).

[0071] The data transfer volume is extracted from the cross-resource data transfer edges, and the data transfer time is calculated by combining it with the bandwidth parameter. The data transfer volume refers to the amount of data transmitted through the cross-resource transfer edges. For example, the data transfer volume from edge 1 to 3 is 64MB, from edge 2 to 4 is 128MB, and from edge 3 to 4 is 96MB. The bandwidth parameter is the effective bandwidth and congestion window size between different computing resources obtained earlier. For example, the effective bandwidth from resource A to resource B is 12GB / s, from resource A to resource C is 8GB / s, and from resource B to resource C is 10GB / s. Based on the data transfer volume and bandwidth parameter, the data transfer time for each edge can be calculated. The transmission time for edge 1 to 3 is 64 MB / 12 GB / s ≈ 5.33 milliseconds, the transmission time for edge 2 to 4 is 128 MB / 8 GB / s = 16 milliseconds, and the transmission time for edge 3 to 4 is 96 MB / 10 GB / s ≈ 9.6 milliseconds.

[0072] The topological level traversed by cross-resource data transmission edges is identified based on data transmission time, and a level penalty factor is generated accordingly. The topological level reflects the distance between two computing resources in the interconnect structure; a higher level indicates a longer communication path and greater communication overhead. For example, if resources A and B are in the same node, the topological level is 1; if resources A and C are in different nodes but in the same rack, the topological level is 2; and if resources B and C are in different nodes but in the same rack, the topological level is also 2. A level penalty factor is generated based on the topological level, with higher levels resulting in larger penalty factors. For example, the penalty factor for topological level 1 is 1.0, for topological level 2 it is 1.5, and for topological level 3 it is 2.0. Therefore, the level penalty factor for edge 1→3 is 1.0, for edge 2→4 it is 1.5, and for edge 3→4 it is 1.5.

[0073] The communication overhead for each cross-resource data transmission edge is calculated based on data transmission time and a hierarchical penalty factor. The communication overhead considers both the pure data transmission time and the additional latency due to the topology, and is typically the product of the two. For example, the communication overhead for edge 1→3 is 5.33 ms × 1.0 = 5.33 ms, for edge 2→4 it is 16 ms × 1.5 = 24 ms, and for edge 3→4 it is 9.6 ms × 1.5 = 14.4 ms.

[0074] From the set of predecessor and successor relationships, identify cross-resource data transmission edges that converge to the same fusion operator and sort them according to communication overhead to identify the maximum communication overhead. A convergence edge refers to a situation where multiple data transmission edges point to the same fusion operator, indicating that the operator needs to wait for multiple input data to arrive before it can start execution. In the previous example, fusion operator 4 has two input edges 2→4 and 3→4, with communication overheads of 24 milliseconds and 14.4 milliseconds, respectively. After sorting, the maximum communication overhead is 24 milliseconds, corresponding to edge 2→4.

[0075] The synchronization overhead is calculated by subtracting the maximum communication overhead from the other communication overheads. Synchronization overhead reflects the waiting time caused by the inconsistent arrival times of different input data. In the example of fusion operator 4, the maximum communication overhead is 24 milliseconds, and the communication overhead of the other input edge is 14.4 milliseconds. The difference between the two is 24 - 14.4 = 9.6 milliseconds, which is the synchronization overhead of edge 3→4. In summary, although the data from edge 3→4 may arrive first, fusion operator 4 needs to wait for the data from edge 2→4 to arrive before it can begin computation, resulting in a 9.6 millisecond idle waiting time.

[0076] The communication overhead for each cross-resource data transfer edge is updated based on the synchronization overhead. After the update, the communication overhead for edge 3→4 becomes the original communication overhead plus the synchronization overhead, i.e., 14.4 milliseconds + 9.6 milliseconds = 24 milliseconds. The updated communication overhead for both input edges of fusion operator 4 is 24 milliseconds, reflecting the actual execution latency. The updated communication overhead is arranged into a communication overhead vector according to the order of the cross-resource data transfer edges. For example, in the order of edge 1→3, edge 2→4, edge 3→4, the communication overhead vector is (5.33, 24, 24) milliseconds.

[0077] The computation time of each fusion operator is extracted from the initial mapping scheme. Computation time refers to the time required for a fusion operator to execute on a specified computing resource, and is related to the computational complexity of the operator and the processing capacity of the resource. For example, fusion operator 1 takes 8 milliseconds to compute on resource A, fusion operator 2 takes 15 milliseconds to compute on resource A, fusion operator 3 takes 12 milliseconds to compute on resource B, fusion operator 4 takes 20 milliseconds to compute on resource C, and fusion operator 5 takes 10 milliseconds to compute on resource C.

[0078] The execution latency of each fusion operator is calculated based on the predecessor-successor relationship set and communication overhead vector, and the global latency distribution is obtained through statistical analysis. The execution latency includes the operator's own computation time, the time spent waiting for the predecessor operator to complete, and the time for data transmission and synchronization. For operators without predecessors, the execution latency equals the computation time. For operators with predecessors, the execution latency equals the sum of the predecessor operator's execution latency, data transmission time, and possible synchronization waiting time, plus the operator's own computation time.

[0079] For example, fusion operator 1 has no predecessor and its execution latency is 8 milliseconds. Fusion operator 2's predecessor is operator 1, and its execution latency is the execution latency of 1 plus the computation time, i.e., 8 + 15 = 23 milliseconds. Fusion operator 3's predecessor is operator 1, but there is cross-resource communication; its execution latency is the execution latency of 1 plus the communication overhead from 1 to 3 plus the computation time, i.e., 8 + 5.33 + 12 = 25.33 milliseconds. Fusion operator 4's predecessors are operators 2 and 3, requiring both predecessors to complete and data transmission to be in place. Its execution latency is the latency of the latest completion of the predecessors plus the corresponding communication overhead plus the computation time. Operator 2's completion time is 23 milliseconds, plus the 24 millisecond communication overhead from 2 to 4, totaling 47 milliseconds; operator 3's completion time is 25.33 milliseconds, plus the 24 millisecond communication overhead from 3 to 4, totaling 49.33 milliseconds. The maximum of these two is 49.33 milliseconds. Adding the 20 millisecond computation time of operator 4, the execution latency is 69.33 milliseconds. The predecessor of fusion operator 5 is operator 4, and the execution delay is the execution delay of 4 plus the computation time, which is 69.33 + 10 = 79.33 milliseconds.

[0080] The global latency distribution was obtained by statistically analyzing the execution latency of all fusion operators, including statistical indicators such as minimum latency, maximum latency, average latency, and median latency. In the example above, the minimum execution latency was 8 milliseconds, the maximum execution latency was 79.33 milliseconds, the average execution latency was (8+23+25.33+69.33+79.33) / 5≈41 milliseconds, and the median execution latency was 25.33 milliseconds.

[0081] In this embodiment, by extracting the data dependencies between fusion operators from the predecessor-successor relationship set and identifying cross-resource data transmission edges in conjunction with the initial mapping scheme, an explicit characterization of cross-computation resource communication paths is achieved. This enables precise localization of cross-resource transmission behavior at the dependency graph level. Furthermore, by combining bandwidth parameters with topology levels for hierarchical modeling, the differences in communication costs under different topologies are more realistically reflected, improving the accuracy of communication overhead assessment. By generating a hierarchical penalty factor based on the topology level and calculating communication overhead in conjunction with data transmission time, the additional latency impact of cross-level transmission is effectively characterized. This avoids underestimating the performance loss of long-distance or multi-hop transmissions, enhances the system's adaptability to complex interconnection structures, and improves the reliability of overall performance prediction. By identifying multiple cross-resource transmission edges converging to the same fusion operator and calculating synchronization overhead based on the maximum communication overhead, the underestimation of critical path latency is effectively avoided, improving the accuracy of global performance assessment.

[0082] Figure 2 This is a flowchart illustrating the communication overhead analysis of the fusion operator in the data stream parallel acceleration method under a heterogeneous computing environment, as described in an embodiment of the present invention.

[0083] In one alternative implementation, Based on the global delay distribution, critical paths are identified, and an optimized mapping scheme is determined. Based on this optimized mapping scheme, the data unit sequence is allocated to corresponding computing resources, and parallel execution is initiated to obtain the output data stream, including: The sequence of fusion operators with the largest execution latency is identified from the global latency distribution as the critical path and the key fusion operators are extracted. The computing resources mapped by the key fusion operators in the initial mapping scheme are subjected to load analysis and the load saturation is calculated. Based on the load saturation, overloaded computing resources are identified and idle computing resources are searched in the heterogeneous computing environment. The key fusion operators are mapped from the overloaded computing resources to the idle computing resources to obtain the corrected mapping relationship. The communication overhead vector is updated based on the corrected mapping relationship, and the corrected global delay distribution is solved based on the updated communication overhead vector. It is determined whether the maximum execution delay in the corrected global delay distribution is less than the maximum execution delay in the global delay distribution. If it is less, the corrected mapping relationship is retained; otherwise, the corrected mapping relationship is rolled back. The mapping adjustment and delay evaluation are repeated until the maximum execution delay converges to obtain the optimized mapping scheme. The computational resources of each fusion operator mapping are extracted from the optimized mapping scheme, the data units in the data unit sequence are allocated to the corresponding computational resources, each computational resource is started to execute the data units in parallel and perform cross-resource data transmission based on the predecessor and successor relationship set, the data output by the computational resources is collected and spliced ​​to obtain the output data stream.

[0084] The sequence of fusion operators with the longest execution latency is identified from the global latency distribution as the critical path, and the critical fusion operators are extracted. The global latency distribution records the execution latency of each fusion operator. By analyzing the latency values, the critical path, i.e., the path with the longest execution time, can be determined. In the example above, the total execution latency of the operator sequence 1→3→4→5 is 79.33 milliseconds, which is greater than the execution latency of the sequence 1→2→4→5, and therefore it is identified as the critical path. The critical fusion operator is the operator on the critical path that contributes the most to the total latency; it is usually the operator with the longest execution time or located at the bottleneck. In the example above, the execution latency of fusion operator 4 is 49.33 milliseconds, which contributes the most to the total latency, and it is identified as the critical fusion operator.

[0085] Load analysis is performed on the computing resources mapped to the key fusion operators in the initial mapping scheme, and load saturation is calculated. Load analysis primarily examines the matching degree between the processing capacity of the computing resources and the current task load. Load saturation refers to the ratio of the current load of a computing resource to its maximum processing capacity, usually expressed as a percentage. For example, fusion operator 4 is mapped to computing resource C, which has a processing capacity of 8 million instructions per second. Since fusion operator 4 needs to process 7.2 million instructions per second, the load saturation is 7.2 million / 8 million = 90%. Simultaneously, resource C also needs to process fusion operator 5, which has 4 million instructions per second, resulting in a total load saturation of (7.2 million + 4 million) million / 8 million = 140% for resource C, significantly exceeding its maximum processing capacity.

[0086] Overloaded computing resources are identified based on load saturation, and idle computing resources are searched in the heterogeneous computing environment. Computing resources with a load saturation exceeding 100% are identified as overloaded resources. In the example above, computing resource C has a load saturation of 140% and is identified as an overloaded resource. Idle computing resources refer to computing resources with low load saturation and the ability to execute key fusion operators. In the heterogeneous computing environment, there is computing resource D with a processing capacity of 10 million instructions per second, a current load of 2 million instructions per second, and a load saturation of 20%, which is identified as an idle resource.

[0087] The corrected mapping relationship is obtained by remapping the key fusion operator from overloaded computing resources to idle computing resources. In the example, fusion operator 4 is remapped from computing resource C to computing resource D, resulting in the corrected mapping relationship. After correction, resource C only needs to process fusion operator 5, and the load saturation is reduced to 50%; resource D needs to process fusion operator 4, and the load saturation is 720 / 1000=72%, both of which are within a reasonable range.

[0088] The communication overhead vector is updated based on the corrected mapping relationship, and the corrected global delay distribution is solved based on the updated communication overhead vector. The corrected mapping relationship changes the execution location of some fusion operators, requiring recalculation of the communication overhead for cross-resource data transmission. For example, the communication overhead for both edge 2→4 and edge 3→4 is 24 milliseconds. After correction, fusion operator 4 is mapped to resource D, and the communication path changes. The bandwidth from resource A to resource D is 15 GB / s, the topology level is 2, and the level penalty factor is 1.5; the bandwidth from resource B to resource D is 20 GB / s, the topology level is 1, and the level penalty factor is 1.0. After correction, the data transmission time for edge 2→4 is 128 MB / 15 GB / s ≈ 8.53 milliseconds, and the communication overhead is 8.53 × 1.5 ≈ 12.8 milliseconds; the data transmission time for edge 3→4 is 96 MB / 20 GB / s = 4.8 milliseconds, and the communication overhead is 4.8 × 1.0 = 4.8 milliseconds. The maximum communication overhead for converging to fusion operator 4 is 12.8 milliseconds, the synchronization overhead is 12.8 - 4.8 = 8 milliseconds, and the updated communication overhead vector is (5.33, 12.8, 12.8).

[0089] The execution latency of each fusion operator is recalculated based on the updated communication overhead vector. The execution latencies of fusion operators 1 and 2 remain unchanged at 8 milliseconds and 23 milliseconds, respectively. The execution latency of fusion operator 3 also remains unchanged at 25.33 milliseconds. The execution latency of fusion operator 4 needs to be recalculated: the completion time of operator 2 is 23 milliseconds, plus the corrected 2→4 communication overhead of 12.8 milliseconds, totaling 35.8 milliseconds; the completion time of operator 3 is 25.33 milliseconds, plus the corrected 3→4 communication overhead of 12.8 milliseconds, totaling 38.13 milliseconds; taking the maximum of the two, 38.13 milliseconds, and adding the computation time of operator 4 on resource D of 18 milliseconds, where the computation time on resource D is shortened due to improved processing power, the execution latency is 56.13 milliseconds. The predecessor of fusion operator 5 is operator 4, but now the two are located on different resources, and cross-resource communication needs to be considered. The bandwidth from resource D to resource C is 25 GB / s, the topology level is 1, the level penalty factor is 1.0, the data transfer volume is 64 MB, the data transfer time is 64 MB / 25 GB / s ≈ 2.56 milliseconds, and the communication overhead is 2.56 × 1.0 = 2.56 milliseconds. The execution latency of fusion operator 5 is the execution latency of operator 4 plus the communication overhead plus the computation time, i.e., 56.13 + 2.56 + 10 = 68.69 milliseconds.

[0090] The process involves determining whether the maximum execution latency in the corrected global latency distribution is less than the maximum execution latency in the original global latency distribution. In the original global latency distribution, the maximum execution latency was 79.33 milliseconds; in the corrected global latency distribution, the maximum execution latency was 68.69 milliseconds, a reduction of 10.64 milliseconds, demonstrating significant optimization. Because the corrected maximum execution latency is less than the original maximum execution latency, the corrected mapping relationship is retained. If the corrected maximum execution latency is greater than or equal to the original maximum execution latency, the corrected mapping relationship is rolled back, i.e., the original mapping scheme is restored. This mapping adjustment and latency evaluation process is repeated until the maximum execution latency converges, resulting in the optimized mapping scheme. In this embodiment, after one adjustment, the maximum execution latency has been significantly reduced, and the load on each computing resource is balanced; therefore, no further adjustment is needed, and the current corrected mapping scheme is the optimized mapping scheme.

[0091] The computational resources mapped to each fusion operator are extracted from the optimized mapping scheme. In the optimized mapping scheme, fusion operators 1 and 2 are mapped to computational resource A, fusion operator 3 is mapped to computational resource B, fusion operator 4 is mapped to computational resource D, and fusion operator 5 is mapped to computational resource C. Data units in the data unit sequence are allocated to their corresponding computational resources. A data unit refers to a specific block of data to be executed, and it is allocated to the corresponding computational resource according to the mapping relationship of the fusion operators. For example, the data unit corresponding to fusion operator 1 contains a feature map with an input tensor size of 32×32×128 and is allocated to computational resource A; the data unit corresponding to fusion operator 2 contains 16 convolutional kernels and is allocated to computational resource A; the data unit corresponding to fusion operator 3 contains 24 convolutional kernels and is allocated to computational resource B; the data unit corresponding to fusion operator 4 contains 8 fully connected layer parameters and is allocated to computational resource D; and the data unit corresponding to fusion operator 5 contains softmax classifier parameters and is allocated to computational resource C.

[0092] The system initiates parallel execution of data units across computing resources and performs cross-resource data transfer based on the predecessor-successor relationship set. Each computing resource begins parallel processing of tasks according to its assigned data units. For example, computing resource A executes fusion operators 1 and 2 to generate intermediate results; computing resource B executes fusion operator 3 to process the output from operator 1 and generate intermediate results; computing resource D executes fusion operator 4 to process the output from operators 2 and 3 and generate intermediate results; and computing resource C executes fusion operator 5 to process the output from operator 4 and generate the final result. During execution, necessary cross-resource data transfers are performed according to the predecessor-successor relationship set. For example, the output of fusion operator 1 needs to be transferred to resource B for use by operator 3; the output of fusion operator 2 needs to be transferred to resource D for use by operator 4; the output of fusion operator 3 needs to be transferred to resource D for use by operator 4; and the output of fusion operator 4 needs to be transferred to resource C for use by operator 5.

[0093] The data output from computing resources is collected and concatenated to obtain the output data stream. After all computing resources have completed their respective processing tasks, the output results, especially the output of the final fusion operator, are collected and concatenated into a complete output data stream according to a predetermined format. In the aforementioned example, the output of fusion operator 5 is a classification probability vector with a dimension of 1×1000, representing the probability distribution of the input image belonging to each category. This output is collected and returned to the user or further processing modules as the final result of the entire heterogeneous computing system.

[0094] In this embodiment, by analyzing the global latency distribution and identifying the fusion operator sequence with the largest execution latency as the critical path, the key computational links affecting the overall execution efficiency can be accurately located during the system operation modeling stage, effectively reducing the overall execution latency. By evaluating the load saturation of the computing resources where the key fusion operators are located and searching for idle computing resources in the heterogeneous computing environment for mapping adjustment, the critical tasks can be reallocated according to the real-time load status, reducing the restriction of overloaded resources on the critical path, improving the overall utilization efficiency of computing resources, and alleviating local resource bottleneck problems. By updating the communication overhead vector and recalculating the corrected global latency distribution, and comparing and backtracking the optimization results, the optimal execution state can be gradually approached during multiple evaluations and adjustments, thereby improving the stability and optimization effect of mapping decisions and avoiding performance degradation caused by unreasonable migration.

[0095] A second aspect of this invention provides a data stream parallel acceleration system for heterogeneous computing environments, comprising: The parsing unit is used to acquire the data stream to be processed and parse it to obtain a data unit sequence, perform dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traverse the data unit sequence and extract unit feature vectors, perform clustering on the unit feature vectors to obtain a computationally intensive subset and a memory-intensive subset, perform operator fusion analysis on the computationally intensive subset to obtain a fusion operator group, and perform memory access pattern recognition on the memory-intensive subset and mark access conflict regions. The analysis unit is used to obtain the processing capability parameters and bandwidth parameters of computing resources in the heterogeneous computing environment, match the feature vectors corresponding to the fusion operator group with the processing capability parameters to obtain the computing affinity matrix, perform path search on the memory-intensive subset based on the bandwidth parameters and the access conflict region and calculate the transmission cost to obtain the memory access affinity matrix, and solve the initial mapping scheme based on the computing affinity matrix and the memory access affinity matrix. The extraction unit is used to identify cross-resource data transmission edges in the initial mapping scheme based on the predecessor-successor relationship set and solve for the communication overhead vector, solve for the global delay distribution based on the communication overhead vector, identify the critical path according to the global delay distribution and determine the optimized mapping scheme, allocate the data unit sequence to the corresponding computing resources based on the optimized mapping scheme and start parallel execution to obtain the output data stream.

[0096] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0097] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0098] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for accelerating data stream parallelism in heterogeneous computing environments, characterized in that, include: The process involves acquiring and parsing the data stream to be processed to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traversing the data unit sequence and extracting unit feature vectors, clustering the unit feature vectors to obtain computationally intensive subsets and memory-intensive subsets, performing operator fusion analysis on the computationally intensive subsets to obtain a fusion operator group, and performing memory access pattern recognition on the memory-intensive subsets and marking access conflict regions. The processing capability parameters and bandwidth parameters of computing resources in a heterogeneous computing environment are obtained. The feature vectors corresponding to the fusion operator group are matched with the processing capability parameters to obtain the computing affinity matrix. Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is obtained by solving the computing affinity matrix and the memory access affinity matrix. Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. The global delay distribution is obtained based on the communication overhead vector. The critical path is identified according to the global delay distribution to determine the optimized mapping scheme. Based on the optimized mapping scheme, the data unit sequence is allocated to the corresponding computing resources and parallel execution is started to obtain the output data stream.

2. The method according to claim 1, characterized in that, The process involves acquiring and parsing the data stream to obtain a sequence of data units, performing dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, and traversing the data unit sequence to extract unit feature vectors, including: The process involves acquiring a data stream to be processed, parsing the data format of the data stream and identifying data boundaries to obtain a set of data segments, performing semantic parsing on each data segment in the set of data segments and marking the operation type to obtain an operation mark sequence, and encapsulating each data segment into a data unit based on the operation mark sequence and arranging them in chronological order to obtain a data unit sequence. Traverse the data unit sequence, extract the input variable identifier and output variable identifier of each data unit and perform intersection matching to identify variable dependencies. Based on the variable dependencies, construct a directed dependency graph and perform topological traversal on the directed dependency graph. Mark the predecessor data unit set and successor data unit set corresponding to each data unit to obtain the predecessor-successor relationship set. The operation type and variable access pattern of each data unit are extracted from the data unit sequence. The operation type is one-hot encoded to obtain the operation type vector. The access frequency and access span are calculated based on the variable access pattern to obtain the access pattern vector. An adjacency matrix is ​​constructed based on the predecessor and successor relationship set, and the adjacency matrix is ​​spectral decomposed to obtain the graph topology feature vector. The operation type vector and the access pattern vector are combined and linearly projected to obtain the unit feature vector.

3. The method according to claim 1, characterized in that, Clustering the unit feature vectors yields computationally intensive and memory-intensive subsets. Operator fusion analysis is performed on the computationally intensive subsets to obtain a fusion operator group. Memory access pattern recognition and access conflict region marking are performed on the memory-intensive subsets, including: Calculate the variance contribution rate of different feature dimensions in the feature vector of the unit and determine the dominant dimension. Construct a feature subspace based on the dominant dimension. Calculate the distance metric between data units in the feature subspace and determine the cluster center. Iteratively cluster the data units based on the distance metric until the cluster center converges. Calculate the memory access ratio based on the converged clusters and divide the clusters into a computationally intensive subset and a memory-intensive subset based on the calculated memory access ratio. Traverse the data units in the computationally intensive subset and extract the data flow relationship to identify matching data unit pairs. Determine the fusion feasibility of the matching data unit pairs and calculate the number of variable eliminations. Based on the number of variable eliminations, perform filtering and fusion operations on the matching data unit pairs to obtain a fusion operator group. Traverse the data units in the memory-intensive subset and extract the variable access sequence. Calculate the step size feature of the access address and the interval feature of the access time in the variable access sequence. Determine the access pattern based on the step size feature. Identify the access window based on the interval feature. Align the variable access sequence with the time axis based on the access pattern and the access window and detect the overlapping interval of the access address. Determine and record the access conflict area based on the overlapping interval.

4. The method according to claim 1, characterized in that, Obtaining the processing capacity and bandwidth parameters of computing resources in a heterogeneous computing environment, and matching the feature vectors corresponding to the fusion operator group with the processing capacity parameters to obtain a computational affinity matrix includes: The microarchitecture pipeline depth and functional unit configuration of each computing resource in the heterogeneous computing environment are obtained and the processing capacity parameters are obtained by performance modeling. The interconnection topology between different computing resources is obtained and a topology adjacency graph is constructed. The topology adjacency graph is traversed and the effective bandwidth and congestion window size are obtained as bandwidth parameters. The fusion operators in the fusion operator group are traversed and aggregated to obtain the fusion operator feature vector. The operation type feature and data scale feature are separated from the fusion operator feature vector. The scheduling feature vector is generated based on the operation type feature and matched with the preset operation mode prototype to determine the dominant operation type. The operation type label is generated based on the dominant operation type. The data throughput requirement is calculated based on the data scale feature and the throughput label is generated. The processing capability parameters are arranged into capability vectors according to computing resources, and a type sensitivity matrix is ​​constructed. A type-weighted capability vector is calculated based on the operation type label and the type sensitivity matrix. A type matching degree is calculated based on the operation type label and the type-weighted capability vector. A capacity matching degree is calculated based on the throughput label and the capability vector, and a weighted sum of the capacity matching degree and the type matching degree is obtained to obtain a comprehensive matching degree. The comprehensive matching degree is filled into a computing affinity matrix based on the fusion operator and the computing resources.

5. The method according to claim 1, characterized in that, Based on the bandwidth parameters and the access conflict region, a path search is performed on the memory-intensive subset, and the transmission cost is calculated to obtain the memory access affinity matrix. The initial mapping scheme is then obtained by solving the calculated affinity matrix and the memory access affinity matrix, including: Extract the physical address range and access port configuration of storage resources from the access conflict region, construct an address space mapping table based on the physical address range, identify mutually exclusive access paths based on the access port configuration, traverse the memory-intensive subset and extract the read operation set and the write operation set, match the read operation set and the write operation set with the address space mapping table, determine potential conflicting memory access operations and group them according to a preset time window, calculate the concurrent access conflict probability based on the grouping results and generate a conflict penalty coefficient; A bandwidth constraint graph is constructed based on the bandwidth parameters and the mutually exclusive access paths. In the bandwidth constraint graph, a multi-path search is performed with the memory access operation of the fusion operator as the starting point and the storage resource as the ending point to obtain candidate paths and calculate the corresponding path capacity. The transmission delay is obtained based on the path capacity and the amount of memory accessed, and the transmission cost is determined by combining the conflict penalty coefficient. The transmission cost is filled into a memory access affinity matrix based on the fusion operator and the storage resource. Align the calculated affinity matrix and the memory access affinity matrix to construct a joint affinity tensor. Perform slice decomposition on the joint affinity tensor to obtain a preference vector. Solve the mapping relationship based on the preference vector and encode it as an initial mapping scheme.

6. The method according to claim 1, characterized in that, Based on the predecessor-successor relationship set, cross-resource data transmission edges in the initial mapping scheme are identified and the communication overhead vector is obtained. Based on the communication overhead vector, the global delay distribution is obtained, including: Data dependencies between fusion operators are extracted from the predecessor-successor relationship set. Fusion operator pairs mapped to different computing resources are identified and cross-resource data transmission edges are marked in conjunction with the initial mapping scheme. Data transmission volume is extracted from the cross-resource data transmission edges and data transmission time is calculated in conjunction with the bandwidth parameter. The topology level crossed by the cross-resource data transmission edge is identified based on the data transmission time and a level penalty factor is generated according to the topology level. The communication overhead of each cross-resource data transmission edge is calculated based on the data transmission time and the level penalty factor. Identify cross-resource data transmission edges that converge to the same fusion operator from the set of predecessor and successor relationships, sort them according to communication overhead, identify the maximum communication overhead, calculate the difference between the maximum communication overhead and other communication overheads to obtain the synchronization overhead, update the communication overhead of each cross-resource data transmission edge based on the synchronization overhead, and arrange the updated communication overhead into a communication overhead vector according to the order of the cross-resource data transmission edges. The computation time of each fusion operator is extracted from the initial mapping scheme. The execution delay of each fusion operator is calculated based on the predecessor-successor relationship set and the communication overhead vector, and statistical analysis is performed to obtain the global delay distribution.

7. The method according to claim 1, characterized in that, Based on the global delay distribution, critical paths are identified, and an optimized mapping scheme is determined. Based on this optimized mapping scheme, the data unit sequence is allocated to corresponding computing resources, and parallel execution is initiated to obtain the output data stream, including: The sequence of fusion operators with the largest execution latency is identified from the global latency distribution as the critical path and the key fusion operators are extracted. The computing resources mapped by the key fusion operators in the initial mapping scheme are subjected to load analysis and the load saturation is calculated. Based on the load saturation, overloaded computing resources are identified and idle computing resources are searched in the heterogeneous computing environment. The key fusion operators are mapped from the overloaded computing resources to the idle computing resources to obtain the corrected mapping relationship. The communication overhead vector is updated based on the corrected mapping relationship, and the corrected global delay distribution is solved based on the updated communication overhead vector. It is determined whether the maximum execution delay in the corrected global delay distribution is less than the maximum execution delay in the global delay distribution. If it is less, the corrected mapping relationship is retained; otherwise, the corrected mapping relationship is rolled back. The mapping adjustment and delay evaluation are repeated until the maximum execution delay converges to obtain the optimized mapping scheme. The computational resources of each fusion operator mapping are extracted from the optimized mapping scheme, the data units in the data unit sequence are allocated to the corresponding computational resources, each computational resource is started to execute the data units in parallel and perform cross-resource data transmission based on the predecessor and successor relationship set, the data output by the computational resources is collected and spliced ​​to obtain the output data stream.

8. A data stream parallel acceleration system for heterogeneous computing environments, used to implement the method of any one of claims 1-7, characterized in that, include: The parsing unit is used to acquire the data stream to be processed and parse it to obtain a data unit sequence, perform dependency analysis on the data unit sequence to obtain a set of predecessor and successor relationships, traverse the data unit sequence and extract unit feature vectors, perform clustering on the unit feature vectors to obtain a computationally intensive subset and a memory-intensive subset, perform operator fusion analysis on the computationally intensive subset to obtain a fusion operator group, and perform memory access pattern recognition on the memory-intensive subset and mark access conflict regions. The analysis unit is used to obtain the processing capability parameters and bandwidth parameters of computing resources in the heterogeneous computing environment, match the feature vectors corresponding to the fusion operator group with the processing capability parameters to obtain the computing affinity matrix, perform path search on the memory-intensive subset based on the bandwidth parameters and the access conflict region and calculate the transmission cost to obtain the memory access affinity matrix, and solve the initial mapping scheme based on the computing affinity matrix and the memory access affinity matrix. The extraction unit is used to identify cross-resource data transmission edges in the initial mapping scheme based on the predecessor-successor relationship set and solve for the communication overhead vector, solve for the global delay distribution based on the communication overhead vector, identify the critical path according to the global delay distribution and determine the optimized mapping scheme, allocate the data unit sequence to the corresponding computing resources based on the optimized mapping scheme and start parallel execution to obtain the output data stream.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Parallel computing method, chip and storage medium based on low-power AI processor

    CN119739488A

  • Matrix calculation adaptive optimization method and system based on ARM architecture

    CN120744299A