Task processing method and system, electronic equipment, storage medium and program product

By setting the remote data cache space in the multi-processor collaborative processing task, the remote data is cached to reduce remote memory access, the problem of remote data access delay in the multi-processor collaborative processing task is solved, and task processing efficiency and performance are improved.

CN120179580AActive Publication Date: 2025-06-20INSPUR (BEIJING) ELECTRONICS INFORMATION IND CO LTD

Patent Information

Application Number
CN202510660290.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the multi-processor collaborative processing task, the processing efficiency is not high due to the delay in remote data access.

Method used

By setting the remote data cache space locally on the processor, the remote data that may be used during the execution of the pending tasks is cached, thereby reducing the number of remote memory accesses, improving local cache utilization, and reducing communication delay.

Benefits of technology

It effectively improves the efficiency and performance of multi-processor collaborative processing tasks, reduces communication delays between different processors, and further improves task processing performance while ensuring data consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179580A_ABST
    Figure CN120179580A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method and system, electronic equipment, a storage medium and a program product, and relates to the technical field of computers. The method comprises the following steps: when multiple processors cooperatively execute a to-be-processed task, determining that the to-be-processed task belongs to a communication dominant task according to memory access behavior characteristics of the to-be-processed task, and synchronizing far-end data of each processor to a local far-end data cache space. Reading target read data from the far-end data cache space when the far-end read operation is carried out; when far-end write operation is carried out, target write data written into the target local cache space is written into the local far-end data cache space, and meanwhile corresponding data in the target local cache space is invalid and synchronized to the target local cache space. According to the invention, the problem of low calculation efficiency of multiprocessor cooperative processing tasks in the prior art can be solved, the communication delay of the multiprocessor cooperative processing tasks can be reduced, the efficiency of the multiprocessor cooperative processing tasks is effectively improved, and the task processing performance of the multiprocessor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a task processing method, system, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] With the continuous growth of the computing power requirements for tasks to be processed, such as high-performance computing tasks and artificial intelligence tasks, the current multi-processor collaboration method is adopted to replace the single processor to execute such high-computing-power-required tasks.

[0003] In the process of using multi-processor collaboration to process tasks in related technologies, due to the interconnection communication between multi-processor nodes, there is a delay in remote data access, resulting in low processing efficiency of multi-processor collaboration in processing tasks. Summary of the Invention

[0004] The present invention provides a task processing method, system, electronic device, computer-readable storage medium, and computer program product, which reduce the communication delay of multi-processor collaboration in processing tasks, effectively improve the efficiency of multi-processor collaboration in processing tasks, and enhance the task processing performance of multi-processors.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: On the one hand, the present invention provides a task processing method, including: When multi-processors collaborate to execute a task to be processed, if it is determined that the task to be processed is a communication-dominated task according to the memory access behavior characteristics during the execution of the task to be processed, the remote data of each processor is respectively synchronized to the corresponding remote data cache space; the local cache space of each processor includes the remote data cache space; when the task to be processed performs a remote read operation, the target read data is read from the remote data cache space; when the task to be processed performs a remote write operation, the target write data written to the target local cache space of the target processor is written to the remote data cache space, and at the same time, the corresponding data in the target local cache space is invalidated, and data synchronization processing is performed on the target local cache space according to the remote data cache space; wherein, the memory access behavior characteristics are determined according to the probability that the data located at the target time and / or target position is accessed again within the target time period, and the communication-dominated task is a task whose remote access data frequency during the task execution process meets the preset communication overhead condition.

[0006] On the other hand, the present invention provides a task processing system, which at least includes a task processor, a first processor and a second processor interconnected; wherein, the first processor includes a first local cache, and the first local cache includes a first remote data cache space, the second processor includes a second local cache, and the second local cache includes a second remote data cache space; the first processor remotely accesses the second local memory, and the first remote data cache space caches the data of the second local cache, the second processor remotely accesses the first local cache, and the second remote data cache space caches the data of the first local cache; when the task processor executes the computer program stored in the memory, the steps of any one of the above task processing methods are implemented.

[0007] The present invention also provides an electronic device, which includes a memory and a processor, and the processor is used to implement the steps of any one of the above task processing methods when executing the computer program stored in the memory.

[0008] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above task processing methods are implemented.

[0009] Finally, the present invention also provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the above task processing methods are implemented.

[0010] The advantages of the technical solution provided by the present invention are as follows: based on the fact that the execution process of the task to be processed is affected by the spatial and temporal locality characteristics, resulting in significant differences in the computational intensity, data access pattern and communication requirements of different types of tasks, the spatial and temporal locality characteristics of the task to be processed can be used to identify whether the task to be processed is a communication-dominated task involving a large amount of remote data access. When the task to be processed belongs to a communication-dominated task, the remote data that may be used during the execution of the task to be processed is cached in the remote data cache space set locally in the processor, so as to reduce the number of remote memory accesses, improve the utilization rate of the local cache, and effectively reduce the communication delay between different processors, effectively improving the efficiency of multi-processor collaborative task processing and the performance of multi-processor collaborative task processing. Further, when modifying local data, the remote data with data synchronization is invalidated, so as to ensure the correctness of the remote data when it is read by other threads or operations when the data in the local cache is modified or not, and the performance of multi-processor collaborative task processing can be further improved on the premise of ensuring data consistency. In addition, the present invention also provides a corresponding implementation system, electronic device, computer-readable storage medium and computer program product for the task processing method, further making the method more practical, and the corresponding system, electronic device, computer-readable storage medium and computer program product have corresponding advantages. Description of the Drawings

[0011] To more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 Schematic diagram of the hardware composition framework applicable to the task processing method provided by the present invention; Figure 2 Schematic flow chart of a task processing method provided by the present invention; Figure 3 Schematic flow chart of a remote read operation provided by the present invention; Figure 4 Schematic flow chart of a remote write operation provided by the present invention; Figure 5 Schematic flow chart of data synchronization provided by the present invention; Figure 6 Another schematic flow chart of a remote read operation provided by the present invention; Figure 7 Another schematic flow chart of a remote write operation provided by the present invention; Figure 8 Structural framework diagram under an exemplary embodiment of the task processing device provided by the present invention; Figure 9 Structural framework diagram under an exemplary embodiment of the task processing system provided by the present invention; Figure 10 Schematic diagram of task recognition model training and deployment in an exemplary application scenario provided by the present invention; Figure 11 Schematic flow chart of the task processing process of the task processing system in an exemplary application scenario provided by the present invention. Detailed implementation manners

[0013] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners. Among them, the terms "first", "second", etc. in the specification and the above drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described here as "exemplary" does not have to be construed as superior to or better than other embodiments.

[0014] Thanks to the continuous optimization of the hardware architecture and the improvement of the software ecosystem, the application scope of GPUs (Graphics Processing Units) in accelerated computing has been greatly expanded and is applied to multiple fields such as high-performance computing, deep learning, and scientific computing to perform corresponding computing tasks. Although the performance of GPUs is constantly improving, as the semiconductor process gradually approaches the physical limit, the slowdown or even failure of Moore's Law has become a consensus in the industry, which makes it difficult for the computing power growth of a single-GPU computing system to continuously meet the growing computing power requirements of tasks. Taking high-computing-power-demand tasks such as high-performance computing and artificial intelligence training tasks as examples, with the continuous growth of data volume and the growth of users' computing performance requirements, the demand for computing power resources for high-computing-power-demand tasks is increasing. The single-GPU computing architecture has problems such as bandwidth bottlenecks, storage access latency, and limited scalability, and can no longer meet the resource requirements of high-computing-power-demand tasks. To solve the problem that a single GPU cannot meet the growing demands for computing and storage, currently, multiple processors are interconnected using high-bandwidth interconnection technologies to improve the overall computing performance. For example, multiple GPUs are interconnected into a multi-GPU computing platform through NVLink (a name of a high-speed interconnection technology), Infinity Fabric (a name of a high-speed interconnection technology), and PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus) to optimize the data transfer efficiency between GPUs. At the programming level of multi-GPU systems, parallel computing frameworks for efficient computing on GPUs, such as OpenCL (Open Computing Language, an open heterogeneous computing framework) and CUDA (Compute Unified Device Architecture, a parallel computing platform and programming model), provide interfaces for programmers to launch thousands of work items in a SPMD (Single Program, Multiple Data) manner on GPUs. Through these frameworks, work items form wavefronts on GPUs and execute in a lock-step manner, and multiple wavefronts form workgroups and are assigned to the same GPU. To simplify the development process and reduce the need to modify legacy code, a unified multi-GPU programming model has emerged. Under this model, programmers no longer need to specify the ID (identification information) of a specific GPU. All GPUs share a unified address space, and the communication process is implicit. Although the interconnection bandwidth is constantly increasing, its performance still lags behind that of local memory. The bandwidth of local memory access is usually 12 times that of remote memory access, which makes the optimization of remote memory access a crucial task in multi-processor systems.In addition, the internal architecture of the GPU is also continuously optimized. Multiple shader engines and compute units can effectively utilize the L1 cache and L2 cache, thereby improving the access efficiency. To reduce the remote access bottleneck, the optimization scheme needs to consider not only the hardware architecture but also improve the speed and efficiency of remote memory access.

[0015] To solve the problem that the performance of multi-processor systems such as multi-GPU computing platforms is limited by the bandwidth bottleneck under the NUMA (Non-Uniform Memory Access) structure, a related technology provides a page migration mechanism for runtime page partitioning, which can evenly distribute data among different GPUs. Page migration realizes data sharing by migrating data pages from one processor such as a GPU to another processor. However, this method will introduce a large performance overhead. Especially when the same page is frequently shared between processors, the "ping-pong effect" may occur, that is, frequent page migration leads to a serious decline in performance. That is to say, for shared pages, cross-GPU access is still required, which will still increase the communication overhead between multi-processors. Another related technology copies remote data to the local memory of each GPU, thereby reducing the cross-NUMA access latency. Although it can reduce the communication cost between different processors, the additional data replication will significantly increase the storage overhead, especially in the application scenario of large-scale working sets. There is also a related technology that uses DCA (Direct Cache Access) to directly access remote memory at the cache line granularity. Although it can avoid the communication overhead of page migration and also reduce the storage overhead, it is still limited by the communication latency of the inter-GPU interconnect, resulting in low computing efficiency for multi-processor collaborative processing tasks.

[0016] In view of this, to solve the problem that the distributed computing efficiency of multi-processors is limited by the communication latency between processors, when the present invention is used for multi-processor collaborative processing tasks, when it is identified that the task to be processed is a communication-dominated task according to the memory access behavior characteristics of the task to be processed, the remote data of each processor is synchronized to the local remote data cache space. When the task to be processed performs a remote read operation, the target read data is read from the local remote data cache space; when the task to be processed performs a remote write operation, the target write data to be written to the remote cache space is first written to the local remote data cache space, and at the same time, the corresponding data in the target local cache space is invalidated. Finally, the corresponding data in the target local cache space is synchronized and updated according to the local remote data cache space. Combining with the specific application environment architecture or specific hardware architecture on which the execution of the task processing method provided by the present invention depends, the specific application environment architecture or specific hardware architecture is described herein. Next, in combination with Figure 1Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content: Multiple GPUs 10 are interconnected through PCIe to form a multi-processor computing platform 11. The multi-processor computing platform can execute the to-be-processed tasks with high computing power requirements through the cooperation of multiple GPUs. In the multi-processor computing platform, each GPU supports remote memory access. For example, it can access the memory resources of other GPUs through the DCA method, and a storage space is pre-opened in the local cache space of each GPU as the remote data cache space. The remote data cache space can pre-store the data of other GPUs or temporarily store the data written to other GPUs.

[0017] The user issues the to-be-processed task to the multi-processor computing platform 11 through the human-computer interaction component of the user terminal 12 and receives the task execution result sent by the multi-processor computing platform 11. If the multi-processor computing platform 11 determines that the to-be-processed task is a communication-dominated task according to the memory access behavior characteristics during the execution of the to-be-processed task, each GPU synchronizes the remote data of other GPUs to the local remote data cache space. When performing a remote read operation, the target read data is first read from the local remote data cache space; when performing a remote write operation, the target write data written to the target local cache space of the target processor can be first written to the local remote data cache space, and the corresponding data in the target local cache space is invalidated at the same time. After the writing is completed, data synchronization processing is performed on the target local cache space according to the local remote data cache space, reducing the communication delay of multiple GPUs in cooperative task processing, effectively improving the task processing efficiency of the multi-processor computing platform 11, and enhancing the task processing performance of the multi-processor computing platform 11.

[0018] It should be noted that the above application scenarios are only shown for the convenience of understanding the idea and principle of the present invention, and the embodiments of the present invention are not restricted in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, the various non-limiting embodiments of the present invention will be described in detail below with reference to the drawings and specific embodiments. First, please refer to Figure 2 , Figure 2 which is a schematic flow chart of a task processing method provided in this embodiment. This embodiment may include the following content: S201: When multiple processors cooperate to execute a to-be-processed task, if it is determined that the to-be-processed task is a communication-dominated task according to the memory access behavior characteristics during the execution of the to-be-processed task, the remote data of each processor is synchronized to the corresponding remote data cache space respectively.

[0019] In this embodiment, during the execution of different tasks on a computing platform composed of multiple processors, due to the influence of the spatio-temporal locality characteristics, there are significant differences in the computational intensity, memory access pattern, and communication requirements. The spatio-temporal locality characteristics refer to the certain regularity of data or instructions during data access in the process of task processing, that is, the probability that the target data at the target time and / or target location is accessed again within the target time period. The spatio-temporal locality characteristics include temporal locality characteristics and spatial locality characteristics. Based on the loop structure, recursive calls, and frequently accessed variables in a computer program, the phenomenon that some data accessed at a certain point in time is very likely to be accessed again in the near future is the temporal locality. Based on the fact that a computer program tends to access memory addresses that are close to each other, if a program accesses a certain storage location or instruction, then it is very likely to access the data or instruction adjacent to its storage location in the near future, that is, the spatial locality. For example, in array operations, once an element is accessed, it is very likely that the adjacent element will be accessed next. The spatio-temporal locality characteristics of the task to be processed in this embodiment are used to identify the task type of the task to be processed. The task types include communication-dominated tasks and computation-dominated tasks. A communication-dominated task is a task in which the frequency of remote access to data during task execution meets a preset communication overhead condition. The preset communication overhead condition means that the frequency of remote access to data is so high that the communication delay caused by the interconnection between different processors will affect the processing performance of the entire task to be processed, or in other words, the communication delay between different processors accounts for a relatively high proportion in the overall delay. Those skilled in the art can determine it according to the actual situation. For example, communication-dominated tasks may be tasks such as matrix multiplication, matrix transpose, SC (Simple Convolution), etc. to be processed. A computation-dominated task is a task in which computation is dominant and the communication overhead between different processors is small, such as tasks like AES (advanced encryption standard), FIR (finite impulse response), etc. Since the task type is caused by the spatio-temporal locality characteristics, correspondingly, in order to identify the task type of the task to be processed, memory access behavior characteristics determined according to the probability that the data at the target time and / or target location is accessed again within the target time period can be extracted. The memory access behavior characteristics at least include cache behavior-related characteristics, remote access-related characteristics, and memory access operation-related characteristics. Cache behavior-related characteristics may be, for example, L1 cache hit rate, L2 cache hit rate, cache line utilization rate. Through the L1 cache hit rate and L2 cache hit rate, it can be reflected whether there is cache access across cores or different processors. The cache line utilization rate refers to the proportion of the cache line that is effectively used during a single memory access. A lower utilization rate may indicate a large number of random accesses or low spatial locality.Remote access related features may include, for example, the ratio of remote cache access, the local / remote memory access ratio. The ratio of remote cache access can determine whether a task highly depends on the remote cache, which may lead to high communication latency. The local / remote memory access ratio is used to reflect whether the task to be processed mainly uses local memory or remote memory. Memory access operation related features may include, for example, the type of memory access operation and the average memory access latency. The type of memory access operation is the ratio of Load (read) to Store (write). Different tasks may exhibit different memory access patterns. The average memory access latency can reflect whether the task to be processed is affected by the memory access bottleneck.

[0020] In this step, when it is determined that the communication overhead between different processors for the task to be processed is large, considering that the cache utilization of the task to be processed directly affects the communication overhead. For example, BS (Bitonic Sort, parallel sorting) and SC exhibit strong spatial locality, resulting in a high L1 cache hit rate, thus reducing the remote memory access demand and lowering the communication latency between different processors. For tasks such as matrix multiplication and matrix transpose, although they have a certain degree of spatial locality, due to the large computational scale, the utilization rate of the L1 cache is low, resulting in more data needing to be directly fetched from the remote memory, thereby exacerbating the communication bottleneck. To improve the overall computing performance of multi-processor collaborative processing tasks, based on accurately characterizing the computing and communication patterns of tasks, the communication latency can be further reduced by effectively utilizing the cache. In this embodiment, a storage space will be pre-allocated in the local cache space of each processor to cache the data of other processors, that is, the remote data of this processor. For the sake of easy description, it is defined as the remote data cache space. For communication-dominated tasks, that is, cache-sensitive tasks, the remote data can be triggered to be cached locally, that is, the remote data of each processor is synchronized to the corresponding local remote data cache space. In this way, during the execution of the task to be processed, the fast local access of data can be maximally ensured, thereby effectively reducing the communication latency of communication-dominated tasks.

[0021] S202: When the task to be processed performs a remote read operation, read the target read data from the remote data cache space; when the task to be processed performs a remote write operation, write the target write data written to the target local cache space of the target processor to the remote data cache space, and at the same time invalidate the corresponding data in the target local cache space, and perform data synchronization processing on the target local cache space according to the remote data cache space.

[0022] In the previous step, after identifying that the task to be processed is a communication-dominated task, the local cache spaces of each processor will cache the data of other processors, and this data is the remote data. In this way, when a processor needs to use the data of other processors, it does not need to perform a remote access across processors and can directly read from the local cache space. During the execution of the task to be processed, it at least includes a data reading operation and a data storage operation required for the intermediate process or the final result. Among them, the target read data is the data that needs to be read for executing the task to be processed, and the target write data is the intermediate data newly generated during the execution of the task to be processed or the calculation result, and these data need to be stored. Moreover, the target write data is stored at a specified position of a specified processor, and the specified processor is the target processor of this step, and the specified position is the target local cache space of this step. To maximize the cache utilization rate, when reading remote data, it is preferred to read from the local remote data cache space first. When writing remote data, it is first written into the local remote data cache space, and then the corresponding remote data is invalidated. A trigger data synchronization condition can be specified in advance, such as when the occupancy of the remote data cache space exceeds a preset threshold or is synchronized every predetermined time, so as to improve the cache utilization rate, reduce remote memory access, and thus reduce the overall calculation latency on the premise of ensuring data consistency.

[0023] Considering that the local storage capacity is limited, and to avoid a large storage overhead and reduce the demand for the memory resources of the processor, after identifying that the task to be processed is a computation-dominated task, there is no need to trigger the local caching of remote data, that is, there is no need to synchronize the remote data of each processor to the corresponding remote data cache space, and it can be executed according to the task execution process recorded in any relevant technology, and the present invention does not make any limitation thereto.

[0024] In the technical solution provided in this embodiment, based on the fact that the execution process of the task to be processed is affected by the spatial and temporal locality characteristics, resulting in significant differences in the computational intensity, data access pattern, and communication requirements among different types of tasks, it is possible to identify whether the task to be processed is a communication-dominated task involving a large amount of remote data access according to the spatial and temporal locality characteristics of the task to be processed. When the task to be processed belongs to the communication-dominated task, the remote data that may be used during the execution of the task to be processed is cached through the remote data cache space set locally in the processor, so as to be able to reduce the number of remote memory accesses, improve the local cache utilization rate, and further effectively reduce the communication latency between different processors, effectively improve the efficiency of multi-processor collaborative task processing, and improve the performance of multi-processor collaborative task processing. Further, when modifying local data, the remote data with data synchronization is invalidated, so as to ensure the correctness of the remote data when it is read by other threads or operations when the data in the local cache is modified or unmodified, and the performance of multi-processor collaborative task processing can be further improved on the premise of ensuring data consistency.

[0025] As can be seen from the above embodiments, there are significant differences in computing load, data access patterns, communication requirements, etc. among different tasks. Accurately identifying the task type of the task to be processed can effectively improve the task processing performance of multi-processors. The above embodiments do not limit how to classify and identify the task to be processed based on the spatio-temporal locality characteristics. This embodiment also provides a more reasonable, simple and effective implementation method for task type identification, which may include the following content: Pre-train the target machine learning model using the task identification sample set until the preset model training stop condition is reached to obtain the task identification model, and convert the task identification model to generate a task identification lookup table; obtain the memory access command stream of the task to be processed, and extract the memory access behavior characteristics to be processed from the memory access command stream; input the memory access behavior characteristics to be processed into the task identification lookup table, and determine whether the task to be processed is a communication-dominated task or a computation-dominated task according to the task identification lookup table.

[0026] Among them, each task identification sample in the task identification sample set is a sample memory access behavior characteristic carrying a communication-dominated task label or a computation-dominated task label. The number of task identification samples can be flexibly determined according to the actual application scenario. The sample access behavior characteristic of the task identification sample refers to the characteristics related to the memory access behavior extracted from the task identification sample, and the type and dimension of the characteristics it contains should be at least the same as those of the memory access behavior characteristics of the task to be processed. That is to say, on the basis of the type of characteristics included in the memory access behavior characteristics of the task to be processed, the type of characteristics included in the sample access behavior characteristics may also include the characteristics of other types of characteristics. The task identification samples can be constructed by obtaining data from the existing database at present. Of course, in order to obtain better model performance, they can be constructed from the historical memory access command stream of the computing platform composed of multi-processors. Exemplarily, the computing platform composed of multi-processors is defined as the hardware end to be deployed. Multiple historical memory access command streams can be obtained from the hardware end to be deployed, and the multi-dimensional historical memory access characteristics of the multiple historical memory access command streams can be extracted. According to the label type corresponding to the multi-dimensional historical memory access characteristics and the sample memory access behavior characteristics, a task identification sample is constructed. The initial memory access characteristics at least include cache behavior-related characteristics, remote access-related characteristics, and memory access operation-related characteristics; the label type is a communication-dominated task or a computation-dominated task. Repeat the above steps continuously according to the number of task identification samples until the total number of task identification samples included in the task identification sample set is reached. The preset model training stop condition can be, for example, that the number of iterations reaches the preset total number of iterations, the model converges, and the prediction accuracy is greater than the preset accuracy threshold.

[0027] Among them, the task identification lookup table is used to represent the corresponding relationship between the output and input of the trained task identification model. In order to improve the overall task execution efficiency, the present invention does not directly use the trained task identification model to identify the task type, but identifies it by means of a lookup table. The index access of the lookup table improves the recognition speed of the type of the task to be identified, and compared with the deployment of the machine learning model, it is more conducive to deploying it on the hardware end constructed by the multi-processor through the hardware writing language. In order to avoid description, the task identification model of the task type used to execute the task to be processed is defined as the task identification model to be deployed. The task identification model to be deployed is a task identification model or a task identification model after pruning. The task identification model after pruning can take into account both model performance and resource requirements. According to the mapping relationship between the input features of the task identification model to be deployed and the predicted output results, a task identification lookup table is generated; the task identification lookup table is encoded, for example, it can be encoded into a binary format to meet the format of the hardware end to be deployed, and the encoded task identification lookup table is deployed on the hardware end to be deployed, and then embedded in the front-end control unit of the instruction pipeline of the multi-processor.

[0028] Furthermore, in order to improve the model prediction performance and improve the model training efficiency, based on the above embodiments, the multi-dimensional historical memory access features of multiple historical memory access command streams can also be processed, including but not limited to important feature screening and / or feature outlier removal and / or feature format standardization processing.

[0029] For ease of description, the multidimensional historical memory access features directly extracted from multiple historical memory access command streams are defined as multidimensional initial historical memory access features. An exemplary implementation of feature screening is: according to the impact of each dimension of the multidimensional initial historical memory access features on the predicted output results of the target machine learning model, the target historical memory access features that meet the preset contribution conditions are selected from the multidimensional initial historical memory access features, and the redundant features of each target historical memory access feature are removed to obtain the sample memory access behavior features; according to the label type corresponding to the multidimensional initial historical memory access features and the sample memory access behavior features, the task identification samples are constructed. An exemplary implementation of feature format standardization and outlier removal is: removing outliers from the multidimensional initial historical memory access features; converting the numerical features in the multidimensional initial historical memory access features into data that conforms to the standard normal distribution; converting the proportional features in the multidimensional initial historical memory access features into a format that meets the input features of the target machine learning model.

[0030] In this embodiment, meeting the preset contribution condition is a screening condition for the feature importance set in advance. For example, the multi-dimensional initial historical memory access features include N initial historical memory access features. Sort their contribution degrees from large to small, and take the first m initial historical memory access features as the features with large contribution degrees, where N > m. That is, the preset contribution condition is the first m features with large contribution degrees. For example, each dimension of the multi-dimensional initial historical memory access features may include: L1 cache hit rate, L2 cache hit rate, cache line utilization rate, remote cache access ratio, local / remote memory access ratio, memory access operation type, and average memory access latency. Calculate the importance or contribution degree of each dimension of the initial historical memory access features respectively. For example, the SHAP value (SHapley Additive exPlanations) of each initial historical memory access feature can be calculated, and the target historical memory access features that meet the preset contribution condition are determined by analyzing the SHAP values of each initial historical memory access feature. The target historical memory access features may be, for example, L1 cache hit rate, L2 cache hit rate, remote cache access ratio, memory access operation type, and average memory access latency. The method based on information gain or mutual information can be used to count whether there is duplicate statistical data, such as whether the L1 cache hit rate and L2 cache hit rate are repeatedly statistically analyzed. The repeatedly statistically analyzed features are redundant features. The 3σ principle can be used to process outliers to identify whether there are outliers. For numerical features, such as cache hit rate and remote access ratio, Z-score (the name of the normalization processing method) normalization processing is performed; the Load / Store ratio is converted into a percentage form.

[0031] Furthermore, the simplification of the task recognition model may include further simplifying the feature dimension and / or further reducing the model scale. For example, deleting the input feature dimension of the task recognition model, such as further selecting more important features. For example, if the input feature dimension of the task recognition model is 8, the feature dimension simplification operation can only retain the top 5 features ranked by importance. Taking the above embodiment as an example, when the input feature dimension of the task recognition model includes L1 cache hit rate, L2 cache hit rate, remote cache access ratio, memory access operation type, and average memory access latency, features with higher importance, such as remote cache access ratio and average memory access latency, can be selected from these input features. The model scale reduction can be achieved by reducing the model structure parameters of the task recognition model. The model structure parameters include, for example, tree depth, number of neural network layers, number of neurons in each neural network layer, depth or width of the encoder / decoder.

[0032] As can be seen from the above, in this embodiment, a task recognition model for classifying and recognizing a task to be processed based on spatio-temporal locality features is constructed in combination with a machine learning model, which can automatically select relatively important features. By screening the feature of the sample data for training the model, it is beneficial to improve the model recognition performance and the model training efficiency. By removing the outliers in the sample data, the model recognition performance is further improved. By performing a normalization process on the input features, it is beneficial to improve the model recognition performance and the model training efficiency, and it is beneficial to accurately recognize the task type of the task to be processed, thereby effectively improving the task processing performance of the multi-processor. Further, the task type recognition efficiency can be improved through the index access of the task recognition lookup table. By simplifying the trained task recognition model, the occupation of hardware computing resources and storage resources can be reduced, and the resource requirements of the task to be processed can be reduced.

[0033] In order to further improve the model performance efficiency, an integerized splitting point can also be used to avoid floating-point calculations. For the sake of description, the simplified task recognition model is defined as the initial task recognition model to be deployed. The floating-point input features of the initial task recognition model to be deployed are converted into integer input features to obtain a task recognition model to be deployed that can be directly deployed on the hardware to be deployed. Of course, in order to improve the task recognition efficiency, the task recognition model to be deployed can be converted into a task recognition lookup table, and the task recognition lookup table is deployed on the hardware to be deployed.

[0034] The above embodiment does not make any limitation on how to train the target machine learning model using the task recognition sample set until the preset model training stop condition is reached. Based on the above embodiment, the present invention also provides a training method for the task recognition model, which may include the following content: When a model parameter setting instruction is received, the basic parameters and hyperparameter search space information are obtained by parsing the model parameter setting instruction; based on the model structure of the random forest model, a target machine learning model is constructed according to the number of decision trees, the maximum tree depth, and the minimum number of leaf node samples; a grid search is performed using a preset model performance evaluation condition to determine the optimal number of trees, the optimal maximum depth, and the optimal minimum number of leaf samples in the hyperparameter search space information; the model parameters of the target machine learning model are configured according to the optimal number of trees, the optimal maximum depth, and the optimal minimum number of leaf samples.

[0035] In this embodiment, since the task recognition type of the task to be processed belongs to a classification problem and the feature dimension is relatively limited, that is, it involves numerical features related to memory access behavior. Considering the non-linear relationship and computational efficiency of memory access features comprehensively, the random forest model can handle mixed numerical features, improve the generalization ability of the model by integrating multiple decision trees, has good generalization ability and interpretability, can automatically select relatively important features, and is suitable for complex feature relationships. Therefore, the random forest model can be used to construct. When constructing the random forest model, basic parameters and hyperparameters need to be set. The basic parameters at least include the number of decision trees, the maximum tree depth, and the minimum number of samples in leaf nodes. For example, the number of decision trees n_estimators = 20, the maximum tree depth max_depth = 4, and the minimum number of samples in leaf nodes min_samples_leaf = 5. The hyperparameter search space information at least includes the range of the number of trees, the range of the maximum depth, and the range of the minimum number of leaf samples. For example, the number of trees (10 - 20), the maximum depth (3 - 5), and the minimum number of leaf samples (3 - 10). The preset model performance evaluation condition can, for example, use 5-fold cross-validation for grid search, and then select the parameter combination with the highest F1 score. For example, n_estimators = 15, max_depth = 4, min_samples_leaf = 5. Of course, other model performance evaluation conditions can also be used, and the present invention does not make any limitation thereto.

[0036] As can be seen from the above, this embodiment combines a random forest model to construct a task recognition model for classifying and recognizing the task to be processed based on spatio-temporal locality features. The task recognition model can not only handle mixed numerical features, but also has good generalization ability and interpretability by integrating multiple decision trees. It can automatically select relatively important features, accurately mine complex feature relationships, and then accurately depict the computing and communication patterns of the task, accurately recognize the task type of the task to be processed, which is beneficial to optimizing the computing and communication efficiency, and thus effectively improving the task processing performance of multi-processors.

[0037] When configuring the model parameters of the target machine learning model according to the optimal values of the number of trees, the maximum depth, and the minimum number of samples in a leaf, a task recognition model to be trained is obtained; the task recognition sample set is divided into a training sample set and a test sample set according to a pre-set ratio, such as an 8:2 ratio. The training sample set is used to train the task recognition model to be trained, and the precision rate, recall rate, and F1 score are calculated in each fold of cross-validation. When the difference in the F1 scores of the communication-dominated task prediction result and the computation-dominated task prediction result output by the task recognition model to be trained is less than a preset prediction threshold, such as 5%, the training of the task recognition model to be trained is ended. The test sample set is used to verify the task recognition model to be trained. If the average accuracy of the prediction results of each test sample is greater than a preset accuracy threshold, such as 90%, the task recognition model to be trained is a well-trained task recognition model. If the average accuracy of the prediction results of each test sample is greater than the preset accuracy threshold, the optimal values corresponding to the number of trees, the maximum depth, and the minimum number of samples in a leaf are determined again in the hyperparameter search space information. According to the construction method of the task recognition samples in the above embodiment, a preset number of task recognition samples are re-obtained, such as 1 / 3 of the number of the original task recognition sample set, and they are randomly replaced with the corresponding number of old task recognition samples in the task recognition sample set as the training sample set for the current training to continue training and adjusting the hyperparameters of the task recognition model to be trained until the average accuracy of the prediction results of each test sample is greater than the preset accuracy threshold.

[0038] As can be seen from the above, during the training process of the task recognition model, by readjusting the optimal values of the hyperparameters in the hyperparameter space and reconstructing new training samples to randomly replace the old training samples, the prediction performance of the task recognition model is effectively improved on the basis of maintaining the training efficiency of the task recognition model.

[0039] The above embodiments do not make any limitations on how to effectively reduce remote memory access and improve cache utilization for the to-be-processed tasks with significant spatio-temporal locality. This embodiment comprehensively considers the characteristics of the to-be-processed tasks and the constraints of the underlying hardware architecture to propose an implementation method for reducing the communication delay between different processors, which may include the following content: For remote read operations, such as Figure 3As shown, for the sake of avoiding description, the processor currently executing the task to be processed is defined as the first processor, and the processor where the remote data read by the first processor is located is defined as the second processor. The implementation process of the remote read operation may include: when the target read data required for the first processor to execute the task to be processed is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space. If the target read data is not in the remote data cache space of the first processor, a remote data read request is generated; according to the remote data read request, the target read data and multiple prefetch data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and each prefetch data are synchronized to the remote data cache space of the first processor. When the target read data is synchronized to the local remote data cache space, the remote read operation is triggered. Similarly, a data synchronization condition can be preset. For example, if the number of remote data read requests exceeds a preset read request quantity threshold, a synchronization operation is executed in batches to reduce communication overhead. The number of prefetch data can be determined according to the total capacity of the actual remote data cache space. Exemplarily, taking the address where the target read data is located as the starting address, several subsequent cache lines are pre-loaded continuously by address as prefetch data. Among them, when the local remote data cache space caches the data of the second processor, it can cache the data in the specified area of the local cache space of the second processor, or it can cache the hot data of the second processor, which does not affect the implementation of the present invention. When reading remote data from the local cache space of other processors, for example, data can be read in a burst (data burst) manner.

[0040] For the remote write operation, such as Figure 4 As shown, for the sake of avoiding description, the processor currently executing the task to be processed is defined as the first processor, and the processor to which the data generated by the first processor is to be written is defined as the target processor. The implementation process of the remote write operation may include: when the first processor generates target write data with a storage address in the target local cache space during the execution of the task to be processed, if the remote data cache space of the first processor caches the data of the target local cache space, then when starting to write the target write data to the remote data cache space of the first processor, such as Figure 5As shown, a status change instruction is sent to the target processor to make the status of the target local cache space an invalid status; when the target write data is successfully written, data synchronization processing is performed on the target local cache space according to the remote data cache space, and after the data synchronization ends, a status change instruction is sent to the target processor again to make the status of the target local cache space return to a valid status; if the remote data cache space of the first processor does not cache the data of the target local cache space, a remote data write request is generated, and according to the remote data write request, the target write data is written into the target local cache space, and after the target write data is successfully written, the target write data and multiple prefetch data adjacent to the address of the target write data are cached into the remote data cache space of the first processor. Among them, the remote data cache space of the first processor caching the data of the target local cache space can be the data in a specified area of the target local cache space that is preset. For this case, when invalidating the data in the target local cache space, only the data in this specified area needs to be invalidated. For example, status information for data invalidation is generated for this specified area and broadcast to all processors, and the data in the invalid status cannot be read. The remote data cache space of the first processor caching the data of the target local cache space can also be hot data. For this case, if the target write data is newly added and has no associated data with the current data, the existing data in the target local cache space may not be invalidated. If the target write data is to modify old data or is related to the current data, only these data can be invalidated. Of course, as a more convenient way, when the target write data needs to be written into the target local cache space, all data in the target local cache space can be directly prohibited from being used to ensure data consistency and the accuracy of task processing. Through this data synchronization method, lightweight consistent data synchronization can be achieved, ensuring data consistency in a multi-accelerator environment while minimizing maintenance overhead as much as possible. When reading prefetch data, exemplarily, the address where the target write data is located can be used as the starting address, and several subsequent cache lines are pre-loaded continuously by address as prefetch data. Similarly, the data synchronization condition of the write data can be preset, such as when the number of remote data write requests exceeds the preset write request quantity threshold, or the total amount of remote write data exceeds the preset write capacity threshold, a synchronization operation is performed in batches.

[0041] As can be seen from the above, in this embodiment, by setting multiple remote data caching methods, the task execution efficiency can be improved while ensuring data consistency. Through the lightweight consistent synchronization method, cache invalidation notifications are sent in a timely manner, and necessary data updates are performed according to the actual situation, which can ensure data consistency in a multi-accelerator environment, guarantee the execution efficiency and execution accuracy of the tasks to be processed, and minimize maintenance overhead at the same time.

[0042] To further improve cache utilization and reduce communication latency between different processors, based on the above embodiments, the present invention also provides that the local cache of the processor has a two-level cache structure, that is, two storage areas can be pre-opened in each local cache space. One large area is used as the remote data cache space, and the other small area is used as the replacement data cache space. The replacement data cache space can be implemented by, for example, Victim Cache. The replacement data cache space is mainly used to store the data blocks replaced within the target time period, such as the data replaced in the local cache space within a week. Under the two-level storage mechanism, the implementation method of the remote read and write operations of the task to be processed is as follows: When a remote read operation is performed on the task to be processed, such as Figure 6 shown, when the target read data required for the first processor to execute the task to be processed is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space. If the target read data is not in the remote data cache space of the first processor, then it reads the target read data from its replacement data cache space. If the target read data is not in the replacement data cache space of the first processor, a remote data read request is generated; according to the remote data read request, the target read data and multiple prefetch data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and each prefetch data are synchronized to the remote data cache space of the first processor, and at the same time, the replacement data cache space of the first processor is updated. When the target read data is synchronized to the local remote data cache space, the remote read operation is triggered again. Similarly, data synchronization conditions can be preset. For example, if the number of remote data read requests exceeds the preset read request number threshold, a batch synchronization operation is executed to reduce communication overhead.

[0043] For the remote write operation, such as Figure 7As shown, when the first processor generates target write data with a storage address in the target local cache space during the execution of a task to be processed, if the remote data cache space of the first processor caches the data in the target local cache space, then when starting to write the target write data to the remote data cache space of the first processor, a status change instruction is sent to the target processor to make the status of the target local cache space an invalid status; when the target write data is successfully written, data synchronization processing is performed on the target local cache space according to the remote data cache space, and after the data synchronization ends, a status change instruction is sent to the target processor again to make the status of the target local cache space return to the valid status; if the remote data cache space of the first processor does not cache the data in the target local cache space, but the replacement data cache space of the first processor caches the data in the target local cache space, then when starting to write the target write data to the replacement data cache space of the first processor, a status change instruction is sent to the target processor to make the status of the target local cache space an invalid status; when the target write data is successfully written, data synchronization processing is performed on the target local cache space according to the replacement data cache space of the first processor, and after the data synchronization ends, a status change instruction is sent to the target processor again to make the status of the target local cache space return to the valid status; if neither the remote data cache space nor the replacement data cache space of the first processor caches the data in the target local cache space, then a remote data write request is generated. According to the remote data write request, the target write data is written to the target local cache space, and after the target write data is successfully written, the target write data and multiple prefetch data adjacent to the address of the target write data are cached in the remote data cache space of the first processor.

[0044] As can be seen from the above, by setting up a two-level cache mechanism in this embodiment, it is possible to ensure that the data to be processed is in the local area to the greatest extent, reduce the number of remote accesses, and ensure the execution efficiency and accuracy of the task to be processed.

[0045] It should be noted that there is no strict order of execution among the steps in the present invention. As long as it conforms to the logical order, these steps can be executed simultaneously or in a certain preset order, which does not affect the implementation of the present invention.

[0046] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. The present invention also provides a corresponding device for the task processing method, which further makes the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The task processing device provided by the present invention will be introduced below, and the task processing device described below can be correspondingly referred to the task processing method described above.

[0047] From the perspective of functional modules, refer to Figure 8 , Figure 8 which is a structural diagram of the task processing device provided in this embodiment under a specific implementation manner. The device may include: A task recognition module 801, configured to determine whether the task to be processed is a communication-dominated task according to the memory access behavior characteristics during the execution of the task to be processed when multiple processors cooperate to execute the task to be processed.

[0048] A task execution module 802, if the task to be processed is a communication-dominated task, synchronize the remote data of each processor to the corresponding remote data cache space respectively; when the task to be processed performs a remote read operation, read the target read data from the remote data cache space; when the task to be processed performs a remote write operation, write the target write data written to the target local cache space of the target processor to the remote data cache space, and at the same time invalidate the corresponding data in the target local cache space, and perform data synchronization processing on the target local cache space according to the remote data cache space. The local cache space of each processor includes the remote data cache space. Among them, the memory access behavior characteristics are determined according to the probability that the data located at the target time and / or target position is accessed again within the target time period, and the communication-dominated task is a task whose remote access data frequency during the task execution process meets the preset communication overhead condition.

[0049] Exemplarily, in some implementation manners of this embodiment, the above task recognition module 801 may further be configured to: pre-train a target machine learning model using a task recognition sample set until a preset model training stop condition is reached to obtain a task recognition model, and convert the task recognition model to generate a task recognition lookup table; where each task recognition sample in the task recognition sample set is a sample memory access behavior characteristic carrying a communication-dominated task label or a computation-dominated task label; obtain the memory access command stream of the task to be processed, and extract the to-be-processed memory access behavior characteristics from the memory access command stream; input the to-be-processed memory access behavior characteristics into the task recognition lookup table, and determine whether the task to be processed is a communication-dominated task or a computation-dominated task according to the task recognition lookup table.

[0050] As an exemplary implementation manner of the above embodiment, the above task recognition module 801 may further be configured to: reduce the input feature dimension of the task recognition model, and / or reduce the model structure parameters of the task recognition model to obtain an initial to-be-deployed task recognition model; convert the floating-point input features of the initial to-be-deployed task recognition model into integer input features to obtain the to-be-deployed task recognition model.

[0051] As another exemplary implementation method of the above embodiment, the above task identification module 801 can be further used to: generate a task identification lookup table according to the mapping relationship between the input features of the task identification model to be deployed and the predicted output results; wherein the task identification model to be deployed is a task identification model or a task identification model after pruning; encode the task identification lookup table to meet the format of the hardware end to be deployed, and deploy the encoded task identification lookup table on the hardware end to be deployed.

[0052] As another exemplary implementation method of the above embodiment, the above-mentioned task identification module 801 can also be further used to: obtain multiple historical memory access command streams from the hardware end to be deployed, and extract multi-dimensional initial historical memory access features of the multiple historical memory access command streams; the initial historical memory access features at least include cache behavior related features, remote access related features, and memory access operation related features; according to the impact of each dimensional initial historical memory access feature of the multi-dimensional initial historical memory access feature on the predicted output result of the target machine learning model, select the target historical memory access feature that meets the preset contribution condition from the multi-dimensional initial historical memory access features, and remove the redundant features of each target historical memory access feature to obtain the sample memory access behavior feature; construct a task identification sample according to the label type corresponding to the multi-dimensional initial historical memory access feature and the sample memory access behavior feature.

[0053] As an exemplary implementation method of the above embodiment, the above task identification module 801 can be further used to: eliminate outliers in the multi-dimensional initial historical memory access features; convert the numerical features in the multi-dimensional initial historical memory access features into data that conform to the standard normal distribution; and convert the proportional features in the multi-dimensional initial historical memory access features into a format that meets the input features of the target machine learning model.

[0054] As another exemplary implementation method of the above embodiment, the above task identification module 801 can also be further used for: when receiving a model parameter setting instruction, obtaining basic parameters and hyperparameter search space information by parsing the model parameter setting instruction; the basic parameters include at least the number of decision trees, the maximum tree depth and the minimum number of leaf node samples; the hyperparameter search space information includes at least the range of the number of trees, the maximum depth range and the minimum number of leaf samples; based on the model structure of the random forest model, constructing a target machine learning model according to the number of decision trees, the maximum tree depth and the minimum number of leaf node samples; using preset model performance evaluation conditions to perform grid search to determine the optimal value of the number of trees, the optimal value of the maximum depth and the optimal value of the minimum number of leaf samples in the hyperparameter search space information; configuring the model parameters of the target machine learning model according to the optimal value of the number of trees, the optimal value of the maximum depth and the optimal value of the minimum number of leaf samples.

[0055] Exemplarily, in some other embodiments of the present embodiment, the above task execution module 802 may further be configured to: when the target read data required for the first processor to execute the to-be-processed task is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space; if the target read data is not in the remote data cache space of the first processor, a remote data read request is generated; according to the remote data read request, the target read data and a plurality of prefetch data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and each prefetch data are synchronized to the remote data cache space of the first processor.

[0056] Exemplarily, in some other embodiments of the present embodiment, the above task execution module 802 may further be configured to: the local cache space of each processor further includes a replacement data cache space for storing data blocks replaced during a target time period. When the target read data required for the first processor to execute the to-be-processed task is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space; if the target read data is not in the remote data cache space of the first processor, the target read data is read from its replacement data cache space; if the target read data is not in the replacement data cache space of the first processor, a remote data read request is generated; according to the remote data read request, the target read data and a plurality of prefetch data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and each prefetch data are synchronized to the remote data cache space of the first processor, and the replacement data cache space of the first processor is updated simultaneously.

[0057] Exemplarily, in some other embodiments of the present embodiment, the above task execution module 802 may further be configured to: when the first processor generates target write data with a storage address of the target local cache space during the execution of the to-be-processed task, if the remote data cache space of the first processor caches the data of the target local cache space, when starting to write the target write data to the remote data cache space of the first processor, a status change instruction is sent to the target processor to make the status of the target local cache space an invalid status; when the target write data is successfully written, data synchronization processing is performed on the target local cache space according to the remote data cache space, and after the data synchronization ends, a status change instruction is sent to the target processor again to make the status of the target local cache space return to the valid status; if the remote data cache space of the first processor does not cache the data of the target local cache space, a remote data write request is generated, and according to the remote data write request, the target write data is written to the target local cache space, and after the target write data is successfully written, the target write data and a plurality of prefetch data adjacent to the address of the target write data are cached in the remote data cache space of the first processor.

[0058] As an exemplary implementation of the above embodiment, the above task execution module 802 may further be configured to: The local cache space of each processor further includes a replacement data cache space for storing data blocks replaced during a target time period. When the first processor generates target write data with a storage address in the target local cache space during the execution of a to-be-processed task; if the remote data cache space of the first processor does not cache the data of the target local cache space, but the replacement data cache space of the first processor caches the data of the target local cache space, then when starting to write the target write data to the replacement data cache space of the first processor, a status change instruction is sent to the target processor to make the status of the target local cache space an invalid status; when the target write data is successfully written, data synchronization processing is performed on the target local cache space according to the replacement data cache space of the first processor, and after the data synchronization ends, a status change instruction is sent to the target processor again to make the status of the target local cache space return to the valid status; if neither the remote data cache space nor the replacement data cache space of the first processor caches the data of the target local cache space, a remote data write request is generated.

[0059] The task processing device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above task processing method embodiments.

[0060] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above task processing method embodiments when running.

[0061] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0062] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above task processing method embodiments.

[0063] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned task processing method embodiments are implemented.

[0064] Finally, the present invention also provides a task processing system. Refer to Figure 9 , the system may at least include a task processor 901, a first processor 902 and a second processor 903 interconnected; the first processor 902 and the second processor 903 can access remote memory resources in a DCA manner, and the first processor 902 and the second processor 903 can be interconnected through NVLink, InfinityFabric, and PCIe. Among them, the first processor 902 includes a first local cache, and the first local cache includes a first remote data cache space. The second processor 903 includes a second local cache, and the second local cache includes a second remote data cache space. The first processor 902 remotely accesses the second local memory, and the first remote data cache space caches the data of the second local cache. The second processor 903 remotely accesses the first local cache, and the second remote data cache space caches the data of the first local cache. The task processor 901 can be deployed in any one of the processors, or on the main CPU of the same computing platform, or on a cloud server, which does not affect the implementation of the present invention. When executing the computer program stored in the memory, the steps of the task processing method recorded in any of the above method embodiments are implemented.

[0065] To make the technical solution of the present invention clearer and more understandable to those skilled in the art, the present invention also provides an exemplary implementation, which may include the following content: A1: Processors 1, 2, 3, and 4 are interconnected through PCIe to build a multi-processor computing platform, and processors 1, 2, 3, and 4 use the DCA method to access the remote data of other processors.

[0066] A2: Build and train a task recognition model on the multi-processor computing platform.

[0067] As Figure 10As shown, the multi-processor computing platform captures the memory access behavior data of each processor core in real time through hardware performance counters or software analysis tools, including L1 / L2 cache hit rate, number of remote memory accesses, cache line utilization rate (proportion of valid data), Load / Store operation ratio, and memory access latency distribution. For every 1000 memory access commands accumulated, a task recognition sample is constructed, and continuous collection is carried out until 1000 task recognition samples are obtained to generate a task recognition sample set. For every 1000 memory access commands extracted or when multi-dimensional initial historical memory access features corresponding to multiple groups of access commands are obtained, the 3σ principle can be used to eliminate outliers from the multi-dimensional initial historical memory access features, and Z-score normalization is performed on numerical features (such as cache hit rate, remote access ratio); the Load / Store ratio is converted into a percentage form. The importance of each initial historical memory access feature after the above processing is evaluated using SHAP value analysis, and the top 8 features with the highest contribution are retained; redundant features are eliminated through mutual information calculation, and a feature vector of dimension 8 is generated as the input of the target machine learning model.

[0068] After the task recognition sample set is generated, a target machine learning model is constructed according to the random forest model, and the number of decision trees (n_estimators = 20), maximum tree depth (max_depth = 4), and minimum number of samples in leaf nodes (min_samples_leaf = 5) of the target machine learning model are set, and automatic calculation of feature importance is enabled; the data set is divided into a training set (80%) and a test set (20%). The hyperparameter search space is defined: number of trees (10 - 20), maximum depth (3 - 5), minimum number of leaf samples (3 - 10); 5-fold cross-validation is used for grid search, and the parameter combination with the highest F1 score is selected as the optimal configuration of the hyperparameters of the target machine learning model.

[0069] After the construction of the target machine learning model is completed, the target machine learning model is trained using the task recognition sample set, and the accuracy of the model is evaluated using cross-validation and confusion matrix to ensure the balanced performance of the model in classifying high communication overhead tasks and high computing-intensive tasks. The precision, recall, and F1 score are calculated in each fold of cross-validation to ensure that the difference in F1 scores between communication-dominated tasks and computing-dominated tasks does not exceed 5%. When the accuracy of the target machine learning model on the test set needs to reach more than 90%, the training process of the target machine learning model is completed, and a task recognition model is obtained.

[0070] A3: The multi-processor computing platform deploys the task recognition model.

[0071] To reduce the resource requirements of the task recognition model for a multi-processor computing platform, the trained task recognition model can also be lightweighted and hardware-adapted: only the top 5 input features with the highest importance in the task recognition model are retained, the depth of the decision tree is compressed to 3 layers, the floating-point segmentation threshold is converted to an integer, for example, the cache line utilization threshold of 0.65 is discretized to 65, reducing the occupation of hardware computing resources. Each decision tree of the lightweighted task recognition model is converted into a corresponding look-up table (LUT), the look-up tables of all decision trees are combined into a task recognition look-up table, and the task recognition look-up table is encoded in binary format and embedded in the front-end control unit of the instruction pipeline of each processor; a dedicated cache is designed to store the LUT index table, supporting the mapping from feature combination to classification result to be completed within a single cycle.

[0072] A4: Deployment of the local cache structure of each processor in the multi-processor computing platform.

[0073] Processors 1, 2, 3, and 4 deploy a two-level cache structure and a data synchronization controller in the local memory: the remote data cache space is used to cache remote data, the replacement data cache space is used to save the data blocks that have been recently replaced, and the data synchronization controller controls the data consistency between the remote data cache space and the corresponding remote cache space based on the lightweight consistency data synchronization protocol.

[0074] A5: Cooperative execution of tasks to be processed by multiple processors in the multi-processor computing platform.

[0075] The application scenario of the multi-processor computing platform is that the application task to be executed currently is determined, such as a matrix multiplication task or an SC task, and then the memory access instruction stream in the entire execution process belongs to the type of this application. As Figure 11 shown, the execution process of the task to be processed is divided into two stages: the task type recognition stage and the task optimized execution stage. The former stage takes a shorter time and only needs to collect 1000 - 10000 memory access instructions to construct the input data required for the task recognition look-up table to realize the discrimination of the task type of the task to be processed. When the discrimination result is obtained, it enters the latter stage. If the task is a cache-sensitive communication-dominated task, the memory access instructions are connected to the task execution module 802 described in the above embodiment, and the hot data in the remote memory of other processors is cached into the local remote data cache space to reduce the high latency of remote access. A small-capacity replacement data cache space is additionally set in the local cache space to store the data blocks that have been recently replaced, further improving the cache hit rate and effectively reducing the performance loss caused by cache misses. To balance efficiency and consistency, an asymmetric lightweight consistency protocol is adopted, and a corresponding data synchronization controller is designed to ensure the correctness of the remote data when it is read by other threads or operations when the data in the local cache is modified or unmodified.

[0076] When receiving a task to be processed, the multi-processor computing platform can initiate a data collection process. The data collection process intercepts the memory access instruction stream within a specified time window, counts a set of feature data once (such as accumulating 1000 instructions), and generates a real-time feature vector. It identifies the task type through a deployed task recognition lookup table, and accelerates cache hit checks and data retrieval through an index structure to ensure that data can be quickly obtained in local memory. When the classification results are the same for three consecutive times, a high-confidence (≥90%) determination is triggered. If the task to be processed is determined to be a communication-dominated task, local caching operations are performed on remote data, and a prefetch mechanism is triggered for unhit data. The remote data is mapped and synchronized to the local cache, and several subsequent cache lines are preloaded in address continuity. At the same time, the cache content of the replacement data cache space is updated. Monitor the write operations of the local cache. When a data block is modified, lightweight consistency maintenance is triggered to invalidate the remote data. When the data block needs to be updated and synchronized, write the remote data to restore the local and remote data to a synchronized state and restore the remote state to a valid state again.

[0077] A6: During the execution of the task to be processed, the average memory access latency, cache hit rate, and consistency protocol overhead are statistically calculated in real time. If the performance improvement does not meet the expected conditions, such as the latency reduction <20%, the task failure model recalibration in A2 is triggered, prompting to re-optimize the model and update the corresponding hardware module.

[0078] As can be seen from the above, in this embodiment, by discriminating the memory access type of the task to be processed on the multi-processor computing platform and using a lightweight machine learning model for efficient and low-overhead real-time recognition, the task scheduling and memory access behavior are optimized, and the overall system performance is improved. Based on the optimization mechanism of spatio-temporal locality, a local two-level cache mechanism and corresponding lightweight consistency data synchronization are constructed, reducing the communication latency of multi-processor collaborative task processing, effectively improving the efficiency of multi-processor collaborative task processing, and enhancing the multi-processor task processing performance. In addition, the cache policy based on spatio-temporal locality can be extended to the edge computing field for data prefetching and synchronization of distributed nodes in low-bandwidth environments, and even provide a low-latency solution for the parameter update mechanism in federated learning. In addition, the deep integration of model lightweighting and hardware acceleration may promote the adaptive instruction set design of artificial intelligence processors and achieve task-aware dynamic allocation of computing power.

[0079] The above has introduced in detail a task processing method, system, electronic device, computer-readable storage medium, and computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A task processing method, characterized in that: include: When multiple processors cooperate to execute a task to be processed, if the task to be processed is determined to be a communication-dominated task based on the memory access behavior characteristics during the execution of the task to be processed, the remote data of each processor is synchronized to the corresponding remote data cache space; the local cache space of each processor includes the remote data cache space; When the task to be processed performs a remote read operation, the target read data is read from the remote data cache space; When the task to be processed performs a remote write operation, the target write data written to the target local cache space of the target processor is written to the remote data cache space, and the corresponding data of the target local cache space is invalidated, and data synchronization processing is performed on the target local cache space according to the remote data cache space; Among them, the memory access behavior characteristics are determined based on the probability of data located at the target time and / or target location being accessed again within the target time period, and the communication-dominated task is a task in which the frequency of remote access to data during task execution meets the preset communication overhead conditions.

2. The task processing method according to claim 1, characterized in that: Determining that the task to be processed is a communication-dominated task according to memory access behavior characteristics during the execution of the task to be processed includes: Pre-training a target machine learning model with a task identification sample set until a preset model training stop condition is reached to obtain a task identification model, and converting the task identification model to generate a task identification lookup table; wherein each task identification sample in the task identification sample set is a sample memory access behavior feature carrying a communication-dominated task label or a computation-dominated task label; Acquire a memory access command stream of the task to be processed, and extract the memory access behavior characteristics to be processed from the memory access command stream; The memory access behavior feature to be processed is input into the task identification lookup table, and it is determined whether the task to be processed is a communication-dominated task or a computation-dominated task according to the task identification lookup table.

3. The task processing method according to claim 2, characterized in that: After obtaining the task identification model, it also includes: Deleting the input feature dimension of the task identification model and / or reducing the model structure parameters of the task identification model to obtain an initial task identification model to be deployed; The floating-point input features of the initial task identification model to be deployed are converted into integer input features to obtain the task identification model to be deployed.

4. The task processing method according to claim 2, characterized in that: Before inputting the to-be-processed memory access behavior feature into the task identification lookup table, the method further includes: Generate a task identification lookup table according to a mapping relationship between input features of the task identification model to be deployed and the predicted output results; wherein the task identification model to be deployed is a task identification model or a task identification model after pruning; The task identification lookup table is encoded to meet the format of the hardware end to be deployed, and the encoded task identification lookup table is deployed on the hardware end to be deployed.

5. The task processing method according to claim 2, characterized in that: The target machine learning model is trained using the task identification sample set until a preset model training stop condition is reached, and further includes: Acquire multiple historical memory access command streams from the hardware to be deployed, and extract multi-dimensional initial historical memory access features of the multiple historical memory access command streams; the multi-dimensional initial historical memory access features at least include cache behavior related features, remote access related features, and memory access operation related features; According to the influence of each dimension of the initial historical memory access features of the multi-dimensional initial historical memory access features on the prediction output result of the target machine learning model, a target historical memory access feature that meets a preset contribution condition is selected from the multi-dimensional initial historical memory access features, and redundant features of each target historical memory access feature are removed to obtain a sample memory access behavior feature; A task identification sample is constructed according to the label type corresponding to the multi-dimensional initial historical memory access feature and the sample memory access behavior feature.

6. The task processing method according to claim 5, characterized in that: After extracting the multi-dimensional initial historical memory access features of the plurality of historical memory access command streams and before selecting the target historical memory access features that meet the preset contribution condition from the multi-dimensional initial historical memory access features, the method further includes: Eliminating abnormal values ​​of the multi-dimensional initial historical access characteristics; Converting the numerical features in the multi-dimensional initial historical access features into data that conforms to a standard normal distribution; The proportional features in the multi-dimensional initial historical access features are converted into a format that satisfies the input features of the target machine learning model.

7. The task processing method according to claim 2, characterized in that: The target machine learning model is trained using the task identification sample set until a preset model training stop condition is reached, and further includes: When a model parameter setting instruction is received, basic parameters and hyperparameter search space information are obtained by parsing the model parameter setting instruction; the basic parameters include at least the number of decision trees, the maximum tree depth and the minimum number of leaf node samples; the hyperparameter search space information includes at least the range of the number of trees, the maximum depth range and the minimum number of leaf samples; Based on the model structure of the random forest model, the target machine learning model is constructed according to the number of decision trees, the maximum tree depth, and the minimum number of leaf node samples; Performing a grid search using a preset model performance evaluation condition to determine an optimal value for the number of trees, an optimal value for the maximum depth, and an optimal value for the minimum number of leaf samples in the hyperparameter search space information; The model parameters of the target machine learning model are configured according to the optimal value of the number of trees, the optimal value of the maximum depth, and the optimal value of the minimum number of leaf samples.

8. The task processing method according to any one of claims 1 to 7, characterized in that: When the task to be processed performs a remote read operation, the target read data is read from the remote data cache space, including: When the target read data required by the first processor to execute the task to be processed is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space, and generates a remote data read request if the target read data is not in the remote data cache space of the first processor; According to the remote data read request, the target read data and a plurality of pre-fetched data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and the pre-fetched data are synchronized to the remote data cache space of the first processor.

9. The task processing method according to any one of claims 1 to 7, characterized in that: The local cache space of each processor also includes a replacement data cache space, and the replacement data cache space is used to store data blocks replaced within a target time period. When the task to be processed performs a remote read operation, the target read data is read from the remote data cache space, including: When the target read data required by the first processor to execute the task to be processed is in the local cache space of the second processor, the first processor first reads the target read data from its remote data cache space, and if the target read data is not in the remote data cache space of the first processor, the first processor then reads the target read data from its replacement data cache space, and if the target read data is not in the replacement data cache space of the first processor, a remote data read request is generated; According to the remote data read request, the target read data and multiple pre-fetched data adjacent to the address of the target read data are read from the local cache space of the second processor, and the target read data and each pre-fetched data are synchronized to the remote data cache space of the first processor, and the replacement data cache space of the first processor is updated at the same time.

10. The task processing method according to any one of claims 1 to 7, characterized in that: The process of performing remote write operation on the task to be processed includes: When the first processor generates target write data whose storage address is a target local cache space of a target processor during the process of executing the task to be processed, if the remote data cache space of the first processor caches the data of the target local cache space, then when starting to write the target write data to the remote data cache space of the first processor, a state change instruction is sent to the target processor to make the state of the target local cache space invalid; When the target write data is written successfully, data synchronization processing is performed on the target local cache space according to the remote data cache space, and after the data synchronization is completed, a state change instruction is sent to the target processor again to restore the state of the target local cache space to a valid state; If the remote data cache space of the first processor does not cache the data of the target local cache space, a remote data write request is generated, and according to the remote data write request, the target write data is written to the target local cache space, and after the target write data is successfully written, the target write data and a plurality of pre-fetched data adjacent to the address of the target write data are cached to the remote data cache space of the first processor.

11. The task processing method according to claim 10, characterized in that: The local cache space of each processor also includes a replacement data cache space, and the replacement data cache space is used to store data blocks replaced within a target time period. The process of performing a remote write operation on the task to be processed includes: When the first processor is executing the task to be processed, the target write data whose storage address is the target local cache space of the target processor is generated; if the remote data cache space of the first processor does not cache the data of the target local cache space, but the replacement data cache space of the first processor caches the data of the target local cache space, then when starting to write the target write data to the replacement data cache space of the first processor, a state change instruction is sent to the target processor to make the state of the target local cache space invalid; When the target write data is written successfully, performing data synchronization processing on the target local cache space according to the replacement data cache space of the first processor, and after the data synchronization is completed, sending a state change instruction to the target processor again to restore the state of the target local cache space to a valid state; If the remote data cache space and the replacement data cache space of the first processor do not cache the data of the target local cache space, a remote data write request is generated.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the task processing method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the task processing method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the task processing method according to any one of claims 1 to 11 are implemented.

15. A task processing system, characterized in that: At least comprising a task processor, a first processor and a second processor interconnected with each other; Among them, the first processor includes a first local cache, the first local cache includes a first remote data cache space, the second processor includes a second local cache, the second local cache includes a second remote data cache space; the first processor remotely accesses the second local memory, the first remote data cache space caches the data of the second local cache, the second processor remotely accesses the first local cache, and the second remote data cache space caches the data of the first local cache; when the task processor executes the computer program stored in the memory, the steps of the task processing method as described in any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Cache access system supporting data prefetching of out-of-order processor

    CN115309453A

  • Data caching method and device, electronic equipment and readable storage medium

    CN117389630A

  • Memory access optimization method and device, equipment, medium and program product

    CN118051189A

  • SLC cache management method and device, equipment and storage medium

    CN118606229A

  • GPU cache management based on locality type detection

    US20200401529A1

Cited By

  • Heterogeneous computing system, cache consistency maintenance method and device, equipment and medium

    CN120353612A

  • Information processing method and device, electronic equipment and storage medium

    CN120353685A