Data transmission acceleration method and system for multi-core heterogeneous ASIC computing motherboard
By schema provisioning and DMA channel generation of the computing core of the multi-core heterogeneous ASIC computing motherboard, the problem of low computing efficiency of traditional data transmission methods is solved, efficient data transmission and computing resource utilization is achieved, and overall computing efficiency is improved.
Patent Information
- Application Number
- CN202510295265.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The traditional multi-core heterogeneous ASIC computing motherboard has inefficient problems in data transmission and computing core resource scheduling, resulting in low overall computing efficiency.
By scheduling the calculation core, switching to the initial allocation mode, using the high-speed transmission protocol to receive the data set to be calculated from the external terminal, obtaining the precedent data part for calculation processing, and monitoring the resource scheduling characteristics. DMA channel is generated according to resource scheduling characteristics, switch the calculation core to the DMA mapping mode, map data directly to the calculation core through the external DMA channel, perform data calculation processing, and realize data exchange between the calculation cores through the interactive DMA channel.
Through DMA technology, reduce data transmission delay, improve computing efficiency, dynamically schedule computing cores, improve resource utilization, avoid uneven loading of computing cores, realize multi-core collaboration and data transmission acceleration, and improve overall computing efficiency.
Smart Images

Figure CN119807122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data transmission, and in particular to a data transmission acceleration method and system for a multi-core heterogeneous ASIC computing motherboard. Background Art
[0002] With the increasing demand for computing, especially in the fields of high-performance computing and artificial intelligence, traditional computing systems often cannot meet the requirements of large-scale data processing and efficient computing. Multi-core heterogeneous ASIC (application-specific integrated circuit) computing motherboards, as an efficient hardware platform, have been widely used in computing tasks that require high throughput and low latency. However, due to the complexity of data transmission and resource scheduling between computing cores, traditional data transmission and scheduling methods often cannot efficiently utilize computing resources, resulting in low overall computing efficiency. Summary of the invention
[0003] The purpose of the present invention is to provide a data transmission acceleration method and system for a multi-core heterogeneous ASIC computing motherboard, aiming to solve the problem of low computing efficiency of traditional data transmission methods in the prior art.
[0004] The present invention is implemented in this way. In a first aspect, the present invention provides a data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard, comprising:
[0005] Performing mode adjustment on a designated computing core on a computing mainboard so that the computing core switches to an initialization allocation mode;
[0006] The calculation core in the initialization allocation mode is driven by a preset high-speed transmission protocol to receive data from the data set to be calculated of the external terminal, so as to obtain the pre-order data part of the data set to be calculated; wherein the data set to be calculated is all the data that needs to be calculated in the external terminal;
[0007] Instructing the computing core in the initialization allocation mode to perform data computing processing on the preceding data portion, and to perform resource scheduling monitoring on the data computing processing of the preceding data portion, so as to obtain the data computing result of the preceding data portion after the data computing processing and the resource scheduling characteristics of the data set to be calculated;
[0008] Generate a DMA channel for each computing core of the computing mainboard according to the resource scheduling characteristics, so that each computing core of the computing mainboard switches to a DMA mapping mode; wherein the computing core in the DMA mapping mode has an external DMA channel and an interactive DMA channel, the external DMA channel is used to directly map the data set to be calculated to the data cache area of the computing core through the DMA technology, and the interactive DMA channel is used to directly map the data processing result of the data set to be calculated to the data cache area of the remaining computing cores through the DMA technology;
[0009] Directly mapping the data set to be calculated to the data cache area of the calculation core through the external DMA channel, so that the calculation core performs data calculation processing on the data set to be calculated to obtain a data calculation result;
[0010] When the data calculation result needs to be processed again by the remaining computing cores, the data calculation result is directly mapped to the data cache area of the remaining computing cores through the interactive DMA channel, so that the remaining computing cores can perform subsequent processing on the data calculation result to obtain a new round of data calculation results;
[0011] When the data calculation result does not need to be processed again by other calculation cores, the data calculation result is output.
[0012] In a second aspect, the present invention provides a data transmission acceleration system for a multi-core heterogeneous ASIC computing motherboard, which is used to implement a data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard as described in any one of the first aspects.
[0013] The present invention provides a data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard, which has the following beneficial effects:
[0014] The present invention switches the designated computing core to the initialization allocation mode, uses a high-speed transmission protocol to receive the data set to be calculated from the external terminal, obtains the preceding data part, the computing core performs calculation processing on the preceding data, and monitors the resource scheduling characteristics, generates a DMA channel according to the resource scheduling characteristics, switches the computing core to the DMA mapping mode, maps the data to the computing core through the external DMA channel, performs data calculation processing, and if the calculation result needs to be processed by other computing cores, performs data exchange and subsequent calculations through the interactive DMA channel. Through the DMA technology, data transmission delay is reduced, efficiency is improved, computing cores are dynamically scheduled, resource utilization is improved, uneven computing core load is avoided, multi-core collaboration and data transmission are accelerated, and overall computing efficiency is improved, thereby solving the problem of low computing efficiency of traditional data transmission methods in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1It is a schematic diagram of the steps of a data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0017] The implementation of the present invention is described in detail below in conjunction with specific embodiments.
[0018] Reference Figure 1 As shown, a preferred embodiment of the present invention is provided.
[0019] In a first aspect, the present invention provides a method for accelerating data transmission of a multi-core heterogeneous ASIC computing motherboard, comprising:
[0020] S1: performing mode adjustment on a computing core designated on a computing mainboard so that the computing core switches to an initialization allocation mode;
[0021] S2: driving the calculation core in the initialization allocation mode to receive data from the data set to be calculated of the external terminal through a preset high-speed transmission protocol to obtain a pre-order data portion of the data set to be calculated; wherein the data set to be calculated is all the data that needs to be calculated in the external terminal;
[0022] S3: Instruct the calculation core in the initialization allocation mode to perform data calculation processing on the preceding data portion, and perform resource scheduling monitoring on the data calculation processing of the preceding data portion, so as to obtain the data calculation result of the preceding data portion after the data calculation processing and the resource scheduling characteristics of the data set to be calculated;
[0023] S4: Generate DMA channels for each computing core of the computing mainboard according to the resource scheduling characteristics, so that each computing core of the computing mainboard switches to a DMA mapping mode; wherein the computing core in the DMA mapping mode has an external DMA channel and an interactive DMA channel, the external DMA channel is used to directly map the data set to be calculated to the data cache area of the computing core through the DMA technology, and the interactive DMA channel is used to directly map the data processing result of the data set to be calculated to the data cache area of the remaining computing cores through the DMA technology;
[0024] S5: directly mapping the data set to be calculated to the data cache area of the calculation core through the external DMA channel, so that the calculation core performs data calculation processing on the data set to be calculated to obtain a data calculation result;
[0025] S6: when the data calculation result needs to be processed again by the remaining computing cores, the data calculation result is directly mapped to the data cache area of the remaining computing cores through the interactive DMA channel, so that the remaining computing cores can perform subsequent processing on the data calculation result to obtain a new round of data calculation results;
[0026] S7: When the data calculation result does not need to be processed again by other computing cores, the data calculation result is output.
[0027] Specifically, in step S1 of the embodiment provided by the present invention, the computing core designated on the computing mainboard is switched to the initialization allocation mode. In this mode, the computing core is in a state of preparing to receive data and perform preliminary computing processing. The computing core switched to the initialization allocation mode is ready to receive the data set to be calculated from the external terminal and make necessary resource reservations for data processing.
[0028] More specifically, the role of the computing core in the initialization allocation mode is to process the data set to be calculated in order to perform processing work on the data set to be calculated, and at the same time monitor the processing work to obtain the data processing method of the data set to be calculated, and obtain the optimal allocation method of the data set to be calculated corresponding to each computing core on the computing main board based on the data processing method, including direct allocation and redistribution of the calculation results.
[0029] Specifically, in step S2 of the embodiment provided by the present invention, a pre-set high-speed transmission protocol is used to drive the computing core in the initialization allocation mode to receive the preamble data portion of the data set to be calculated sent by the external terminal. The preamble data portion is a subset of the data set to be calculated and contains key information related to subsequent data calculations or data required for initialization calculations.
[0030] More specifically, the data set to be calculated is composed of a series of data parts. When some of the data parts are transmitted to the computing core in the initialization allocation mode, the computing core can perform data processing on them, and analyze the overall picture of the data set to be calculated by monitoring various aspects such as work resource scheduling and work effect of the data processing work, and judge the work item corresponding to the data set to be calculated, and how the computing main board corresponds to this work item, including which computing core the work of the data set to be calculated is assigned to, and whether the computing core needs to transfer the data calculation results to the remaining computing cores for subsequent calculations after processing the data part, and which computing core is the best to receive the transferred data calculation results.
[0031] More specifically, the high-speed transmission protocol ensures that data can be quickly and reliably transmitted from the external terminal to the computing core, reducing transmission delays and improving data exchange efficiency.
[0032] Specifically, in step S3 of the embodiment provided by the present invention, the computing core performs data calculation processing on the received preamble data part, which involves mathematical calculations, data analysis, model reasoning, etc. During the calculation process, the computing core performs resource scheduling monitoring, that is, monitors the usage of the computing resources (such as computing units, memory bandwidth, cache, etc.) of the computing core to ensure reasonable allocation of resources, and dynamically adjusts resource allocation as needed. Through calculation processing, data calculation results are generated, and resource scheduling characteristics are extracted, such as resource utilization efficiency, computing task completion time, delay and other indicators.
[0033] It can be understood that the preamble data part belongs to a part of the data set to be calculated, so the preamble data part also needs to be processed by data calculation, and the data calculation result obtained by the preamble data part after data calculation is also part of the work.
[0034] Furthermore, the resource scheduling feature is an allocation plan obtained by monitoring the data calculation and processing of the previous data part, that is, how the data set to be calculated needs to be allocated and transmitted. This transmission work will be achieved by DMA technology. DMA technology can directly map data to the memory of the computing core, so that the computing core can directly process data without receiving data, greatly speeding up work efficiency.
[0035] Specifically, in step S4 of the embodiment provided by the present invention, based on the resource scheduling characteristics generated by the computing core, the system decides how to further allocate computing tasks to different computing cores. On this basis, the system generates appropriate DMA channels for each computing core of the computing mainboard, and switches the computing core to DMA mapping mode. The computing core in DMA mapping mode can directly map the data set to be calculated to the data cache area through an external DMA channel, and realize data transmission between different computing cores through interactive DMA channels.
[0036] More specifically, the DMA channel includes an external DMA channel and an interactive DMA channel. The external DMA channel is used to map the data set to be calculated directly to the data cache area of the calculation core through DMA technology, and the interactive DMA channel is used to map the data processing results of the data set to be calculated directly to the data cache areas of the remaining calculation cores through DMA technology.
[0037] It can be understood that generating DMA channels based on resource scheduling characteristics enables data computing tasks to be optimally allocated according to actual load and resource conditions. In DMA mapping mode, data can be directly transferred from external storage or other computing cores to the data cache area of the computing core, reducing data access latency and improving computing efficiency.
[0038] Specifically, in steps S5, S6, and S7 of the embodiment provided by the present invention, the data set to be calculated is directly mapped to the data cache area of the computing core through the external DMA channel, so that the computing core can perform data calculation and processing. The computing core uses its local computing unit to process the data mapped to the cache area to generate data calculation results. The use of external DMA channels greatly improves data transmission efficiency, avoids repeated copying or access of data, and reduces bandwidth pressure. By directly mapping data, the computing core can obtain processed data more quickly, thereby improving the response time and throughput of computing processing.
[0039] More specifically, if the calculation results need further processing, the computing core transfers the data calculation results directly to the data cache areas of other computing cores through the interactive DMA channel. Other computing cores will perform subsequent calculations on the data calculation results to generate new data calculation results. The interactive DMA channel enables computing cores to share data efficiently, avoiding traditional memory access delays and improving the system's computing efficiency and task scheduling capabilities. This collaborative processing mode between computing cores enhances the performance of multi-core computing and can quickly respond to complex computing tasks.
[0040] It should be noted that the DMA channel is constructed based on the resource scheduling characteristics obtained in the previous steps. That is to say, by calculating and monitoring the preceding data part, a processing solution for the data set to be calculated is obtained, and DMA memory mapping is performed based on this.
[0041] More specifically, if the data calculation result does not require further calculation or processing, the final result can be output directly to an external device, storage system or other computing module. If the data calculation result does not require further processing, the output operation can directly return the result, reducing unnecessary calculation steps and delays, and ensuring a quick system response.
[0042] The present invention provides a data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard, which has the following beneficial effects:
[0043] The present invention switches the designated computing core to the initialization allocation mode, uses a high-speed transmission protocol to receive the data set to be calculated from the external terminal, obtains the preceding data part, the computing core performs calculation processing on the preceding data, and monitors the resource scheduling characteristics, generates a DMA channel according to the resource scheduling characteristics, switches the computing core to the DMA mapping mode, maps the data to the computing core through the external DMA channel, performs data calculation processing, and if the calculation result needs to be processed by other computing cores, performs data exchange and subsequent calculations through the interactive DMA channel. Through the DMA technology, data transmission delay is reduced, efficiency is improved, computing cores are dynamically scheduled, resource utilization is improved, uneven computing core load is avoided, multi-core collaboration and data transmission are accelerated, and overall computing efficiency is improved, thereby solving the problem of low computing efficiency of traditional data transmission methods in the prior art.
[0044] Preferably, the step of performing mode adjustment on a designated computing core on a computing mainboard so that the computing core switches to an initialization allocation mode includes:
[0045] S11: Obtain computing performance indicators of each computing core on the computing mainboard, and evaluate the task balance of each computing core according to the computing performance indicators, so as to obtain a computing core with the best task balance characteristics as an initialization object;
[0046] S12: performing external data interface configuration based on a high-speed transmission protocol on the initialization object so that the initialization object maintains a communication connection state with an external terminal; wherein the initialization object in a communication connection state with the external terminal is used to receive data from the external terminal;
[0047] S13: Perform resource scheduling monitoring and mapping channel construction program configuration on the initialization object, so that the initialization object switches to the initialization allocation mode.
[0048] Specifically, through hardware monitoring mechanisms or built-in sensors, the performance indicators of each computing core on the computing motherboard are collected, such as computing power (number of operations per second), memory bandwidth, cache usage, load, latency, power consumption, etc. Performance indicators may include CPU clock frequency, computing resource utilization, memory bandwidth usage, temperature, etc. These data help determine the load capacity and performance of each computing core. Collecting and analyzing the performance indicators of each computing core helps to understand the distribution of computing resources, ensure that tasks can be reasonably allocated to the computing cores with the best performance, and avoid overloading of certain computing cores or waste of resources.
[0049] More specifically, by analyzing the performance indicators of each computing core and evaluating the load balancing of the computing core, it is ensured that the task load of each computing core is as balanced as possible. The evaluation content includes factors such as computing power, memory bandwidth, cache utilization, etc., and comprehensively judges which computing cores are suitable for taking on more tasks and which computing cores should take on fewer tasks. When evaluating task balance, the system can use load balancing algorithms, such as minimizing response time algorithms, maximizing throughput algorithms, etc. Task balance evaluation enables computing cores to be reasonably allocated according to actual load and performance indicators, avoiding excessive concentration or dispersion of tasks, and improving the overall computing performance and response efficiency of the system. By optimizing task allocation, resource competition between computing cores can be avoided, computing bottlenecks can be reduced, and the parallel processing capability of the entire system can be improved.
[0050] More specifically, based on the task balance evaluation results, the computing core with the best performance and the most balanced load is selected as the initialization object. This computing core will serve as the dominant processing unit for subsequent tasks. The selection criteria of the computing core usually include computing performance, memory bandwidth, current load, resource idleness, etc. Selecting the computing core with the best task balancing characteristics as the initialization object can ensure that its processing performance is the most superior, thereby improving the computing efficiency of subsequent tasks. Optimizing the computing core selection helps to improve the overall load balance of the system and avoid computing bottlenecks and resource conflicts.
[0051] More specifically, the external data interface of the selected initialization object is configured, and the communication connection with the external terminal is usually achieved through a high-speed transmission protocol (such as PCIe, RDMA, Infiniband, etc.). The configuration includes setting protocol parameters (such as data transmission rate, connection stability, data packet size, etc.) to ensure that the initialization object can quickly and stably receive the data set to be calculated from the external terminal. The high-speed transmission protocol configuration ensures that the initialization object can efficiently and seamlessly exchange data with the external terminal, thereby improving the speed and stability of data reception. The configuration of the external data interface optimizes the data transmission path, reduces the possible delay and packet loss risk during data transmission, and enhances the real-time processing capability of the system.
[0052] More specifically, after the external data interface is configured, a stable communication connection is established between the initialization object and the external terminal to ensure that the computing core can continuously receive data in the initialization allocation mode, and the computing core maintains real-time synchronization with the external terminal to receive continuously transmitted data and ensure the timeliness and integrity of data transmission. By maintaining the communication connection, the initialization object can receive and process the data to be calculated from the external terminal in real time to ensure that there is no delay or data loss during the data exchange process.
[0053] More specifically, the resource scheduling monitoring function of the computing core is configured. The role of the resource scheduling monitoring program is to generate resource scheduling features in subsequent steps. Similarly, the role of the mapping channel construction program is to construct the DMA channel in subsequent steps.
[0054] Preferably, the step of driving the computing core in the initialization allocation mode to receive data from the data set to be calculated of the external terminal through a preset high-speed transmission protocol to obtain the pre-order data part of the data set to be calculated includes:
[0055] S21: driving the computing core in the initialization allocation mode to continuously receive data from the data set to be calculated of the external terminal through a preset high-speed transmission protocol, and time-marking the data parts of the data set to be calculated received at different times to obtain a plurality of data parts arranged in sequence according to time sequence;
[0056] S22: Substitute each data part received continuously at different times into the task execution sequence in the computing core in the initialization allocation mode in turn, and evaluate the task integrity of the task execution sequence. When the evaluation result shows that each data part in the task execution sequence constitutes an executable computing task, convert the task execution sequence to obtain the preceding data part.
[0057] Specifically, the computing core is in the initialization allocation mode, and establishes a connection with the external terminal through a pre-configured high-speed transmission protocol (such as PCIe, RDMA, etc.), and starts the data receiving process. The data receiving process is continuous, and the computing core continuously receives the data set to be calculated from the external terminal. The received data is usually divided into multiple parts, which may be segments of a large data set. The high-speed transmission protocol ensures that the data can be transmitted to the computing core in an efficient and low-latency manner, reducing the bottleneck of data transmission and improving the real-time and accuracy of data processing. In the initialization allocation mode, the computing core can efficiently receive and process a large amount of data streams and prepare the required data for subsequent computing tasks.
[0058] More specifically, each received data portion of the data set to be calculated is time-stamped. The time tagging usually includes a timestamp of the received data, which is used to ensure that the various data portions are arranged in the order in which they are received. These data portions are received continuously, so each data portion can be accompanied by a timestamp to indicate its relative order in the data receiving process. The time tagging ensures that the data portions can be accurately arranged in the order in which they are received, which is crucial for subsequent calculations and task scheduling. The time tagging process avoids data disorder or loss, so that during task execution, data can be transmitted and processed in the correct order, ensuring the correctness of the task logic.
[0059] More specifically, the data parts received at different times and arranged in chronological order are added to the task execution sequence in the computing core in sequence. This means that as the data arrives, the execution order of the tasks is constantly changing. The system will process each data part in chronological order during execution. The task execution sequence is a list of computing tasks constructed in the order in which the data is received. Each data part is an element in the task sequence. The generation of the task execution sequence ensures that the data is processed in the order in which it is received, avoiding confusion in task execution. The data parts are added to the task sequence in sequence, so that the tasks can be scheduled according to the chronological order and logical order during the processing process, avoiding the influence of the order in which the data arrives on the computing results.
[0060] More specifically, the integrity of the constructed task execution sequence is evaluated. The purpose of the evaluation is to check whether each data part in the task sequence constitutes an executable computing task. The basis for the task integrity evaluation can include whether the data part is complete (for example, whether data is missing), whether the data meets the input requirements of the computing task (format, type, sequence, etc.), and whether there are errors or data inconsistencies. The task integrity evaluation ensures that each data part has been correctly verified before entering the task execution sequence and can smoothly participate in subsequent computing tasks. By evaluating the integrity of the task, the data parts that do not meet the conditions can be discovered and excluded in a timely manner to avoid incomplete or erroneous data affecting the execution of subsequent computing tasks.
[0061] More specifically, when the task integrity assessment shows that all data parts in the task execution sequence meet the requirements and can be executed, the computing core will perform a conversion process. The conversion process combines multiple data parts into an overall data set to obtain the so-called "previous data part", which refers to the previous data that has been checked and is ready to perform calculations. This part of the data will be used as the input for subsequent computing tasks. The conversion process allows multiple received data parts to be merged into a complete computing task data set in the correct order and format, preparing for the subsequent processing of the computing task. The obtained previous data part can be used by the subsequent computing core to perform specific computing tasks, thereby ensuring the smooth progress of the computing process.
[0062] Preferably, the step of causing the computing core in the initialization allocation mode to perform data computing processing on the preceding data portion, and performing resource scheduling monitoring on the data computing processing of the preceding data portion, so as to obtain the data computing result of the preceding data portion after the data computing processing and the resource scheduling characteristics of the data set to be calculated includes:
[0063] S31: instructing the calculation core in the initialization allocation mode to perform data calculation processing on the preceding data portion to obtain a data calculation result of the preceding data portion after the data calculation processing;
[0064] S32: when the computing core performs data processing on the preceding data portion, data collection and feature analysis are performed on the computing resource scheduling status of the computing core to obtain computing resource scheduling parameters for the computing core to perform data computing processing on the preceding data portion;
[0065] S33: performing a reverse calculation of the overall form of the data set to be calculated on the preceding data portion according to the computing resource scheduling parameters, so as to obtain the overall resource requirement form of the data set to be calculated;
[0066] S34: The overall resource demand form is allocated and processed according to the computing performance indicators of each computing core of the computing mainboard to obtain resource scheduling characteristics of each computing core corresponding to the data set to be calculated; wherein the resource scheduling characteristics are used to describe the allocation method between each data part in the data set to be calculated and each computing core, as well as the computing processing method of the data part allocated to each computing core.
[0067] Specifically, the computing core starts to perform data calculation processing on the pre-order data part in the initialization allocation mode. The pre-order data part is the data part that has been received and arranged in sequence from the data set to be calculated. The specific content of the data calculation processing may include the execution of computing tasks, data analysis, model training, etc. The specific task depends on the application scenario. Using the processing power of the computing core to process the data for computing tasks can efficiently complete specific computing tasks. The initialization allocation mode provides efficient data processing preparation, so that the computing core can start quickly and perform accurate calculations when processing data.
[0068] More specifically, while the computing core is performing data computing and processing on the preceding data, the system needs to monitor the resource scheduling status of the computing core. Specifically, the system will collect data on the resource usage of the computing core, including the usage of resources such as CPU, memory, storage, and network bandwidth. Resource scheduling monitoring also involves real-time tracking of the computing core's load, resource allocation, and computing status to evaluate the resource requirements of the current computing task. Real-time resource monitoring can ensure that the computing core's resource usage is optimized when performing data computing and processing, avoiding resource overload or idle waste. Data collection and feature analysis help optimize resource allocation strategies so that subsequent computing tasks can be performed more efficiently.
[0069] More specifically, based on data collection and feature analysis, the computing core calculates parameters regarding resource scheduling. These parameters describe the resource scheduling status and requirements of the computing core during data calculation and processing. These parameters include, but are not limited to, the computing load, memory usage, I / O bottleneck, bandwidth requirements, etc. of each computing core, and can provide the real-time requirements of computing resources during task processing. The resource scheduling parameters help the system understand the resource requirements of each computing core, thus providing a strong basis for subsequent resource allocation. Through the analysis of the resource scheduling status, the load of the computing core can be effectively predicted, avoiding the situation of over-concentration or imbalance of resources, and ensuring the efficient completion of computing tasks.
[0070] More specifically, according to the resource scheduling parameters of the computing core, a reverse calculation of the overall resource requirements of the data set to be calculated is performed. The reverse calculation process analyzes the resource usage of the computing core and reversely deduces the overall requirements of the computing task to obtain the form of the overall resource requirements of the data set to be calculated. This step is actually to estimate the resource consumption of the entire data set in order to provide data support for subsequent resource scheduling and task allocation. Through the reverse calculation, a prediction of the resource requirements for the entire data set to be calculated can be obtained, helping the system to effectively plan future computing tasks. This process improves the prediction accuracy of resource utilization, making subsequent resource scheduling more forward-looking and avoiding delays in computing tasks caused by insufficient resources.
[0071] More specifically, according to the computing performance indicators of each computing core on the computing motherboard, the overall resource requirements of the data set to be calculated are allocated. Specifically, the system reasonably allocates the resource requirements of the data set to be calculated according to the performance of the computing core (such as computing power, load-bearing capacity, memory capacity, etc.), divides the computing task among the appropriate computing cores, and generates resource scheduling characteristics between each data part of the data set to be calculated and each computing core through the allocation process. These characteristics describe the computing tasks allocated to each computing core and their processing methods. The resource scheduling characteristics can accurately describe the resource allocation relationship between the data set to be calculated and the computing core, helping to achieve load balancing of tasks in a multi-computing core environment. This process ensures that each computing core can be allocated computing tasks according to demand through performance analysis and resource prediction, maximizing the utilization rate of computing resources, avoiding task congestion or resource waste, and the generated resource scheduling characteristics provide a detailed basis for subsequent task scheduling and computing core cooperation, thus ensuring the efficient progress of the computing process.
[0072] Preferably, the step of performing a reverse calculation of the overall form of the data set to be calculated on the previous data part according to the computing resource scheduling parameters to obtain the form of the overall resource requirements of the data set to be calculated includes:
[0073] S331: performing a task descriptive analysis of the computing task on the preceding data portion to obtain computing task features of the preceding data portion;
[0074] S332: performing an inductive analysis of correlation characteristics on the computing task characteristics according to the computing resource scheduling parameters to obtain computing correlation characteristics between the computing resource scheduling parameters and the computing task characteristics;
[0075] S333: Identify and process the calculation-related features according to the task identification standards summarized by the machine learning model in the preset database to obtain the overall task form corresponding to the calculation-related features and the overall resource demand form corresponding to the overall task form.
[0076] Specifically, a task descriptive analysis is performed on the preceding data. The purpose of the task descriptive analysis is to extract relevant features of the computing task by analyzing the preceding data. These features include the scale of the task, computational complexity, data dependency, computational time, memory requirements, input and output characteristics, etc. This step actually captures the important parameters and features of the task during execution from the perspective of the computing task, and provides basic data for subsequent resource scheduling and task allocation. The task descriptive analysis can fully understand the basic characteristics of the computing task, and help the system to more accurately schedule resources and optimize computing tasks in subsequent steps. Through the descriptive analysis of the preceding data, the system can identify the potential complexity of the computing task and evaluate the required resource types and amounts in advance.
[0077] More specifically, after obtaining the characteristics of the computing task, it is necessary to perform an inductive analysis of the correlation characteristics of these task characteristics based on the computing resource scheduling parameters. Through the correlation characteristic analysis, the system matches the characteristics of the computing task with the parameters of the computing resource scheduling in order to identify the intrinsic connection between resource scheduling and task characteristics. This analysis usually includes: analyzing the processing power of the computing core, resource usage, task input and output requirements, memory and bandwidth limitations, etc., to find out the relationship between these characteristics and task execution efficiency. Through inductive analysis, the system can identify the intrinsic connection between different task characteristics and resource scheduling parameters, which not only helps to further clarify the resource requirements of each computing task, but also provides optimized resource scheduling strategies in multi-core computing environments. Correlation analysis can reveal bottlenecks that may occur during task execution, and then provide decision-making basis for task allocation and scheduling, thereby improving the utilization of computing resources.
[0078] More specifically, the aforementioned computing-related features are identified and processed using a machine learning model in a preset database. The machine learning model is trained on historical data to summarize the standards and rules for task identification. These standards include identification of task types, prediction of task computing requirements, and recommendation of resource scheduling modes. The model is trained on the correlation features between computing task features and resource scheduling features to identify the overall task form of a task and the resource features required. The application of machine learning models can greatly improve recognition efficiency and accuracy, and automatically identify different types of tasks and their corresponding resource requirements. Based on training on historical data, the model can provide accurate predictions, reducing the workload of manual analysis and guesswork, and as data accumulates, the model's prediction accuracy will gradually improve. The machine learning model helps realize the intelligence and automation of task scheduling, enabling the system to optimize resource allocation for computing tasks based on historical rules and real-time data.
[0079] More specifically, after identifying the computing-related features through the machine learning model, the system can obtain the overall task form of each computing task. These overall task forms refer to the total demand for computing tasks, including the required resource types, computing time, memory requirements, etc. According to the overall form of the task, the system further derives the overall resource demand form, that is, the total amount of resources required for the task during execution and the allocation ratio of each resource. The overall resource demand form usually takes into account the performance indicators of different computing cores, resource scheduling strategies, etc. The technical effect of this step is to clarify the resource allocation plan required for each task during execution. This resource demand form is based on the previous task feature analysis, correlation analysis and machine learning model, which can help the system to plan resources in advance during the calculation process. The overall resource demand form obtained can provide clear guidance for subsequent computing task allocation, ensuring that the execution of tasks is not limited by insufficient resources, thereby improving computing efficiency and the overall performance of the system.
[0080] Preferably, the step of generating a DMA channel for each computing core of the computing mainboard according to the resource scheduling feature so that each computing core of the computing mainboard switches to a DMA mapping mode includes:
[0081] S41: parsing the data allocation standard for each computing core of the computing mainboard according to the resource scheduling characteristics, so as to obtain the data allocation standard for each computing core corresponding to the data set to be calculated;
[0082] S42: Taking the computing core as a data mapping object and the external terminal as a data mapping subject, constructing a framework for performing direct memory mapping on the data mapping object and the data mapping subject according to the data allocation standard to generate an external DMA channel of the computing core;
[0083] S43: parsing the data processing flow of each computing core of the computing mainboard according to the resource scheduling characteristics to obtain a transfer path of data calculation results between each computing core; wherein the transfer path includes a starting node, an end node and a data transfer standard, the starting node is a computing core to which the data calculation result is to be transferred, the end node is a computing core to which the data calculation result is transferred, and the data transfer standard is used to determine the data calculation result that needs to be transferred from the starting node to the end node;
[0084] S44: Taking the starting node as the data mapping subject and the ending node as the data mapping object, constructing a framework for direct memory mapping of the data mapping subject and the data mapping object according to the data transfer standard to generate an interactive DMA channel of the computing core.
[0085] Specifically, according to the resource scheduling characteristics, we first analyze the data allocation standards of each computing core on the computing motherboard. The data allocation standard refers to how to allocate the data of computing tasks to different computing cores. Specifically, we analyze the amount of data, data type, processing speed requirements, etc. required by each computing core when executing tasks, and derive the corresponding relationship between each computing core and the data set to be calculated.
[0086] The analysis of data allocation standards includes: determining which computing cores need which data, and determining how the data is allocated between the computing cores (for example, some computing cores are responsible for processing a certain part of the data, while other computing cores are responsible for another part of the data). By analyzing the data allocation standards, it can be ensured that each computing core obtains the data it needs, thereby improving the parallelism and efficiency of computing tasks. Reasonable data allocation can avoid resource conflicts between computing cores and optimize the efficiency of data transmission and storage. This step helps to provide a basis for data mapping for subsequent DMA channel generation, ensuring that the computing cores can access the required data efficiently and accurately.
[0087] More specifically, each computing core is taken as a data mapping object, and the external terminal (which may be an external storage device or other computing resource) is taken as a data mapping subject. According to the aforementioned data allocation standard, a direct memory mapping framework is created between the computing core and the external terminal. Specifically, the computing core directly exchanges data with the external device through the DMA channel, avoiding the overhead of CPU interruption and improving the efficiency of data transmission. A DMA channel is set between the computing core and the external terminal, allowing data to be transferred directly from the memory to the computing core or from the computing core to the external storage without passing through the central processing unit (CPU). The generation of the external DMA channel enables data to be transferred between the computing core and the external device efficiently and with low latency without passing through the CPU, thereby reducing the bottleneck of data transmission. In this way, the system can better utilize the parallel processing capabilities of the computing motherboard and the high bandwidth of the external storage, significantly improving data processing efficiency.
[0088] More specifically, based on the resource scheduling characteristics, the internal data processing flow of the computing core is analyzed to obtain the transfer path of the data calculation results between the various computing cores. The data transfer path involves the transfer rules of data from one computing core to another, especially how the data is transferred from the output of the computing core to other cores after processing is completed, or how the data is returned from the computing core to the external storage or terminal.
[0089] Each transfer path includes three elements: starting node: the computing core where the data calculation results come from, end node: the target computing core that receives the data calculation results, data transfer standard: defines the type, quantity, transfer conditions, etc. of the data calculation results to be transferred. The analysis of the data transfer path helps the system understand the data flow rules and dependencies between computing cores and optimize the flow of data between cores. It avoids unnecessary data copying and transmission, reduces data access conflicts and delays, and provides a basis for the next step of creating interactive DMA channels between computing cores.
[0090] More specifically, based on the data processing flow analyzed above, the starting node and the end node of the data are used as the subject and object of the DMA channel respectively to construct a direct memory mapping framework. Specifically, the data is transferred from the starting node (computing core) to the end node (computing core) through the DMA channel and processed in accordance with the data transfer standard. The generation of interactive DMA channels enables direct data transmission between computing cores through DMA without CPU intervention, ensuring efficient data transmission between different computing cores. The generation of interactive DMA channels makes data transmission between computing cores more efficient and low-latency, reduces CPU intervention, and improves parallel processing capabilities. Through DMA's automated data transmission, the system can more efficiently share data across cores and coordinate computing tasks. This step makes the collaborative computing between computing cores no longer limited to traditional memory access methods, but is optimized through efficient data flow and resource scheduling mechanisms, improving the overall performance of the system.
[0091] Preferably, the step of directly mapping the data set to be calculated to the data cache area of the computing core through the external DMA channel so that the computing core performs data calculation processing on the data set to be calculated, and obtaining the data calculation result includes:
[0092] S51: performing data matching on the data set to be calculated in the external terminal through the external DMA channels of each of the computing cores, and marking the content of each data part of the data set to be calculated in the external terminal according to the result of the data matching;
[0093] S52: performing direct memory mapping of the data portion marked with the corresponding content to the data cache area of each computing core according to the DMA channel of each computing core, and assigning a process number to the data portion according to the content mark to obtain the data portion with the process number;
[0094] S53: Instruct the computing core to perform data computing processing on the data portion of the data set to be calculated that is in its own data cache area to obtain a data calculation result, and generate a process number for the data calculation result according to the process number of the data portion to obtain a data calculation result with a process number; wherein the process number is used to describe whether the data calculation result needs to be processed again by the remaining computing cores.
[0095] Specifically, before loading the data set to be calculated from the external terminal to the data cache area of the computing core through the external DMA channel of each computing core, the data set to be calculated is first matched. This process will check the different data parts of the data set to be calculated in the external terminal, and match each data part with the data required by the computing core. The result of the data matching will determine which data parts should be mapped to the cache area of the computing core. Each data part to be calculated will be marked so that subsequent memory mapping operations can accurately allocate these parts to the appropriate computing core. Through data matching, it can be ensured that the computing core only receives the required relevant data parts, avoiding the transmission of irrelevant data, saving memory and bandwidth resources, and the data marking mechanism makes data management more orderly, reduces the complexity of data management and scheduling, and improves the flexibility and response speed of the system.
[0096] More specifically, after the data part is marked, direct memory mapping is performed according to the external DMA channel of each computing core. Specifically, the marked data part will be transmitted and mapped to the data cache area of the computing core through DMA technology. This process does not require CPU intervention. The data is directly transmitted from the external terminal to the memory area of the target computing core through the DMA channel, realizing efficient data loading. During the mapping process, the system will assign a "process number" to each data part according to the marking information of the data part. The process number is an important identifier used to track the execution order and dependencies of data calculations. Direct memory mapping reduces CPU intervention, improves data transmission efficiency, and enables data to quickly reach the computing core cache area. The process number gives each data part a unique identification, which is convenient for subsequent processing and scheduling, ensuring that the order and processing dependencies of the data are maintained.
[0097] More specifically, after the data part is successfully mapped to the cache area of the computing core, the computing core will perform data calculation processing on it. During this process, the computing core will perform specific computing tasks (such as matrix operations, image processing, data analysis, etc.) to generate data calculation results. While calculating, the computing core generates a process number for the data calculation result based on the process number of the data part, ensuring that the process number of each calculation result can correctly reflect the order and dependency of its calculation processing. The computing core can calculate the data independently and efficiently, and through the process number, it can accurately track and manage the processing process of each data part. The generated calculation results can be subsequently processed and scheduled according to the order and dependency of the process numbers, ensuring that the data calculation results are passed to other computing cores or terminals as needed.
[0098] More specifically, after the data calculation is completed, each data calculation result will be assigned a process number. The process number is used to describe whether the calculation result needs to be processed again by other computing cores. If the calculation result of a computing core needs to be further processed by other cores (for example, it depends on the calculation results of other computing cores), the process number of the calculation result will be marked as "pending" or "waiting". If a data calculation result does not need to be processed again, its process number will be marked as "completed" or "final result", and then the result can be output to an external terminal or used by other computing cores. The introduction of process numbers provides the system with an efficient result management and scheduling mechanism. Through the process number, the system can clearly know which data calculation results have been completed and which need to be further processed by other cores, so as to better schedule the computing tasks. The process number helps reduce the dependency conflicts of data calculation results, ensure the coordination of parallel computing processes, and avoid waiting and resource waste between computing cores.
[0099] Preferably, the step of directly mapping the data calculation result to the data cache area of the remaining computing cores through the interactive DMA channel includes:
[0100] S61: performing format verification on the data calculation result according to the interactive DMA channel, and performing data compression processing on the data calculation result according to the verification result to obtain a compressed data packet;
[0101] S62: Directly map the compressed data packet to the data cache area of the corresponding computing core through the interactive DMA channel.
[0102] Specifically, before the data calculation results are transmitted through the interactive DMA channel, the system first performs format verification on the data calculation results. This process includes verifying whether the data meets the data format requirements of the target computing core, whether the data is complete, whether it conforms to the predetermined data structure, etc. If there are format problems with the data (such as mismatch, missing fields or damage), the system will take corresponding measures (such as discarding invalid data or marking it as erroneous data). Format verification can ensure that the calculation results maintain data consistency during the transmission and reception process, and avoid calculation errors or crashes due to data format errors. This step can improve the robustness and fault tolerance of the system, ensure the correctness of the data calculation results, and prevent format-error data from affecting data transmission between computing cores.
[0103] More specifically, once the data calculation results pass the format verification, the system will perform data compression processing on them. Data compression can use a variety of algorithms (such as Huffman coding, LZ compression, etc.) to reduce bandwidth consumption during data transmission. The compressed data forms a compressed data packet, which occupies less memory space than the original data and can improve data transmission efficiency in subsequent transmission processes. Data compression reduces the size of the data, reduces memory usage and transmission bandwidth consumption, especially in high-performance computing environments. Data compression can greatly increase the data transmission speed. The compressed data packet is easier to transmit efficiently between multiple computing cores, reducing data bottleneck problems, thereby improving the system's parallel computing capabilities and resource utilization.
[0104] More specifically, the compressed data packet is transmitted to the target computing core through the interactive DMA channel. Interactive DMA (Direct Memory Access) allows data to be directly transferred from the source address to the target memory area without passing through the main processor. During the transmission process, the DMA channel will ensure that the compressed data packet is accurately mapped to the data cache area of the target computing core. The DMA controller of the computing core will decompress the received data packet and directly put it into the data cache area of the corresponding computing core for subsequent processing or calculation by the computing core. The use of the DMA channel avoids CPU intervention and improves the efficiency of data transmission. DMA can perform high-speed, low-latency data transmission, avoiding the heavy burden of the CPU in traditional data transmission. Through DMA direct mapping, the time of memory copy operation can be saved, and the intermediate links in data transmission can be reduced, thereby improving the data transmission speed. Direct mapping to the data cache area of the target computing core ensures data consistency and synchronization, and avoids data duplication or repeated transmission.
[0105] Preferably, the step of performing format verification on the data calculation result according to the interactive DMA channel, and performing data compression processing on the data calculation result according to the verification result to obtain a compressed data packet includes:
[0106] S611: Obtain computing performance indicators and current task execution status of the remaining computing cores corresponding to the interactive DMA channel, and perform real-time evaluation of data processing capabilities of the remaining computing cores according to the current task execution status and the computing performance indicators to obtain real-time data processing performance of the remaining computing cores;
[0107] S612: Analyzing the most suitable format for the data calculation result according to the real-time data processing performance to obtain a theoretical compression format of the data calculation result;
[0108] S613: Perform format verification on the data calculation result according to the theoretical compression format to obtain a verification result. When the verification result shows that the data calculation result does not conform to the theoretical compression format, perform data compression processing on the data calculation result according to the theoretical compression format to obtain a compressed data packet.
[0109] Specifically, the system first obtains the computing performance indicators (such as computing power, memory bandwidth, load status, etc.) of other computing cores corresponding to the interactive DMA channel, as well as the current task execution status (such as the progress of the current computing task, load distribution, etc.). These data provide real-time computing core performance status, including whether each computing core is in an idle state, the processing progress of the computing task, and whether the load is balanced. Obtaining the performance indicators and task execution status of the computing core enables the system to dynamically evaluate the load status of different computing cores, helping to optimize data allocation and transmission strategies. By real-time monitoring of the computing core status, bottlenecks caused by uneven performance or overload can be avoided to ensure effective data scheduling.
[0110] More specifically, based on the performance indicators obtained above and the current task execution status, the system evaluates the data processing capabilities of each computing core in real time to obtain real-time data processing performance. This step evaluates the actual processing capabilities of each computing core in the current task by combining factors such as computing power and memory bandwidth, ensuring that the resource usage of the computing core during task execution is not over-congested or idle. The real-time data processing performance evaluation can intelligently and dynamically adjust task allocation according to the load of the computing core and optimize resource usage. This evaluation ensures that each computing core can participate in computing tasks with optimal performance and avoids performance waste due to improper resource allocation.
[0111] More specifically, based on the real-time evaluation of the computing core data processing capabilities, the system performs the most suitable format analysis on the data calculation results and selects the most suitable compression format. This analysis optimizes the data compression strategy based on the real-time performance of different computing cores. Specifically, the system considers the computing core's processing power, memory limitations, and task characteristics, and selects a theoretically optimal compression format to maximize data transmission efficiency and reduce transmission delays. The most suitable format analysis ensures that the selected compression format can achieve the best performance under given hardware resources and task requirements. This analysis helps the system dynamically select the most suitable compression algorithm based on the current resource status, thereby maximizing the compression effect and ensuring the efficiency of data transmission.
[0112] More specifically, after determining the theoretical compression format, the system performs format verification on the data calculation results to check whether the data meets the requirements of the selected theoretical compression format. If the format verification result shows that the data does not conform to the predetermined compression format, the system performs data compression processing on the data according to this theoretical compression format to generate a compressed data packet. The compressed data packet will be further transmitted to the target computing core, reducing the bandwidth requirement and memory occupancy during transmission. Format verification ensures that important information is not lost during the compression process and conforms to the standards of the theoretical compression format, guaranteeing the correctness and efficiency of the data after compression. Data compression processing optimizes the data transmission process, reduces the bandwidth burden, and improves the overall performance of the system. Especially in parallel computing tasks that require processing a large amount of data, the efficient data packets generated after compression not only reduce memory usage but also avoid possible communication bottlenecks between computing cores, improving the efficiency of data transmission and processing.
[0113] In a second aspect, the present invention provides a data transmission acceleration system for a multi-core heterogeneous ASIC computing motherboard, which is used to implement the data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard according to any one of the first aspects.
[0114] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard, characterized in that: include: Performing mode adjustment on a designated computing core on a computing mainboard so that the computing core switches to an initialization allocation mode; The calculation core in the initialization allocation mode is driven by a preset high-speed transmission protocol to receive data from the data set to be calculated of the external terminal, so as to obtain the pre-order data part of the data set to be calculated; wherein the data set to be calculated is all the data that needs to be calculated in the external terminal; Instructing the computing core in the initialization allocation mode to perform data computing processing on the preceding data portion, and to perform resource scheduling monitoring on the data computing processing of the preceding data portion, so as to obtain the data computing result of the preceding data portion after the data computing processing and the resource scheduling characteristics of the data set to be calculated; Generate a DMA channel for each computing core of the computing mainboard according to the resource scheduling characteristics, so that each computing core of the computing mainboard switches to a DMA mapping mode; wherein the computing core in the DMA mapping mode has an external DMA channel and an interactive DMA channel, the external DMA channel is used to directly map the data set to be calculated to the data cache area of the computing core through the DMA technology, and the interactive DMA channel is used to directly map the data processing result of the data set to be calculated to the data cache area of the remaining computing cores through the DMA technology; Directly mapping the data set to be calculated to the data cache area of the calculation core through the external DMA channel, so that the calculation core performs data calculation processing on the data set to be calculated to obtain a data calculation result; When the data calculation result needs to be processed again by the remaining computing cores, the data calculation result is directly mapped to the data cache area of the remaining computing cores through the interactive DMA channel, so that the remaining computing cores can perform subsequent processing on the data calculation result to obtain a new round of data calculation results; When the data calculation result does not need to be processed again by other calculation cores, the data calculation result is output; The step of driving the computing core in the initialization allocation mode to receive data from the data set to be calculated of the external terminal through a preset high-speed transmission protocol to obtain the pre-order data part of the data set to be calculated includes: The computing core in the initialization allocation mode is driven by a preset high-speed transmission protocol to continuously receive data from the data set to be calculated of the external terminal, and time-marks the data parts of the data set to be calculated received at different times to obtain a plurality of data parts arranged in time sequence; The various data parts continuously received at different times are sequentially substituted into the task execution sequence in the computing core in the initialization allocation mode, and the task integrity of the task execution sequence is evaluated. When the evaluation result shows that the various data parts in the task execution sequence constitute an executable computing task, the task execution sequence is converted to obtain the preceding data part.
2. The data transmission acceleration method of a multi-core heterogeneous ASIC computing motherboard according to claim 1, characterized in that: The step of performing mode adjustment on a designated computing core on a computing mainboard so that the computing core switches to an initialization allocation mode includes: Obtaining computing performance indicators of each computing core on the computing mainboard, and evaluating the task balance of each computing core according to the computing performance indicators, so as to obtain a computing core with the best task balance characteristics as an initialization object; Performing an external data interface configuration based on a high-speed transmission protocol on the initialization object so that the initialization object maintains a communication connection state with an external terminal; wherein the initialization object in a communication connection state with the external terminal is used to receive data from the external terminal; The program configuration of resource scheduling monitoring and mapping channel construction is performed on the initialization object, so that the initialization object is switched to the initialization allocation mode.
3. The data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard according to claim 1, characterized in that: The steps of causing the computing core in the initialization allocation mode to perform data computing processing on the preceding data portion, and performing resource scheduling monitoring on the data computing processing of the preceding data portion, so as to obtain the data computing result of the preceding data portion after the data computing processing and the resource scheduling characteristics of the data set to be calculated include: Instructing the calculation core in the initialization allocation mode to perform data calculation processing on the preceding data portion to obtain a data calculation result of the preceding data portion after the data calculation processing; When the computing core performs data processing on the preceding data portion, data collection and feature analysis are performed on the computing resource scheduling status of the computing core to obtain computing resource scheduling parameters for the computing core to perform data computing processing on the preceding data portion; Performing a reverse calculation of the overall form of the data set to be calculated on the preceding data portion according to the computing resource scheduling parameters, so as to obtain the overall resource requirement form of the data set to be calculated; The overall resource demand form is allocated and processed according to the computing performance indicators of each computing core of the computing mainboard to obtain resource scheduling characteristics of each computing core corresponding to the data set to be calculated; wherein the resource scheduling characteristics are used to describe the allocation method between each data part in the data set to be calculated and each computing core, as well as the computing processing method of the data part allocated to each computing core.
4. The data transmission acceleration method of a multi-core heterogeneous ASIC computing motherboard as claimed in claim 3, characterized in that: The step of performing a reverse calculation of the overall form of the data set to be calculated on the preceding data portion according to the computing resource scheduling parameters to obtain the overall resource requirement form of the data set to be calculated comprises: Performing a task descriptive analysis of the computing task on the preceding data portion to obtain computing task characteristics of the preceding data portion; Performing an inductive analysis of correlation characteristics on the computing task characteristics according to the computing resource scheduling parameters to obtain computing correlation characteristics between the computing resource scheduling parameters and the computing task characteristics; The calculation-related features are identified and processed according to the task identification standards summarized by the machine learning model in the preset database to obtain the overall task form corresponding to the calculation-related features and the overall resource demand form corresponding to the overall task form.
5. The data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard according to claim 1, characterized in that: The step of generating a DMA channel for each computing core of the computing mainboard according to the resource scheduling feature so that each computing core of the computing mainboard switches to a DMA mapping mode includes: Parsing the data allocation standard for each computing core of the computing mainboard according to the resource scheduling characteristics to obtain the data allocation standard for each computing core corresponding to the data set to be calculated; Taking the computing core as a data mapping object and the external terminal as a data mapping subject, constructing a framework for performing direct memory mapping on the data mapping object and the data mapping subject according to the data allocation standard to generate an external DMA channel of the computing core; According to the resource scheduling characteristics, the data processing flow of each computing core of the computing mainboard is analyzed to obtain a transfer path of data calculation results between each computing core; wherein the transfer path includes a starting node, an end node and a data transfer standard, the starting node is a computing core to which the data calculation result is to be transferred, the end node is a computing core to which the data calculation result is transferred, and the data transfer standard is used to determine the data calculation result that needs to be transferred from the starting node to the end node; The starting node is used as the data mapping subject, the end node is used as the data mapping object, and a framework for direct memory mapping of the data mapping subject and the data mapping object is constructed according to the data transfer standard to generate an interactive DMA channel of the computing core.
6. The data transmission acceleration method of a multi-core heterogeneous ASIC computing motherboard according to claim 1, characterized in that: The step of directly mapping the data set to be calculated to the data cache area of the computing core through the external DMA channel so that the computing core performs data calculation processing on the data set to be calculated, and obtaining the data calculation result includes: Performing data matching on the data set to be calculated in the external terminal through the external DMA channels of each of the computing cores, and marking the content of each data part of the data set to be calculated in the external terminal according to the result of the data matching; Directly memory-map the data portion marked with the corresponding content to the data cache area of each computing core according to the DMA channel of each computing core, and assign a process number to the data portion according to the content mark to obtain the data portion with the process number; The computing core is instructed to perform data computing processing on the data portion of the data set to be calculated that is in its own data cache area to obtain a data computing result, and a process number is generated for the data computing result according to the process number of the data portion to obtain a data computing result with a process number; wherein the process number is used to describe whether the data computing result needs to be processed again by the remaining computing cores.
7. The data transmission acceleration method of a multi-core heterogeneous ASIC computing motherboard according to claim 1, characterized in that: The step of directly mapping the data calculation result to the data cache area of the remaining computing cores through the interactive DMA channel comprises: Performing format verification on the data calculation result according to the interactive DMA channel, and performing data compression processing on the data calculation result according to the verification result to obtain a compressed data packet; The compressed data packet is directly mapped to the data cache area of the corresponding computing core through the interactive DMA channel.
8. The data transmission acceleration method of a multi-core heterogeneous ASIC computing motherboard as claimed in claim 7, characterized in that: The steps of formatting the data calculation result according to the interactive DMA channel and performing data compression processing on the data calculation result according to the verification result to obtain a compressed data packet include: Obtaining computing performance indicators and current task execution status of the remaining computing cores corresponding to the interactive DMA channel, and performing real-time evaluation of data processing capabilities of the remaining computing cores according to the current task execution status and the computing performance indicators to obtain real-time data processing performance of the remaining computing cores; Analyzing the most suitable format for the data calculation result according to the real-time data processing performance to obtain a theoretical compression format of the data calculation result; The data calculation result is format-verified according to the theoretical compression format to obtain a verification result. When the verification result shows that the data calculation result does not conform to a predetermined compression format, the data calculation result is data-compressed according to the theoretical compression format to obtain a compressed data packet.
9. A data transmission acceleration system for a multi-core heterogeneous ASIC computing motherboard, characterized in that: A data transmission acceleration method for a multi-core heterogeneous ASIC computing motherboard is used to implement any one of claims 1-8.
Citation Information
Patent Citations
In-memory calculation method and chip design
CN118349212A
Multi-channel DMA data processing method and electronic equipment
CN119166570A