Multi-core heterogeneous architecture chip task processing method, system and device and storage medium

By building the resource allocation matrix and dual-loop feedback mechanism of multi-core heterogeneous architecture chips, dynamically adjusting task allocation, the problem of unbalanced power consumption of heterogeneous cores is solved, and the chip's high-efficiency energy consumption balance and performance optimization are achieved.

CN120407156APending Publication Date: 2025-08-01GUANGZHOU KETENG INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510417997.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

During the task execution, unbalanced power consumption of multi-core heterogeneous architecture chips leads to chip heating, performance degradation and resource waste, and it is difficult to balance performance and energy consumption without real-time closed-loop feedback.

Method used

By collecting short-term and long-term operation data of heterogeneous cores, building a resource allocation matrix, dynamically adjusting task allocation, combining the dual-ring feedback mechanism to optimize power consumption and performance, and using three-dimensional resource allocation matrix and inter-core communication path optimization.

Benefits of technology

It realizes that while ensuring task execution performance, reduce chip power consumption, improve resource utilization and energy efficiency ratio, and improve task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407156A_ABST
    Figure CN120407156A_ABST
Patent Text Reader

Abstract

The invention provides a multi-core heterogeneous architecture chip task processing method, system and device and a storage medium, and relates to the technical field of computers. Comprising the steps that in response to a task allocation instruction, short-term operation data of a plurality of heterogeneous cores are collected, a first resource allocation matrix is constructed according to the short-term operation data, the first resource allocation matrix comprises vector data and vector weights of the heterogeneous cores, the vector data are used for representing performance resources of the heterogeneous cores, and the performance resources of the heterogeneous cores are allocated through the first resource allocation matrix; the task in the task allocation instruction is allocated to the target heterogeneous core, long-term operation data of the multiple heterogeneous cores are collected, the first resource allocation matrix is updated according to the long-term operation data, a second resource allocation matrix is obtained, the long-term operation data comprise heterogeneous core power consumption, and secondary allocation is conducted on the task according to the second resource allocation matrix. According to the method, task execution and resource allocation can be dynamically adjusted, the energy efficiency ratio is improved, and the power consumption of the heterogeneous core is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, system, device, and storage medium for task processing of a multi-core heterogeneous architecture chip. Background Art

[0002] Multi-core heterogeneous architecture chips are widely used in fields such as industrial control, autonomous driving, and edge computing. In a multi-core heterogeneous architecture chip, as tasks are executed, the power consumption of different heterogeneous cores will vary. Excessive power consumption will not only cause the chip to heat up, affecting its performance and lifespan, but also result in waste of chip resources. Related technologies lack real-time closed-loop feedback on core power consumption and it is difficult to balance performance indicators and energy consumption. Summary of the Invention

[0003] The main objective of the embodiments of the present disclosure is to propose a method, system, device, and storage medium for task processing of a multi-core heterogeneous architecture chip, which can improve the energy efficiency ratio of chip task processing and reduce chip power consumption.

[0004] To achieve the above objective, on the one hand, an embodiment of the present application proposes a method for task processing of a multi-core heterogeneous architecture chip, including the following steps:

[0005] In response to a task allocation instruction, collect short-term operation data of multiple heterogeneous cores, and construct a first resource allocation matrix according to the short-term operation data. The first resource allocation matrix includes vector data and vector weights of multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores;

[0006] Through the first resource allocation matrix, allocate the tasks in the task allocation instruction to the target heterogeneous core;

[0007] Collect long-term operation data of multiple heterogeneous cores, and update the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes heterogeneous core power consumption;

[0008] Perform secondary allocation of the tasks according to the second resource allocation matrix.

[0009] In some embodiments, the short-term operation data and the long-term operation data of the heterogeneous cores are obtained through the following steps:

[0010] Collect the short-term operation data through the first monitoring module of the heterogeneous core;

[0011] Collect the long-term operation data through the second monitoring module of the heterogeneous core.

[0012] In some embodiments, the short-term operation data includes core computing data, remaining capacity of the storage unit, and communication channel delay data, and the vector data includes a first dimension, a second dimension, and a third dimension;

[0013] The first dimension is used to characterize the core computing data;

[0014] The second dimension is used to characterize the remaining capacity of the storage unit;

[0015] The third dimension is used to characterize the communication channel delay data.

[0016] In some embodiments, the vector weights include weight coefficients of multiple dimensions. Assigning the task in the task assignment instruction to the target heterogeneous core through the first resource allocation matrix includes the following steps:

[0017] Parse the task instruction to obtain the identifier of the task type and the corresponding resource requirement parameters, where the resource requirement parameters include the storage capacity threshold and the upper limit of data transmission delay;

[0018] Determine the weight coefficient of the first dimension of each heterogeneous core according to the identifier;

[0019] Determine the target heterogeneous core according to the weight coefficient of the first dimension of each heterogeneous core;

[0020] Determine the weight coefficient of the second dimension of each heterogeneous core according to the storage capacity threshold;

[0021] Determine the weight coefficient of the third dimension of each heterogeneous core according to the upper limit of data transmission delay;

[0022] Determine the performance requirement data for the target heterogeneous core to process the task according to the weight coefficients of the second dimension and the third dimension of each heterogeneous core;

[0023] Generate a first allocation instruction according to the performance requirement data and the task in the task assignment instruction, and send the first allocation instruction to the target heterogeneous core.

[0024] In some embodiments, sending the first allocation instruction to the target heterogeneous core includes the following steps:

[0025] Use the instruction adaptation conversion interface to convert the first allocation instruction into a second allocation instruction;

[0026] Send the second allocation instruction to the target heterogeneous core.

[0027] In some embodiments, updating the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix includes the following steps:

[0028] Obtain the energy efficiency ratio of the heterogeneous cores according to the power consumption of the heterogeneous cores and the performance resources consumed in the long-term operation data;

[0029] Taking maximizing the energy efficiency ratio of the heterogeneous cores as the goal, correct the vector weights of the first resource allocation matrix to obtain a second resource allocation matrix.

[0030] In some embodiments, performing secondary allocation of the tasks according to the second resource allocation matrix includes the following steps:

[0031] Analyze the vector weights of the second resource allocation matrix to obtain priority coding information;

[0032] Use the priority coding information to find and modify the entries of the routing table of the communication system to obtain target entries;

[0033] Determine a new inter-core communication path according to the target entries, and perform secondary allocation of the tasks according to the new inter-core communication path.

[0034] On the other hand, an embodiment of the present invention provides a task processing system for a multi-core heterogeneous architecture chip, including:

[0035] A first module, configured to collect short-term operation data of multiple heterogeneous cores in response to a task allocation instruction, and construct a first resource allocation matrix according to the short-term operation data, where the first resource allocation matrix includes vector data and vector weights of multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores;

[0036] A second module, configured to allocate the tasks in the task allocation instruction to target heterogeneous cores through the first resource allocation matrix;

[0037] A third module, configured to collect long-term operation data of multiple heterogeneous cores, and update the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes the power consumption of the heterogeneous cores;

[0038] A sixth module, configured to perform secondary allocation of the tasks according to the second resource allocation matrix.

[0039] On the other hand, an embodiment of the present invention provides an electronic device, including:

[0040] At least one processor;

[0041] At least one memory, configured to store at least one program;

[0042] When the at least one program is executed by the at least one processor, such that at least one of the processors implements the multi-core heterogeneous architecture chip task processing method as described in the previous embodiments.

[0043] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the multi-core heterogeneous architecture chip task processing method as described in the previous embodiments.

[0044] At least one of the above technical solutions of the present invention has the following advantages or beneficial effects:

[0045] Based on short-term operation data, a resource allocation matrix is constructed, which can quickly allocate tasks to heterogeneous cores for processing, achieving an accurate match between tasks and heterogeneous core resources. During the task execution process, according to the power consumption in the long-term operation data, the resource allocation matrix is updated and the tasks are re-allocated, which can flexibly adapt to the resource requirements of the tasks and the power consumption of the heterogeneous cores, timely adjust the task execution and resource allocation, balance the performance index and power consumption of the heterogeneous cores, that is, reduce the power consumption of the heterogeneous cores while ensuring the performance requirements of task execution, combine short-term operation data and long-term operation data for task allocation, realize reasonable allocation of tasks on different time scales, give full play to the advantages of each heterogeneous core, and improve the task processing efficiency of the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart of a multi-core heterogeneous architecture chip task processing method provided by an embodiment of the present application;

[0047] Figure 2 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.

[0049] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0051] Please refer to Figure 1 , Figure 1 which is an optional flowchart of a method for processing tasks of a multi-core heterogeneous architecture chip provided by some embodiments of this application. Embodiments of the present invention provide a method for processing tasks of a multi-core heterogeneous architecture chip, including but not limited to steps S100 to S400:

[0052] Step S100, in response to a task allocation instruction, collect short-term operation data of multiple heterogeneous cores, and construct a first resource allocation matrix according to the short-term operation data. The first resource allocation matrix includes vector data and vector weights of multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores;

[0053] Step S200, through the first resource allocation matrix, allocate the task in the task allocation instruction to the target heterogeneous core;

[0054] Step S300, collect long-term operation data of multiple heterogeneous cores, and update the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes the power consumption of the heterogeneous cores;

[0055] Step S400, perform secondary allocation of the task according to the second resource allocation matrix.

[0056] In step S100 of some embodiments, the driver interface of the task parsing module of the main control core receives an external task instruction. To process the task in the task instruction, first, the main control core establishes a communication connection with the monitoring modules of multiple heterogeneous cores. The monitoring modules obtain the operation status data of their respective heterogeneous cores, and the operation status data includes short-term operation data and long-term operation data. The short-term operation data refers to the performance-related data generated by each heterogeneous core during operation within a relatively short time range, such as the occupancy rate of the core computing unit of the heterogeneous core, the remaining capacity of the storage unit, and the communication channel delay data, etc. The main control core can construct a first resource allocation matrix based on the short-term operation data. This matrix includes vector data and vector weights of multiple heterogeneous cores. The vector data can describe the performance resources of each heterogeneous core. It can be a multi-dimensional vector, and each dimension represents a performance index of the heterogeneous core. Different dimensions can correspond to different short-term operation data. The vector weights reflect the adaptation degree of different performance indexes in task allocation.

[0057] In some embodiments, the short-term operation data and long-term operation data of the heterogeneous cores are obtained through steps S110 to S120:

[0058] Step S110: Collect short-term operation data through the first monitoring module of the heterogeneous core.

[0059] Step S120: Collect long-term operation data through the second monitoring module of the heterogeneous core.

[0060] In steps S110 to S120 of some embodiments, in a multi-core heterogeneous architecture chip, each heterogeneous core is usually equipped with a monitoring module. The main function of the monitoring module is to monitor the operation status data of the heterogeneous core in real time and send the operation status data to the main control core as the basis for task allocation. The monitoring module can be divided into a first monitoring module and a second monitoring module. The second monitoring module has a longer data collection period and a lower frequency compared to the first monitoring module. The first monitoring module is used to collect short-term operation data.

[0061] The short-term operation data reflects the hardware status data of the heterogeneous core within a relatively short period of time. For example, the computing load within a short cycle, the remaining capacity of the storage unit, and the communication channel delay data of the heterogeneous core, etc. The main control core can quickly evaluate the current performance resource status of each heterogeneous core based on these data, and thus make a reasonable task allocation decision.

[0062] The main task of the second monitoring module is to collect the long-term operation data of the heterogeneous core. The long-term operation data is a comprehensive record of the operation status of the heterogeneous core within a relatively long period of time, including the power consumption information of the heterogeneous core.

[0063] The long-term operation data can help the system understand the behavior patterns of the heterogeneous core under different workloads and time scales, especially the characteristics in terms of power consumption. Based on these long-term operation data, the main control core can update the resource allocation matrix to achieve a more optimized and energy-saving task allocation strategy, and avoid problems such as overheating or excessive power consumption of some heterogeneous cores caused by long-term high-load operation.

[0064] In some embodiments, the monitoring module adopts a double-loop feedback mechanism. The inner loop can collect hardware status data with a period of 10 ms, and the outer loop will calculate the system energy efficiency ratio index with a period of 100 ms.

[0065] In some embodiments, the short-term operation data includes core computing data, the remaining capacity of the storage unit, and communication channel delay data. The vector data includes a first dimension, a second dimension, and a third dimension;

[0066] The first dimension is used to represent the core computing data;

[0067] The second dimension is used to represent the remaining capacity of the storage unit;

[0068] The third dimension is used to represent the communication channel delay data.

[0069] Through the vector data in these three dimensions, the main control core can comprehensively consider the computing power, storage resources, and communication efficiency of heterogeneous cores, and combine the vector weights to reasonably allocate the tasks in the task allocation instructions, so as to achieve the efficient operation of the multi-core heterogeneous architecture chip.

[0070] In some embodiments, the resource allocation matrix uses three-dimensional coordinate modeling, including the X-axis, Y-axis, and Z-axis. The X-axis represents the type of computing core, the Y-axis represents the storage unit capacity, and the Z-axis represents the communication channel bandwidth. The coordinate points on each axis correspond to the allocated weight values.

[0071] In step S200 of some embodiments, after constructing the first resource allocation matrix, the main control core can use this matrix to complete task allocation. It will comprehensively evaluate the performance and resource status of each heterogeneous core according to the vector data and vector weights of each heterogeneous core in the matrix, so as to reasonably allocate the tasks in the task allocation instructions to the most suitable target heterogeneous core.

[0072] In some embodiments, step S200 may include but is not limited to steps S210 to S270:

[0073] Step S210, parse the task instruction to obtain the identifier of the task type and the corresponding resource requirement parameters, where the resource requirement parameters include the storage capacity threshold and the upper limit of data transmission delay;

[0074] Step S220, determine the weight coefficient of the first dimension of each heterogeneous core according to the identifier;

[0075] Step S230, determine the target heterogeneous core according to the weight coefficient of the first dimension of each heterogeneous core;

[0076] Step S240, determine the weight coefficient of the second dimension of each heterogeneous core according to the storage capacity threshold;

[0077] Step S250, determine the weight coefficient of the third dimension of each heterogeneous core according to the upper limit of data transmission delay;

[0078] Step S260, determine the performance requirement data for the target heterogeneous core to process the task according to the weight coefficients of the second and third dimensions of each heterogeneous core;

[0079] Step S270, generate the first allocation instruction according to the performance requirement data and the task in the task allocation instruction, and send the first allocation instruction to the target heterogeneous core.

[0080] In steps S210 to S270 of some embodiments, the task parsing module of the main control core parses the received task instruction to obtain the identifier of the task type and resource requirement parameters such as the storage capacity threshold and the upper limit of data transmission delay.

[0081] Different types of tasks have different requirements for the core computing power. For example, for compute-intensive tasks, heterogeneous cores with strong core computing power should be assigned higher weight coefficients; while for other types of tasks, the weight coefficients may be correspondingly reduced. The master core determines the weight coefficients of the first dimension (core computing data) of each heterogeneous core according to the identifier. The core computing data includes the core computing type, computing power value, etc., so as to select the target heterogeneous core. Usually, heterogeneous cores with higher weight coefficients are more likely to be selected as the target heterogeneous core.

[0082] Similarly, the master core then determines the weight coefficients of the second dimension (remaining capacity of the storage unit) of each heterogeneous core according to the storage capacity threshold, and determines the weight coefficients of the third dimension (communication channel delay data) of each heterogeneous core according to the upper limit of data transmission delay.

[0083] After that, by synthesizing the weight coefficients of the second and third dimensions, the performance requirement data for the target heterogeneous core to process the task is determined. The performance requirement data includes performance resources such as the storage capacity and communication channel used by the target heterogeneous core to process the task. Finally, a first allocation instruction is generated according to the performance requirement data and the task in the task allocation instruction, and is sent to the target heterogeneous core.

[0084] The master core comprehensively considers the computing, storage, and communication capabilities of the heterogeneous cores according to the specific requirements of the task, realizes the reasonable allocation of the task, and improves the overall performance and efficiency of the multi-core heterogeneous architecture chip.

[0085] In some embodiments, the task types recognized by the task parsing module for external task instructions include compute-intensive tasks, IO-intensive tasks, and hybrid tasks. The resource requirement parameters at least include the expected computing power value, memory bandwidth threshold, and upper limit of data transmission delay.

[0086] In some embodiments, step S270 may include but is not limited to steps S271 to S272:

[0087] Step S271, using the instruction adaptation conversion interface, convert the first allocation instruction into a second allocation instruction;

[0088] Step S272, send the second allocation instruction to the target heterogeneous core.

[0089] In steps S271 to S272 of some embodiments, since different heterogeneous cores may have different instruction sets and operating rules, the first allocation instruction may not be directly recognized and executed by the target heterogeneous core. The instruction adaptation and conversion interface converts the first allocation instruction into a second allocation instruction. The role of the instruction adaptation and conversion interface is to adjust and convert the instruction in terms of format, etc., to meet the requirements of the target heterogeneous core. The second allocation instruction after conversion will be sent to the target heterogeneous core, enabling the target heterogeneous core to accurately understand and execute tasks according to the instruction requirements. Through this process of instruction conversion and sending, the compatibility and effectiveness of task allocation in a heterogeneous core environment are ensured.

[0090] In some embodiments, according to the handshake protocol and arbitration rules pre-stored in the inter-core communication protocol library, the driver interface generation module of the master core converts the multi-level scheduling policy (the first allocation instruction) into a hardware abstraction layer instruction set (the second allocation instruction) including synchronous clock domain control signals and asynchronous packet routing information. The hardware abstraction layer instruction set is written into the configuration registers of each heterogeneous core through a standardized bus interface; the synchronous clock domain control signal is used to ensure data consistency between cores, the routing information is the rules and policies used by network devices to determine how to forward packets, the transmission path is the actual transmission path that the packet passes from the source core to the target core, and the asynchronous packet routing information is used to determine the burst asynchronous transmission path, and the burst asynchronous transmission path is used to process real-time interrupt requests. The hardware abstraction layer instruction set adopts a variable-length coding format, including an opcode field, a target core address field, and a data payload field, where the opcode field supports 32 basic operation type extensions.

[0091] In step S300 of some embodiments, during the task processing of a multi-core heterogeneous architecture chip, the power consumption performance of different heterogeneous cores will vary. Excessive power consumption will not only cause the chip to heat up, affecting its stability and lifespan, but also result in energy waste.

[0092] After the master core collects long-term operation data, it will update the first resource allocation matrix based on this data. For example, if a certain heterogeneous core has a consistently high power consumption during long-term operation, during the update, the vector weight corresponding to the heterogeneous core in the first resource allocation matrix may be reduced, thereby reducing the priority of this heterogeneous core in resource allocation and the probability of subsequent tasks being assigned to this core. Through such an update operation, a second resource allocation matrix is obtained. This new matrix can more accurately reflect the comprehensive factors such as the current performance, resource status, and power consumption of the heterogeneous core, providing a more reasonable basis for the secondary allocation of subsequent tasks, thereby achieving energy conservation and stable operation of the chip while ensuring task processing efficiency.

[0093] In some embodiments, step S300 may include but is not limited to steps S310 to S320:

[0094] Step S310: Obtain the energy efficiency ratio of the heterogeneous cores according to the power consumption of the heterogeneous cores and the performance resources consumed in the long-term operation data.

[0095] Step S320: With the goal of maximizing the energy efficiency ratio of the heterogeneous cores, correct the vector weights of the first resource allocation matrix to obtain the second resource allocation matrix.

[0096] In steps S310 and S320 of some embodiments, the energy efficiency ratio is an important indicator for measuring the relationship between the performance and power consumption of the chip, reflecting the performance output that the heterogeneous cores can provide under a certain power consumption. For example, when a heterogeneous core completes the same computing task, the lower the power consumption, the higher the energy efficiency ratio.

[0097] The main control core corrects the vector weights of the first resource allocation matrix with the goal of maximizing the energy efficiency ratio of the heterogeneous cores. The vector weights determine the importance of different performance indicators of each heterogeneous core during the task allocation process. By adjusting these weights, the system can make tasks more likely to be allocated to heterogeneous cores with high energy efficiency ratios, and avoid over-allocation of tasks to those heterogeneous cores with strong performance but high power consumption. After such correction, the second resource allocation matrix is obtained. The second resource allocation matrix can better balance the performance and power consumption of the chip, so that during the subsequent secondary task allocation, on the premise of meeting the resource requirements for processing tasks, the overall power consumption of the chip can be effectively reduced, and the comprehensive performance and energy utilization efficiency of the chip can be improved.

[0098] In some embodiments, the main control core uses the gradient descent method with constraints to process the operation state data, with the power consumption-performance optimization function as the objective function, so as to change the weight coefficients of the parameters in the resource allocation matrix.

[0099] In step S400 of some embodiments, after obtaining the second resource allocation matrix, the main control core will perform secondary allocation of tasks according to the second resource allocation matrix. Since the second resource allocation matrix takes into account the long-term operation data of the heterogeneous cores, the secondary allocation can more optimally distribute tasks among the heterogeneous cores, improve the overall performance and energy efficiency ratio of the chip, and reduce the overall power consumption of the chip.

[0100] In some embodiments, step S400 may include but is not limited to steps S410 to S430:

[0101] Step S410: Analyze the vector weights of the second resource allocation matrix to obtain priority coding information.

[0102] Step S420: Use the priority coding information to search for and modify the entries in the routing table of the communication system to obtain the target entries.

[0103] Step S430: Determine a new inter-core communication path based on the target table entry, and reallocate the tasks according to the new inter-core communication path.

[0104] In steps S410 and S430 of some embodiments, the priority encoder in the master core analyzes the vector weights in the second resource allocation matrix, re-evaluates the priorities of each heterogeneous core based on the vector weights, and priority encoding information can be obtained after the analysis.

[0105] The routing table of the communication interaction system is used to record the path information of data transmission between various cores and modules inside the chip. Each entry in the routing table specifies the specific path that the data should pass through from the source address to the destination address. The priority encoder will modify the routing table entry according to the priority encoding information. For example, if a heterogeneous core has a higher priority, then the relevant table entry in the routing table will be adjusted so that the data can be transmitted to this core more efficiently, avoiding unnecessary delays and data conflicts, and finally obtaining the target table entry.

[0106] The new communication path can ensure the efficiency of data transmission between cores. Then, the master core will reallocate the tasks according to this new inter-core communication path.

[0107] The reallocation can enable the tasks to more reasonably utilize the chip resources during the transmission and processing, improve the overall performance and task processing efficiency of the multi-core heterogeneous architecture chip, and at the same time help reduce power consumption and improve the power consumption efficiency ratio.

[0108] The embodiments of the present invention have at least the following beneficial effects:

[0109] Improved resource utilization: Based on the dynamic scheduling of the three-dimensional resource allocation matrix, the computing power utilization rate of compute-intensive tasks is increased by more than 40%, and the latency of IO tasks is reduced by 30%;

[0110] Enhanced real-time performance: The multi-level communication protocol supports 10-μs-level interrupt response, and the communication failure recovery time is shortened to within 50 ms through redundant channels;

[0111] Energy efficiency optimization: The dual-loop feedback mechanism controls the system load balancing deviation within 5%, and the overall power consumption is reduced by 15%-20%;

[0112] Compatibility extension: The standardized hardware abstraction layer instruction set can adapt to heterogeneous cores of multiple architectures such as ARM and RISC-V, reducing the secondary development cost.

[0113] In some embodiments, in the security protection of the power monitoring system

[0114] Security task classification and resource allocation:

[0115] The main control core (ARM Cortex-A72) analyzes the power monitoring data stream, identifies the encrypted communication tasks (computation-intensive) and anomaly detection tasks (real-time requirement < 5ms), and allocates them to the NPU core and the real-time coprocessor (RISC-V core) respectively.

[0116] Build a three-dimensional resource matrix: the X-axis divides the security cores (ARM / RISC-V), the Y-axis configures the secure memory area (128MB ECC encrypted storage), and the Z-axis allocates dual-redundant MU communication channels (main channel 32Gbps, backup channel 16Gbps).

[0117] Inter-core secure communication mechanism:

[0118] The encrypted data is transmitted through the TXVring buffer, encrypted in real time using the AES-256 algorithm, and a CRC32 checksum is appended to each frame of data. When the check fails, it switches to the backup channel for retransmission.

[0119] The anomaly detection core (RISC-V) sends an alarm signal to the main control core through register interrupt, and the interrupt response time is controlled within 2ms. The priority preemption mechanism is used to handle emergency events.

[0120] Trusted execution environment construction:

[0121] An independent security domain is divided in the NPU core, and the trusted execution environment (TEE) runs to physically isolate the key management and identity authentication modules and block malicious code injection.

[0122] The main control core sends a heartbeat packet to the security domain every 30 seconds. If there is no response after the timeout, it triggers an inter-core reset, resets the communication channel, and reconstructs the resource allocation matrix.

[0123] Dynamic defense strategy:

[0124] The monitoring module collects the core load and communication delay data in real time. When abnormal traffic is detected (such as the packet rate exceeding the threshold by 50%), the traffic shaping algorithm is automatically enabled to limit the bandwidth of non-critical tasks.

[0125] Adopt a power consumption-security linkage mechanism: when encountering a DDoS attack, the NPU voltage is dynamically increased from 1.0V to 1.2V, and the computing power is increased by 25% to accelerate encryption processing. At the same time, unnecessary peripherals are turned off to reduce the overall power consumption.

[0126] In some embodiments, in the edge-side protection of smart meters

[0127] Lightweight security protocol deployment:

[0128] Run the lightweight encryption protocol (LWC-SHA3) on the Cortex-M4 core to hash the electricity meter measurement data and generate a 128-bit digital fingerprint for storage in the secure Flash area.

[0129] Dynamic key distribution: Synchronize and update the SM4 algorithm key with the master station core (ARM Cortex-A53) through the MU module every 15 minutes, and the key transmission is protected by a physically unclonable function (PUF).

[0130] Multi-level intrusion detection system:

[0131] First-level detection (hardware layer): Monitor the GPIO port level jump through the RISC-V core to identify abnormal physical access (such as short-circuit attack) and trigger the IO port locking mechanism.

[0132] Second-level detection (protocol layer): Deploy a Modbus protocol whitelist in the main control core to filter illegal function code requests, and record illegal operations in the secure log area and encrypt and upload them.

[0133] Anti-side-channel attack design:

[0134] Clock randomization processing: The CPU main frequency dynamically jitters within the range of 48 - 72 MHz to interfere with power consumption analysis attacks, and at the same time, the idle core power supply is turned off through the gated clock module.

[0135] Memory protection: The key data storage uses the XOR scattering write technology to disperse a single data block to 8 physical addresses to block buffer overflow attacks.

[0136] Fault self-healing mechanism:

[0137] Dual-core mutual inspection mechanism: The main control core and the coprocessor exchange check codes every 10 ms. If the check fails continuously for 3 times, the secure startup process is triggered, and the original firmware is loaded from the read-only secure area.

[0138] Communication link self-repair: When the CRC error rate of the MU channel is detected to be > 0.1%, automatically switch to the backup vring channel and reconstruct the Z-axis bandwidth allocation weight to 0.863.

[0139] The embodiment of the present invention also provides a multi-core heterogeneous architecture chip task processing system, including:

[0140] The first module is used to collect the short-term operation data of multiple heterogeneous cores in response to a task allocation instruction, and construct a first resource allocation matrix according to the short-term operation data. The first resource allocation matrix includes vector data and vector weights of multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores;

[0141] A second module, configured to allocate tasks in the task allocation instruction to target heterogeneous cores through a first resource allocation matrix;

[0142] A third module, configured to collect long-term operation data of multiple heterogeneous cores, and update the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes heterogeneous core power consumption;

[0143] A fourth module, configured to perform secondary allocation of tasks according to the second resource allocation matrix.

[0144] It can be understood that the content in the above embodiments of the task processing method for a multi-core heterogeneous architecture chip is applicable to the embodiments of this system. The functions specifically implemented by the embodiments of this system are the same as those of the above embodiments of the collaborative operation training method for an equipment carrier platform, and the beneficial effects achieved are also the same as those of the above embodiments of the task processing method for a multi-core heterogeneous architecture chip.

[0145] Next, in conjunction with Figure 2 the electronic device of the embodiments of the present application will be introduced in detail.

[0146] As Figure 2 , Figure 2 shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0147] A processor 1100, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;

[0148] A memory 1200, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1200 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1200 and are called by the processor 1100 to execute the task processing method for a multi-core heterogeneous architecture chip of the present disclosure;

[0149] An input / output interface 1300, configured to implement information input and output;

[0150] A communication interface 1400 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0151] A bus 1500 transmits information between various components of the device (such as a processor 1100, a memory 1200, an input / output interface 1300, and a communication interface 1400);

[0152] Among them, the processor 1100, the memory 1200, the input / output interface 1300, and the communication interface 1400 achieve communication connections with each other inside the device through the bus 1500.

[0153] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned multi-core heterogeneous architecture chip task processing method.

[0154] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0155] The embodiments described in the embodiments of the present disclosure are to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0156] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.

[0157] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0159] As used in the description of the present application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0160] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0161] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0162] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] In addition, in each embodiment of this application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0164] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0165] The preferred embodiments of the embodiments of the present disclosure have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present disclosure. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present disclosure shall be within the scope of the rights of the embodiments of the present disclosure.

Claims

1. A method for task processing of a multi-core heterogeneous architecture chip, characterized in that Including the following steps: In response to a task assignment instruction, collect short-term operation data of multiple heterogeneous cores, and construct a first resource allocation matrix based on the short-term operation data. The first resource allocation matrix includes vector data and vector weights of multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores; Through the first resource allocation matrix, assign the task in the task assignment instruction to a target heterogeneous core; Collect long-term operation data of multiple heterogeneous cores, and update the first resource allocation matrix based on the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes heterogeneous core power consumption; Perform secondary allocation of the task according to the second resource allocation matrix.

2. The task processing method for a multi-core heterogeneous architecture chip according to claim 1, wherein The short-term operation data and the long-term operation data of the heterogeneous cores are obtained through the following steps: Collect the short-term operation data through a first monitoring module of the heterogeneous core; Collect the long-term operation data through a second monitoring module of the heterogeneous core.

3. The method for processing tasks of a multi-core heterogeneous architecture chip according to claim 1, wherein, The short-term operation data includes core computing data, remaining storage unit capacity, and communication channel delay data, and the vector data includes a first dimension, a second dimension, and a third dimension; The first dimension is used to characterize the core computing data; The second dimension is used to characterize the remaining storage unit capacity; The third dimension is used to characterize the communication channel delay data.

4. The method for processing tasks of a multi-core heterogeneous architecture chip according to claim 3, wherein The vector weights include weight coefficients of multiple dimensions. Assigning the task in the task assignment instruction to a target heterogeneous core through the first resource allocation matrix includes the following steps: Parse the task instruction to obtain an identifier of the task type and corresponding resource requirement parameters, where the resource requirement parameters include a storage capacity threshold and a data transmission delay upper limit; According to the identifier, determine the weight coefficients of the first dimension of each heterogeneous core; Determine a target heterogeneous core according to the weight coefficients of the first dimension of each heterogeneous core; According to the storage capacity threshold, determine the weight coefficients of the second dimension of each heterogeneous core; According to the data transmission delay upper limit, determine the weight coefficients of the third dimension of each heterogeneous core; According to the weight coefficients of the second dimension and the third dimension of each heterogeneous core, determine the performance requirement data for the target heterogeneous core to process the task; Generate a first allocation instruction according to the performance requirement data and the task in the task assignment instruction, and send the first allocation instruction to the target heterogeneous core.

5. The method for processing tasks of a multi-core heterogeneous architecture chip according to claim 4, wherein Sending the first allocation instruction to the target heterogeneous core includes the following steps: Use an instruction adaptation conversion interface to convert the first allocation instruction into a second allocation instruction; Send the second allocation instruction to the target heterogeneous core.

6. The method for processing tasks of a multi-core heterogeneous architecture chip according to claim 1, wherein Updating the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix includes the following steps: According to the heterogeneous core power consumption and the consumed performance resources in the long-term operation data, obtain the energy efficiency ratio of the heterogeneous core; Taking maximizing the energy efficiency ratio of the heterogeneous core as the goal, correct the vector weights of the first resource allocation matrix to obtain a second resource allocation matrix.

7. The task processing method for a multi-core heterogeneous architecture chip according to claim 6, wherein Performing secondary allocation of the tasks according to the second resource allocation matrix includes the following steps: Analyzing the vector weights of the second resource allocation matrix to obtain priority coding information; Using the priority coding information to search for and modify the entries of the routing table of the communication system to obtain target entries; Determining a new inter-core communication path according to the target entries, and performing secondary allocation of the tasks according to the new inter-core communication path.

8. A task processing system for a multi-core heterogeneous architecture chip, characterized in that, Including: A first module, configured to collect short-term operation data of multiple heterogeneous cores in response to a task allocation instruction, and construct a first resource allocation matrix according to the short-term operation data, where the first resource allocation matrix includes vector data and vector weights of the multiple heterogeneous cores, and the vector data is used to characterize the performance resources of the heterogeneous cores; A second module, configured to allocate the tasks in the task allocation instruction to target heterogeneous cores through the first resource allocation matrix; A third module, configured to collect long-term operation data of the multiple heterogeneous cores, and update the first resource allocation matrix according to the long-term operation data to obtain a second resource allocation matrix, where the long-term operation data includes heterogeneous core power consumption; A fourth module, configured to perform secondary allocation of the tasks according to the second resource allocation matrix.

9. An electronic device, characterized in that, Including: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, enabling the at least one processor to implement the multi-core heterogeneous architecture chip task processing method according to any one of claims 1-7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to implement the multi-core heterogeneous architecture chip task processing method according to any one of claims 1-7.

Citation Information

Cited By

  • Task allocation method and device, equipment and medium

    CN121210150A

  • BMC (Baseboard Management Controller) coprocessing method and device, electronic equipment and storage medium

    CN121807654A