Inter-core communication method, system and device based on heterogeneous asymmetric processor and medium

Through task feature extraction and dynamic resource management, the inter-core communication of heterogeneous asymmetric processors is optimized, which solves the problems of low task scheduling efficiency, high communication redundancy and insufficient security isolation, improves the system's computing resource utilization and energy efficiency ratio, and enhances the security of cross-core data transmission.

CN120277022APending Publication Date: 2025-07-08GUANGZHOU KETENG INFORMATION TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510350205.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing heterogeneous multi-core processors have low task scheduling efficiency, high inter-core communication redundancy, insufficient security isolation and coarse-grained energy efficiency control, resulting in insufficient computing power utilization, high communication delay, poor security and large energy efficiency fluctuations, which cannot meet the needs of industrial control scenarios.

Method used

By extracting and classifying task features of heterogeneous asymmetric processors, dynamically adjusting voltage frequency, dividing trusted execution environments, optimizing resource allocation using dynamic session keys and three-dimensional resource matrix, and synchronizing cross-core data using a dual-channel redundant transmission protocol.

Benefits of technology

The computing resource utilization, security and energy efficiency ratio of heterogeneous asymmetric processors is improved, efficient mapping and stable migration of tasks on different computing units are realized, and the security of cross-core data transmission and system resource scheduling capabilities are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277022A_ABST
    Figure CN120277022A_ABST
Patent Text Reader

Abstract

The invention discloses an inter-core communication method, system and device based on a heterogeneous asymmetric processor and a medium, and the method comprises the steps: carrying out the classification and feature extraction of a target task, constructing a task feature vector, and mapping the target task to a processor core; the load is detected in real time, when the real-time load exceeds a threshold value, a task migration strategy is executed through atomic operation and a breakpoint management mechanism, and the voltage and frequency of a processor core are adjusted according to a target task; a trusted execution environment is divided in a processor core, a dynamic session key is generated, the dynamic session key is transmitted through a physical isolation bus, cross-core communication data is encrypted, resources are allocated according to a three-dimensional dynamic resource matrix, and cross-core data synchronization is achieved through a two-channel redundancy transmission protocol. According to the method, the software collaboration framework of the heterogeneous asymmetric processor is optimally designed, so that the performance of the heterogeneous asymmetric processor is improved, and the method can be widely applied to the technical field of inter-core communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of inter-core communication, and in particular, to an inter-core communication method, system, device and medium based on heterogeneous asymmetric processors. Background Art

[0002] The existing software cooperation framework of heterogeneous multi-core processors has the following technical defects:

[0003] Low task scheduling efficiency: The traditional static task allocation strategy cannot adapt to the dynamic load differences of asymmetric processors (such as CPU / GPU / NPU), resulting in a computing power utilization rate of less than 60%, and the cross-core migration latency is as high as more than 50 μs; High inter-core communication redundancy: The shared memory mechanism does not distinguish between real-time / non-real-time data, and frequent data copying and lock competition result in a communication bandwidth waste of more than 30%, and the fault tolerance recovery time exceeds 5 ms; Insufficient security isolation: Key management and sensitive data are not physically isolated from heterogeneous cores, and side-channel attacks can steal critical instructions, making it difficult for the system security to meet the requirements of industrial control scenarios; Coarse-grained energy efficiency control: Voltage and frequency regulation rely on global policies and do not dynamically optimize in combination with task types, and the energy efficiency ratio (TOPS / W) fluctuates within a range of more than ±20%. Summary of the Invention

[0004] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0005] To this end, an object of an embodiment of the present invention is to provide an inter-core communication method based on heterogeneous asymmetric processors, which improves the data utilization rate, security and energy efficiency ratio of heterogeneous asymmetric processors through an optimized design of the software cooperation framework of heterogeneous asymmetric processors.

[0006] Another object of an embodiment of the present invention is to provide an inter-core communication system based on heterogeneous asymmetric processors.

[0007] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:

[0008] In the first aspect, an embodiment of the present invention provides an inter-core communication method based on heterogeneous asymmetric processors, including:

[0009] Classify and extract features of a target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector;

[0010] Detect the real-time load of the processor core. When the real-time load exceeds a threshold, execute a task migration strategy through an atomic operation and a breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task;

[0011] Partition a trusted execution environment in the processor core, generate a dynamic session key through a physically unclonable function, and transmit the dynamic session key between the processor cores through a physically isolated bus. The dynamic session key is used to encrypt cross-core communication data;

[0012] Allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighted algorithm, and achieve cross-core data synchronization through a dual-channel redundant transmission protocol.

[0013] Further, the target tasks include compute-intensive tasks, control-intensive tasks, and I / O-intensive tasks. The classification and feature extraction of the target tasks, constructing a task feature vector, and mapping the target tasks to the processor cores according to the task feature vector include:

[0014] Judge the type of the target task;

[0015] When the target task is the compute-intensive task, extract the floating-point operation volume and cache hit rate of the target task, and generate the task feature label through a lightweight hash function according to the floating-point operation volume and cache hit rate;

[0016] When the target task is the control-intensive task, extract the interrupt response time and context switching frequency of the target task, and generate the task feature label through the lightweight hash function according to the interrupt response time and context switching frequency;

[0017] When the target task is the I / O-intensive task, extract the DMA channel occupancy rate and throughput of the target task, and generate the task feature label through the lightweight hash function according to the DMA channel occupancy rate and throughput;

[0018] Map the target task to the corresponding processor core according to the task feature vector through a clustering algorithm.

[0019] Further, the three-dimensional dynamic resource matrix is updated through the following steps:

[0020] Collect the load data of the processor core according to a preset period, predict the future resource requirements according to the load data, and then update the three-dimensional dynamic resource matrix;

[0021] Detect the memory bandwidth utilization rate. When the memory bandwidth utilization rate exceeds the threshold, expand the virtual memory pool, and then update the three-dimensional dynamic resource matrix.

[0022] Further, the dual-channel includes a main channel and a secondary channel, the cross-core data includes cross-core batch data and cross-core control instructions, and the implementation of cross-core data synchronization through the dual-channel redundant transmission protocol includes:

[0023] Transmit the batch cross - core data through the main channel using the Remote Direct Memory Access protocol;

[0024] Transmit the cross - core control instructions through the secondary channel using an interrupt - driven mechanism, and set the priority of the secondary channel higher than that of the main channel;

[0025] Verify and re - transmit the cross - core data through a cyclic redundancy check mechanism.

[0026] Furthermore, execute the task migration strategy through the atomic operation and breakpoint management mechanism, including:

[0027] Use atomic operations to lock the memory addresses of the source core and the target core;

[0028] Generate snapshot data of the current task and store it in the shared memory;

[0029] Store the breakpoint status data of the current task through a differential compression algorithm;

[0030] Manage the multi - core concurrent access to the breakpoint status data through an atomic counter;

[0031] Re - construct the current task on the target core according to the snapshot data and the breakpoint status data.

[0032] Furthermore, adjust the voltage and frequency of the processor core according to the target task, including:

[0033] Judge the type of the target task;

[0034] When the target task is a compute - intensive task, adjust the voltage and frequency of the processor core through a dynamic voltage and frequency regulation unit;

[0035] When the target task is an I / O - intensive task, adjust the voltage and frequency of the processor core through a gated - clock technique.

[0036] Furthermore, transmit the dynamic session key between the processor cores through a physically isolated bus, including:

[0037] Set an access control mechanism in the trusted execution environment to detect unauthorized access of the dynamic session key by the processor core. When the unauthorized access is detected, trigger a processor reset signal;

[0038] Set a secure isolation area in the storage unit of the processor core, and store the dynamic session key through the secure isolation area.

[0039] Second aspect, an embodiment of the present invention provides an inter-core communication system based on heterogeneous asymmetric processors, including:

[0040] A task scheduling module, configured to classify and extract features of a target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector;

[0041] A task migration module, configured to detect the real-time load of the processor core. When the real-time load exceeds a threshold, execute a task migration strategy through an atomic operation and a breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task;

[0042] A communication security module, configured to divide a trusted execution environment in the processor core, generate a dynamic session key through a physically unclonable function, and transmit the dynamic session key between the processor cores through a physically isolated bus. The dynamic session key is used to encrypt cross-core communication data;

[0043] A resource allocation module, configured to allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighting algorithm, and implement cross-core data synchronization through a dual-channel redundant transmission protocol.

[0044] Third aspect, an embodiment of the present invention provides a device, including:

[0045] At least one processor;

[0046] At least one memory, configured to store at least one program;

[0047] When the at least one program is executed by the at least one processor, the at least one processor implements an inter-core communication method based on heterogeneous asymmetric processors as described above.

[0048] Fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a processor-executable program is stored. The processor-executable program, when executed by a processor, is used to execute an inter-core communication method based on heterogeneous asymmetric processors as described above.

[0049] The advantages and beneficial effects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention:

[0050] In the embodiments of the present invention, through a dynamic task scheduling strategy based on task feature extraction, efficient mapping of tasks on different computing units is achieved, improving the utilization rate of computing resources; through a three-dimensional dynamic resource matrix and an asymmetric weighting algorithm, computing, memory, and communication resources are dynamically adjusted, improving the resource scheduling ability and resource utilization rate of the system; through a communication mechanism based on a trusted execution environment and dynamic keys, the security of cross-core data transmission is improved; through an atomic operation and breakpoint management mechanism, the stability of task migration is ensured; by dynamically adjusting the voltage and frequency of the processor core according to the task type, the energy efficiency ratio of the system is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 FIG. is a schematic diagram of the steps of an inter-core communication method based on a heterogeneous asymmetric processor provided by an embodiment of the present invention;

[0052] Figure 2 FIG. is a schematic diagram of an inter-core communication system based on a heterogeneous asymmetric processor provided by an embodiment of the present invention;

[0053] Figure 3 FIG. is a schematic diagram of the structure of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0055] In the description of the present invention, the meaning of "a plurality of" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art of this technology.

[0056] The English abbreviations used in the present invention include:

[0057] FPGA: Field Programmable Gate Array, field programmable gate array;

[0058] GPU: Graphics Processing Unit, the graphics processing unit;

[0059] CPU: Central Processing Unit, the central processing unit.

[0060] Figure 1 It is a schematic diagram of the steps of an inter-core communication method based on a heterogeneous asymmetric processor provided by an embodiment of the present invention. Refer to Figure 1 , an embodiment of the present invention provides an inter-core communication method based on a heterogeneous asymmetric processor, including:

[0061] S101. Classify and extract features of the target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector;

[0062] Specifically, by dividing the type of the target task and extracting key features to construct a task feature vector, and then allocating the task to a suitable processor core for processing, thereby improving the computing efficiency and resource utilization rate of the system.

[0063] In some alternative embodiments, the target tasks include compute-intensive tasks, control-intensive tasks, and IO-intensive tasks. Classifying and extracting features of the target task, constructing a task feature vector, and mapping the target task to a processor core according to the task feature vector includes:

[0064] A1. Determine the type of the target task;

[0065] A2. When the target task is a compute-intensive task, extract the floating-point operation volume and cache hit rate of the target task, and generate a task feature label through a lightweight hash function according to the floating-point operation volume and cache hit rate;

[0066] A3. When the target task is a control-intensive task, extract the interrupt response time and context switching frequency of the target task, and generate a task feature label through a lightweight hash function according to the interrupt response time and context switching frequency;

[0067] A4. When the target task is an IO-intensive task, extract the DMA channel occupancy rate and throughput of the target task, and generate a task feature label through a lightweight hash function according to the DMA channel occupancy rate and throughput;

[0068] A5. Map the target task to the corresponding processor core according to the task feature vector through a clustering algorithm.

[0069] Specifically, its type can be classified and judged according to information such as the execution instructions and call frequencies of the target tasks. For compute-intensive tasks, their floating-point operation amounts and cache hit rates are used as key features. The former characterizes the computing scale of the task, and the latter characterizes the stability of the task's computing process; for control-intensive tasks, their interrupt response times and context switching frequencies are used as key features. The former characterizes the response speed of the task to external interrupt signals, and the latter can reflect the scheduling frequency of the task on the processor. Frequent context switching will affect the execution efficiency of real-time tasks; for I / O-intensive tasks, their DMA (Direct Memory Access) channel occupancy rates and throughputs are used as key features. The former reflects the occupancy of the task on the channel, and the latter determines the bandwidth requirements of the task; in this implementation, after these characteristic parameters are extracted, they will be input into a lightweight hash function to generate a unique task characteristic label, and these tasks will be mapped to the most suitable processing core through the K-means clustering algorithm, so as to maximize the performance of the system and reduce the waste of computing resources.

[0070] S102. Detect the real-time load of the processor core. When the real-time load exceeds the threshold, execute the task migration strategy through the atomic operation and breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task.

[0071] Specifically, in this embodiment, when the real-time load of the processor core exceeds the threshold, the atomic operation and breakpoint management mechanism will be used to safely migrate the task from the source core to the target core; in addition, the voltage and frequency of the processor core will be dynamically adjusted according to different types of tasks to improve the energy consumption ratio of the system.

[0072] In some alternative embodiments, executing the task migration strategy through the atomic operation and breakpoint management mechanism includes:

[0073] B1. Use atomic operations to lock the memory addresses of the source core and the target core.

[0074] B2. Generate snapshot data of the current task and store it in the shared memory.

[0075] B3. Store the breakpoint status data of the current task through the differential compression algorithm.

[0076] B4. Manage the multi-core concurrent access to the breakpoint status data through an atomic counter.

[0077] B5. Reconstruct the current task on the target core according to the snapshot data and the breakpoint status data.

[0078] Specifically, an atomic operation is a mechanism to prevent race conditions and is used to ensure data consistency during migration. A task snapshot is used to record the execution status, register values, memory data, etc. of a task to ensure that the task can be resumed on the target core. Shared memory is an area that can be accessed by multiple cores. Since storing the complete task snapshot may consume a large amount of storage resources, in this embodiment, sampling differential compression technology is used to only store the critical state change part, thereby reducing the occupied storage space and accelerating the migration speed. Since breakpoint state data needs to be shared among multiple cores, to prevent conflicts caused by multiple cores modifying data simultaneously, an atomic counter is used to manage access permissions to ensure data consistency. Finally, the target core reads the task snapshot and breakpoint state data to restore the execution state of the task and ensure that the task can run seamlessly.

[0079] In some alternative embodiments, the voltage and frequency of the processor core are adjusted according to the target task, including:

[0080] C1. Determine the type of the target task;

[0081] C2. When the target task is a compute-intensive task, adjust the voltage and frequency of the processor core through a dynamic voltage and frequency regulation unit;

[0082] C3. When the target task is an I / O-intensive task, adjust the voltage and frequency of the processor core through a gated clock technique.

[0083] Specifically, this embodiment can adjust the voltage and frequency of the processor core according to the task type to optimize the computing efficiency and improve the energy efficiency ratio. For compute-intensive tasks, dynamic voltage and frequency regulation is used to ensure that its frequency and voltage are within an appropriate range to balance the computing efficiency and power consumption. For I / O-intensive tasks, which usually involve a large amount of data transmission rather than high-intensity computing, the gated clock technique can be used to reduce its voltage and frequency and turn off some clock signals when the task is idle, thereby improving the energy efficiency ratio.

[0084] S103. Divide a trusted execution environment in the processor core, generate a dynamic session key through a physically unclonable function, and transmit the dynamic session key between processor cores through a physically isolated bus. The dynamic session key is used to encrypt cross-core communication data.

[0085] Specifically, the trusted execution environment is a secure computing area inside the processor core isolated from ordinary applications, used to store and manage keys to prevent unauthorized access. The physically unclonable function uses hardware characteristics to generate a unique and non-replicable key, which has high security. The physically isolated bus ensures that the key data is transmitted between different processor cores without passing through the shared bus, further guaranteeing the security of the key.

[0086] In some alternative embodiments, a dynamic session key is transmitted between processor cores through a physically isolated bus, including:

[0087] D1. Set an access control mechanism in the trusted execution environment to detect unauthorized access by a processor core to the dynamic session key. When unauthorized access is detected, trigger a processor reset signal;

[0088] D2. Set a secure isolation area in the storage unit of the processor core, and store the dynamic session key through the secure isolation area.

[0089] Specifically, the access control mechanism of the trusted execution environment monitors which processor cores attempt to access the key and identifies whether it is an authorized access. If an unauthorized processor core is detected attempting to access the key, the system will immediately trigger a processor reset signal to prevent potential attacks. At the same time, the memory protection mechanism of the storage unit is specifically used to store the key, which can further prevent unauthorized access or tampering.

[0090] S104. Allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighting algorithm, and implement cross-core data synchronization through a dual-channel redundant transmission protocol;

[0091] Specifically, in this embodiment, the three-dimensional dynamic resource matrix is used to comprehensively manage the states of computing resources, memory resources, and communication resources. When allocating resources, different weights are assigned to different task types to achieve more efficient resource scheduling. At the same time, cross-core data synchronization is performed through a dual-channel redundant transmission protocol to improve system efficiency and the stability of task execution.

[0092] In some alternative embodiments, the three-dimensional dynamic resource matrix is updated through the following steps:

[0093] E1. Collect the load data of the processor core according to a preset period, predict the future resource requirements based on the load data, and then update the three-dimensional dynamic resource matrix;

[0094] E2. Detect the memory bandwidth utilization rate. When the memory bandwidth utilization rate exceeds the threshold, expand the virtual memory pool, and then update the three-dimensional dynamic resource matrix.

[0095] Specifically, the future resource requirements can be predicted through algorithms such as the sliding window algorithm or regression analysis based on the load data, and then the resource weights for task allocation are adjusted according to the prediction results to update the three-dimensional dynamic resource matrix;

[0096] When the memory bandwidth utilization is too high, the system will expand the virtual memory pool to avoid memory resource bottlenecks, allocate more storage resources to reduce memory pressure, and at the same time adjust the task scheduling strategy to reduce the occupancy of the bus by high-bandwidth tasks. These operations will trigger the update of the three-dimensional dynamic resource matrix to ensure that the system's resource scheduling strategy always runs in an optimal state.

[0097] In some alternative embodiments, the dual-channel includes a main channel and a secondary channel, and the cross-core data includes cross-core batch data and cross-core control instructions. Cross-core data synchronization is achieved through a dual-channel redundancy transmission protocol, including:

[0098] F1. Transmit the batch cross-core data through the main channel using the Remote Direct Memory Access protocol;

[0099] F2. Transmit the cross-core control instructions through the secondary channel using an interrupt-driven mechanism, and set the priority of the secondary channel higher than that of the main channel;

[0100] F3. Check and retransmit the cross-core data through a cyclic redundancy check mechanism.

[0101] Specifically, Remote Direct Memory Access (RDMA) is an efficient data transmission protocol that allows data to be directly transmitted between the memories of different computing cores without going through CPU processing. Transmitting batch cross-core data through this protocol can effectively reduce transmission latency and improve data throughput. The main role of the secondary channel is to transmit key control signals such as task status synchronization and error notification, and use an interrupt-driven mechanism to ensure that control instructions can be transmitted immediately when critical tasks need to communicate. At the same time, the high priority of the secondary channel can ensure that control instructions will not be blocked by large-scale data transmission. The cyclic redundancy check mechanism can guarantee the integrity of cross-core data and ensure the accurate transmission of data.

[0102] The following describes the inter-core communication method of the present invention in conjunction with a specific embodiment.

[0103] In this embodiment, for the security protection application scenario of the power grid monitoring system, methods such as heterogeneous task scheduling, data security protection, and energy efficiency optimization are used to improve the security and computing efficiency of the system. The main control processor analyzes the monitoring data stream and classifies tasks using the K-means clustering algorithm, scheduling computationally intensive encrypted communication tasks and anomaly detection tasks with high real-time requirements to appropriate computing units respectively. By constructing a three-dimensional dynamic resource matrix (such as computing power weight 0.4, memory occupancy rate 85%, communication bandwidth 32 Gbps) to manage computing, memory, and communication resources, and using an asymmetric weighted algorithm to allocate bandwidth, ensuring that the main channel transmits encrypted data and the secondary channel transmits control signals. In terms of data security, the system uses AES-256 encryption and appends a 32-bit CRC (Cyclic Redundancy Check) check code to each frame. If the check fails, the backup channel is enabled for retransmission. At the same time, a trusted execution environment is partitioned in the CPU core, a key management module is deployed, interacting with the NPU core through a physically isolated bus, and a read-only isolation zone is set in the video memory area to prevent unauthorized access and block side-channel attacks. In addition, to reduce the impact of abnormal traffic, the system restricts the bandwidth occupancy of non-critical tasks through a traffic control mechanism and dynamically adjusts the voltage of the computing unit to optimize energy efficiency. The main control processor periodically sends encrypted heartbeat packets to the trusted execution environment. If no response is received after a timeout, the processor is reset and tasks are restored from the encrypted snapshot in the shared memory to ensure the stable operation of the power grid monitoring system.

[0104] The inter-core communication method of the present invention will be described below in combination with another specific embodiment.

[0105] In this embodiment, for the edge security protection application scenario of smart meters, lightweight security protocols, intrusion detection mechanisms, and anti-attack optimization strategies are adopted to improve the security and self-healing ability of the device. The LWC-SHA3 lightweight protocol is deployed in the processor core to generate a 128-bit hash fingerprint, and the key is dynamically updated every 15 minutes through PUF (Physically Unclonable Function) to reduce the risk of key leakage. At the same time, the scattering writing technique is used to disperse and store critical data in multiple physical addresses, combined with the dynamic expansion mechanism of the memory controller to prevent buffer overflow attacks. In terms of intrusion detection, at the hardware layer, the RISC-V core monitors the GPIO (General-Purpose Input / Output) level changes, and immediately locks the port when a short-circuit attack is detected; at the protocol layer, the Modbus protocol whitelist is deployed through the main control core to intercept illegal function code requests, and the violation logs are encrypted and uploaded to the cloud through the edge-cloud collaboration architecture. In terms of anti-attack and self-healing ability, the system uses dynamic clock jitter technology to interfere with power analysis attacks, and the gated clock technology is used to turn off idle processor cores to reduce energy consumption. In addition, the main control core and the coprocessor regularly exchange check codes through the dual-core mutual inspection mechanism. If the failures are consecutive, the firmware is loaded from the read-only security area to ensure the system quickly recovers and maintains stable operation.

[0106] It can be recognized that in the embodiment of the present invention, through the dynamic task scheduling strategy based on task feature extraction, the efficient mapping of tasks on different computing units is realized, and the utilization rate of computing resources is improved; through the three-dimensional dynamic resource matrix and the asymmetric weighting algorithm, the computing, memory, and communication resources are dynamically adjusted, and the resource scheduling ability and resource utilization rate of the system are improved; through the communication mechanism based on the trusted execution environment and dynamic keys, the security of cross-core data transmission is improved; the stability of task migration is ensured through the atomic operation and breakpoint management mechanism; the voltage and frequency of the processor core are dynamically adjusted according to the task type, effectively improving the energy efficiency ratio of the system.

[0107] Referring to Figure 2 , the embodiment of the present invention provides an inter-core communication system based on a heterogeneous asymmetric processor, including:

[0108] A task scheduling module, configured to classify and extract features of a target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector;

[0109] A task migration module, configured to detect the real-time load of the processor core. When the real-time load exceeds a threshold, execute a task migration strategy through an atomic operation and a breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task;

[0110] A communication security module is used to divide a trusted execution environment in a processor core, generate a dynamic session key through a physically unclonable function, transmit the dynamic session key between processor cores through a physically isolated bus, and the dynamic session key is used to encrypt cross-core communication data;

[0111] A resource allocation module is used to allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighting algorithm, and achieve cross-core data synchronization through a dual-channel redundant transmission protocol.

[0112] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0113] Referring to Figure 3 , an embodiment of the present invention provides a device, including:

[0114] At least one processor;

[0115] At least one memory for storing at least one program;

[0116] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above-mentioned inter-core communication method based on heterogeneous asymmetric processors.

[0117] An embodiment of the present invention also provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above-mentioned inter-core communication method based on heterogeneous asymmetric processors when executed by the processor.

[0118] A computer-readable storage medium of an embodiment of the present invention can execute an inter-core communication method based on heterogeneous asymmetric processors provided by the method embodiments of the present invention, can execute any combination of implementation steps of the method embodiments, and has the corresponding functions and beneficial effects of the method.

[0119] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of the device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the device executes Figure 1 the inter-core communication method based on a multi-core heterogeneous architecture shown.

[0120] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0121] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0122] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods of the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0123] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.

[0124] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above-mentioned program can be printed, because the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0125] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0126] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0127] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0128] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. An inter-core communication method based on heterogeneous asymmetric processors, characterized in that, Including: Classify and extract features of the target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector; Detect the real-time load of the processor core. When the real-time load exceeds the threshold, execute a task migration policy through an atomic operation and a breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task; Divide a trusted execution environment in the processor core, generate a dynamic session key through a physically unclonable function, and transmit the dynamic session key between the processor cores through a physically isolated bus. The dynamic session key is used to encrypt cross-core communication data; Allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighted algorithm, and achieve cross-core data synchronization through a dual-channel redundant transmission protocol.

2. The inter-core communication method based on a heterogeneous asymmetric processor according to claim 1, wherein The target task includes a compute-intensive task, a control-intensive task, and an I / O-intensive task. The classifying and extracting features of the target task, constructing a task feature vector, and mapping the target task to a processor core according to the task feature vector include: Judge the type of the target task; When the target task is the compute-intensive task, extract the floating-point operation volume and cache hit rate of the target task, and generate the task feature label through a lightweight hash function according to the floating-point operation volume and cache hit rate; When the target task is the control-intensive task, extract the interrupt response time and context switching frequency of the target task, and generate the task feature label through the lightweight hash function according to the interrupt response time and context switching frequency; When the target task is the I / O-intensive task, extract the DMA channel occupancy rate and throughput of the target task, and generate the task feature label through the lightweight hash function according to the DMA channel occupancy rate and throughput; Map the target task to the corresponding processor core through a clustering algorithm according to the task feature vector.

3. The inter-core communication method based on a heterogeneous asymmetric processor according to claim 1, wherein The three-dimensional dynamic resource matrix is updated through the following steps: Collect the load data of the processor core according to a preset period, predict the future resource requirements according to the load data, and then update the three-dimensional dynamic resource matrix; Detect the memory bandwidth utilization rate. When the memory bandwidth utilization rate exceeds the threshold, expand the virtual memory pool, and then update the three-dimensional dynamic resource matrix.

4. An inter-core communication method based on a heterogeneous asymmetric processor according to claim 1, characterized in that The dual-channel includes a main channel and a secondary channel. The cross-core data includes cross-core batch data and cross-core control instructions. The realizing cross-core data synchronization through the dual-channel redundant transmission protocol includes: Transmit the batch cross-core data through the main channel using the remote direct memory access protocol; Transmit the cross-core control instructions through the secondary channel using an interrupt-driven mechanism, and set the priority of the secondary channel higher than that of the main channel; Verify and retransmit the cross-core data through a cyclic redundancy check mechanism.

5. The inter-core communication method based on a heterogeneous asymmetric processor according to claim 1, wherein The executing the task migration policy through the atomic operation and the breakpoint management mechanism includes: Lock the memory addresses of the source core and the target core using an atomic operation; Generate snapshot data of the current task and store it in the shared memory; Store the breakpoint status data of the current task through a differential compression algorithm; Manage the multi-core concurrent access to the breakpoint status data through an atomic counter; Reconstruct the current task on the target core according to the snapshot data and the breakpoint status data.

6. The inter-core communication method based on a heterogeneous asymmetric processor according to claim 2, characterized in that The adjusting the voltage and frequency of the processor core according to the target task includes: Judge the type of the target task; When the target task is a compute-intensive task, adjust the voltage and frequency of the processor core through a dynamic voltage and frequency regulation unit; When the target task is an I / O-intensive task, adjust the voltage and frequency of the processor core through a gated clock technique.

7. A method for inter-core communication based on a heterogeneous asymmetric processor according to claim 1, characterized in that The transmitting the dynamic session key between the processor cores through a physically isolated bus includes: Set an access control mechanism in the trusted execution environment to detect unauthorized access by a processor core to the dynamic session key, and when the unauthorized access is detected, trigger a processor reset signal; Set a secure isolation area in the storage unit of the processor core, and store the dynamic session key through the secure isolation area.

8. An inter-core communication system based on heterogeneous asymmetric processors, characterized in that, Comprising: A task scheduling module, configured to classify and extract features of a target task, construct a task feature vector, and map the target task to a processor core according to the task feature vector; A task migration module, configured to detect the real-time load of the processor core, and when the real-time load exceeds a threshold, execute a task migration strategy through an atomic operation and a breakpoint management mechanism, and adjust the voltage and frequency of the processor core according to the target task; A communication security module, configured to divide a trusted execution environment in the processor core, generate a dynamic session key through a physically unclonable function, transmit the dynamic session key between the processor cores through a physically isolated bus, and the dynamic session key is used to encrypt cross-core communication data; A resource allocation module, configured to allocate computing resources, memory resources, and communication resources according to a three-dimensional dynamic resource matrix and an asymmetric weighting algorithm, and implement cross-core data synchronization through a dual-channel redundant transmission protocol.

9. A device, characterized in that, Comprising: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements an inter-core communication method based on a heterogeneous asymmetric processor according to any one of claims 1-7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute an inter-core communication method based on a heterogeneous asymmetric processor according to any one of claims 1-7.

Citation Information

Cited By

  • Malicious access information detection method and device

    CN120880779A

  • Signal processor access method and computer equipment

    CN120892095A

  • Cloud edge collaborative adaptation method for supervision multi-heterogeneous system

    CN121098887A

  • Reliable reasoning execution method and device for NPU heterogeneous scene

    CN121333571A