Chip design method, system and device based on multi-core heterogeneous architecture and medium

Through dynamic configuration of heterogeneous computing core and hybrid topological structure design, combined with power management and post-silicon verification, the problems of low resource utilization, poor real-time performance and energy efficiency imbalance in traditional multi-core heterogeneous chip designs are solved, and efficient computing resource allocation and energy efficiency optimization are achieved.

CN120278090APending Publication Date: 2025-07-08GUANGZHOU KETENG INFORMATION TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510350214.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In traditional multi-core heterogeneous chip design, there are problems such as static configuration solutions that are difficult to adapt to dynamic computing load changes, low task scheduling efficiency, rigid on-chip interconnect architecture, extensive energy efficiency optimization, and lag in verification and correction, resulting in low resource utilization, poor real-time performance and imbalance in energy efficiency.

Method used

By dynamically configuring heterogeneous computing core types and quantities, designing on-chip interconnect structures using a hybrid topology, configuring power management units to achieve dynamic task scheduling and energy efficiency optimization, and firmware updates are performed in the post-silicon verification stage to support continuous optimization after chip deployment.

Benefits of technology

It realizes efficient allocation of computing resources, low latency and high bandwidth utilization of computing core communications, improves energy efficiency ratio, and improves the long-term reliability and security of the chip through continuous optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278090A_ABST
    Figure CN120278090A_ABST
Patent Text Reader

Abstract

The invention discloses a chip design method, system and device based on a multi-core heterogeneous architecture and a medium. The method comprises the steps that the type and number of heterogeneous computing cores are determined according to a target application scene, and a task scheduling unit is configured; designing an on-chip interconnection structure of a heterogeneous computing core by adopting a hybrid topological structure, and verifying and adjusting a dynamic task scheduling strategy and the on-chip interconnection structure through hardware simulation; configuring a power management unit according to the energy efficiency optimization target, wherein the power management unit is used for controlling the working state of the heterogeneous computing core; operating data of the heterogeneous computing core are collected in the post-silicon verification stage, a firmware updating scheme is generated according to the operating data, and firmware updating is conducted according to the firmware updating scheme. By optimizing the chip design process under the multi-core heterogeneous architecture, the performance, the energy efficiency ratio and the reliability of the chip are improved, and the method can be widely applied to the technical field of chip design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and particularly to a chip design method, system, device and medium based on a multi-core heterogeneous architecture. Background Art

[0002] With the rapid development of technologies such as artificial intelligence and edge computing, the traditional homogeneous multi-core chip architecture has been difficult to meet the differentiated requirements for computing efficiency, real-time performance and energy efficiency in diverse scenarios. For example, in scenarios such as autonomous driving, intelligent security, and industrial control, computing tasks include both general computing tasks with high parallelism (such as data processing, logical operations, etc.) and specially optimized computing tasks (such as image processing, AI inference, 3D modeling, cryptographic computing, etc.). In order to complete different types of computing tasks more efficiently, the current industry widely adopts a multi-core heterogeneous architecture that integrates multiple types of computing cores on a single chip to design chips. The design of multi-core heterogeneous chips involves aspects such as the selection and configuration of computing cores, task scheduling strategies and on-chip interconnection structures for different cores, energy efficiency optimization, and verification and debugging. The traditional multi-core heterogeneous chip design methods have the following problems:

[0003] 1) The static configuration scheme has defects: Existing schemes usually pre-fix the types and quantities of heterogeneous cores (such as the CPU+GPU combination), and it is difficult to adapt to the dynamic changes of computing loads in complex application scenarios;

[0004] 2) Low task scheduling efficiency: Existing task scheduling strategies are mostly based on static rules (such as round-robin scheduling, fixed-priority scheduling, etc.), and fail to combine task feature analysis and real-time load prediction, which easily leads to uneven load distribution among computing cores and affects the overall performance;

[0005] 3) Rigid on-chip interconnection architecture: Traditional multi-core chips often adopt a single topology structure (such as a bus or Mesh), and do not perform channel partitioning according to data types and computing priorities, resulting in high-priority tasks (such as real-time control instructions) being blocked by low-priority data streams, affecting the real-time performance of the system.

[0006] 4) Coarse energy efficiency optimization: Existing energy efficiency control methods usually adjust the voltage and frequency as a whole for each core, and fail to perform refined control for different types of computing cores;

[0007] 5) Lag in verification and correction: The traditional design process relies on front-end simulation and lacks a post-silicon dynamic verification mechanism, resulting in difficult and quick repair of core cooperation problems after actual deployment. Summary of the Invention

[0008] The purpose of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0009] To this end, an object of an embodiment of the present invention is to provide a chip design method based on a multi-core heterogeneous architecture, which optimizes the chip design process to solve problems such as low resource utilization, poor real-time performance, and energy efficiency imbalance caused by rigid architecture, inefficient scheduling, and extensive energy efficiency control in the prior art.

[0010] Another object of an embodiment of the present invention is to provide a chip design system based on a multi-core heterogeneous architecture.

[0011] In order to achieve the above technical objectives, the technical solutions adopted in the embodiments of the present invention include:

[0012] In a first aspect, an embodiment of the present invention provides a chip design method based on a multi-core heterogeneous architecture, including:

[0013] Determine the types and quantities of heterogeneous computing cores according to the target application scenario and configure a task scheduling unit. The types of the heterogeneous computing cores include at least two of a general computing core, a dedicated accelerator core, and a reconfigurable logic unit. The task scheduling unit is used to generate a dynamic task scheduling strategy;

[0014] Design the on-chip interconnect structure of the heterogeneous computing cores using a hybrid topology structure, and verify and adjust the dynamic task scheduling strategy and the on-chip interconnect structure through hardware simulation;

[0015] Configure a power management unit according to the energy efficiency optimization target. The power management unit is used to control the working states of the heterogeneous computing cores;

[0016] Collect the operation data of the heterogeneous computing cores in the post-silicon verification stage, generate a firmware update plan according to the operation data, and perform firmware update according to the firmware update plan.

[0017] Further, the determining the types and quantities of heterogeneous computing cores according to the target application scenario includes:

[0018] Conduct a requirements analysis on the target application scenario to obtain the task characteristics of the target application scenario;

[0019] Determine the types and quantities of the heterogeneous computing cores according to the task characteristics and constraint conditions through a multi-objective optimization algorithm.

[0020] Further, the generating the dynamic task scheduling strategy includes:

[0021] Classify the computing tasks of the target application scenario to obtain task types;

[0022] Monitor the load conditions of the heterogeneous computing cores to obtain real-time load data, and perform real-time load prediction based on the real-time load data and historical load data to obtain real-time load prediction results;

[0023] Generate the dynamic task scheduling policy according to the task type and the real-time load prediction results.

[0024] Furthermore, designing the on-chip interconnection structure of the heterogeneous computing cores by adopting a hybrid topology structure includes:

[0025] Connect the homogeneous computing cores through a star topology;

[0026] Connect the heterogeneous computing cores through a ring topology;

[0027] Divide the data transmission channels according to the task type.

[0028] Furthermore, the hardware simulation includes load perturbation testing and data interaction testing. Verifying and adjusting the dynamic task scheduling policy and the on-chip interconnection structure through hardware simulation includes:

[0029] Evaluate the dynamic task scheduling policy through the load perturbation testing to obtain the first hardware simulation verification result, and adjust the dynamic task scheduling policy according to the first hardware simulation verification result;

[0030] Evaluate the on-chip interconnection architecture through the data interaction testing to obtain the second hardware simulation verification result, and adjust the on-chip interconnection structure according to the second hardware simulation verification result.

[0031] Furthermore, the power management unit includes a dynamic voltage and frequency control circuit and a gate-level clock gating circuit. The working states of the heterogeneous computing cores include working voltage, working frequency, and start / stop state. Controlling the working states of the heterogeneous computing cores includes:

[0032] Adjust the working voltage and working frequency of the general computing cores through the dynamic voltage and frequency control circuit according to the load prediction results;

[0033] Control the start / stop state of the dedicated accelerator cores through the gate-level clock gating circuit according to the task type.

[0034] Furthermore, the operation data includes data conflict events, task scheduling information, and resource allocation information among the heterogeneous computing cores. The firmware update plan includes a collaborative working mode correction strategy, a task scheduling optimization strategy, and a resource allocation adjustment strategy for the heterogeneous computing cores. Generating the firmware update plan according to the operation data includes:

[0035] Collect the data conflict event, the task scheduling information, and the resource allocation information through the reserved debugging interface;

[0036] Generate the collaborative work mode correction strategy according to the data conflict event, generate the task scheduling optimization strategy according to the task scheduling information, and generate the resource allocation adjustment strategy according to the resource allocation information.

[0037] In a second aspect, an embodiment of the present invention provides a chip design system based on a multi-core heterogeneous architecture, including:

[0038] A heterogeneous core configuration module, configured to determine the types and quantities of heterogeneous computing cores according to a target application scenario and configure a task scheduling unit, where the types of the heterogeneous computing cores include at least two of a general computing core, a dedicated accelerator core, and a reconfigurable logic unit, and the task scheduling unit is configured to generate a dynamic task scheduling strategy;

[0039] A network-on-chip design module, configured to design an on-chip interconnection structure of the heterogeneous computing cores by using a hybrid topology structure, and verify and adjust the dynamic task scheduling strategy and the on-chip interconnection structure through hardware simulation;

[0040] An energy efficiency optimization module, configured to configure a power management unit according to an energy efficiency optimization target, where the power management unit is configured to control the working states of the heterogeneous computing cores;

[0041] A post-silicon verification module, configured to collect operation data of the heterogeneous computing cores in a post-silicon verification phase, generate a firmware update plan according to the operation data, and perform firmware update according to the firmware update plan.

[0042] In a third aspect, an embodiment of the present invention provides a device, including:

[0043] At least one processor;

[0044] At least one memory, configured to store at least one program;

[0045] When the at least one program is executed by the at least one processor, the at least one processor implements a chip design method based on a multi-core heterogeneous architecture as described above.

[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute a chip design method based on a multi-core heterogeneous architecture as described above when being executed by the processor.

[0047] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention:

[0048] In the embodiments of the present invention, the conflict between "static design - dynamic requirements" in the traditional solution is solved through the dynamic configuration of heterogeneous core types; the efficient allocation of computing resources is achieved through dynamic task scheduling and load balancing; the low latency and high bandwidth utilization of the communication between computing cores are achieved through the on-chip interconnection structure and channel division of the hybrid topology design, and the working state of the computing cores is controlled through the dynamic voltage and frequency adjustment unit and the gate-level clock gating circuit, effectively improving the energy efficiency ratio of the computing cores; the continuous optimization after chip deployment is supported through the post-silicon verification and firmware update mechanism, improving the long-term reliability and security of the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the steps of a chip design method based on a multi-core heterogeneous architecture provided by the embodiments of the present invention;

[0050] Figure 2 It is a schematic diagram of a chip design system based on a multi-core heterogeneous architecture provided by the embodiments of the present invention;

[0051] Figure 3 It is a schematic diagram of the structure of a device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.

[0053] In the description of the present invention, the meaning of "a plurality" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0054] The English abbreviations used in the present invention include:

[0055] FPGA: Field Programmable Gate Array, field programmable gate array;

[0056] GPU: Graphics Processing Unit, Graphics Processing Unit;

[0057] CPU: Central Processing Unit, Central Processing Unit.

[0058] Figure 1 This is a schematic diagram of the steps of a chip design method based on a multi-core heterogeneous architecture provided by an embodiment of the present invention. Refer to Figure 1 An embodiment of the present invention provides a chip design method based on a multi-core heterogeneous architecture, including:

[0059] S101. Determine the types and quantities of heterogeneous computing cores according to the target application scenario and configure a task scheduling unit. The types of heterogeneous computing cores include at least two of a general computing core, a dedicated accelerator core, and a reconfigurable logic unit. The task scheduling unit is used to generate a dynamic task scheduling strategy;

[0060] Specifically, the types of heterogeneous computing cores include a general computing core, a dedicated accelerator core, and a reconfigurable logic unit. The general computing core is mainly used to execute general instructions, such as CPUs (Central Processing Units) with ARM architecture and x86 architecture; the dedicated accelerator core is mainly used to provide hardware support for specific algorithm optimizations, such as GPUs (Graphics Processors) suitable for graphics rendering and NPUs (Neural Network Processors) suitable for neural network inference acceleration, etc.; the reconfigurable logic unit is mainly used to reconstruct the computing logic according to different tasks to adapt to the changes of dynamic tasks, such as FPGAs (Field Programmable Gate Arrays), etc. In this embodiment, the multi-core heterogeneous chip includes at least two of the above three computing cores.

[0061] In some alternative embodiments, determining the types and quantities of heterogeneous computing cores according to the target application scenario includes:

[0062] A1. Perform a requirements analysis on the target application scenario to obtain the task characteristics of the target application scenario;

[0063] A2. Determine the types and quantities of heterogeneous computing cores through a multi-objective optimization algorithm according to the task characteristics and constraint conditions.

[0064] Specifically, the requirement analysis of the target application scenario includes extracting the computing tasks to be processed in the target application scenario, obtaining the task characteristics in the target application scenario according to the computational intensity, data parallelism, and real-time requirements of these tasks. The constraint conditions include area constraints, cost constraints, power consumption constraints, etc. Multi-objective optimization methods such as evolutionary algorithms, reinforcement learning, and Monte Carlo simulation can be used to find the optimal core combination under these task characteristics and constraint conditions. For example, Monte Carlo simulation is used to randomly sample possible core combinations, and the core combinations with a success rate of completing the computing tasks exceeding the threshold under the current task characteristics and constraint conditions are screened out, and further the Pareto optimal solution of the core type combination is selected.

[0065] In some alternative embodiments, generating a dynamic task scheduling policy includes:

[0066] B1. Classify the computing tasks of the target application scenario to obtain the task types;

[0067] B2. Monitor the load conditions of heterogeneous computing cores to obtain real-time load data, and perform real-time load prediction based on the real-time load data and historical load data to obtain real-time load prediction results;

[0068] B3. Generate a dynamic task scheduling policy according to the task types and real-time load prediction results.

[0069] Specifically, in this embodiment, to generate a dynamic task scheduling policy through the task scheduling unit, first, the computing tasks of the target application scenario need to be classified, which can be divided into compute-intensive, control-intensive, and data-intensive. The realization of real-time load prediction can be to train a long short-term memory network through historical load data to obtain a load prediction model, and then combine the real-time load data and the load prediction model to obtain real-time load prediction results. The dynamic task scheduling policy can be divided into two parts. One part is the priority scheduling policy based on task types, which preferentially executes high-priority tasks such as real-time control. The other part is the load balancing policy based on real-time load prediction results, which migrates computing tasks among heterogeneous processors according to the load prediction results to achieve load balancing.

[0070] S102. Design the on-chip interconnection structure of heterogeneous computing cores using a hybrid topology structure, and verify and adjust the dynamic task scheduling policy and the on-chip interconnection structure through hardware simulation;

[0071] Specifically, the on-chip interconnection structure is a communication structure that connects different computing cores in a heterogeneous multi-core architecture. Its topology design determines key metrics such as data transmission bandwidth, latency, and energy efficiency. The hybrid topology design refers to a topology structure that combines two or more topology methods. Compared with a single topology structure, it can better adapt to the needs of different computing cores. Common topology structures include star topology, ring topology, Mesh topology, and tree topology, etc.

[0072] In some alternative embodiments, a hybrid topology design is adopted for the on-chip interconnection structure of heterogeneous computing cores, including:

[0073] C1. Connecting homogeneous computing cores through star topology;

[0074] C2. Connecting heterogeneous computing cores through ring topology;

[0075] C3. Dividing data transmission channels according to task types.

[0076] Specifically, in this embodiment, homogeneous computing cores are connected through star topology, and a central control communication method is adopted among the same type of computing cores, which can achieve low-latency high-speed data exchange. Ring topology is used to connect heterogeneous computing cores, which can improve bandwidth utilization and reduce congestion. The priorities of computing tasks are divided according to task types, and then data transmission channels are divided according to task priorities, which can prevent low-priority tasks from blocking high-priority tasks and implement a differential service quality allocation strategy.

[0077] In some alternative embodiments, hardware simulation includes load perturbation testing and data interaction testing. The dynamic task scheduling strategy and on-chip interconnection structure are verified and adjusted through hardware simulation, including:

[0078] D1. Evaluating the dynamic task scheduling strategy through load perturbation testing to obtain the first hardware simulation verification result, and adjusting the dynamic task scheduling strategy according to the first hardware simulation verification result;

[0079] D2. Evaluating the on-chip interconnection architecture through data interaction testing to obtain the second hardware simulation verification result, and adjusting the on-chip interconnection structure according to the second hardware simulation verification result.

[0080] Specifically, the purpose of the load disturbance test is to evaluate the effectiveness of the dynamic task scheduling strategy. This is achieved by constructing different load patterns and injecting load disturbances into the simulation environment. During the test, it is monitored whether the computing tasks are mapped to appropriate cores for processing according to the task type, priority, and load conditions, and the monitoring results are recorded. Based on the monitoring results, the task scheduling strategy is adjusted. For example, when task backlogs occur, the preemption weight of high-priority tasks is increased. The purpose of the data interaction test is to verify the throughput, bandwidth utilization, and latency of the on-chip interconnection structure. Packets are generated under different task scenarios, and various indicators during data transmission on the communication link are monitored, and the monitoring results are recorded. Based on the monitoring results, the on-chip interconnection structure is adjusted.

[0081] S103. Configure the power management unit according to the energy efficiency optimization goal, where the power management unit is used to control the operating states of heterogeneous computing cores;

[0082] Specifically, the energy efficiency optimization goal refers to minimizing the power consumption of the chip while ensuring its computing performance, mainly achieved by the power management unit adjusting the operating states such as voltage, frequency, startup and shutdown states of the computing cores.

[0083] In some alternative embodiments, the power management unit includes a dynamic voltage and frequency control unit and a gate-level clock gating circuit. The operating states of the heterogeneous computing cores include operating voltage, operating frequency, and startup and shutdown states. Controlling the operating states of the heterogeneous computing cores includes:

[0084] E1. Adjust the operating voltage and operating frequency of the general computing core according to the load prediction result through the dynamic voltage and frequency control unit;

[0085] E2. Control the startup and shutdown states of the dedicated accelerator core according to the task type through the gate-level clock gating circuit.

[0086] Specifically, the dynamic voltage and frequency control unit is used to adjust the voltage and frequency of the general computing core. Its working principle is to monitor the load of the computing core in combination with the load prediction result. When the load is low, the voltage and frequency of the computing core are reduced, thus significantly reducing its power consumption. When the load is high, the voltage and frequency of the computing core are increased to improve its performance to meet the load requirements. The gate-level clock gating circuit is used to dynamically start and stop the dedicated accelerator core by controlling the clock signal. Its working principle is to monitor the current computing task type and control the clock gating according to the task type. When the current task type does not require the dedicated accelerator core, the clock is turned off to make the dedicated accelerator core enter the low-power mode. When the dedicated accelerator core needs to execute a task, the clock signal is restored, so that the dedicated accelerator core can operate intermittently as needed, greatly reducing the overall power consumption and improving the energy efficiency ratio.

[0087] S104. During the post-silicon verification phase, collect the operation data of the heterogeneous computing cores, generate a firmware update plan based on the operation data, and perform firmware update according to the firmware update plan.

[0088] Specifically, post-silicon verification refers to, after the chip is manufactured, collecting the actual operation data to discover and correct problems in the hardware design and scheduling strategy. The firmware update plan is a set of optimization strategies generated based on the post-silicon verification data. Through update methods such as Bootloader, write new scheduling strategies or co-optimization algorithms to the chip to achieve firmware update.

[0089] In some alternative embodiments, the operation data includes data conflict events between heterogeneous computing cores, task scheduling information, and resource allocation information. The firmware update plan includes a collaborative working mode correction strategy for heterogeneous computing cores, a task scheduling optimization strategy, and a resource allocation adjustment strategy. Generating a firmware update plan based on the operation data includes:

[0090] F1. Collect data conflict events, task scheduling information, and resource allocation information through the reserved debugging interface;

[0091] F2. Generate a collaborative working mode correction strategy based on the data conflict events, generate a task scheduling optimization strategy based on the task scheduling information, and generate a resource allocation adjustment strategy based on the resource allocation information.

[0092] Specifically, in this embodiment, collect the operation data through the debugging interface reserved in the chip design phase, including data conflict events (competition or errors generated during data access by different computing cores), task scheduling information (the allocation of tasks on different cores, the load conditions of each core, etc.), and resource allocation information (the computing resource allocation of each core, etc.). The collaborative working mode correction strategy includes access priority control strategies, data sharing method adjustment strategies, etc. designed for specific data conflict events. The task scheduling optimization strategy includes load balancing strategies, etc. designed for task scheduling information. The resource allocation adjustment strategy includes cache allocation strategies and power management optimization strategies, etc. designed for resource allocation information. It can be understood that post-silicon verification performs actual tests on the manufactured chip and collects relevant data to generate a firmware update plan, which can continuously improve the performance of the chip in the actual operating environment.

[0093] It can be recognized that the present application solves the conflict of "static design - dynamic requirements" in the traditional solution through the dynamic configuration of heterogeneous core types; realizes the efficient allocation of computing resources through dynamic task scheduling and load balancing; realizes low latency and high bandwidth utilization of computing core communication through the on-chip interconnection structure and channel division of the hybrid topology design, controls the working state of the computing core through the dynamic voltage and frequency adjustment unit and the gate-level clock gating circuit, effectively improves the energy efficiency ratio of the computing core; supports the continuous optimization after chip deployment through the post-silicon verification and firmware update mechanism, and improves the long-term reliability and security of the chip.

[0094] The chip design method of the present invention will be described below in conjunction with a specific embodiment.

[0095] This embodiment is for the application scenario of the special chip design for power grid security protection, and is described in combination with the requirements of real-time monitoring and network attack protection in the power system. The real-time data acquisition-encryption-decision requirements of intelligent substations and distribution automation terminals require the processor to meet the following conditions: multi-source heterogeneous data processing (SCADA telemetry data, relay protection signals, encrypted communication messages); anti-network attack ability (resisting DDoS, malicious instruction injection, etc.); strong real-time performance (protection trip instruction delay ≤ 2 ms); high reliability (MTBF ≥ 100,000 hours). Through Monte Carlo simulation, with a task success rate of 99.99% as the reliability constraint, the type and number of heterogeneous computing cores are determined to be 2 real-time control cores (ARM Cortex-R52 with lockstep redundancy design), 1 cryptographic security accelerator (supporting national cryptographic SM4 / SM9 algorithms, encryption and decryption throughput ≥ 10 Gbps), and 1 reconfigurable logic unit (which can be dynamically switched to a protocol parsing engine or an intrusion detection module). The computing resources are allocated according to the task priority through dynamic task scheduling, which can be divided into priority 1: relay protection trip instruction (hard real-time task); priority 2: encrypted communication message processing (soft real-time task); priority 3: device status monitoring data cleaning (non-real-time task). The time window segmentation strategy is adopted to reserve a dedicated time slice for hard real-time tasks. When network attack traffic is detected, the reconfigurable unit switches to the intrusion detection mode to parallelly process malicious feature matching. The control core and the security accelerator adopt a double-star redundant topology, and the reconfigurable unit and the I / O interface are connected through an independent ring channel. The production data and control instructions are physically isolated. The QoS (Quality of Service) strategy is adopted to allocate an exclusive channel for the trip instruction to ensure that the end-to-end delay meets the threshold. The CRC check retransmission mechanism is adopted to transmit encrypted messages to ensure that the bit error rate meets the threshold. The task scheduling strategy and the on-chip interconnection structure are verified and optimized through hardware simulation. The dynamic voltage and frequency regulation and the gate-level clock gating circuit are used to intelligently adjust the power consumption of the computing core to meet the energy efficiency optimization goal. The operation data is collected through the debugging interface in the post-silicon verification stage. Combining with the functional safety (IEC61508 SIL3) and network security (GB / T 22239-2019) standards, it supports dynamic defense and online optimization to ensure the stable operation of the power grid system.

[0096] The chip design method of the present invention will be described below in combination with another specific embodiment.

[0097] This embodiment is for the application scenario of intelligent security edge computing chip design, which needs to meet the real-time analysis requirements of multi-modal data in smart parks. Taking the energy efficiency ratio as the optimization goal and the area as the constraint condition, the type and number of heterogeneous computing cores are determined by the Pareto optimization algorithm as 2 low-power RISC-V cores, 1 video coding accelerator (H.265 4K@60fps), and 1 reconfigurable logic unit (supporting dynamic switching to AI inference or encryption engine). A dynamic task scheduling strategy is adopted, and the peak period of the video stream is predicted based on historical data to allocate accelerator resources in advance. When the load of the RISC-V core > 85%, the face feature comparison task is migrated to the reconfigurable unit. The control plane (RISC-V + security module) adopts a tree topology, and the data plane (accelerator + reconfigurable unit) adopts a crossbar topology. QoS channel division is used to ensure that control instructions and video streams do not interfere with each other. Power management adopts hierarchical dynamic voltage and frequency control and module-level gating. The video preprocessing module is automatically turned off under low load, and the power supply strategy of the reconfigurable unit is adjusted. During the post-silicon verification stage, the DMA channel contention situation is monitored, and the task scheduling strategy is optimized by combining online learning.

[0098] Referring to Figure 2 , an embodiment of the present invention provides a chip design system based on a multi-core heterogeneous architecture, including:

[0099] A heterogeneous core configuration module, which is used to determine the type and number of heterogeneous computing cores according to the target application scenario and configure a task scheduling unit. The types of heterogeneous computing cores include at least two of general computing cores, dedicated accelerator cores, and reconfigurable logic units. The task scheduling unit is used to generate a dynamic task scheduling strategy;

[0100] An on-chip network design module, which is used to design the on-chip interconnection structure of heterogeneous computing cores using a hybrid topology structure, and verify and adjust the dynamic task scheduling strategy and the on-chip interconnection structure through hardware simulation;

[0101] An energy efficiency optimization module, which is used to configure a power management unit according to the energy efficiency optimization goal. The power management unit is used to control the working state of heterogeneous computing cores;

[0102] A post-silicon verification module, which is used to collect the operation data of heterogeneous computing cores during the post-silicon verification stage, generate a firmware update plan according to the operation data, and perform firmware update according to the firmware update plan.

[0103] The content in the above method embodiments is applicable to this system embodiment. The functions specifically implemented by this system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0104] Referring to Figure 3, an embodiment of the present invention provides a device, including:

[0105] At least one processor;

[0106] At least one memory for storing at least one program;

[0107] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above-mentioned chip design method based on a multi-core heterogeneous architecture.

[0108] An embodiment of the present invention also provides a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to execute the above-mentioned chip design method based on a multi-core heterogeneous architecture when executed by a processor.

[0109] A computer-readable storage medium according to an embodiment of the present invention can execute a chip design method based on a multi-core heterogeneous architecture provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0110] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the device executes Figure 1 The chip design method based on a multi-core heterogeneous architecture shown.

[0111] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown can actually be executed substantially simultaneously, or the above blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical processes presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0112] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0113] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0114] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0115] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the above programs can be printed, because the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, deciphering, or otherwise processing as appropriate, and then stored in a computer memory.

[0116] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0117] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0118] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0119] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A chip design method based on a multi-core heterogeneous architecture, characterized in that Including: Determine the types and quantities of heterogeneous computing cores according to the target application scenario and configure a task scheduling unit. The types of the heterogeneous computing cores include at least two of a general computing core, a dedicated accelerator core, and a reconfigurable logic unit. The task scheduling unit is used to generate a dynamic task scheduling strategy; Design the on-chip interconnection structure of the heterogeneous computing cores using a hybrid topology structure, and verify and adjust the dynamic task scheduling strategy and the on-chip interconnection structure through hardware simulation; Configure a power management unit according to the energy efficiency optimization goal. The power management unit is used to control the operating states of the heterogeneous computing cores; Collect the operation data of the heterogeneous computing cores during the post-silicon verification phase, generate a firmware update plan according to the operation data, and perform firmware update according to the firmware update plan.

2. A chip design method based on a multi-core heterogeneous architecture according to claim 1, characterized in that The determining the types and quantities of heterogeneous computing cores according to the target application scenario includes: Conduct a requirements analysis on the target application scenario to obtain the task characteristics of the target application scenario; Determine the types and quantities of the heterogeneous computing cores through a multi-objective optimization algorithm according to the task characteristics and constraint conditions.

3. A chip design method based on a multi-core heterogeneous architecture according to claim 1, characterized in that The generating the dynamic task scheduling strategy includes: Classify the computing tasks of the target application scenario to obtain task types; Monitor the load conditions of the heterogeneous computing cores to obtain real-time load data, and perform real-time load prediction based on the real-time load data and historical load data to obtain a real-time load prediction result; Generate the dynamic task scheduling strategy according to the task types and the real-time load prediction result.

4. A chip design method based on a multi-core heterogeneous architecture according to claim 3, characterized in that The designing the on-chip interconnection structure of the heterogeneous computing cores using a hybrid topology structure includes: Connect homogeneous computing cores through a star topology; Connect the heterogeneous computing cores through a ring topology; Divide data transmission channels according to the task types.

5. A chip design method based on a multi-core heterogeneous architecture according to claim 1, characterized in that The hardware simulation includes a load perturbation test and a data interaction test. The verifying and adjusting the dynamic task scheduling strategy and the on-chip interconnection structure through hardware simulation includes: Evaluate the dynamic task scheduling strategy through the load perturbation test to obtain a first hardware simulation verification result, and adjust the dynamic task scheduling strategy according to the first hardware simulation verification result; Evaluate the on-chip interconnection architecture through the data interaction test to obtain a second hardware simulation verification result, and adjust the on-chip interconnection structure according to the second hardware simulation verification result.

6. A chip design method based on a multi-core heterogeneous architecture according to claim 3, characterized in that The power management unit includes a dynamic voltage and frequency control circuit and a gate-level clock gating circuit. The operating states of the heterogeneous computing cores include operating voltage, operating frequency, and start-stop state. The controlling the operating states of the heterogeneous computing cores includes: Adjust the operating voltage and operating frequency of the general computing core through the dynamic voltage and frequency control circuit according to the load prediction result; Control the start-stop state of the dedicated accelerator core through the gate-level clock gating circuit according to the task types.

7. A chip design method based on a multi-core heterogeneous architecture according to claim 1, characterized in that, The operation data includes data conflict events, task scheduling information, and resource allocation information among the heterogeneous computing cores. The firmware update scheme includes a collaborative working mode correction strategy, a task scheduling optimization strategy, and a resource allocation adjustment strategy for the heterogeneous computing cores. Generating the firmware update scheme based on the operation data includes: Collecting the data conflict events, the task scheduling information, and the resource allocation information through a reserved debugging interface; Generating the collaborative working mode correction strategy based on the data conflict events, generating the task scheduling optimization strategy based on the task scheduling information, and generating the resource allocation adjustment strategy based on the resource allocation information.

8. A chip design system based on a multi-core heterogeneous architecture, characterized in that, Includes: A heterogeneous core configuration module, configured to determine the types and quantities of heterogeneous computing cores according to a target application scenario and configure a task scheduling unit. The types of the heterogeneous computing cores include at least two of a general computing core, a dedicated accelerator core, and a reconfigurable logic unit. The task scheduling unit is configured to generate a dynamic task scheduling strategy; An on-chip network design module, configured to design an on-chip interconnect structure of the heterogeneous computing cores by adopting a hybrid topology structure, and verify and adjust the dynamic task scheduling strategy and the on-chip interconnect structure through hardware simulation; An energy efficiency optimization module, configured to configure a power management unit according to an energy efficiency optimization target. The power management unit is configured to control the working states of the heterogeneous computing cores; A post-silicon verification module, configured to collect operation data of the heterogeneous computing cores in a post-silicon verification phase, generate a firmware update scheme based on the operation data, and perform firmware update according to the firmware update scheme.

9. A device, characterized in that, Includes: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, enabling the at least one processor to implement a chip design method based on a multi-core heterogeneous architecture according to any one of claims 1-7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute a chip design method based on a multi-core heterogeneous architecture according to any one of claims 1-7.

Citation Information

Cited By

  • Layout design rule checking method, system, equipment, medium and product

    CN120803676A

  • Cooperative scheduling and optimization method for heterogeneous resources of algorithm training platform

    CN120872535A

  • Performance level coding method, performance level decoding method and system-level chip

    CN121255725A

  • Performance monitoring driving type heterogeneous computing task hardware scheduling system and method

    CN121722575A

  • Object-oriented network-on-chip verification environment construction method and verification system

    CN121814580A