Performance test and verification method, system and device for multi-core heterogeneous chip and medium

By dynamically configuring test parameters and generating test cases covering heterogeneous collaborative working modes, combining reconfigurable monitoring and multi-dimensional evaluation models, the static and singularization problems in multi-core heterogeneous chip performance testing are solved, and efficient performance verification and optimization are achieved.

CN120294534APending Publication Date: 2025-07-11GUANGZHOU KETENG INFORMATION TECH
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510350218.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The performance testing methods of existing multi-core heterogeneous chips have problems such as static test parameters, insufficient coverage of test cases, fragmentation of monitoring means and single analysis dimensions, resulting in significant deviations from the actual operation, and lack of systematic verification of heterogeneous core collaborative working mode.

Method used

By configuring the dynamic test parameter set, test cases covering the collaborative working mode of heterogeneous computing are generated, feature data is collected using the reconfigurable monitoring module, and quantitative analysis is used for multi-dimensional evaluation model to generate performance deviation indicators and architectural optimization suggestions.

Benefits of technology

It improves the verification accuracy and reliability of multi-core heterogeneous chips, realizes efficient allocation of computing resources and improves energy efficiency ratio, supports continuous optimization after chip deployment, and improves long-term reliability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120294534A_ABST
    Figure CN120294534A_ABST
Patent Text Reader

Abstract

The invention discloses a performance test and verification method, system and device for a multi-core heterogeneous chip and a medium, and the method comprises the steps: configuring a test parameter set according to a target application scene, the test parameter set comprising a processor core type recognition parameter, a communication topology weight parameter and a load balancing threshold parameter; generating a test case covering the heterogeneous computing core cooperative working mode; the method comprises the following steps: acquiring characteristic data during testing through a reconfigurable monitoring sub-module, wherein the characteristic data comprises inter-core communication delay information, shared cache hit rate information and power consumption distribution thermodynamic information; and performing quantitative analysis on the feature data through a multi-dimensional evaluation model to obtain a performance deviation index and an architecture optimization suggestion parameter. By optimizing the chip performance verification process under the multi-core heterogeneous architecture, the verification precision is effectively improved, and the chip performance verification method can be widely applied to the technical field of chip design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and particularly to a method, system, device and medium for performance testing and verification of multi-core heterogeneous chips. Background Art

[0002] With the rapid development of fields such as artificial intelligence and autonomous driving, multi-core heterogeneous architecture chips have been widely used due to their high-performance computing characteristics. However, the existing performance testing methods face the following technical bottlenecks:

[0003] Static test parameters: Traditional tests use fixed parameter configurations, which cannot adapt to the dynamic load characteristics of chips in different application scenarios, resulting in a significant deviation between the test results and the actual operation; Insufficient test case coverage: Existing test schemes are mostly designed for the interaction between homogeneous cores, lacking systematic verification of the collaborative working modes of heterogeneous cores such as CPU / GPU / DSP / accelerators; Fragmented monitoring means: Traditional monitoring modules rely on independent probes to collect local data, making it difficult to reconstruct the complete execution link of cross-core tasks, and hardware probes introduce additional performance losses; Single analysis dimension: Existing evaluation models mostly focus on single indicators such as computing throughput, lacking the ability to jointly analyze key parameters such as communication latency and energy consumption efficiency ratio. Summary of the Invention

[0004] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0005] To this end, an object of an embodiment of the present invention is to provide a method for performance testing and verification of multi-core heterogeneous chips, which improves the accuracy and reliability of verification by optimizing the chip performance verification process under a multi-core heterogeneous architecture.

[0006] Another object of an embodiment of the present invention is to provide a system for performance testing and verification of multi-core heterogeneous chips.

[0007] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:

[0008] In the first aspect, an embodiment of the present invention provides a method for performance testing and verification of multi-core heterogeneous chips, including:

[0009] Configuring a test parameter set according to a target application scenario, where the test parameter set includes processor core type identification parameters, communication topology weight parameters, and load balancing threshold parameters;

[0010] Generating test cases that cover the collaborative working modes of heterogeneous computing cores, where the test cases include CPU-GPU heterogeneous computing verification cases, DSP-accelerator data flow verification cases, and mixed-precision operation verification cases;

[0011] Collect the characteristic data during testing through the reconfigurable monitoring sub-module, where the characteristic data includes inter-core communication delay information, shared cache hit rate information, and power consumption distribution heat map information;

[0012] Perform quantitative analysis on the characteristic data through a multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters

[0013] Furthermore, the configuration of the test parameter set according to the target application scenario includes:

[0014] Determine the physical layout of the computing cores, storage cores, and communication cores according to the chip layout information, and configure the processor core type identification parameters according to the physical layout;

[0015] Adjust the communication topology weight parameters and bus arbitration strategy according to the historical test data;

[0016] Constrain the test boundary conditions according to the temperature-frequency coupling model and determine the load balancing threshold parameters.

[0017] Furthermore, the generation of test cases covering the collaborative working mode of heterogeneous computing cores includes:

[0018] Generate a basic test vector through a randomization algorithm;

[0019] Generate a basic test case according to the basic test vector and the typical application mode matching template;

[0020] Generate the test case according to the basic test case and the bursty load disturbance factor.

[0021] Furthermore, the reconfigurable monitoring sub-module includes a hardware-level probe unit, a system-level tracing unit, and an application-level analysis unit. The collection of the characteristic data during testing through the reconfigurable monitoring sub-module includes:

[0022] Collect the execution pipeline state with clock-level accuracy through the hardware-level probe unit to obtain the inter-core communication delay information;

[0023] Record the task scheduling queue and interrupt response events through the system-level tracing unit to obtain the shared cache hit rate information;

[0024] Construct a cross-core function call relationship graph through the application-level analysis unit to obtain the power consumption distribution heat map information.

[0025] Further, the multi-dimensional evaluation model includes a computing unit utilization rate evaluation sub-model, a bandwidth saturation rate evaluation sub-model, and an energy efficiency ratio evaluation sub-model. The performance deviation index includes a load balancing deviation, a data transportation deviation, and an energy efficiency deviation. Quantitatively analyzing the feature data through the multi-dimensional evaluation model to obtain the performance deviation index and architecture optimization suggestion parameters, including:

[0026] Analyzing the power consumption distribution thermal information through the computing unit utilization rate evaluation sub-model to obtain the load balancing deviation;

[0027] Analyzing the inter-core communication delay information and the shared cache hit rate information through the bandwidth saturation rate evaluation sub-model to obtain the data transportation deviation;

[0028] Analyzing the power consumption distribution thermal information through the energy efficiency ratio evaluation sub-model to obtain the energy efficiency deviation;

[0029] Obtaining the architecture optimization suggestion parameters according to the load balancing deviation, the data transportation deviation, and the energy efficiency deviation.

[0030] Further, the method further includes:

[0031] Generating the visualization report according to the performance deviation index;

[0032] Generating a verification meta data packet according to the visualization report.

[0033] Further, the method further includes:

[0034] Creating a virtual test environment through virtualization technology;

[0035] Executing the test case in the virtual test environment in shadow mode to obtain test data;

[0036] Calibrating the deviation compensation coefficient between the test environment and the target application scenario according to the test data.

[0037] In a second aspect, an embodiment of the present invention provides a performance test and verification system for a multi-core heterogeneous chip, including:

[0038] A parameter configuration module, configured to configure a test parameter set according to a target application scenario, where the test parameter set includes a processor core type identification parameter, a communication topology weight parameter, and a load balancing threshold parameter;

[0039] A test case generation module, generating test cases covering the cooperative working mode of heterogeneous computing cores, where the test cases include CPU-GPU heterogeneous computing verification test cases, DSP-accelerator data stream verification test cases, and mixed-precision operation verification test cases;

[0040] A monitoring module, configured to collect characteristic data during testing through a reconfigurable monitoring sub-module, where the characteristic data includes inter-core communication delay information, shared cache hit rate information, and power consumption distribution thermal information;

[0041] An analysis and evaluation module, configured to perform quantitative analysis on the characteristic data through a multi-dimensional evaluation model to obtain a performance deviation degree index and architecture optimization suggestion parameters.

[0042] In a third aspect, an embodiment of the present invention provides a device, including:

[0043] At least one processor;

[0044] At least one memory, configured to store at least one program;

[0045] When the at least one program is executed by the at least one processor, the at least one processor implements the performance test and verification method for a multi-core heterogeneous chip as described above.

[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the performance test and verification method for a multi-core heterogeneous chip as described above when executed by the processor.

[0047] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or can be understood through the practice of the present invention:

[0048] The embodiment of the present invention solves the conflict of "static design - dynamic requirements" in the traditional solution by dynamically configuring heterogeneous core types; realizes efficient allocation of computing resources through dynamic task scheduling and load balancing; realizes low latency and high bandwidth utilization of computing core communication through a hybrid topology on-chip interconnection structure and channel division, controls the working state of the computing core through a dynamic voltage and frequency adjustment unit and a gate-level clock gating circuit, effectively improves the energy efficiency ratio of the computing core; supports continuous optimization after chip deployment through post-silicon verification and firmware update mechanisms, and improves the long-term reliability and security of the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a step schematic diagram of the performance test and verification method for a multi-core heterogeneous chip provided by an embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of the performance test and verification system for a multi-core heterogeneous chip provided by an embodiment of the present invention;

[0051] Figure 3Schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners

[0052] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0053] In the description of the present invention, "a plurality of" means two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0054] The English abbreviations used in the present invention include:

[0055] FPGA: Field Programmable Gate Array, Field Programmable Gate Array;

[0056] GPU: Graphics Processing Unit, Graphics Processing Unit;

[0057] CPU: Central Processing Unit, Central Processing Unit;

[0058] DSP: Digital Signal Processor, Digital Signal Processor.

[0059] Figure 1 Schematic diagram of the steps of the performance test and verification method for a multi-core heterogeneous chip provided by an embodiment of the present invention. Refer to Figure 1 , an embodiment of the present invention provides a performance test and verification method for a multi-core heterogeneous chip, including:

[0060] S101. Configure a test parameter set according to the target application scenario. The test parameter set includes processor core type identification parameters, communication topology weight parameters, and load balancing threshold parameters;

[0061] Specifically, the processor core type recognition parameter is a parameter for distinguishing different types of computing cores inside the chip, which is used to ensure that the test tool can correctly identify various computing cores (such as the general-purpose computing core CPU or the dedicated computing core GPU, etc.) so as to formulate targeted test cases. The communication topology weight parameter is a parameter for measuring the importance and bandwidth occupancy of the communication links between different computing cores inside the chip, which is used to determine the task scheduling priority between different computing cores. The load balancing threshold parameter is a threshold for determining how tasks are allocated between computing cores, which is used to control the distribution of computing loads and avoid overloading some cores while other cores are idle.

[0062] In some alternative embodiments, the test parameter set is configured according to the target application scenario, including:

[0063] A1. Determine the physical layout of the computing core, storage core, and communication core according to the chip layout information, and configure the processor core type recognition parameter according to the physical layout;

[0064] A2. Adjust the communication topology weight parameter and the bus arbitration strategy according to the historical test data;

[0065] A3. Constrain the test boundary conditions according to the temperature-frequency coupling model, and determine the load balancing threshold parameter.

[0066] Specifically, the chip layout information is the physical design diagram of the chip, including the distribution of cores, caches, and interconnection structures. According to the physical design data of the chip layout information, the computing capabilities and interconnection methods of different cores can be determined, and then the processor core type to which each computing core belongs can be identified. The historical test data includes the access conflict situations of each computing core to the bus. The communication topology weight parameter is adjusted according to the conflict frequency and data throughput. When the communication link of a high-priority task is blocked, its bandwidth allocation ratio is increased. The bus arbitration strategy determines how to preferentially allocate bandwidth when multiple cores request the bus simultaneously. The temperature-frequency coupling model is used to predict the temperature change of the computing core under different load conditions based on the coupling relationship between the operating frequency and temperature of the core, and can be constructed by collecting the relationship data between the chip temperature and frequency. The test boundary conditions are used to set the temperature and frequency safety ranges for performance testing, and the computing frequency is reduced when the core temperature exceeds the threshold to simulate the dynamic regulation mechanism in the real operating environment.

[0067] S102. Generate test cases covering the collaborative working modes of heterogeneous computing cores. The test cases include CPU-GPU heterogeneous computing verification cases, DSP-accelerator data flow verification cases, and mixed-precision operation verification cases;

[0068] Specifically, the heterogeneous computing core collaborative working mode refers to the working method in which different types of computing cores (including CPUs, GPUs, DSPs, and dedicated accelerators) collaborate to execute computing tasks. Since the architectures and data processing methods of different computing cores are different, test cases need to cover the collaborative working conditions of these cores in actual applications. Among them, the CPU-GPU heterogeneous computing verification test cases are used to verify data transfer, computing scheduling, and task offloading strategies between the CPU and the GPU; the DSP-accelerator data stream verification test cases are used to test the data transfer capabilities between the DSP and the dedicated accelerator; the mixed-precision arithmetic verification test cases are used to test the correctness of computations with different precisions (such as FP32, FP16, INT8, etc.) among multiple cores.

[0069] In some alternative embodiments, the test parameter set is configured according to the target application scenario, including:

[0070] B1. Generate a basic test vector through a randomization algorithm;

[0071] B2. Generate basic test cases according to the basic test vector and the typical application pattern matching template;

[0072] B3. Generate test cases according to the basic test cases and the burst load disturbance factor.

[0073] Specifically, the basic test vector is a set of test input parameters obtained by randomly generating task types and data scales through a randomization algorithm. To make the randomly generated test vector more valuable as a reference and closer to the real situation in the current application scenario, it is necessary to compare the randomly generated test vector with the typical application pattern matching template, find the closest pattern, and adjust the test parameters accordingly to obtain the basic test cases. Finally, add a burst load dynamically on top of the basic test cases to further test the load fluctuation of the system in extreme situations and verify the stability of the system.

[0074] S103. Collect feature data during testing through a reconfigurable monitoring module. The feature data includes inter-core communication latency information, shared cache hit rate information, and power consumption distribution heat map information;

[0075] Specifically, the inter-core communication latency information reflects the data transfer performance between different computing cores and can be used to analyze whether there are bottlenecks in task scheduling. The shared cache hit rate information is used to measure the effectiveness of the cache hierarchy and reflects the efficiency of data access. The power consumption distribution heat map information indicates the energy consumption characteristics of different computing cores. By using a reconfigurable monitoring module to adjust the collection strategy according to different test requirements, it is possible to adapt to different test scenarios under various chip architectures.

[0076] In some alternative embodiments, the reconfigurable monitoring module includes a hardware-level probe unit, a system-level tracing unit, and an application-level analysis unit. Feature data during testing is collected through the reconfigurable monitoring module, including:

[0077] C1. Collect the execution pipeline status with clock-level precision through the hardware-level probe unit to obtain inter-core communication delay information;

[0078] C2. Record the task scheduling queue and interrupt response events through the system-level tracing unit to obtain shared cache hit rate information;

[0079] C3. Construct a cross-core function call relationship graph through the application-level analysis unit to obtain power consumption distribution heat map information.

[0080] Specifically, the hardware-level probe unit is used to monitor the pipeline execution of the computing cores, record the delays of instruction execution and data transmission, and calculate the communication delays between different cores. The shared cache hit rate is a key indicator to measure data access efficiency. The system-level tracing unit is used to record the task scheduling queue, track the migration process of computing tasks on different cores, and analyze the shared cache hit rate in combination with interrupt response events. The application-level analysis unit constructs the flow path of computing tasks between different cores by analyzing the cross-core function call relationship, and generates a power consumption distribution heat map in combination with power consumption data to visually display high-power consumption computing areas.

[0081] S104. Quantitatively analyze the feature data through a multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters;

[0082] Specifically, the multi-dimensional evaluation model is used to comprehensively analyze the chip performance, quantify the key performance deviations in three dimensions of computing unit load, data flow, and energy efficiency. Among them, the performance deviation index is used to measure the deviation between the current operating state of the chip and the ideal state, and the architecture optimization suggestion parameters are the optimization schemes calculated according to the deviation.

[0083] In some alternative embodiments, the multi-dimensional evaluation model includes a computing unit utilization rate evaluation sub-model, a bandwidth saturation rate evaluation sub-model, and an energy efficiency ratio evaluation sub-model. The performance deviation index includes a load balance deviation, a data transportation deviation, and an energy efficiency deviation. Quantitatively analyze the feature data through the multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters, including:

[0084] D1. Analyze the power consumption distribution heat map information through the computing unit utilization rate evaluation sub-model to obtain a load balance deviation;

[0085] D2. Analyze the inter-core communication delay information and the shared cache hit rate information through the bandwidth saturation rate evaluation sub-model to obtain a data transportation deviation;

[0086] D3. Analyze the thermal information of the power consumption distribution through the energy efficiency ratio evaluation sub-model to obtain the energy efficiency deviation degree.

[0087] D4. Obtain the architecture optimization suggestion parameters according to the load balancing deviation degree, data transportation deviation degree, and energy efficiency deviation degree.

[0088] Specifically, according to the thermal information of the power consumption distribution, it can be judged whether the core load is balanced. If some cores are highly loaded for a long time while other cores are idle, a relatively high load balancing deviation degree can be calculated; according to the communication delay between cores and the shared cache hit rate, it can be evaluated whether the data transportation is efficient. If the shared cache hit rate is low and the communication delay between cores is high, indicating that a large amount of data transfer needs to pass through high-delay paths, a relatively high data transportation deviation degree can be calculated; according to the thermal information of the power consumption distribution, the energy utilization rate of the computing cores can be evaluated. If the power consumption of some cores is high but the load is low, a relatively high energy efficiency deviation degree can be calculated. Combining the above performance deviation degree indicators, architecture optimization suggestion parameters such as adjusting the task scheduling parameters of the computing cores, cache allocation strategy, and power consumption management strategy can be obtained.

[0089] In some alternative embodiments, the method further includes:

[0090] E1. Generate a visualization report according to the performance deviation degree indicators;

[0091] E2. Generate a verification meta data packet according to the visualization report.

[0092] Specifically, the visualization report includes a hot code annotation diagram, a resource contention analysis tree, an architecture improvement suggestion matrix, etc., which can intuitively display the content of the performance deviation degree analysis result. According to the visualization report, key test data can be extracted to generate a verification meta data packet that complies with the IEEE P2851 standard, ensuring the standardization, portability, and reusability of the test results.

[0093] In some alternative embodiments, the method further includes:

[0094] F1. Create a virtual test environment through virtualization technology;

[0095] F2. Execute test cases in the virtual test environment in shadow mode to obtain test data;

[0096] F3. Calibrate the deviation compensation coefficient between the test environment and the target application scenario according to the test data.

[0097] Specifically, virtualization technology can be used to build a controllable and repeatable virtual test environment to reproduce the chip architecture and the actual operating conditions of the computing core. In the virtual test environment, the test cases can be executed in parallel in shadow mode to compare the performance of different computing solutions, better evaluate the chip performance, calculate and apply the deviation compensation coefficient based on the test data, and calibrate the differences between the virtual test environment and the target application scenario, which can ensure the accuracy of the test results.

[0098] The chip design method of the present invention will be described below in conjunction with a specific embodiment.

[0099] In this embodiment, for the application scenario of real-time detection of cyber attacks in smart substations, a test environment is constructed and the security protection strategy is optimized to cope with the hybrid attack of False Data Injection Attack (FDIA) and DDoS (Distributed Denial of Service). First, the system determines the core type identification parameters according to the chip layout information of different computing cores to ensure the collaborative work of the security core, control core, and communication core. At the same time, the communication topology weight is adjusted to improve the priority of GOOSE (Generic Object Oriented Substation Event) messages, reduce the data occupancy of the MMS protocol, and set that when the CPU occupancy rate of the SCADA (Supervisory Control and Data Acquisition) system exceeds 75%, the attack mitigation mode is automatically triggered. Subsequently, a virtual test environment is built based on the IEC61850 protocol and test cases are generated, including FDIA attacks that simulate tampering with bus voltage measurement data and collaborative attacks that induce misoperation of protection devices when DDoS causes communication blockage. At the same time, a stress test scenario in which 20% of the smart meter nodes are controlled by a botnet is introduced as a sudden load disturbance factor to evaluate the load-bearing capacity of the system.

[0100] During the testing process, the system deploys an FPGA-level message verification module in the merging unit through hardware probes to monitor communication latency, and uses the system tracing unit to record the action logic of the protection device and the network traffic pattern to analyze attack behaviors. In addition, the application analysis unit constructs a cross-device data consistency map to identify abnormal measurement nodes. Based on the collected data, the attack feature deviation is calculated through a multi-dimensional evaluation model to detect abnormal fluctuations in the GOOSE message period and evaluate the effectiveness of the security protection strategy. Finally, the system proposes optimization suggestions according to the analysis results, such as adjusting the task scheduling period of the encryption core to reduce the bit error rate and improve data integrity. It can be recognized that this testing scheme can effectively optimize the network defense ability of intelligent substations and improve the detection and response speed to malicious attacks.

[0101] The chip design method of the present invention will be described below in conjunction with another specific embodiment.

[0102] In this embodiment, for the application scenario of dynamic security protection of a new energy high-penetration power grid, a test environment is constructed and the frequency modulation strategy is optimized to address the frequency instability problem caused by the fluctuations of photovoltaic and wind power. First, the system identifies parameters according to the chip layout information and core type of different computing cores to ensure the coordinated operation of the frequency modulation core, prediction core, and regulation core. At the same time, the communication weight parameter is adjusted, and the transmission priority of the PMU (Phasor Measurement Unit) data stream is set to the highest level (coefficient 0.9), and the emergency frequency modulation mode is automatically activated when the grid frequency deviation exceeds 0.5 Hz. Subsequently, a digital twin model of the power grid with a new energy proportion of 30% is constructed based on RT-LAB (Real-Time Laboratory), and test cases are generated, including a working condition where the simulated photovoltaic power generation drops suddenly by 80% within 5 seconds due to cloud occlusion, a disturbance scenario where the grid cascading fault propagation is caused by the disconnection of the wind turbine, and a stress test with the disordered charging of an electric vehicle cluster as the sudden load disturbance factor to evaluate the stability and regulation ability of the power grid.

[0103] During the test, the system monitors the harmonic changes under grid disturbances by deploying a high-speed harmonic component acquisition module (sampling rate 1MHz) for the energy storage converter through a hardware probe. At the same time, the system uses the system tracking unit to synchronously record the PMU phasor data of 12 nodes and the instruction response delay of the AGC (Power Conversion System, energy storage converter), and reconstructs the whole-network power angle stability situation map through the application analysis unit to locate potential unstable units. Based on the collected data, the evaluation model calculates the performance deviation degree, evaluates the response delay of the energy storage system, measures the safety margin of the power grid, analyzes the improvement range of voltage stability, and controls the frequency deviation within ±0.2Hz. Finally, the system proposes optimization suggestions according to the analysis results. For example, adjust the wind power prediction algorithm of the prediction kernel to increase the weight of the LSTM (Long Short-Term Memory) model by 30%, so as to improve the prediction accuracy and regulation ability. It can be recognized that this test scheme can effectively optimize the dynamic stability of the new energy high-penetration power grid, improve the power grid fault recovery speed, and enhance the new energy consumption capacity.

[0104] It can be recognized that in the embodiment of the present invention, by dynamically configuring the test parameter set and generating test cases covering heterogeneous cooperation, the coverage rate of the verification scenario for the target application scenario can be effectively improved; through the reconfigurable monitoring module and the multi-dimensional evaluation model to hierarchically collect test data, the full-dimensional tracking from the clock-level pipeline state to the system-level task scheduling can be realized, and the architecture optimization suggestion parameters with high reference value are output, so as to play a guiding role in the design and improvement of the chip and reduce the number of repeated verification iterations.

[0105] Referring to Figure 2 , the embodiment of the present invention provides a performance test and verification system for a multi-core heterogeneous chip, including:

[0106] A parameter configuration module, configured to configure a test parameter set according to the target application scenario, and the test parameter set includes processor core type identification parameters, communication topology weight parameters for communication, and load balancing threshold parameters;

[0107] A test case generation module, generating test cases covering the cooperative working mode of heterogeneous computing cores, and the test cases include CPU-GPU heterogeneous computing verification cases, DSP-accelerator data flow verification cases, and mixed-precision operation verification cases;

[0108] A monitoring module, configured to collect characteristic data during the test through a reconfigurable monitoring sub-module, and the characteristic data includes inter-core communication delay information, shared cache hit rate information, and power consumption distribution heat map information;

[0109] An analysis and evaluation module for quantitatively analyzing feature data through a multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters.

[0110] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0111] Referring to Figure 3 , an embodiment of the present invention provides a device, including:

[0112] At least one processor;

[0113] At least one memory for storing at least one program;

[0114] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the performance test and verification method of the above multi-core heterogeneous chip.

[0115] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the performance test and verification method of the above multi-core heterogeneous chip when executed by the processor.

[0116] A computer-readable storage medium according to an embodiment of the present invention can execute the performance test and verification method of the multi-core heterogeneous chip provided by the method embodiment of the present invention, can execute any combination of the implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0117] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the device executes Figure 1 The chip design method based on a multi-core heterogeneous architecture shown.

[0118] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order noted in the operational illustrations. For example, depending on the functionality / operation involved, two blocks shown in succession may actually be executed substantially concurrently or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are envisioned in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0119] In addition, although the present invention has been described in the context of functional modules, it should be understood that one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module unless otherwise stated to the contrary, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of such modules would be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0120] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product stored in a storage medium, including several instructions for causing a device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0121] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0122] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above-mentioned program can be printed, because the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, when necessary, other suitable processing, and then storing it in a computer memory.

[0123] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0124] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0125] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0126] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for performance testing and verification of a heterogeneous multi-core chip, characterized in that Including: Configuring a test parameter set according to a target application scenario, where the test parameter set includes processor core type identification parameters, communication topology weight parameters, and load balancing threshold parameters; Generating test cases covering the collaborative working mode of heterogeneous computing cores, where the test cases include CPU-GPU heterogeneous computing verification cases, DSP-accelerator data flow verification cases, and mixed-precision operation verification cases; Collecting feature data during testing through a reconfigurable monitoring sub-module, where the feature data includes inter-core communication delay information, shared cache hit rate information, and power consumption distribution heat map information; Performing quantitative analysis on the feature data through a multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters.

2. The performance testing and verification method of the multi-core heterogeneous chip according to claim 1, characterized in that The configuring the test parameter set according to the target application scenario includes: Determining the physical layout of computing cores, storage cores, and communication cores according to chip layout information, and configuring the processor core type identification parameters according to the physical layout; Adjusting the communication topology weight parameters and bus arbitration strategy according to historical test data; Constraining test boundary conditions according to a temperature-frequency coupling model and determining the load balancing threshold parameters.

3. The performance testing and verification method of the multi-core heterogeneous chip according to claim 1, characterized in that, The generating the test cases covering the collaborative working mode of heterogeneous computing cores includes: Generating a basic test vector through a randomization algorithm; Generating basic test cases according to the basic test vector and a typical application mode matching template; Generating the test cases according to the basic test cases and a sudden load disturbance factor.

4. The performance testing and verification method of the multi-core heterogeneous chip according to claim 1, characterized in that The reconfigurable monitoring sub-module includes a hardware-level probe unit, a system-level tracing unit, and an application-level analysis unit. The collecting the feature data during testing through the reconfigurable monitoring sub-module includes: Collecting the execution pipeline state with clock-level precision through the hardware-level probe unit to obtain the inter-core communication delay information; Recording the task scheduling queue and interrupt response events through the system-level tracing unit to obtain the shared cache hit rate information; Constructing a cross-core function call relationship graph through the application-level analysis unit to obtain the power consumption distribution heat map information.

5. The performance testing and verification method of the multi-core heterogeneous chip according to claim 1, characterized in that The multi-dimensional evaluation model includes a computing unit utilization rate evaluation sub-model, a bandwidth saturation rate evaluation sub-model, and an energy efficiency ratio evaluation sub-model. The performance deviation index includes a load balancing deviation, a data transportation deviation, and an energy efficiency deviation. The performing quantitative analysis on the feature data through the multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters includes: Analyzing the power consumption distribution heat map information through the computing unit utilization rate evaluation sub-model to obtain the load balancing deviation; Analyzing the inter-core communication delay information and the shared cache hit rate information through the bandwidth saturation rate evaluation sub-model to obtain the data transportation deviation; Analyzing the power consumption distribution heat map information through the energy efficiency ratio evaluation sub-model to obtain the energy efficiency deviation; Obtaining the architecture optimization suggestion parameters according to the load balancing deviation, the data transportation deviation, and the energy efficiency deviation.

6. The performance testing and verification method for the multi-core heterogeneous chip according to claim 1, characterized in that The method further includes: Generating the visualization report according to the performance deviation index; Generating a verification meta data packet according to the visualization report.

7. The performance testing and verification method for the multi-core heterogeneous chip according to claim 1, characterized in that The method further includes: creating a virtual test environment through virtualization technology; executing the test case in the virtual test environment in shadow mode to obtain test data; calibrating the deviation compensation coefficient between the test environment and the target application scenario according to the test data.

8. A performance testing and verification system for a multi-core heterogeneous chip, characterized in that, It includes: a parameter configuration module, configured to configure a test parameter set according to a target application scenario, where the test parameter set includes a processor core type identification parameter, a communication topology weight parameter, and a load balancing threshold parameter; a test case generation module, generating test cases covering the collaborative working mode of heterogeneous computing cores, where the test cases include CPU-GPU heterogeneous computing verification cases, DSP-accelerator data flow verification cases, and mixed-precision operation verification cases; a monitoring module, configured to collect feature data during testing through a reconfigurable monitoring sub-module, where the feature data includes inter-core communication delay information, shared cache hit rate information, and power consumption distribution heat map information; an analysis and evaluation module, configured to perform quantitative analysis on the feature data through a multi-dimensional evaluation model to obtain a performance deviation index and architecture optimization suggestion parameters.

9. A device, characterized in that, It includes: at least one processor; at least one memory, configured to store at least one program; when the at least one program is executed by the at least one processor, enabling the at least one processor to implement the performance test and verification method of the multi-core heterogeneous chip according to any one of claims 1-7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute the performance test and verification method of the multi-core heterogeneous chip according to any one of claims 1-7.

Citation Information

Cited By

  • Processor test method, device, medium and program product

    CN120670239A

  • Chip prototype verification method and device, medium and product

    CN120723558A

  • Automatic testing method and platform system for computing performance of domain control chip

    CN120973608A

  • Parameter adjustment method and device of test software, equipment, medium and program product

    CN121349901A

  • Adaptive evaluation method and device for core demand of control chip based on vehicle application scene

    CN121681393A