Performance detection method and device, equipment and storage medium
By simulating the collaborative performance detection of AI software architecture and chip architecture, the software and hardware collaboration problems in ultra-large-scale processor design are solved, early problem discovery and resource optimization are achieved, and the design efficiency and performance of AI products are improved.
Patent Information
- Application Number
- CN202510396890.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-25
AI Technical Summary
In ultra-large-scale processor design, it is difficult to achieve reasonable coordination between software and hardware, resulting in high design indicators but far below expectations, and lack of effective performance detection methods, resulting in waste of resources and declining market competitiveness.
Provide a performance detection method, by simulating the initial AI software architecture and the initial AI chip architecture to run target tasks in different simulation scenarios, obtain target performance results, and obtain collaborative detection results based on this, identify collaborative problems, and guide the adjustment of software and hardware.
In the early stages of the project, software and hardware collaboration problems were discovered, design efficiency was improved, resource waste wasted, and the overall performance and market competitiveness of AI products were improved.
Smart Images

Figure CN120373224A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to technologies such as artificial intelligence, large models, and chips. Background Art
[0002] In the design process of ultra-large-scale processors (such as artificial intelligence (AI) chips), the current challenge is that it is difficult to truly achieve reasonable software-hardware cooperation, and there may even be a problem that the design indicators of the processor are high but the actual performance is far from the expectation. Therefore, there is an urgent need for a method for detecting software-hardware cooperation performance. Summary of the Invention
[0003] The present disclosure provides a performance detection method, device, equipment, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a performance detection method, including:
[0005] Determine an initial artificial intelligence (AI) software architecture and an initial AI chip architecture for simulating the operation of the initial AI software architecture;
[0006] Based on the initial AI software architecture and the initial AI chip architecture, simulate the operation of a target task to obtain a target performance result, where the target performance result includes performance data of the initial AI chip architecture under different simulation scenarios;
[0007] Based on the target performance result, obtain a collaborative detection result for the initial AI software architecture and the initial AI chip architecture, where the collaborative detection result characterizes the collaborative performance of the initial AI software and the initial AI chip when working together.
[0008] According to another aspect of the present disclosure, there is provided a performance detection device, including:
[0009] A data determination unit for determining an initial artificial intelligence (AI) software architecture and an initial AI chip architecture for simulating the operation of the initial AI software architecture;
[0010] A task operation unit for simulating the operation of a target task based on the initial AI software architecture and the initial AI chip architecture to obtain a target performance result, where the target performance result includes performance data of the initial AI chip architecture under different simulation scenarios;
[0011] A performance detection unit for obtaining a collaborative detection result for the initial AI software architecture and the initial AI chip architecture based on the target performance result, where the collaborative detection result characterizes the collaborative performance of the initial AI software and the initial AI chip when working together.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any of the methods in the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any of the methods in the embodiments of the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a computer program product, comprising a computer program which, when executed by a processor, implements any of the methods in the embodiments of the present disclosure.
[0018] The solution of the present disclosure can simulate the operation of a target task in different simulation scenarios based on an initial AI software architecture and an initial AI chip architecture to obtain a target performance result, and further obtain a co-detection result of the initial AI software architecture and the initial AI chip architecture. In other words, the solution of the present disclosure can simulate the co-performance of the initial AI software architecture and the initial AI chip architecture when working together. Therefore, it provides data support for discovering possible co-problems between software and hardware, facilitates timely adjustment of software and / or hardware, effectively improves the design efficiency of AI products, and at the same time, avoids waste of resources caused by poor software-hardware cooperation.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0021] Figure 1 is a schematic flowchart of a performance detection method according to an embodiment of the present application Figure 1 ;
[0022] Figure 2 is a schematic flowchart of a performance detection method according to an embodiment of the present application Figure 2 ;
[0023] Figure 3It is a schematic flowchart of software and hardware performance collaborative optimization using a software and hardware collaborative performance model according to an embodiment of the present disclosure;
[0024] Figure 4 It is a schematic diagram of the hardware abstraction features of each model according to an embodiment of the present disclosure;
[0025] Figure 5 It is a schematic diagram of a software and hardware collaborative optimization framework according to an embodiment of the present disclosure;
[0026] Figure 6 It is a structural schematic of a performance detection device 600 according to an embodiment of the present disclosure Figure 1 ;
[0027] Figure 7 It is a structural schematic of a performance detection device 600 according to an embodiment of the present disclosure Figure 2 ;
[0028] Figure 8 It shows a schematic block diagram of an example electronic device 800 that can be used to implement the embodiments of the present disclosure. Detailed Embodiments
[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0030] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C. The terms "first" and "second" in this article represent referring to multiple similar technical terms and distinguishing them, and do not mean limiting the order, or limiting to only two. For example, the first feature and the second feature refer to two types / two features. The first feature can be one or more, and the second feature can also be one or more.
[0031] In addition, to better illustrate the present disclosure, numerous specific details are provided in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0032] The related technologies of the embodiments of the present disclosure are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as optional solutions, and they all fall within the protection scope of the embodiments of the present disclosure.
[0033] In the practice of contemporary very-large-scale processor design, achieving reasonable software-hardware co-design faces a series of challenges. Specifically, there are the following pain points:
[0034] (1) Since it is impossible to optimize software from a global perspective at the beginning of the design, it is difficult to provide clear optimization direction guidance for the processor;
[0035] (2) How to plan the overall performance of the processor at the beginning of the design, and how to find the design balance among various hardware resources;
[0036] (3) Before the processor design is completed, it is often impossible to accurately judge the degree of fit between the architecture design and the current software, and it is difficult to reasonably estimate key indicators such as the theoretical computing power of the processor in the initial stage of design. As a result, although high standards are set for the computing power performance, the final performance is far from the expectation.
[0037] For example, in the field of AI chip design, with the rapid development of algorithms and applications, the performance requirements for chips are increasing day by day. However, before the final design of the AI chip is completed, it is difficult to accurately detect its overall performance when working with AI software. This is not only because of the diversity and complexity of AI applications, making prediction and optimization difficult, but also because of the tight coupling relationship between software and hardware, making any minor change in either party may have a significant impact on the overall performance.
[0038] Furthermore, due to the high complexity and long cycle of AI chip design, once it is found that the performance does not meet the standard in the later stage of design, it will bring huge time and cost losses. More seriously, even if the design indicators seem high, due to the lack of effective software-hardware performance co-detection methods, the actual running performance is often much lower than expected. This not only leads to waste of resources, but also seriously affects the market competitiveness of AI chips.
[0039] Based on this, the solution of the present disclosure provides a performance detection method, which can detect the co-performance of software and hardware when working together, and thus provides strong support for realizing software-hardware co-design.
[0040] Specifically, Figure 1 FIG. is a schematic flowchart of a performance detection method according to an embodiment of the present application. Figure 1 This method is optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0041] Furthermore, this method at least includes at least part of the following content. As Figure 1 shown, it includes:
[0042] Step S101: Determine the initial artificial intelligence (AI) software architecture and the initial AI chip architecture for running the initial AI software architecture.
[0043] Step S102: Based on the initial AI software architecture and the initial AI chip architecture, simulate running the target task to obtain the target performance result.
[0044] Here, the target performance result includes the performance data of the initial AI chip architecture under different simulation scenarios.
[0045] Furthermore, in an example, the performance data of the initial AI chip architecture may include multiple metrics such as data processing speed and resource utilization rate. The specific performance metrics are not limited in the present disclosure solution.
[0046] In an example, the target task is a task that the initial AI software architecture can execute, such as, specifically, AI tasks such as intelligent question answering, intelligent reasoning, and intelligent image generation. The specific tasks are not limited in the present disclosure solution.
[0047] Step S103: Based on the target performance result, obtain the collaborative detection result (also referred to as the collaborative performance detection result) for the initial AI software architecture and the initial AI chip architecture.
[0048] Here, the collaborative detection result characterizes the collaborative performance of the initial AI software architecture and the initial AI chip architecture when working together.
[0049] In this way, the present disclosure solution can simulate running the target task under different simulation scenarios based on the initial AI software architecture and the initial AI chip architecture to obtain the target performance result, and then obtain the collaborative detection result of the initial AI software architecture and the initial AI chip architecture. In other words, the present disclosure solution can simulate and obtain the collaborative performance of the initial AI software architecture and the initial AI chip architecture when working together. Therefore, it provides data support for discovering possible collaborative problems between software and hardware, facilitates timely adjustment of software and / or hardware, effectively improves the design efficiency of AI products, and at the same time, avoids resource waste caused by poor software-hardware collaboration.
[0050] In addition, the disclosed solution can perform simulation runs and collaborative detection in the early stage of a project, enabling the discovery of performance collaboration issues earlier, facilitating adjustments to the software and hardware designs. Compared with existing performance testing, the disclosed solution does not rely on detection data in the later stage of the project, providing greater space and time for collaborative performance optimization of software and hardware architectures. Moreover, the disclosed solution can provide more realistic simulation data, which can be used not only for exploring and evaluating the hardware architecture but also for optimizing and guiding the software layer (i.e., the application layer).
[0051] It should be noted that the above-mentioned initial AI software architecture can be specifically a model (such as, a large model, etc.), an operator, etc. In other words, the disclosed solution can be specifically applied to at least one of the following: the collaborative design scenario of a model (such as, a large model, etc.) and an AI chip architecture, the collaborative design scenario of an operator and an AI chip architecture, the performance optimization scenario of a large model, the optimization of an AI chip architecture, etc. In other words, the disclosed solution can provide multi-dimensional optimization guidance, realizing software and hardware collaborative design with a global perspective at the architecture design stage, facilitating the on-demand abstraction of the hardware level during the continuous deepening of the design, and achieving the on-demand balance of simulation speed and simulation accuracy.
[0052] For example, in one example, the AI chip architecture can specifically adopt forms of multi-chip interconnection such as Die-to-Die (D2D), Chip-to-Chip (C2C), etc. This multi-chip interconnected AI chip architecture, through an efficient inter-chip communication and data transmission mechanism, can support more complex and larger-scale AI computing tasks, meeting diverse application requirements such as high performance and low power consumption. In other words, the disclosed solution can also provide a more complete performance evaluation solution for contemporary AI large model training chips.
[0053] In addition, the disclosed solution can also be used for optimizing compilers (for code compilation of the initial AI software architecture). For example, based on the collaborative performance detection results, analyze and utilize the code structure and corresponding characteristics of the initial AI software architecture obtained by the editor to optimize the code, such as adjusting the execution order of the code to reduce data dependencies, or performing loop unrolling or loop merging, etc., to improve the execution efficiency of the loop body. By optimizing the code of the initial AI software architecture, the performance of the compilation result can be improved, thereby accelerating the execution speed of the software and reducing resource consumption.
[0054] It should be noted that, in one example, a scripting language such as Python can be used as the development language for the present disclosure solution, and coroutine behavior can be implemented through a generator to achieve a multi-task concurrent execution framework in hardware modeling. In this way, the collaborative performance detection between the initial AI software architecture and the initial AI chip architecture can be achieved.
[0055] Compared with the cycle-accurate model based on C++, the present disclosure solution uses the Python language and can shorten the development cycle in the performance detection method, accelerating the iteration process of the software and hardware architectures. Moreover, Python provides rich third-party libraries in data analysis and visualization, facilitating data processing and performance analysis, and making the exploration of software and hardware architectures and data demonstration more efficient.
[0056] In addition, Python shows higher simplicity and convenience in cross-platform development and prototype verification compared with other programming languages. For projects that require frequent and rapid iteration, the interpreted execution method of Python provides significant advantages. There is no need to compile into machine code for a specific platform in advance like C++, thus promoting the flexibility and efficiency of the development process.
[0057] In summary, using Python syntax has the following advantages:
[0058] First, Python syntax is concise. The amount of code required to implement the same function is less than that of other languages. Moreover, using the Python language greatly shortens the development cycle and facilitates code maintenance. It meets the requirements for the rapid implementation of architectural innovation and performance exploration at the beginning of the design, and has faster iteration of performance feedback.
[0059] Second, coroutines are implemented through Python generators to achieve lightweight concurrency. In addition, the lazy evaluation mode of Python generators improves the response performance of the system-level model (program) (generating values only when needed), reducing the development cost while improving the running simulation efficiency.
[0060] Third, using the Python language realizes the integration of a complete architecture modeling flow and performance analysis flow, greatly simplifying the steps of architecture exploration and data demonstration, enabling developers to devote more energy to the architecture exploration itself (such as hardware architecture, software architecture, etc.), and quickly iterating an efficient architecture for software and hardware collaboration.
[0061] Further, in a specific example, the following method can be used to obtain the co-detection result for the initial AI software architecture and the initial AI chip architecture; specifically, obtaining the co-detection result for the initial AI software architecture and the initial AI chip architecture based on the target performance result (such as step S103) can specifically include:
[0062] Step S103-1: Based on the target performance result, obtain the operating performance of each hardware unit in the initial AI chip architecture, and obtain the performance bottlenecks of each hardware unit in the initial AI chip architecture.
[0063] Step S103-2: Based on the operating performance of each hardware unit in the initial AI chip architecture and the performance bottlenecks of each hardware unit in the initial AI chip architecture, obtain the co-detection result.
[0064] Here, the solution of the present disclosure can use performance counters to obtain the operating data of each hardware unit (such as processor cores, cache systems, memory controllers, input / output interfaces, etc.) in the initial AI chip architecture during operation. For example, use performance counters to record performance data such as the workload, processing speed, and latency time of each hardware unit in the initial AI chip architecture.
[0065] For example, in an example, based on the target performance result and using performance counters, record the workload of each hardware unit in the initial AI chip architecture when simulating at least some subtasks of the target task, or, alternatively, record the running speed, etc. In this way, the actual operating conditions of each hardware unit can be known, which is convenient for identifying problems such as unbalanced resource utilization, providing data support for design optimization.
[0066] It should be noted that in practical applications, the required performance metrics can be selected based on design requirements, and the present disclosure does not limit specific performance metrics either.
[0067] Further, in another example, the solution of the present disclosure can use performance analysis tools to analyze and accurately identify the performance bottlenecks encountered by each hardware unit in the initial AI chip architecture during operation. These performance bottlenecks may manifest as insufficient processing speed, unbalanced resource occupancy rate, or abnormal power consumption increase, etc.
[0068] It should be noted that in practical applications, the performance bottlenecks to be analyzed can be determined based on design requirements, and the present disclosure does not limit specific bottleneck metrics either.
[0069] Further, after obtaining the running performance of each hardware unit in the initial AI chip architecture and the performance bottlenecks of each hardware unit, a collaborative detection result that can characterize the collaborative performance between the initial AI software architecture and the initial AI chip architecture can be obtained. In this way, it is convenient to reveal possible collaborative problems between the initial AI software architecture and the initial AI chip architecture through the collaborative detection result, such as problems that the software algorithm fails to fully utilize the performance of the hardware acceleration unit, and the memory management strategy leads to frequent cache misses, etc., providing data support for subsequent software and hardware optimization.
[0070] In this way, the proposed solution of the present disclosure utilizes the obtained target performance results of the simulation, and can not only obtain the running performance and performance bottlenecks of each hardware unit in the initial AI chip architecture, but also obtain the collaborative performance between the initial AI software architecture and the initial AI chip architecture. In this way, it provides data support for obtaining the optimization direction and optimization strategy of the initial AI software architecture and / or the initial AI chip architecture, and further lays a foundation for jointly promoting the improvement of the performance of the AI system.
[0071] Further, in a specific example, in order to implement the performance detection method proposed in the solution of the present disclosure, it is also necessary to determine the target task before performance simulation. Specifically, the following steps can be used to obtain the target task:
[0072] Obtain the software feature data of the initial AI software architecture, and obtain the architecture feature data of the initial AI chip architecture;
[0073] Generate a target task that can detect the collaborative performance between the initial AI software architecture and the initial AI chip architecture based on the software feature data and the architecture feature data.
[0074] Here, in an example, the software feature data of the initial AI software architecture may include execution information such as the start time, end time, and execution order of each abstraction layer (such as Grid, Group, Block, Warp, Thread, etc.) in the initial AI software architecture.
[0075] It should be noted that the software feature data of the initial AI software architecture can be obtained after analysis based on the initial software data of the initial AI software architecture (such as trace data). For example, in one example, the initial software data of the initial AI software architecture can be obtained by using a behavior model. Here, the behavior model focuses on algorithm and functional completeness detection, and has characteristics such as a high level of abstraction, concise code volume, and fast running speed. Therefore, the initial software data with rich data dimensions can be quickly obtained by using this behavior model. Further, the functional model can be used to analyze the initial software data obtained by the behavior model, and then the software feature data can be obtained. Here, the functional model can focus on signal integrity detection, and then obtain the software feature data after performing signal integrity detection on the initial software data.
[0076] In another example, the architecture feature data of the initial AI chip architecture may include the number of processing cores, clock speed (i.e., the basic frequency at which the AI chip operates), cache size, memory bandwidth (i.e., the speed at which the AI chip accesses the main memory and the data transfer capacity), etc.
[0077] It should be noted that in practical applications, the required software feature data and architecture feature data can be determined based on design requirements, and the present disclosure solution does not make specific limitations in this regard.
[0078] Further, based on the software feature data and architecture feature data, a target task for detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture can be generated.
[0079] It can be understood that in practical applications, the target task may include processing of data sets of different scales, diverse computationally intensive tasks, etc., to ensure the comprehensiveness and accuracy of the test results.
[0080] In this way, based on the obtained feature data, such as software feature data and architecture feature data, a target task matching the initial AI software architecture and the initial AI chip architecture can be generated. Moreover, since the target task is obtained based on the software feature data and architecture feature data, in other words, it is obtained after fully considering the characteristics of the initial AI software architecture and the initial AI chip architecture, it provides support for effectively detecting the performance of the initial AI software architecture and the initial AI chip architecture when working together subsequently, and thus lays a foundation for discovering potential optimization points for the initial AI software architecture and / or the initial AI chip architecture.
[0081] Further, in a specific example, after the collaborative detection result for the initial AI software architecture and the initial AI chip architecture, the method further includes:
[0082] Based on the collaborative detection result, adjust the initial AI software architecture and / or the initial AI chip hardware.
[0083] For example, after analyzing the collaborative detection results, targeted fine-tuning can be performed on the initial AI software architecture and / or the initial AI chip architecture.
[0084] Specifically, in one example, the adjustment of the initial AI software architecture may include algorithm optimization, adjustment of the data processing flow, improvement of the memory management strategy, etc. These adjustments aim to improve the execution efficiency of the software, reduce unnecessary resource consumption, and make full use of the computing power provided by the hardware.
[0085] Furthermore, in another example, the adjustment of the initial AI chip architecture may include configuration optimization of the processing core, adjustment of the cache hierarchy, expansion of the memory bandwidth, etc. These hardware-level adjustments aim to match the software workload and improve the overall performance of the software-hardware collaboration.
[0086] In this way, fine-tuning the software architecture and / or the hardware architecture based on the collaborative detection results to optimize the performance metrics found in the collaborative detection provides strong support for improving the overall performance of the AI system. In addition, since the solution of the present disclosure can provide data support from a global perspective in the design stage, it provides guidance for the optimization direction of the software architecture and / or the hardware architecture, and thus provides strong support for more reasonable allocation and utilization of resources and avoidance of resource waste.
[0087] Specifically, Figure 2 is a schematic flowchart of a performance detection method according to an embodiment of the present application. Figure 2 This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, and other electronic devices. It can be understood that the relevant content of the method shown above Figure 1 can also be applied to this example, and the relevant associated content will not be elaborated in this example.
[0088] Furthermore, this method at least includes at least part of the following content. As Figure 2 shown, it includes:
[0089] Step S201: Determine the initial artificial intelligence (AI) software architecture and the initial AI chip architecture for running the initial AI software architecture.
[0090] Step S202: Based on the initial AI software architecture and the initial AI chip architecture, simulate the operation of the target task to obtain at least two of the following: the first performance data of the initial AI chip architecture in the first simulation scenario, the second performance data of the initial AI chip architecture in the second simulation scenario, and the third performance data of the initial AI chip architecture in the third simulation scenario.
[0091] Here, the simulation granularities of the first simulation scenario, the second simulation scenario, and the third simulation scenario are different from each other.
[0092] Here, it should be noted that the simulation granularity can refer to the degree of detail and complexity of the real-world entities represented by each independent unit or component in the simulation scenario. For example, in one example, the simulation granularity can specifically represent the degree of proximity between the environment where the hardware unit in the simulation scenario is located and the preset conditions (such as actual limiting conditions, or what can be called the actual environment or real environment). In other words, the degrees of proximity between the simulation environments where the hardware units simulated in the first simulation scenario, the second simulation scenario, and the third simulation scenario are located and the real environment are different. That is to say, the corresponding abstraction levels of the first simulation scenario, the second simulation scenario, and the third simulation scenario are different. Here, the simulation environments where the hardware units are located at different abstraction levels are different. In this way, by constructing different simulation environments, performance data under different simulation environments can be obtained.
[0093] For example, in one example, the target task can be simulated and run in the first simulation scenario to obtain first performance data, and the target task can be simulated and run in the second simulation scenario to obtain second performance data. At this time, the first performance data and the second performance data are the target performance results.
[0094] In another example, the target task can be simulated and run in the first simulation scenario to obtain first performance data, and the target task can be simulated and run in the third simulation scenario to obtain third performance data. At this time, the first performance data and the third performance data are the target performance results.
[0095] In another example, the target task can be simulated and run in the first simulation scenario to obtain first performance data, the target task can be simulated and run in the second simulation scenario to obtain second performance data, and the target task can be simulated and run in the third simulation scenario to obtain third performance data. At this time, the first performance data, the second performance data, and the third performance data are the target performance results.
[0096] Here, the relevant descriptions regarding the initial AI software architecture, the initial AI chip architecture, and the target task can refer to the above descriptions and will not be elaborated here.
[0097] Step S203: Based on the target performance results, obtain the collaborative detection results for the initial AI software architecture and the initial AI chip architecture.
[0098] It should be noted that relevant examples of the collaborative detection results can be seen in the above descriptions and will not be elaborated here.
[0099] In this way, since the target task is simulated and run under different simulation scenarios, the running performance of the initial AI chip architecture under different conditions can be comprehensively detected. And since the target task is simulated and run based on the initial AI software architecture deployed on the initial AI chip architecture, the running performance of the initial AI software architecture under different conditions can also be comprehensively detected while detecting the running performance of the initial AI chip architecture under different conditions, thereby detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture. Further, by analyzing the collaborative performance, the bottleneck of the collaborative performance between the initial AI software architecture and the initial AI chip architecture can be effectively identified, thus guiding subsequent software and hardware optimization and design improvement to enhance the overall performance of the AI product.
[0100] In addition, compared with the architecture evaluation scheme guided by the performance test of the previous generation, the proposed solution in this disclosure provides an optimization scheme for the software and hardware collaborative architecture from a global perspective, avoiding local optimization limited to modules or within self-modules, and thus providing strong support for achieving comprehensive performance improvement of the AI product from the whole to the part.
[0101] Further, in a specific example, in order to obtain the first performance data of the initial AI chip architecture, the following method can also be used to construct the first simulation scenario. Specifically, the first simulation scenario can be obtained through the following steps:
[0102] Utilize the hardware abstraction features of the first simulation model to construct the first simulation scenario.
[0103] Here, the first simulation scenario is used to simulate the running scenarios of each hardware unit in the initial AI chip architecture under theoretical limit conditions (or can be called high abstraction levels).
[0104] Here, in one example, the first simulation model, for example, can be specifically a performance analysis model, which can be used to define the abstraction level of the hardware, so as to simulate the running environment (such as architecture parameters, data processing flow, etc.) of each hardware unit in the initial AI chip architecture.
[0105] Further, in this example, the first simulation scenario is a specific environment simulated and constructed by the first simulation model. Under this first simulation scenario, the running conditions of each hardware unit in the initial AI chip architecture under theoretical limit conditions or can be called high abstraction levels are simulated. For example, the running conditions of each hardware unit in the initial AI chip architecture are simulated under the ideal condition of not considering actual manufacturing deviations, physical limitations or external environmental interferences, so as to obtain the first performance data.
[0106] In this way, the present disclosure provides a refined solution for constructing a first simulation environment to simulate the operation of each hardware unit in the initial AI chip architecture under theoretical constraint conditions, thus providing strong support for efficiently obtaining the first performance data of the initial AI chip architecture.
[0107] Further, in a specific example, the following method can be used to obtain the first performance data of the initial AI chip architecture; specifically, by simulating the running of a target task to obtain the first performance data, it can specifically include:
[0108] Obtain the performance data of at least some subtasks of the target task when each hardware unit in the initial AI chip architecture runs under the theoretical architecture parameters corresponding to the first simulation scenario;
[0109] Based on the performance data of each hardware unit, obtain the first performance data.
[0110] Here, it should be noted that in this example, the simulation of theoretical constraint conditions can be achieved by setting or simulating theoretical architecture parameters. This simulation method is simple and efficient, and at the same time, it is convenient to adjust the simulation scenario, providing strong support for the iteration and optimization of subsequent software and hardware co-design.
[0111] Further, in this example, in the first simulation scenario, the specific tasks that the hardware units in the initial AI chip architecture need to execute can also be specifically simulated. For example, in an example, the target task can be decomposed into a series of subtasks. At this time, at least some subtasks that each hardware unit needs to execute can be simulated, so as to observe and analyze the performance data of each hardware unit, and then obtain the first hardware performance-to-price ratio.
[0112] In this way, based on the performance data of each hardware unit when processing subtasks, the first performance data of the initial AI chip architecture is obtained. This first performance data reflects the ideal performance level of the initial AI chip architecture under theoretical constraint conditions, thus providing reference information for subsequent software and hardware optimization design or performance tuning.
[0113] Further, in a specific example, in order to obtain the second performance data of the initial AI chip architecture, the following method can also be used to construct a second simulation scenario. Specifically, the following steps can be used to obtain the second simulation scenario:
[0114] Utilize the hardware abstraction features of the second simulation model to construct the second simulation scenario.
[0115] Here, the second simulation scenario is used to simulate the operation scenario of each hardware unit in the initial AI chip architecture under mixed constraint conditions (or can be called mixed abstraction levels).
[0116] Among them, the mixed constraint condition is a constraint condition between the theoretical constraint condition and the actual constraint condition. That is to say, the mixed constraint condition is the constraint condition corresponding to the mixed abstraction level, where the mixed abstraction level is an abstraction level between the high abstraction level and the low abstraction level.
[0117] Here, in one example, the second simulation model, for example, can be specifically a mixed-precision model, which can be used to define the abstraction level of the hardware, so as to simulate the operating environment (such as architecture parameters, data processing flow, etc.) of each hardware unit in the initial AI chip architecture.
[0118] Furthermore, in this example, the second simulation scenario is a specific environment simulated and constructed by the second simulation model. Under this second simulation scenario, the operating conditions of each hardware unit in the initial AI chip architecture under the mixed constraint condition are simulated. For example, the operating conditions of each hardware unit in the initial AI chip architecture are simulated in a state between the ideal state and the real state, so as to obtain the second performance data.
[0119] Here, it should be noted that the mixed constraint condition can be adjusted based on the simulation requirements. For example, it is closer to the actual constraint condition, or in the middle position between the ideal constraint condition and the actual constraint condition, etc. The solution of the present disclosure does not make specific limitations on this.
[0120] In this way, the solution of the present disclosure provides a refined solution for constructing the second simulation environment, so as to use this second simulation environment to simulate the operating conditions of each hardware unit in the initial AI chip architecture under the mixed constraint condition, thus providing strong support for efficiently obtaining the second performance data of the initial AI chip architecture.
[0121] Furthermore, in a specific example, the following method can be used to obtain the second performance data of the initial AI chip architecture; specifically, the target task is simulated to obtain the second performance data, which can specifically include:
[0122] Obtain the performance data of at least some sub-tasks of the target task when at least some hardware units in the initial AI chip architecture operate under the multi-task concurrency corresponding to the mixed constraint condition;
[0123] Based on the performance data of each hardware unit, obtain the second performance data.
[0124] Here, it should be noted that in this example, the mixed constraint condition is a non-ideal condition. At this time, the mixed constraint condition is used to simulate the operating conditions of the hardware unit under multi-concurrent tasks, so as to obtain the second performance data of the initial AI chip architecture. This simulation method is simple and efficient. At the same time, it is convenient to adjust the task concurrency, providing strong support for the iteration and optimization of subsequent software and hardware co-design.
[0125] Further, in this example, in the second simulation scenario, the specific tasks that the hardware units in the initial AI chip architecture need to execute can also be specifically simulated. For example, in one example, the target task can be decomposed into a series of subtasks. At this time, at least some of the subtasks that each hardware unit needs to execute under concurrent tasks can be simulated, so as to observe and analyze the performance data of each hardware unit, and further obtain the second performance data.
[0126] In this way, based on the performance data of each hardware unit when processing subtasks, the second performance data of the initial AI chip architecture is obtained. This second performance data reflects the performance level of the initial AI chip architecture under multi-concurrent tasks in the mixed constraint conditions. Thus, it provides reference information for subsequent software and hardware optimization design or performance tuning.
[0127] Further, in a specific example, in order to obtain the third performance data of the initial AI chip architecture, the following method can also be used to construct the third simulation scenario. Specifically, the following steps can be used to obtain the third simulation scenario:
[0128] Utilize the hardware abstraction characteristics of the third simulation model to construct the third simulation scenario.
[0129] Here, the third simulation scenario is used to simulate the operating scenarios of each hardware unit in the initial AI chip architecture under actual constraint conditions (or can be called the low abstraction level).
[0130] Here, in one example, the third simulation model, for example, can be specifically a cycle-accurate model, which can be used to define the abstraction level of the hardware, so as to simulate the operating environment (such as architecture parameters, data processing flow, etc.) of each hardware unit in the initial AI chip architecture.
[0131] Further, in this example, the third simulation scenario is a specific environment simulated and constructed by the third simulation model. Under this third simulation scenario, the operating conditions of each hardware unit in the initial AI chip architecture under real conditions are simulated. For example, the operating conditions of each hardware unit in the initial AI chip architecture are simulated under the conditions of considering actual manufacturing deviations, physical limitations, or external environmental interferences. Thus, the third performance data is obtained.
[0132] In this way, the solution of the present disclosure provides a refined solution for constructing the third simulation environment to utilize this third simulation environment to simulate the operating conditions of each hardware unit in the initial AI chip architecture under actual constraint conditions (also called real constraint conditions). Thus, it provides strong support for efficiently obtaining the third performance data of the initial AI chip architecture.
[0133] Further, in a specific example, the following method can be used to obtain the third performance data of the initial AI chip architecture; specifically, simulate the operation of the target task to obtain the third performance data, which can specifically include:
[0134] Obtain the performance data of at least some of the subtasks in the target task when at least some of the hardware units in the initial AI chip architecture operate under the multi-task concurrency corresponding to the actual limiting conditions;
[0135] Based on the performance data of each hardware unit, obtain the third performance data.
[0136] Here, it should be noted that in this example, the multi-task concurrency corresponding to the actual limiting conditions is different from the multi-task concurrency simulated by the mixed limiting conditions. It can be understood that the multi-task concurrency simulated by the mixed limiting conditions tends to be an ideal state, or it can be said that the multi-task concurrency simulated by the mixed limiting conditions is weaker than the multi-task concurrency corresponding to the actual limiting conditions. In this way, it is convenient to obtain the performance data at different abstraction levels, providing strong support for the iteration and optimization of subsequent software and hardware co-design.
[0137] Furthermore, in this example, in the third simulation scenario, the specific tasks that the hardware units in the initial AI chip architecture need to execute can also be specifically simulated. For example, in one example, the target task can be decomposed into a series of subtasks. At this time, at least some of the subtasks that each hardware unit needs to execute under concurrent tasks can be simulated, so as to observe and analyze the performance data of each hardware unit, and then obtain the third performance data.
[0138] It should be noted that in this example, the simulation granularity of the theoretical limiting conditions, the mixed limiting conditions, and the actual limiting conditions gradually becomes finer. In this way, it is convenient to obtain the performance data at different simulation granularities, providing strong support for the iteration and optimization of subsequent software and hardware co-design.
[0139] In this way, based on the performance data of each hardware unit when processing subtasks, the third performance data of the initial AI chip architecture is obtained. This third performance data reflects the true performance level of the initial AI chip architecture under multi-concurrent tasks in the real environment. In this way, it provides reference information for subsequent software and hardware optimization design or performance tuning.
[0140] The following will further elaborate on the present disclosure solution with specific examples. Specifically, to address the pain points in existing architecture exploration / performance analysis, the present disclosure solution provides a fast and lightweight architecture or micro-architecture modeling solution based on the Python language. This solution can provide a fast, easy-to-use (easy to configure / easy to modify), reliable (simulation), and trustworthy software-hardware co-design performance model for the team at the beginning of the design. This software-hardware co-design performance model integrates a performance analysis model, a mixed-precision model, and a cycle-accurate model to simulate and model each major hardware unit and conduct corresponding analysis, so as to provide high-precision and global perspective analysis results while ensuring simulation efficiency, thereby providing strong support for the iteration and optimization of software-hardware co-design.
[0141] Figure 3 is a schematic flowchart of software-hardware performance co-optimization using a software-hardware co-design performance model according to an embodiment of the present disclosure. As Figure 3 shown, the software-hardware co-optimization method may include the following steps:
[0142] S301: Obtain the initial software data of the initial AI software architecture.
[0143] Here, in one example, the initial software data may be log data (also referred to as trace data) generated during the operation of the initial AI software architecture. For example, in one example, the initial software data of the initial AI software architecture can be obtained using a behavior model.
[0144] S302: Determine the software feature data of the initial AI software architecture from the initial software data.
[0145] For example, in one example, a functional model can be used to analyze the initial software data obtained by the behavior model, and then the software feature data of the initial AI software architecture can be obtained.
[0146] It should be noted that for relevant examples of software feature data, please refer to the above description and will not be elaborated here.
[0147] S303: Obtain the architecture feature data of the initial AI chip architecture.
[0148] In one example, for more convenient architecture exploration, for example, to explore and optimize the architecture of AI products (such as those including AI software and AI chips), the architecture feature data of this AI chip architecture may also include configuration information of the architecture and micro-architecture, for example, including configuration information at each level from the macro architecture to the micro architecture, so that it is convenient to flexibly adjust the hardware architecture using this configuration information during the optimization stage.
[0149] It should be noted that for relevant examples of the architecture feature data of the AI chip architecture, please refer to the above description and will not be elaborated here.
[0150] S304: Obtain the target task corresponding to the software feature data and the architecture feature data; simultaneously perform hardware initialization.
[0151] In one example, the target task can detect the collaborative performance of the initial AI software architecture and the initial AI chip architecture. For example, using a simulator, generate a target task for detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture, so as to simulate the initial AI software architecture deployed on the initial AI chip architecture to run the target task.
[0152] In addition, in one example, the ways of hardware initialization include but are not limited to: configuring the parameters of the initial AI chip architecture, establishing connections between the initial AI chip architecture and other hardware, performing self-check and calibration of the initial AI chip architecture, etc. The details of the hardware initialization in the present disclosure solution are not specifically limited.
[0153] S305: Under theoretical constraint conditions (i.e., high abstraction levels), model the workload of each hardware unit of the initial AI chip architecture, and simulate the running of the target task to obtain the first performance data. For example, obtain the Figure 4 ideal performance as shown.
[0154] Here, a performance analysis model can be used to model the workload of each hardware unit of the initial AI chip architecture under theoretical constraint conditions.
[0155] For example, in one example, by comprehensively considering the parameters of the initial AI chip architecture (such as bandwidth information, latency information, etc.) and the workload information provided by the software feature data, the performance data of the initial AI chip architecture under ideal conditions can be simulated. For example, under theoretical constraint conditions, using an analytical calculation method to analyze the internal relationship between the parameters of the initial AI chip architecture and the workload information, so as to simulate and obtain the theoretical time required for a specific hardware unit in the initial AI chip architecture to complete the corresponding task (or, can be called the theoretical limit performance). At this time, the time required for this hardware unit to simulate the running of the target task can be used as the first performance data of the initial AI chip architecture.
[0156] S306: Under mixed constraint conditions, model the concurrent behavior of each hardware unit of the initial AI chip architecture, and simulate the running of the target task to obtain the second performance data. For example, obtain the Figure 4 mixed performance as shown.
[0157] Here, a mixed-precision model can be utilized to model the concurrent behaviors of the hardware units of the initial AI chip architecture under mixed constraints. For example, in one example, synchronization logic similar to that of a Register Transfer Level (RTL) model can be implemented in the form of a global synchronous clock to ensure the accuracy and consistency of the modeling.
[0158] Furthermore, in one example, during the modeling process, a series of synchronous logic hardware including flip-flops (Flop), First Input First Output (FIFO) queues, Random Access Memory (RAM), and pipeline queues (Pipe) can also be constructed. The introduction of this hardware not only accelerates the management and modeling of Starve / Stall behaviors but also narrows the abstract relationship between the initial AI chip architecture and the microarchitecture, providing reliable hardware behavior support for the exploration of the macroarchitecture / microarchitecture, enabling AI system designers to obtain the actually achievable performance peak.
[0159] S307: Under actual constraints, model the concurrent behaviors of the hardware units of the initial AI chip architecture and simulate the running of the target task to obtain third performance data. For example, obtain the actual performance as shown in Figure 4 the figure.
[0160] Here, a cycle-accurate model can be utilized to model the concurrent behaviors of the hardware units of the initial AI chip architecture under actual constraints.
[0161] In one example, the modeling method of S307 can be similar to that of step S306, except that the conditions for the two are different.
[0162] It should be noted that in the present disclosure solution, the abstraction levels of the performance analysis model, the mixed-precision model, and the cycle-accurate model lay the foundation for realizing performance modeling of the initial AI chip architecture under different conditions.
[0163] In addition, it should be noted that compared with the currently commonly used event-based (asynchronous) cycle modeling method, the method mentioned in S306 - S307 in the present disclosure solution provides more accurate parallelism control. This improvement not only significantly simplifies the debugging and development process of the modeling model but also facilitates the improvement of the modeling accuracy of the mixed-precision model and the cycle-accuracy model, and at the same time promotes the calibration work with the RTL model.
[0164] Figure 4 is a schematic diagram of the hardware abstraction characteristics of each model according to an embodiment of the present disclosure. Here, in Figure 4The hardware abstraction features of the performance analysis model, mixed-precision model, and cycle-accurate model are described. As Figure 4 shown, the horizontal axis represents the performance accuracy of each model (i.e., from left to right, the performance cycle accuracy gradually increases), and the vertical axis represents the hardware signal integrity (i.e., from top to bottom, the signal integrity gradually enhances).
[0165] According to Figure 4 what is shown, from the performance analysis model, mixed-precision model to the cycle-accurate model, the requirements for performance accuracy and signal integrity gradually increase. In other words, the abstraction level gradually decreases, and the cycle-accurate model also partially overlaps with the RTL model. Thus, it provides reliable performance data for subsequent iteration and optimization of the hardware-software co-design in a real scenario.
[0166] S308: Determine the operating performance of each hardware unit in the initial AI chip architecture. For example, analyze the above-obtained first performance data, second performance data, and third performance data to obtain the operating performance of each hardware unit.
[0167] It should be noted that the solution of the present disclosure can utilize a complete set of performance counters to obtain the first performance data, second performance data, and third performance data. Thus, it provides a top-down (or, which can also be called from theory to practice) global performance perspective for architecture exploration and provides a clear performance profiling path for the performance prediction of future AI chips.
[0168] S309: Based on the operating performance of each hardware unit, obtain the performance bottlenecks of each hardware unit in the initial AI chip architecture.
[0169] For example, in one example, the solution of the present disclosure analyzes and identifies the performance bottlenecks encountered by each hardware unit in the initial AI chip architecture during operation under different performance modeling conditions through a performance analysis tool. It should be noted that since the target task is simulated and run based on the initial AI software architecture deployed on the initial AI chip architecture, when determining the operating performance and performance bottlenecks of the initial AI chip architecture, the operating performance and performance bottlenecks of the initial AI software can also be indirectly obtained, thereby detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture.
[0170] S310: Obtain the collaborative detection result.
[0171] Here, after obtaining the operating performance of each hardware unit in the initial AI chip architecture and the performance bottlenecks of each hardware unit, the collaborative detection result that can characterize the collaborative performance of the initial AI software architecture and the initial AI chip architecture can be obtained.
[0172] Furthermore, based on the collaborative detection results, software optimization guidance can be provided for the initial AI software architecture, and / or hardware optimization guidance can be provided for the initial AI hardware architecture.
[0173] Figure 5 It is a schematic diagram of a software and hardware collaborative optimization framework according to an embodiment of the present disclosure.
[0174] As Figure 5 shown, in an example, first, the software and hardware collaborative performance model 501 is used to perform modeling under theoretical limit conditions, mixed limit conditions, and actual limit conditions respectively. Secondly, based on the performance counter, the software and hardware collaborative system performance counter data is obtained (for example, the operation performance data of each hardware unit in the initial AI chip architecture under theoretical limit conditions, mixed limit conditions, and actual limit conditions). At the same time, based on the software and hardware collaborative performance model 501, the system execution log data of the software and hardware system under theoretical limit conditions, mixed limit conditions, and actual limit conditions is obtained (for example, trace data). Finally, based on the software and hardware collaborative system performance counter data, the performance analyzer 502 is used to obtain the collaborative detection results of the initial AI software architecture and the initial AI hardware architecture.
[0175] Here, in practical applications, the data visualization module 503 can be used to display the collaborative detection results and the system execution log data of the software and hardware system, so as to perform corresponding optimization guidance on the initial AI software architecture and / or the initial AI hardware architecture based on the displayed collaborative detection results and the system execution log data of the software and hardware system.
[0176] The solution of the present disclosure also provides a performance detection device 600, as Figure 6 shown, including:
[0177] A data determination unit 601, configured to determine an initial artificial intelligence (AI) software architecture and an initial AI chip architecture for simulating the operation of the initial AI software architecture;
[0178] A task running unit 602, configured to simulate the operation of a target task based on the initial AI software architecture and the initial AI chip architecture to obtain a target performance result, where the target performance result includes the performance data of the initial AI chip architecture under different simulation scenarios;
[0179] A performance detection unit 603, configured to obtain a collaborative detection result for the initial AI software architecture and the initial AI chip architecture based on the target performance result, where the collaborative detection result characterizes the collaborative performance of the initial AI software and the initial AI chip when working together.
[0180] In a specific example of the solution of the present disclosure, the task running unit 602 is specifically configured to:
[0181] Based on the initial AI software architecture and the initial AI chip architecture, simulate the operation of the target task to obtain at least two of the following: the first performance data of the initial AI chip architecture in the first simulation scenario, the second performance data of the initial AI chip architecture in the second simulation scenario, and the third performance data of the initial AI chip architecture in the third simulation scenario;
[0182] Among them, the simulation granularities of the first simulation scenario, the second simulation scenario, and the third simulation scenario are different from each other.
[0183] In a specific example of the present disclosure solution, as Figure 7 shown, the performance detection device 600 may further include a scenario construction unit 604, where the scenario construction unit 604 is used for:
[0184] Construct a first simulation scenario by using the hardware abstraction features of the first simulation model; where the first simulation scenario is used to simulate the operation scenarios of each hardware unit in the initial AI chip architecture under theoretical limiting conditions.
[0185] In a specific example of the present disclosure solution, the task running unit 602 is specifically used for:
[0186] Obtain the performance data of at least some subtasks of the target task when each hardware unit in the initial AI chip architecture operates under the theoretical architecture parameters corresponding to the first simulation scenario;
[0187] Based on the performance data of each hardware unit, obtain the first performance data.
[0188] In a specific example of the present disclosure solution, the scenario construction unit 604 is further used for:
[0189] Construct a second simulation scenario by using the hardware abstraction features of the second simulation model; where the second simulation scenario is used to simulate the operation scenarios of each hardware unit in the initial AI chip architecture under mixed limiting conditions; where the mixed limiting conditions are limiting conditions between theoretical limiting conditions and actual limiting conditions.
[0190] In a specific example of the present disclosure solution, the task running unit 602 is specifically used for:
[0191] Obtain the performance data of at least some subtasks of the target task when at least some hardware units in the initial AI chip architecture operate under multi-task concurrency corresponding to the mixed limiting conditions;
[0192] Based on the performance data of each hardware unit, obtain the second performance data.
[0193] In a specific example of the present disclosure solution, the scenario construction unit 604 is further used for:
[0194] Construct a third simulation scenario by using the hardware abstraction features of the third simulation model, where the third simulation scenario is used to simulate the operating scenarios of each hardware unit in the initial AI chip architecture under actual constraint conditions.
[0195] In a specific example of the present disclosure solution, the task running unit 602 is specifically configured to:
[0196] Obtain the performance data of at least some of the sub-tasks in the target task when at least some of the hardware units in the initial AI chip architecture operate under the multi-task concurrency corresponding to the actual constraint conditions;
[0197] Based on the performance data of each hardware unit, obtain the third performance data.
[0198] In a specific example of the present disclosure solution, the performance detection unit 603 is specifically configured to:
[0199] Based on the target performance result, obtain the operating performance of each hardware unit in the initial AI chip architecture, and obtain the performance bottlenecks of each hardware unit in the initial AI chip architecture;
[0200] Based on the operating performance of each hardware unit in the initial AI chip architecture and the performance bottlenecks of each hardware unit in the initial AI chip architecture, obtain a collaborative detection result.
[0201] In a specific example of the present disclosure solution, the data determination unit 601 is further configured to:
[0202] Obtain the software feature data of the initial AI software architecture and obtain the architecture feature data of the initial AI chip architecture;
[0203] Based on the software feature data and the architecture feature data, generate a target task capable of detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture.
[0204] In a specific example of the present disclosure solution, as Figure 7 shown, the performance detection device 600 may further include an architecture adjustment unit 605, and the architecture adjustment unit 605 is configured to:
[0205] Based on the collaborative detection result, adjust the initial AI software architecture and / or the initial AI chip architecture.
[0206] For the specific functions and example descriptions of the units of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0207] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0208] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0209] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0210] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0211] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0212] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a performance detection method. For example, in some embodiments, a performance detection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the performance detection method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute a performance detection method by any other suitable means (e.g., by means of firmware).
[0213] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0214] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0215] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0216] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0217] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0218] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0219] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0220] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A performance detection method, comprising: Determine an initial artificial intelligence (AI) software architecture and an initial AI chip architecture for simulating the operation of the initial AI software architecture; Based on the initial AI software architecture and the initial AI chip architecture, simulate the operation of a target task to obtain a target performance result, wherein the target performance result includes performance data of the initial AI chip architecture under different simulation scenarios; Based on the target performance result, obtain a collaborative detection result for the initial AI software architecture and the initial AI chip architecture, wherein the collaborative detection result characterizes the collaborative performance of the initial AI software and the initial AI chip when working together.
2. The method according to claim 1, wherein, The step of, based on the initial AI software architecture and the initial AI chip architecture, simulating the operation of a target task to obtain a target performance result, includes: Based on the initial AI software architecture and the initial AI chip architecture, simulate the operation of a target task to obtain at least two of the following: first performance data of the initial AI chip architecture in a first simulation scenario, second performance data of the initial AI chip architecture in a second simulation scenario, and third performance data of the initial AI chip architecture in a third simulation scenario; Wherein, the simulation granularities of the first simulation scenario, the second simulation scenario, and the third simulation scenario are different from each other.
3. The method according to claim 2, further comprising: Construct a first simulation scenario by using the hardware abstraction features of a first simulation model; wherein the first simulation scenario is used to simulate the operation scenarios of each hardware unit in the initial AI chip architecture under theoretical limit conditions.
4. The method according to claim 3, wherein The step of simulating the operation of a target task to obtain first performance data includes: Obtain performance data of at least some subtasks of the target task when each hardware unit in the initial AI chip architecture operates under the theoretical architecture parameters corresponding to the first simulation scenario; Based on the performance data of each hardware unit, obtain the first performance data.
5. The method according to claim 2, further comprising: Construct a second simulation scenario by using the hardware abstraction features of a second simulation model; wherein the second simulation scenario is used to simulate the operation scenarios of each hardware unit in the initial AI chip architecture under mixed limit conditions; wherein the mixed limit conditions are limit conditions between theoretical limit conditions and actual limit conditions.
6. The method according to claim 5, wherein, The step of simulating the operation of a target task to obtain second performance data includes: Obtain performance data of at least some subtasks of the target task when at least some hardware units in the initial AI chip architecture operate under multi-task concurrency corresponding to the mixed limit conditions; Based on the performance data of each hardware unit, obtain the second performance data.
7. The method according to claim 2, further comprising: Construct a third simulation scenario by using the hardware abstraction features of a third simulation model; wherein the third simulation scenario is used to simulate the operation scenarios of each hardware unit in the initial AI chip architecture under actual limit conditions.
8. The method according to claim 7, wherein The step of simulating the operation of a target task to obtain third performance data includes: Obtain performance data of at least some subtasks of the target task when at least some hardware units in the initial AI chip architecture operate under multi-task concurrency corresponding to the actual limit conditions; Based on the performance data of each hardware unit, obtain the third performance data.
9. The method according to any one of claims 1-8, wherein, Based on the target performance results, obtaining a collaborative detection result for the initial AI software architecture and the initial AI chip architecture, including: Based on the target performance results, obtaining the running performance of each hardware unit in the initial AI chip architecture, and obtaining the performance bottlenecks of each hardware unit in the initial AI chip architecture; Based on the running performance of each hardware unit in the initial AI chip architecture and the performance bottlenecks of each hardware unit in the initial AI chip architecture, obtaining the collaborative detection result.
10. The method according to any one of claims 1-8, further comprising: Obtaining software feature data of the initial AI software architecture, and obtaining architecture feature data of the initial AI chip architecture; Based on the software feature data and the architecture feature data, generating a target task capable of detecting the collaborative performance of the initial AI software architecture and the initial AI chip architecture.
11. The method according to any one of claims 1-8, further comprising: Based on the collaborative detection result, adjusting the initial AI software architecture and / or the initial AI chip architecture.
12. A performance detection device, comprising: A data determination unit for determining an initial artificial intelligence (AI) software architecture and an initial AI chip architecture for simulating the operation of the initial AI software architecture; A task running unit for simulating the operation of a target task based on the initial AI software architecture and the initial AI chip architecture to obtain target performance results, wherein the target performance results include performance data of the initial AI chip architecture under different simulation scenarios; A performance detection unit for obtaining a collaborative detection result for the initial AI software architecture and the initial AI chip architecture based on the target performance results, wherein the collaborative detection result characterizes the collaborative performance of the initial AI software and the initial AI chip when working together.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11.