Processor design method and related product

By simulating the processor and adjusting the software and hardware until the design improvement is achieved, and by combining the stream timeline information to optimize the processor design, the problem of low processor design efficiency in the existing technology is solved, and a more efficient processor design is realized.

CN121009852APending Publication Date: 2025-11-25SHANGHAI CAMBRICON INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410629502.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies lack sufficient methods for analyzing the competitiveness of processor products, making it impossible to effectively assess the impact of the configuration of each unit in the chip on performance, resulting in resource waste and inefficiency.

Method used

The design employs a simulated processor, adjusting initial software and hardware specifications or implementation strategies until performance data achieves the design improvement multiple, and utilizing streaming timeline information to optimize the processor design.

Benefits of technology

It improves the efficiency and accuracy of processor design, avoids the waste of resources in hardware physical design, and can intuitively show the processing of instructions inside the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009852A_ABST
    Figure CN121009852A_ABST
Patent Text Reader

Abstract

The invention relates to a method for designing a processor, which comprises the following steps of: acquiring performance data after running by utilizing an analog processor, and enabling initial hardware simulated by the analog processor to have a design improvement multiple relative to reference hardware, when the performance improvement of the performance data relative to the reference performance data of the reference hardware does not reach the design improvement multiple, adjusting initial software of the analog processor, and updating the analog processor according to the adjusted initial software; and repeating the steps until the performance improvement of the performance data relative to the reference performance data reaches the design improvement multiple. According to the method for designing the processor, resource waste caused by the fact that a hardware entity processor is adopted for design can be avoided, specification parameters or implementation strategies of hardware can be adjusted more conveniently, and the design efficiency of the processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of chips, and in particular to a method for designing a processor and related products. BACKGROUND

[0002] With the continuous progress of artificial intelligence technology and the increasing demand for computing performance, new products of artificial intelligence chips are launched at an increasingly fast pace. Accordingly, the design requirements for the core component processor of artificial intelligence chips are also increasingly high, and the competitiveness analysis of processor products has already started in the product design stage.

[0003] In the prior art, the main means of processor product competitiveness analysis include changing the configuration on the existing hardware board, sometimes needing to use two chips to equivalent the performance of the target chip, etc., which cannot evaluate the influence of the settings of each unit in the chip on the performance change, can only see the total running time and cannot see the performance bottleneck, when the computing power bandwidth of the target chip is increased, it is limited by the original board computing power bandwidth and cannot be evaluated. The existing means also include changing important operators on the software side, such as using EXCEL for evaluation, listing the implementation steps of the operator, rearranging the operator according to the required specifications, etc., which needs to be manually implemented, the evaluation of a single operator is time-consuming and labor-intensive, and only operator-level evaluation can be achieved and the overall network cannot be evaluated. It can be seen that how to improve the competitiveness analysis means in the processor design stage and improve the input-output ratio and work efficiency in the chip design field has always been a problem to be solved in the chip design field. SUMMARY

[0004] In order to at least partially solve the technical problems mentioned in the background art, the scheme of the present application provides a method for designing a processor.

[0005] In one aspect, the present application provides a method for designing a processor, the design method comprising: obtaining performance data by running a simulation processor, wherein the initial hardware simulated by the simulation processor has a design improvement multiple relative to the benchmark hardware, and the design improvement multiple includes a computing power improvement multiple or a bandwidth improvement multiple; when the performance improvement of the performance data relative to the benchmark performance data of the benchmark hardware does not reach the design improvement multiple, adjusting the initial software of the simulation processor, and updating the simulation processor according to the adjusted initial software; repeating the above steps until the performance improvement of the performance data relative to the benchmark performance data reaches the design improvement multiple.

[0006] In an aspect, the design method further comprises: when the performance data reaches the design improvement multiple relative to the benchmark performance data, determining whether there is a stream whose time proportion is greater than a proportion threshold based on the stream time axis information output by the simulation processor; when there is no stream whose time proportion is greater than the proportion threshold, adjusting the initial software, otherwise adjusting the initial hardware, and updating the simulation processor according to the adjusted initial hardware or initial software until the performance data of the simulation processor reaches the design target.

[0007] In an aspect, the adjusting the initial hardware comprises: adjusting an implementation strategy of the simulation processor hardware, the implementation strategy comprising at least one of the following strategies: changing parallelism of an operation stream, an IO stream or a MOVE stream, adjusting a hardware delay parameter, adjusting a number of operation units, adjusting a component work efficiency.

[0008] In an aspect, the adjusting the initial hardware further comprises: adjusting a specification parameter of the simulation processor hardware, the specification parameter comprising at least one of a working frequency, a computing power, a read-write bandwidth, an instruction startup overhead, an operation unit efficiency coefficient, a read-write RAM efficiency coefficient.

[0009] In an aspect, the adjusting the initial software comprises at least one of the following modifications of an implementation strategy of a software operator: reducing a synchronization instruction overhead, increasing a number of parallel streams, and increasing a fused instruction.

[0010] In an aspect, the performance data comprises any one of the following: a number of batch data processed per second by the simulation processor, an execution time of a set task completed by the simulation processor, and a running efficiency of an execution unit in the simulation processor.

[0011] In an aspect, the simulation processor is a dynamic library compiled by a code, used to simulate hardware processing components of a processor, access initial software running and output stream time axis information, the stream comprising at least one of an IO stream, a MOVE stream and an operation stream, and the stream time axis information comprising a running start time and a running end time of a stream.

[0012] In an aspect, the hardware processing component comprises a control unit and at least one operation unit, the control unit being configured to read and decode instructions, and distribute decoded operation instructions to the operation unit; the operation unit being configured to execute the operation instructions and return an execution duration to the control unit, so that the control unit outputs the stream time axis information according to the returned execution duration and decoded synchronization instructions.

[0013] In one aspect, the synchronization instruction comprises a producer-consumer mode, if the producer and the consumer are in the same stream, the control unit outputs the stream timeline information according to the returned execution duration and the decoded synchronization instruction, comprising: the control unit determines the end time according to the execution duration of the stream, and adds the synchronization duration in the synchronization instruction to the end time to determine the running end time of the stream related to the synchronization instruction.

[0014] In one aspect, the synchronization instruction comprises a producer and a consumer, if the producer and the consumer are not in the same stream, the producer stream and / or the consumer stream is at least one, the control unit outputs the stream timeline information according to the execution duration of the stream and the decoded synchronization instruction, comprising: the control unit determines the end time of the first consumer stream as the latest time in the end time of the producer stream plus the synchronization duration in the synchronization instruction, and determines the end time of the second consumer stream as its own end time plus the synchronization duration in the synchronization instruction, the first consumer stream is the consumer stream whose end time is earlier than the latest time in the end time of the producer stream, and the second consumer stream is the consumer stream whose end time is later than the latest time in the end time of the producer stream.

[0015] In one aspect, the present application provides a computer program product, the computer program is executed by a processor to realize the steps of the method in any of the above aspects.

[0016] In one aspect, the present application provides a computer device, comprising a memory, a processor and a computer program stored in the memory, the processor executes the computer program to realize the steps of the method in any of the above aspects.

[0017] The method for designing a processor provided by the embodiment of the present application adopts the simulation processor obtained by code compilation to process the processor design, which can avoid the resource waste caused by the design using a hardware entity processor, can more conveniently adjust the hardware specification parameters or implementation strategies, and improves the efficiency of the processor design. Meanwhile, the simulation processor provides the stream timeline information of different streams, which can intuitively present the processing process of the instructions in the processor, and is convenient for the analysis and adjustment in the processor design stage. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0019] Figure 1 is a timeline schematic diagram of one instruction / task according to an embodiment of the present application;

[0020] Figure 2 is a flow chart showing a method of designing a processor according to the present application;

[0021] Figure 3 is a flow chart showing a method of designing a processor according to the present application.

[0022] The specific embodiments of the present application will be described in detail below with reference to the drawings. The drawings and the following description are not intended to limit the scope of the present application in any way, but merely to illustrate the concepts of the present application. DETAILED DESCRIPTION

[0023] The exemplary embodiments will be described in detail with reference to the drawings. The following description is not intended to limit the scope of the present application in any way, but merely to illustrate the concepts of the present application. The following exemplary embodiments are described in detail with reference to the drawings.

[0024] The design target of a processor product can be set according to the performance data of a competitor or the performance data of a previous generation product, according to certain design rules, business considerations or other factors. The design process of a processor product is a process of constantly adjusting the software and hardware of the processor so that the performance data of the processor reaches the design target.

[0025] To avoid the troubles caused by using physical hardware for processor design, the present application uses a simulation processor for design: the function and internal structure of the hardware processor are simulated by software, so that the behavior of the hardware processor can be simulated, and the running state and result of the hardware processor can be shown. By adjusting various specification parameters or implementation strategies in the simulation processor, the change of the specification parameters or implementation strategies of the hardware processor can be simulated, so that the performance change of the hardware processor can be measured, and the effect of adjusting the hardware of the processor in the processor design process can be achieved.

[0026] In an embodiment of the present application, the simulation processor can be equivalent to a dynamic library compiled by code, which is used to simulate the hardware processing components of the processor, and can be accessed by the initial software to run and output the flow timeline information of each internal flow. The hardware processing components simulated in the simulation processor include a control unit and at least one operation unit, which can be used to run the instructions executed on the processor.

[0027] In this embodiment, the simulation processor in the present application can be a dynamic library compiled by different programming languages, for example, it can be compiled by C++ code. Taking the convolution operation executable by the operation unit as an example, an example of the code of the control unit and the convolution operation unit is as follows:

[0028]

[0029] The simulation processor can be equivalent to a hardware board card, and corresponding internal components are set according to a design target, and the simulation processor is built based on an existing project or initial setting. The simulation processor can access initial software, simulate the running of the simulation processor product, and output results. The initial software can include a runtime library, a driver, and an operator library, etc., for supporting the processor hardware and supporting the upper application. The initial software can also be obtained based on the existing project or the initial setting.

[0030] In the embodiment, after the simulation processor runs the initial software, the instruction set of the existing project can be used as input, and after the implementation strategy and the specification parameter are modified according to the design target of the processor product, the simulation processor is run in the running environment of the initial software or offline, performance data and stream time axis information of different output streams in the simulation processor are output, and the execution of each stream is visually displayed. In the process of executing instructions, the simulation processor includes at least one of an IO stream, a MOVE stream, and a calculation stream, for example. The simulation processor can output stream time axis information of each stream, and the stream time axis information includes a running start time and a running end time of one stream when executing instructions.

[0031] In the embodiment, in the process of executing instructions, the simulation processor can record stream time axis information of at least one of an IO stream, a MOVE stream, and a calculation stream. The IO stream is data transfer related to off-chip storage (for example, DDR storage), the MOVE stream is data transfer between on-chip storage (for example, RAM), and the calculation stream is operation related to calculation. The stream time axis information of the three streams includes a running start time and a running end time of the IO stream, a running start time and a running end time of the MOVE stream, and a running start time and a running end time of the calculation stream. According to the stream time axis information of each stream, the simulation processor can output stream time axis information of instructions / tasks, that is, output time information of each stream in the simulation processor when processing instructions / tasks.

[0032] Figure 1 is a schematic diagram of a time axis of one instruction / task according to an embodiment of the application, as Figure 1 shown, the running time of each stream in the processor in the whole instruction running can be intuitively displayed through the stream time axis information of different streams. The proportions of the IO stream, the MOVE stream, and the calculation stream in the whole instruction running time can be intuitively displayed through Figure 1

[0033] ​In the embodiment, the processor is designed by using the simulation processor obtained by code compiling, so that resource waste caused by using a hardware entity processor for design can be avoided, the specification parameters or implementation strategies of the hardware can be adjusted more conveniently, and the efficiency of processor design is improved. Meanwhile, the simulation processor provides flow time axis information of different flows, so that the processing process of instructions in the processor can be intuitively presented, and analysis and adjustment in the processor design stage is facilitated.

[0034] In one embodiment of the application, the hardware processing component includes a control unit and at least one operation unit, the control unit is configured to read instructions and decode, and distribute the decoded operation instructions to the operation unit; the operation unit is configured to execute the operation instructions and return the execution duration to the control unit, so that the control unit outputs the flow time axis information according to the returned execution duration and the decoded synchronization instruction.

[0035] In the embodiment, the simulation processor includes a control unit and at least one operation unit, the control unit reads instructions, and distributes the decoded instructions to the operation unit for execution and records the running start time, the plurality of operation units can run in parallel, and each operation unit returns the execution end time of the instruction to the control unit after execution, so that the control unit records the execution time returned by each operation unit.

[0036] In the process of running the simulation processor, in order to support parallel operation of a plurality of operation units, parallel operation of a plurality of simulation processors, support the interdependence between a plurality of instructions, and support parallel processing between data transfer and data operation when executing instructions in the simulation processor, there are a large number of synchronization instructions of different levels in the instructions processed by the simulation processor. For example, the instructions read by the simulation processor can include synchronization instructions between simulation processors, and other instructions read by the simulation processor can also be decoded to obtain synchronization instructions for synchronization between a plurality of operation units. Whether the synchronization instructions read by the simulation processor or the decoded synchronization instructions, the time duration information required for synchronization is included. The control unit can output the flow time axis information of each flow according to the execution duration returned by the operation unit and the time duration information in the decoded synchronization instruction corresponding thereto. The time duration information in the synchronization instruction can be set according to an experience value or preset according to project requirements.

[0037] In one embodiment, the instruction received by the simulation processor is a barrier instruction (a kind of synchronization instruction), and the code is as follows

[0038]

[0039] In the above code, mainTimeline is the overall running timeline of the simulation processor to execute instructions / tasks, and otherTimeline is the timeline of each stream.

[0040] In this example, the simulation processor can obtain accurate stream timeline information of each stream according to the duration information in the synchronization instruction and the execution time of each operation unit, and can provide reliable support for the design of the processor.

[0041] In an embodiment of the present application, the synchronization instruction includes a producer-consumer mode, if the producer and the consumer are in the same stream, the control unit outputs the stream timeline information according to the returned execution duration and the decoded synchronization instruction, including, the control unit determines the end time according to the execution duration of the stream, and adds the synchronization duration in the synchronization instruction to the end time to determine the running end time of the stream related to the synchronization instruction.

[0042] In an embodiment of the present application, the synchronization instruction includes synchronization of the producer-consumer mode, which is commonly used for process synchronization. In the synchronization instruction executed by the simulation processor, the producer and the consumer can be one or more, and can belong to the same stream or belong to different streams. For example, the IO stream is the producer and the consumer; the MOVE stream is the producer and the consumer; the operation stream is the producer and the consumer; the IO stream is the producer and the operation stream is the consumer; the operation stream is the producer and the IO stream is the consumer; the IO stream is the producer and the MOVE stream is the consumer; the MOVE stream is the producer and the IO stream is the consumer, etc. There are different combinations of producers and consumers according to actual needs.

[0043] As shown in Figure 1 Each stream is in an execution or idle state in stages during the execution of the instruction, and the control unit records the execution start time and the execution end time of each stream, and adds the synchronization duration in the synchronization instruction to the corresponding execution stage when the execution reaches the synchronization instruction. It can be understood that if the producer and the consumer are in the same stream, the end time in the stream timeline information of this stream is equal to the time obtained by adding the synchronization duration in the synchronization instruction to the execution end time of the stream itself.

[0044] In this embodiment, when the producer and the consumer are in the same stream, the control unit determines the end time according to the execution duration of the stream, and adds the synchronization duration in the synchronization instruction to the end time to determine the running end time of the stream related to the synchronization instruction. The simulation processor can more accurately simulate the working state of the processor through the synchronization instruction, and improve the accuracy and efficiency of the processor design.

[0045] In one embodiment of the present application, the synchronization instruction comprises a producer and a consumer, if the producer and the consumer are not in the same stream, the producer stream and / or the consumer stream is at least one, and the control unit outputs the stream timeline information according to the execution duration of the stream and the decoded synchronization instruction, comprising: the control unit determines the end time of the first consumer stream as the latest time in the end time of the producer stream plus the synchronization duration in the synchronization instruction according to the execution duration of the stream, and determines the end time of the second consumer stream as its own end time plus the synchronization duration in the synchronization instruction, the first consumer stream is the consumer stream whose end time is earlier than the latest time in the end time of the producer stream, and the second consumer stream is the consumer stream whose end time is later than the latest time in the end time of the producer stream.

[0046] In one embodiment, the producer and the consumer are not in the same stream, comprising: the producer stream is multiple consumer streams, the producer stream is one consumer stream, or the producer stream and the consumer stream are both multiple. Specifically:

[0047] In one embodiment, when the producer stream is multiple consumer streams, the latest end time of the producer stream is compared with the end time of the consumer stream, if the latest time in the end time of the producer stream is later than the end time of the consumer stream, the latest time in the end time of the producer stream is taken as the end time of the consumer stream, and the execution end time of the consumer stream is obtained by adding the synchronization duration in the synchronization instruction to the latest time in the end time of the producer stream, and the stream timeline information of the consumer stream is updated.

[0048] In one embodiment, when the producer stream is one consumer stream, the end time of the producer stream is compared with the end time of each consumer stream, if the end time of one consumer stream is earlier than the end time of the producer stream, the time obtained by adding the synchronization duration in the synchronization instruction to the end time of the producer stream is taken as the end time of this consumer stream, and the stream timeline information is updated; if the end time of one consumer stream is later than the end time of the producer stream, the time obtained by adding the synchronization duration in the synchronization instruction to the end time of this consumer stream is taken as the end time of this consumer stream, and the stream timeline information is updated.

[0049] In one of the embodiments, when there are multiple producer streams and multiple consumer streams, the latest end time among the end times of the multiple producer streams is compared with the end times of the multiple consumer streams, if the end time of one of the consumer streams is earlier than the latest end time, the end time of the producer stream is added with the synchronization time length in the synchronization instruction, and the obtained time is taken as the end time of the consumer stream, and the stream timeline information is updated; if the end time of one of the consumer streams is later than the end time of the producer stream, the end time of the producer stream is added with the synchronization time length in the synchronization instruction, and the obtained time is taken as the end time of the consumer stream, and the stream timeline information is updated.

[0050] In the above embodiments, the multiple simulation processors can form a multi-core processor and run in parallel, and each simulation processor serves as a core. The multiple streams involved in the synchronization instruction can be an intra-core synchronization instruction, involving multiple streams in one core, or an inter-core synchronization instruction of multiple cores, involving multiple streams in different cores. When the simulation processors are running, in addition to the stream timeline showing the running time of each stream, an instruction timeline showing the overall time of the instruction / task is also maintained. The end time of the instruction timeline of each core is updated according to the latest value of the end time of each stream in the core. If the inter-core synchronization instruction of multiple cores needs to be executed, the intra-core stream timeline information of different cores is determined according to the end time determination method of each producer stream and consumer stream in the above embodiments. When the multiple simulation processors execute a task or program together, the latest end time of the instruction timeline of each core is the execution end time of the task or program.

[0051] In the embodiments, the simulation processor can execute the synchronization instruction and update the stream timeline information of each stream in the processor according to the synchronization time length in the synchronization instruction. In the processor design process, the execution of the instruction in the simulation processor can be more clearly and accurately given, and the efficiency and accuracy of the processor design can be improved.

[0052] In the processor design, the simulation processor is run based on the initial software and the initial hardware, and is continuously adjusted in the design stage until the design target is reached. In the process of adjusting the hardware or the software, the specific components of the hardware or the software that need to be adjusted, accurately and efficiently judging the adjustment direction of the hardware and the software, giving the adjustment basis, and avoiding manual adjustment based on experience are the key problems that need to be solved in the processor design stage.

[0053] Figure 2 is a flow chart showing a method for designing a processor according to the present application, as shown in Figure 2 In one of the embodiments of the present application, a method for designing a processor is provided, and the design method comprises:

[0054] Step S100: obtaining performance data after running the simulation processor, wherein the initial hardware of the simulation processor has a design improvement multiple relative to the benchmark hardware, and the design improvement multiple includes a computing power improvement multiple or a bandwidth improvement multiple;

[0055] Step S200: adjusting the initial software of the simulation processor when the performance improvement of the performance data relative to the benchmark performance data of the benchmark hardware does not reach the design improvement multiple, and updating the simulation processor according to the adjusted initial software;

[0056] Step S300: repeating the above steps until the performance improvement of the performance data relative to the benchmark performance data reaches the design improvement multiple.

[0057] In the embodiment, the implementation strategy or specification parameter of the hardware during the processor design is usually improved by a multiple relationship based on the last generation product or the competitive product. The hardware of the last generation product or the competitive product serving as the improvement benchmark is the benchmark hardware. The initial hardware of the simulation processor given at the initial design stage has a design improvement multiple relative to the implementation strategy or specification parameter of the benchmark hardware, for example, the computing power (implementation strategy) is improved by 2 times or the bandwidth (specification parameter) is improved by 3 times. The performance data can be any data of the running result of the simulation processor, for example, the performance data can be the running time of processing given data, the result accuracy, etc. Since the performance data is used to measure the performance of the simulation processor, theoretically, the improvement of the hardware configuration or the hardware implementation will inevitably bring the improvement of the performance, and the performance improvement is equivalent to the improvement multiple of the hardware implementation. In short, assuming that the computing power is improved by 2 times in the hardware implementation, it is expected that the performance data is also improved by 2 times, and when the bandwidth is improved by 3 times in the hardware implementation, it is expected that the performance data is also improved by 3 times. It can be understood that in the actual engineering implementation, there can be a coefficient between the improvement multiple of the hardware implementation and the multiple of the performance improvement, and the coefficient can be a value greater than 1 or less than 1. The performance data is obtained after the simulation processor runs based on the initial software and the initial hardware. The to-be-processed data can be any data type or application field, and the present application does not limit this.

[0058] When the performance data obtained by the initial hardware of the simulation processor does not reach the design performance improvement multiple relative to the benchmark performance data of the benchmark hardware, it indicates that the initial software supporting the operation of the simulation processor has problems, so that the simulation processor cannot fully exert the performance of the initial hardware. The initial software of the simulation processor needs to be adjusted and optimized, and the simulation processor is updated according to the adjusted initial software. The simulation processor is run again to obtain updated performance data, and the performance data is compared with the benchmark performance data again. The above steps are repeated until the performance data reaches the design performance improvement multiple relative to the benchmark performance data. At this time, the software of the simulation processor is in a state of being more suitable for the hardware, and the supporting software can fully exert the designed performance of the hardware.

[0059] In the embodiment, when the performance data obtained after the simulation processor is run cannot reach the performance improvement multiple, the initial software of the simulation processor is adjusted until the performance data reaches the design performance improvement multiple relative to the benchmark performance data.

[0060] The method provided in the embodiment gives a clear adjustment direction at the initial stage of processor design. When the performance improvement multiple cannot reach the design performance improvement multiple, the software is first adjusted so that the software can fully exert the performance of the hardware, and then the hardware is further adjusted. The method provided in the embodiment gives a simple and clear judgment method and adjustment direction in the design process, and improves the overall efficiency of processor design.

[0061] In an embodiment of the present application, the performance data includes any one of the following: the number of batch data processed per second by the simulation processor, the execution time of the simulation processor for completing a set task, and the running efficiency of an execution unit in the simulation processor.

[0062] In the embodiment, the performance data can be the number of batch data processed per second by the simulation processor, which can be used to measure the throughput of the processor, and the unit is pieces per second. For example, in image processing, how many batch data the simulation processor completes per second, or in natural speech processing, how many token data the simulation processor completes per second. The performance data can also be the execution time of the simulation processor for completing a set task, such as the execution time of a certain instruction, a certain software operator, or the running of a trained neural network model, etc. The unit is us / s. The performance data can also be the running efficiency of an execution unit in the simulation processor, for example, in a network mainly using convolution operation, the longer the processing resources or processing time of the operation unit responsible for convolution operation, the better the performance of the processor. The performance data can also be any dimension of any measurement index, as long as the performance of the processor can be compared. The present application does not limit this.

[0063] In the embodiment, the simulation processor can select different performance data to measure performance improvement, which provides more possibilities for the design process of the processor and improves the applicability of the processor design method.

[0064] Figure 3 is a flow chart illustrating a method for designing a processor according to the present application, as shown in the embodiment of the present application, the method for designing a processor provided by the present application further comprises: Figure 3

[0065] Step S400: When the performance improvement of the performance data relative to the benchmark performance data reaches the design improvement multiple, determine whether there is a stream whose time proportion is greater than the proportion threshold according to the stream time axis information output by the simulation processor;

[0066] Step S500: When there is no stream whose time proportion is greater than the proportion threshold, adjust the initial software,

[0067] Step S600: When there is a stream whose time proportion is greater than the proportion threshold, adjust the initial hardware,

[0068] Step S700: Update the simulation processor according to the adjusted initial hardware or initial software until the performance data of the simulation processor reaches the design target.

[0069] In the field of artificial intelligence, due to the particularity of the algorithm, some operations implemented by the processor will occupy most of the resources during the running of the processor. For example, in a task including convolution operation, the time length of executing convolution operation can account for more than half of the entire task running time; for example, when the processor processes a convolution instruction, the time length of the operation stream can account for more than half of the entire instruction execution time.

[0070] In the embodiment, when the performance data obtained by the simulation processor has reached the design improvement multiple, further determine whether there is a stream whose time proportion is greater than the proportion threshold according to the stream time axis information output by the simulation processor. The proportion threshold can be set according to the demand, for example, the proportion threshold can be set to eighty percent, that is, determine whether there is a stream whose processing time accounts for eighty percent of the total processing time of the entire instruction or task. If so, it can be considered that the software of the simulation processor can develop the potential of the hardware, otherwise, the software needs to be further adjusted until there is a stream whose time proportion is greater than the proportion threshold, and the simulation processor is updated at the same time.

[0071] ​In the embodiment, when the performance data of the simulation processor reaches the design improvement multiple, and the time proportion of one stream is greater than the proportion threshold, it is indicated that the software does not have a design bottleneck, and the full performance of the hardware can be exerted, and the initial hardware needs to be further adjusted, and the simulation processor is updated until the performance data of the simulation processor reaches the design target.

[0072] In an embodiment of the application, adjusting the initial hardware comprises: adjusting an implementation strategy of the simulation processor hardware, and the implementation strategy comprises at least one of the following strategies: changing parallelism of an operation stream, an IO stream or a MOVE stream, adjusting a hardware delay parameter, adjusting an operation unit quantity, and adjusting a component work efficiency.

[0073] In the embodiment, when the performance of the simulation processor does not reach the design target, the hardware needs to be adjusted, and the implementation strategy of the simulation processor hardware can be adjusted, and the implementation strategy can be regarded as a strategy for implementing various functions based on the hardware structure of the processor. During processor running, the parallelism of the operation stream, the IO stream and the MOVE stream affects the hardware performance, the parallelism of each stream can be improved by adjusting the number of the operation stream, the IO stream or the MOVE stream, and the parallelism can also be improved by changing the IO stream or the MOVE stream into read-write split flow, that is, the hardware performance is improved by changing the parallelism of the operation stream, the IO stream or the MOVE stream. The software and hardware performance alignment parameters related to the hardware implementation can also be adjusted, for example, the number of operation units and the hardware delay parameter including reading delay and writing delay, so that the software and hardware of the processor are better adapted to work together to fully exert the performance of the hardware. In addition, the component work efficiency can also be adjusted, for example, the efficiency of calculation depends on whether the read data received by the operation unit is continuous and whether the write data can be not back-pressured, and the implementation strategy of the hardware is adjusting the bank conflict rate and the on-chip RAM arbitration priority.

[0074] The hardware implementation strategy that needs to be adjusted can be located in combination with the stream time axis information output by the simulation processor. For example, when the stream time axis information proportion of the IO stream is too large, and the performance data of the simulation processor does not reach the set target, the method provided in the embodiment can adjust the number or parallelism of the IO stream in the hardware implementation strategy to improve the performance of the simulation processor.

[0075] In the embodiment, various means for adjusting the hardware implementation strategy are provided, which can be used as needed during processor design to improve the flexibility and design efficiency of the processor design.

[0076] When the performance data of the simulation processor does not reach the set target, the work efficiency of each component in the processor hardware is preferentially tried to be improved, and if the work efficiency is improved to the upper limit (100% work efficiency or the upper limit of hardware implementation, for example, 80%), if the performance data still cannot be improved to the set target, the specification parameter is considered to be improved.

[0077] In an embodiment of the present application, adjusting the initial hardware further comprises: adjusting a specification parameter of the analog processor hardware, the specification parameter comprising at least one of a working frequency, a computing power, a read-write bandwidth, an instruction startup overhead, an efficiency coefficient of an operator, and an efficiency coefficient of a read-write RAM.

[0078] In the embodiment, adjusting the specification parameter of the hardware is equivalent to modifying a configuration of the processor hardware architecture, and when the performance of the analog processor does not meet the design target, the performance of the analog processor can be improved by adjusting the processor hardware architecture. The specification parameter can comprise a working frequency of the processor (for example, 1.6 GHz), a convolution computing power under various data types (INT8, 512 TFLOPS), or a vector computing power (INT8, 64 TFLOPS), a read-write line width of an on-chip RAM (64 Byte), a read-write bandwidth of a DDR (220 GB / s), a startup overhead of various types of instructions (30 cycles), and an efficiency of each operator and each RAM (an integer less than 1: read-write DDR efficiency 0.8, read-write CT efficiency 0.95).

[0079] The specification parameter of the hardware that needs to be adjusted can be located in combination with the stream timeline information output by the analog processor. For example, when the execution time of the analog processor for completing a set task does not meet a set target, the method provided in the embodiment can improve the performance of the analog processor by adjusting the working frequency and the like in the specification parameter of the hardware.

[0080] In the embodiment, various means of adjusting the specification parameter of the hardware are provided, which can be used as needed in the process of designing the processor, thereby improving the flexibility and design efficiency of the processor.

[0081] In an embodiment of the present application, adjusting the initial software comprises at least one of the following modifications of an implementation strategy of a software operator: reducing a synchronization instruction overhead, increasing a number of parallel streams, and increasing a fused instruction.

[0082] In the embodiment, if the performance data of the analog processor does not meet the design target due to a software bottleneck, and the stream arrangement manner of the software needs to be improved according to the stream timeline information output by the analog processor, the parallel degree needs to be improved from the perspective of the software operator, so as to fully use the hardware resources. Moreover, if only the computing power of a certain stream is improved in each stream of the analog processor, or the stream timeline information between the streams is not proportional, for example, the proportion of the time length of an IO stream to the time length of a computing stream is abnormal, it is indicated that there is a large bottleneck in the software.

[0083] The implementation strategy of the modified software operator provided in the embodiment includes: reducing the synchronization instruction overhead to improve the execution efficiency of the instruction, increasing the number of parallel streams to improve the parallelism of the hardware, increasing the fusion instruction so that the originally independent hardware running units can be executed in parallel, etc. The present application does not limit this.

[0084] In the embodiment, various means of implementation strategy of the modified software operator are given, which can be used as needed in the processor design process to improve the flexibility and design efficiency of the processor design.

[0085] In an embodiment of the present application, the present application provides a computer program product, wherein the computer program is executed by a processor to implement the steps of any of the above-mentioned methods.

[0086] In an embodiment of the present application, the present application provides a computer device comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods.

[0087] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action order described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0088] It should be further noted that, although each step in the flowchart is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless explicitly stated in this paper, the execution of these steps has no strict order limitation, and these steps can be executed in other order. Moreover, at least part of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or sub-steps or stages of other steps.

[0089] It should be understood that the above-mentioned device embodiments are only schematic, and the device of the present application can also be realized by other ways. For example, the division of units / modules in the above-mentioned embodiments is only a logical function division, and another division mode can be used in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0090] In addition, each functional unit / module in each embodiment of the present application can be integrated in one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated unit / module can be realized in the form of hardware or in the form of a software program module.

[0091] If the integrated unit / module is realized in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor can be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, and an ASIC, etc. Unless otherwise specified, the storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc.

[0092] If the integrated unit / module is realized in the form of a software program module and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0093] In the above embodiments, the description of each of the embodiments has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any manner. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, it should be considered that it is within the scope of the description.

Claims

1. A method for designing a processor, characterized in that, The design method includes: acquiring performance data after running a simulated processor, wherein the initial hardware simulated by the simulated processor has a design improvement factor relative to the benchmark hardware, the design improvement factor including a computing power improvement factor or a bandwidth improvement factor; when the performance improvement of the performance data relative to the benchmark performance data of the benchmark hardware does not reach the design improvement factor, adjusting the initial software of the simulated processor, and updating the simulated processor according to the adjusted initial software; repeating the above steps until the performance improvement of the performance data relative to the benchmark performance data reaches the design improvement factor.

2. The method as described in claim 1, characterized in that, The method further includes: when the performance improvement of the performance data relative to the baseline performance data reaches the design improvement multiple, using the stream time axis information output by the simulation processor to determine whether the time proportion of a stream is greater than the proportion threshold; when the time proportion of no stream is greater than the proportion threshold, adjusting the initial software; otherwise, adjusting the initial hardware, and updating the simulation processor according to the adjusted initial hardware or initial software, until the performance data of the simulation processor reaches the design target.

3. The method as described in claim 1, characterized in that, The adjustment of the initial hardware includes: adjusting the implementation strategy of the simulated processor hardware, wherein the implementation strategy includes at least one of the following strategies: changing the parallelism of the arithmetic flow, IO flow or MOVE flow, adjusting hardware latency parameters, adjusting the number of arithmetic units, and adjusting the working efficiency of components.

4. The method as described in claim 3, characterized in that, The adjustment of the initial hardware further includes: adjusting the specifications of the simulated processor hardware, wherein the specifications include at least one of the following: operating frequency, computing power, read / write bandwidth, instruction startup overhead, arithmetic unit efficiency coefficient, and RAM read / write efficiency coefficient.

5. The method as described in claim 1, characterized in that, The adjustment of the initial software includes at least one of the following strategies for modifying the implementation of software operators: reducing synchronization instruction overhead, increasing the number of parallel streams, and increasing fusion instructions.

6. The method as described in claim 2, characterized in that, The performance data includes any one of the following: the number of batches of data processed per second by the simulation processor, the execution time for the simulation processor to complete the set task, and the operating efficiency of the execution unit in the simulation processor.

7. The method as described in claim 1, characterized in that, The simulated processor is a dynamic library compiled with code, used to simulate the hardware processing components of a processor, to access the initial software operation and output stream timeline information. The stream includes at least one of IO stream, MOVE stream and computation stream, and the stream timeline information includes the start time and end time of a stream.

8. The method as described in claim 7, characterized in that, The hardware processing component includes a control unit and at least one arithmetic unit. The control unit is used to read and decode instructions, and distribute the decoded arithmetic instructions to the arithmetic unit. The arithmetic unit is used to execute the arithmetic instructions and return the execution duration to the control unit, so that the control unit outputs the stream timeline information according to the returned execution duration and the decoded synchronization instructions.

9. The method as described in claim 8, characterized in that, The synchronization instruction includes a producer-consumer pattern. If the producer and consumer are in the same stream, the control unit outputs the stream timeline information based on the returned execution duration and the decoded synchronization instruction. This includes the control unit determining the end time based on the execution duration of the stream, and adding the synchronization duration in the synchronization instruction to the end time to determine the end time of the stream associated with the synchronization instruction.

10. The method as described in claim 8, characterized in that, The synchronization instruction includes producers and consumers. If the producers and consumers are not in the same stream, and the producer stream and / or consumer stream is at least one, then the control unit outputs the stream timeline information according to the execution duration of the stream and the decoded synchronization instruction. This includes: the control unit determining the end time of the first consumer stream as the latest end time among the end times of the producer stream plus the synchronization duration in the synchronization instruction, and determining the end time of the second consumer stream as its own end time plus the synchronization duration in the synchronization instruction, wherein the first consumer stream is a consumer stream whose end time is earlier than the latest end time among the end times of the producer stream, and the second consumer stream is a consumer stream whose end time is later than the latest end time among the end times of the producer stream.

11. A computer program product, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

12. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.