Computer-aided method for designing a specialised processor
The method for designing specialized computer-aided processors automates the optimization of performance, power consumption, and area parameters, addressing the complexity and time-consuming nature of custom processor design, and making it accessible to non-experts while achieving efficient and tailored processor design.
Patent Information
- Application Number
- PCT/FR2024/051579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
The design of custom processors for specific applications is a complex and time-consuming process that requires specialized engineering expertise and involves a lengthy optimization of performance, power consumption, and surface area, making it inaccessible to non-experts and inefficient in terms of processing time.
A method for designing a specialized computer-aided processor that involves accessing the source code of an application and a software test bench, along with design parameters and user-defined performance requests, to generate an optimal instruction set and corresponding performance, power consumption, and area parameters, allowing for automated optimization and generation of a final binary file compatible with the specialized processor.
This method enables the rapid design of specialized processors tailored to specific applications without the need for specialized engineering teams, allowing for significant reductions in processing time and making the design process accessible to non-expert users while achieving optimized performance, power consumption, and area parameters.
Smart Images

Figure FR2024051579_05062025_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] TITLE: METHOD FOR DESIGNING A SPECIALIZED COMPUTER-ASSISTED PROCESSOR
[0003] Technical field
[0004] The technical field of the invention is the design of processors, and more particularly the automatic design of processors.
[0005] Previous techniques
[0006] Today's processors include a large collection of processing elements to support a variety of applications. These processors are typically designed and tested in multi-stage processes involving multiple design stages and multiple specialized engineers.
[0007] A general-purpose processor can include a complex instruction set. This is called a CISC architecture (an acronym for "complex instruction-set computer").
[0008] When an application is run on such a general-purpose processor, many instructions included in this processor are not used. Indeed, such processors are designed to be able to run any type of application on a minimal memory footprint and must therefore include a wide selection of instructions adapted to these different types of applications.
[0009] This means that the processor's occupancy is higher than that of a specialized processor that only includes the instructions needed by the application being executed. This also has an impact on the power consumed by the processor, which is more important to keep these unused instructions accessible. This unnecessarily higher processor occupancy also implies a larger critical path and naturally a limitation of the frequency of use and therefore degraded performance.
[0010] It is on this observation that the RISC architecture (an English acronym for "Reduced Instruction Set Computer") was developed, in which a number of instructions from the CISC architecture were removed to be replaced by instructions emulated in software from simpler instructions.
[0011] In the case of a RISC-V architecture, instructions are grouped into extensions, each extension supporting certain operations such as multiplication, addition, code compression, or bit manipulation, to name a few. All instructions can be emulated from a few I instructions and two A instructions.
[0012] Designing a processor involves a step of optimization between performance, power consumption and surface area occupied.
[0013] A processor's performance is a measure of its processing or computing capacity, and is typically expressed in floating-point operations per second (FLOPS). It also depends on the efficiency of the instructions and / or the type of instructions included in the processor's instruction set (ISA, Instruction Set Architecture).
[0014] The power consumption determines the energy required for the operation of the processor as well as its cooling requirements and is generally expressed in Watts.
[0015] The surface area represents the size of the processor, and indirectly its cost. Processors are etched on silicon wafers. The manufacturing cost of processors, for a given manufacturing process, is calculated per silicon wafer. However, the number of processors per silicon wafer and therefore the cost per processor depends on the surface area occupied by each processor. The smaller the surface area of the processor, the greater the number of processors per silicon wafer and the lower the unit cost.
[0016] These three quantities are dependent on each other. For example, power can be increased by adding instructions. However, this increases the power consumption and the area occupied.
[0017] In contrast, when some instructions are emulated in software, the space occupied can be reduced, and in some cases, the power consumed can be reduced. However, such emulation consumes computing cycles, which reduces the processor's power.
[0018] Several hundred different instruction sets can exist for a single algorithm, making the choice of instruction combination particularly complicated in itself. When combined with power consumption and area requirements, the number of possibilities becomes colossal.
[0019] Designing custom processors for at least one application generally requires particularly long periods of time for teams of specialized engineers to determine the compromise between performance, power consumption and surface area, known as the PPA compromise (an acronym for "Performance, Power, Area").
[0020] There is therefore a need for a processor design suitable for a software application that does not require a team of specialized engineers and has significantly shorter processing times.
[0021] There is also a need for processor design to be made accessible to a non-expert user.
[0022] Finally, there is a need for fine-tuning of a processor during design.
[0023] Statement of the invention
[0024] The invention relates to a method for designing a computer-aided specialized processor, said specialized processor comprising an optimized instruction set and being designed according to a performance request defined by a user to execute a specific application in an optimized manner, comprising the following steps: a. accessing the source code of the specific application, b. accessing the source code of the software test bench, c. accessing design parameters of the specialized processor, d. accessing the performance request defined by the user, e.generating a report comprising an optimal instruction set to be implemented in the specialized processor for the application under consideration as well as the performance, power consumption and area parameters of the specialized processor corresponding to this optimal instruction set as a function of the design parameters of the specialized processor and the performance, power consumption and area requirements of the specialized processor defined in the performance query defined by the user, f. if the report on the performance parameters is satisfactory, the optimal instruction set can be considered as a final optimal instruction set, g. if the report on the performance parameters is not satisfactory, the method resumes at the step of accessing the performance, power consumption and area requirements of the specialized processor defined by the user, with new requirements, h.a compilation chain and a conversion tool are then generated to convert the instructions of an application intended to be executed on a general-purpose processor into instructions understandable by the specialized processor, i. the specialized processor is designed according to the final optimal instruction set, and j. a final binary file compatible with the specialized processor is generated in order to execute applications.
[0025] An initial binary file can be accessed instead of the source code, and a linker file can then be accessed at the same time as the initial binary file is accessed, with the final binary file then being generated based on the initial binary file, the linker file, and the conversion tool.
[0026] To generate a report, the following sub-steps can be carried out: a. a dictionary is generated associating with an instruction set the performance, the power consumption, the occupied surface determined by simulation and the logical elements shared with another instruction, the dictionary being initialized with values combining as input at least two instruction sets distributed to contain both almost complete and almost empty instruction families, b. using the dictionary, a linear optimization is carried out to determine locations of the areas with the best instruction sets giving performance, power consumption and surface parameters of the specialized processor best satisfying the performance request defined by the user, c. the dictionary is enriched with additional entries determined at the end of the linear optimization and located in the areas with the best instruction sets, d.machine learning is performed on the enriched dictionary in order to estimate the optimal instruction set, then a simulation of the performance, power consumption and occupied area parameters is performed as a function of the optimal instruction set and the dictionary is enriched with the determined performance, power consumption and occupied area parameters, e. if only one optimal instruction set has been determined, the method resumes at the dictionary enrichment step, if this is not the case it is determined whether the at least one optimal instruction set has sufficiently converged with respect to at least one predetermined convergence threshold as a function of a criterion to be maximized among the performance, the power consumption and the occupied area whose constraint is to be favored over the others, the criterion to be maximized being included in the performance query, and f.If so, the instruction set is retained as the final optimal instruction set and a report is generated including the final optimal instruction set and the performance, power consumption and area parameters of the specialized processor adjusted by accurate simulation.
[0027] To determine whether the optimal instruction sets in the dictionary have sufficiently converged in view of the user's performance query, the following steps can be performed: the current criterion to be maximized is compared with the at least one criterion to be maximized of the at least one previous occurrence, the deviation of the current criterion to be maximized from the criterion to be maximized of the previous occurrence is determined, and the derivative of the deviation of the criterion to be maximized with respect to the iteration number is determined. Then, the optimal instruction set is considered to have been found when the deviation of the current criterion to be maximized from the criterion to be maximized of the previous occurrence is less than a predetermined convergence threshold and when the derivative of the deviation of the criterion to be maximized with respect to the iteration number is negative or zero.
[0028] The user-defined performance request may include at least one of a performance requirement, a power consumption requirement, and / or an occupied area requirement, each requirement being expressed either as a boundary value defining an open range of values bounded on one side by the boundary value, or as a range of values bounded by two boundary values.
[0029] The user-defined performance query includes a mathematical relationship that is a function of at least two of a performance requirement, a power requirement, and / or an occupied area requirement.
[0030] After determining a final optimal instruction set and before generating a compilation chain and a conversion tool, at least one alternative optimal instruction set can be determined, by inserting at least one variation into the user-defined performance query, and by re-executing the steps of the method for automatic design of a specialized processor up to the step of generating a compilation chain and a conversion tool, this step not being carried out, the determination of the final optimal instruction set taking into account the same criterion to be maximized as when determining the final optimal instruction set, then the final optimal instruction set and the at least one alternative optimal instruction set are submitted to the user for choice,the user is then able to decide whether a compromise on one of the requirements of his performance query allows to obtain an unexpected result preferable to the result obtained with the final optimal instruction set.
[0031] Brief description of the drawings
[0032] Other aims, characteristics and advantages of the invention will appear on reading the following description, given solely by way of non-limiting example and made with reference to the appended drawings in which:
[0033] - figure [Fig. 1] illustrates the main steps of a method for designing a computer-aided specialized processor according to the invention, and
[0034] - Figure [Fig. 2] illustrates the main sub-steps involved in generating a report including an optimal instruction set to be implemented in a specialized processor for the application under consideration as well as the PPA parameters corresponding to this optimal instruction set.
[0035] Detailed Description Figure [Fig. 1 ] illustrates a computer-aided specialized processor design method for generating a binary file compatible with an optimized instruction set defining a specialized processor and designed to execute a specific application based on a user performance request including performance, power consumption and area requirements of the specialized processor.
[0036] The user-defined performance query includes at least one of a performance requirement, a power requirement, and / or an occupied area requirement. Each requirement included in the query is expressed either as a limit value defining an open range of values bounded on one side by the limit value, or as a range of values bounded by two limit values. This allows the power consumed by the specialized processor to be limited, but leaves the computational performance and occupied area requirements free to take any value.
[0037] A range of values for at least one of the requirements can also be defined.
[0038] For example, one can define a range of values for power consumption and occupied area, while allowing performance to take any value. It is obviously understood that, when a value is not associated with a requirement, the design process will attempt to obtain the best possible value. For power consumption and occupied area, this means obtaining the lowest possible value. For computational performance, it will be a question of obtaining the highest possible value.
[0039] The performance query may also include a mathematical function having as a parameter at least two of a processing performance requirement of the specialized processor, a power requirement consumed by the specialized processor, and a surface area requirement occupied by the specialized processor. When the performance query includes such a mathematical function, it may itself be associated with a limit value or a range of values.
[0040] During a first step 101, we access the source code of an application for which we seek to design a specialized processor.
[0041] In a second step 102, the software test bench source code is accessed, the execution of which provides the key performance indicators. The software test bench is a C / C++ code for evaluating the performance of the processor. This software can evaluate, for example, the number of clock cycles required to execute a predefined heavy processing.
[0042] During a third step 103, design parameters of the specialized processor are accessed. The design parameters of the specialized processor may include the cache protocol, the cache size, the number of cache levels. Each of these parameters has an impact on the performance, the power consumed or the occupied surface area of the corresponding processor.
[0043] In a fourth step 104, user-defined performance requests are accessed. User-defined performance requests are understood to mean PPA (Performance, Power Consumption, Occupied Area) requests, i.e., a request comprising a performance requirement, a power consumption requirement, and / or an occupied area requirement. The PPA requests may relate to at least one of a performance requirement, a power consumption requirement, and / or an occupied area requirement. Each requirement may be expressed as a limit value defining an open range of values bounded on one side by the limit value. Each requirement may also be expressed as a range of values bounded by two limit values.
[0044] Finally, PPA queries can also be defined as a mathematical relationship between at least two requirements among a performance requirement, a power requirement and / or an occupied area requirement.
[0045] In all cases, we also define a criterion to maximize among performance, power consumed and occupied surface area, the constraint of which is to be favored over the others.
[0046] Steps 101 to 104 may be carried out partially or completely sequentially and / or in parallel.
[0047] The method then continues with a fifth step 1 10, during which a report is generated comprising an optimal instruction set to be implemented in the specialized processor for the application considered as well as the PPA parameters corresponding to this optimal instruction set. Step 1 10 is illustrated by the figure [Fig. 2] and described in more detail later in this description. During a step 1 15, the user determines whether the report is satisfactory by comparing the PPA parameters corresponding to this optimal instruction set with the performance request defined by the user. In a particular embodiment, in light of the results (report obtained at the end of step 1 10) already provided, the user can quickly evaluate another PPA compromise.
[0048] If the report is not satisfactory, the method resumes at step 104 with the submission of a new performance request by the user.
[0049] If the ratio is satisfactory, the method continues to step 120 during which the optimal instruction set determined in step 110 is considered satisfactory. The optimal instruction set is then considered as the final instruction set.
[0050] During a step 130, a compilation chain and a conversion tool are then generated for converting the instructions of an application intended to be executed on a general-purpose processor into instructions understandable by the specialized processor. In other words, the conversion tool makes it possible to match the instructions of the application to equivalent instructions included in the final instruction set.
[0051] During a step 150, the final binary file defining the specialized processor is generated from the source code provided in step 101 and the compilation chain generated in step 130.
[0052] In a particular embodiment, an initial binary file is used instead of the source code during the first step 101. A linker file is then accessed at the same time as the initial binary file is accessed during the first step 101.
[0053] The final binary file is then generated based on the initial binary file, the linker file, and the conversion tool generated in step 130. The specialized processor is considered to be designed based on the final optimal instruction set. The specialized processor can be manufactured based on this final optimal instruction set.
[0054] Figure [Eig. 2] illustrates the main sub-steps included in step 110 of generating a report. During a first sub-step 220, a dictionary is generated in the form of an associative table with a set of instructions as input or key.
[0055] The input of the associative array (its key) is an instruction set. The PPA parameters, i.e. the performance, the power consumed, the area occupied by the processor and the identifiers of the common elementary logical resources used by the instructions of the same family are the values (outputs) associated with each key.
[0056] Elementary logic resources are, for example, an adder, a comparator, a shift register. They are very often common to several instructions. For example, the adder is used by all addition instructions but also by branch instructions.
[0057] For each input corresponding to an instruction set, the PPA parameters of performance, area, and consumption resulting from a precise simulation with an external tool are associated at the output. For each input, the PPA parameters at the output are completed with a list identifying the generic processing units used, such as an adder or a shift register. The dictionary does not contain an exhaustive set of all possible instruction sets. Indeed, obtaining such an exhaustive set would represent an extremely large quantity of data requiring a computation time significantly longer than the marketing life of the custom processors obtained, given current computational resources. The choice of instruction sets is made in order to cover all instruction sets in a deliberately sparse manner.Instead of proposing all possible combinations down to the instruction, for each similar instruction (such as multiplications, divisions, etc.), we will choose a very fragmentary subset of instructions and an almost complete subset of instructions. We will fill the dictionary with all possible combinations of the presence of these subsets of instructions.
[0058] With a reduced instruction set as the key, we will have a small surface area, low performance and low consumption, whereas for a complete instruction set, we will have the opposite result (large surface area, large performance and large consumption).
[0059] During a second sub-step 240, a linear optimization is carried out to determine the best combinations of instructions giving PPA parameters that best satisfy the performance request defined by the user accessed during step 104. Using the PPA parameters associated with the instruction sets included in the dictionary, and the list identifying the generic processing units used for each instruction set in the dictionary, it is possible, with a linear regression, to estimate the PPA parameters of all the possible combinations of instruction sets not found in the dictionary.
[0060] With a classical linear or simplex optimization algorithm, one can quickly evaluate the areas in which the PPA parameters will be optimal according to the performance requirements defined by the user. Since this is done by approximation without performing more precise simulations, these results will be approximate and will have to be adjusted by simulations with an external tool.
[0061] During a third sub-step 260, the dictionary is enriched with additional entries found during sub-step 240.
[0062] During a fourth sub-step 270, machine learning is carried out in order to determine the optimal instruction set. To achieve this, the learning uses the dictionary enriched at the end of the third sub-step 260 and estimates the criterion to be maximized, more precisely than a linear optimization. The dictionary is completed with the new entries found by machine learning by carrying out simulations with an external tool.
[0063] During a sub-step 275, it is determined whether the optimal instruction sets of the dictionary in view of the user's performance query have sufficiently converged. To achieve this, the current criterion to be maximized is compared with the at least one criterion to be maximized of the at least one previous occurrence. The optimal instruction set is considered to have been found when the deviation of the criterion to be maximized is less than a predetermined convergence threshold and the derivative of the deviation of the criterion to be maximized with respect to the number of iterations is negative or zero.
[0064] If this is not the case, the process then resumes after the enrichment step 260.
[0065] If so, the method continues with a sixth sub-step 280, during which the instruction set is retained as the final optimal instruction set and a report is generated comprising the final optimal instruction set and the adjusted PPA parameters determined in the third sub-step 260.
[0066] In a particular embodiment, after determining a final optimal instruction set, at least one alternative optimal instruction set is determined by inserting at least one variation into the user-defined performance query by re-executing all of the steps of the method. Determining the final optimal instruction set takes into account the same criterion to be maximized as when determining the final optimal instruction set. The final optimal instruction set and the at least one alternative optimal instruction set are then submitted to the user for selection. The user is then able to decide whether a compromise on one of the requirements of his performance query makes it possible to obtain an unexpected result that is preferable to the result obtained with the final optimal instruction set.
Claims
CLAIMS 1. A method for computer-aided design of a specialized processor, said specialized processor comprising an optimized instruction set and being designed according to a performance query defined by a user to execute a specific application in an optimized manner, comprising the following steps: a. accessing (101) the source code of the specific application, b. accessing (102) the source code of a software test bench, c. accessing (103) design parameters of the specialized processor, d. accessing (104) the performance query defined by the user, e.generating (1 10) a report comprising an optimal instruction set to be implemented in the specialized processor for the considered application as well as the performance, power consumption and area parameters of the specialized processor corresponding to this optimal instruction set according to the design parameters of the specialized processor and the performance, power consumption and area requirements of the specialized processor defined in the performance query defined by the user, f. if the report on the performance parameters satisfies the performance query defined by the user, the optimal instruction set is considered as a final optimal instruction set, g.if the performance parameter report does not satisfy the user-defined performance query, the method resumes at the step of accessing the user-defined performance, power consumption and specialized processor area requirements, with new requirements, h. a compilation chain and a conversion tool are then generated (150) for converting the instructions of a. applications intended to be executed on a general-purpose processor in instructions understandable by the specialized processor, i. the specialized processor is designed according to the final optimal instruction set, and j. a final binary file compatible with the specialized processor is generated in order to execute applications.
2. A method of designing a computer-aided specialized processor according to claim 1, wherein an initial binary file is accessed in place of the source code of the application, and a linker file is then accessed at the same time as the initial binary file is accessed, the final binary file then being generated on the basis of the initial binary file, the linker file and the conversion tool.
3. A method for designing a computer-aided specialized processor according to claim 1 or 2, wherein, to generate a report, the following sub-steps are carried out: a. a dictionary is generated (220) associating with a set of instructions the performance, the power consumed, the occupied surface determined by simulation and the logical elements shared with another instruction, the dictionary being initialized with values combining as input at least two sets of instructions distributed to contain both families of instructions comprising at least one software-emulated instruction, and families of instructions comprising at least one instruction which is not software-emulated, b.using the dictionary, a linear optimization is carried out (240) to determine locations of the areas with instruction sets giving parameters of performance, power consumption and surface area of the specialized processor whose deviation from the performance request defined by the user is the smallest, said instruction sets having the smallest deviation being considered as the best instruction sets. c. enriching (260) the dictionary with additional entries determined at the end of the linear optimization and located in the areas with the best instruction sets, d. performing (270) machine learning on the enriched dictionary in order to estimate the optimal instruction set, then performing a simulation of the performance, power consumption and occupied surface area parameters as a function of the optimal instruction set and enriching the dictionary with the determined performance, power consumption and occupied surface area parameters, e.if only one optimal instruction set has been determined, the method resumes at the dictionary enrichment step, if this is not the case, it is determined (275) whether the at least one optimal instruction set is considered to have converged with respect to at least one predetermined convergence threshold depending on a criterion to be maximized included in the performance query and whose constraint is to be favored over the others, said constraint being chosen from performance, power consumed and occupied surface area, and f. if this is the case, the instruction set is retained (280) as the final optimal instruction set and a report is generated comprising the final optimal instruction set and the performance, power consumed and surface area parameters of the specialized processor adjusted by a precise simulation.
4. A method for designing a specialized processor according to claim 3, wherein, to determine whether the optimal instruction sets of the dictionary are considered to have converged in view of the user's performance query, the current criterion to be maximized is compared with the at least one criterion to be maximized of the at least one previous occurrence, the deviation of the current criterion to be maximized with the criterion to be maximized of the previous occurrence is determined and the derivative of the deviation of the criterion to be maximized with respect to the number of iterations is determined, then the optimal instruction set is considered to have been found when the deviation of the current criterion to be maximized with the criterion to be maximized of the previous occurrence is less than a threshold of predetermined convergence and when the derivative of the deviation of the criterion to be maximized with respect to the iteration number is negative or zero.
5. A method of designing a specialized processor according to any one of claims 1 to 4, wherein the user-defined performance request comprises at least one of a performance requirement, a power consumption requirement and / or an occupied surface area requirement, each requirement being expressed either as a limit value defining an open range of values bounded on one side by the limit value, or as a range of values delimited by two limit values.
6. A method of designing a specialized processor according to any one of claims 1 to 5, wherein the user-defined performance request comprises a mathematical relationship depending on at least two requirements from a performance requirement, a power requirement and / or an occupied area requirement, the user-defined performance request.
7. Method for designing a specialized processor according to any one of claims 3 to 6, wherein, after determining a final optimal instruction set and before generating a compilation chain and a conversion tool, at least one alternative optimal instruction set is determined, by inserting at least one variation into the performance query defined by the user, and by re-executing the steps of the method for automatic design of a specialized processor up to the step of generating a compilation chain and a conversion tool, this step not being carried out, the determination of the final optimal instruction set taking into account the same criterion to be maximized as when determining the final optimal instruction set, then the final optimal instruction set and the at least one alternative optimal instruction set are submitted to the user for choice,the user then being able to decide whether a compromise on one of the requirements of his performance query allows to obtain an unexpected result preferable to the result obtained with the final optimal instruction set.
Citation Information
Patent Citations
Architectural level power-aware optimization and risk mitigation
US20120017189A1
Synthesis system for pipelined digital circuits with multithreading
WO2011156741A1