Method for automatically designing processors

The automatic processor design method using machine learning to generate processor topologies from source code or binary executables addresses the inefficiencies and expertise requirements of current design methods, achieving faster, cheaper, and more accessible processor design.

WO2025114674A1PCT designated stage expired Publication Date: 2025-06-05KEYSOM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/FR2024/051578
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current processor design methods require extensive time and expertise from specialized engineers, making it difficult to design custom processors efficiently and accessible to non-experts, while also being costly and time-consuming.

Method used

A method for automatically designing processors using machine learning to generate processor topologies from source code or binary executables, allowing for rapid iteration and fine-tuning of design parameters without requiring specialized expertise.

Benefits of technology

This approach significantly reduces processor design time and costs, enabling rapid prototyping and making processor design accessible to non-expert users while maintaining performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2024051578_05062025_PF_FP_ABST
    Figure FR2024051578_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for automatically designing a specialised processor for executing at least one program, wherein the method comprises the following steps: a. accessing (110) the source code or a binary executable of the program; b. analysing (120) the source code or the binary executable to obtain processor design parameters; c. encoding (140) the design parameters into a design parameter data structure and recording (150) the design parameter data structure; and d. generating (160) a processor topology from the design parameter data structure via a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] TITLE: METHOD FOR AUTOMATICALLY DESIGNING PROCESSORS

[0003] Technical field

[0004] The technical field of the invention is the design of processors, and more particularly the automatic design of processors.

[0005] Previous techniques

[0006] Today's processors include a large collection of processing elements to support a variety of applications. These processors are typically designed and tested in multi-stage processes involving multiple design stages and multiple specialized engineers.

[0007] Complete modularity of processor design is generally difficult to achieve without involving a variety of design elements, either proprietary or open source.

[0008] Processor designers typically implement an instruction set architecture (ISA) in a low-level hardware description language (HDL) or a high-level hardware construction language (HCL).

[0009] It is recalled that a hardware description language HDL is a specialized computer language used to describe the structure and behavior of electronic circuits, and in particular digital logic circuits.

[0010] Despite the availability of both commercial and open source EDA (Electronic Design Automation) software, custom processor design for at least one application typically requires a significant amount of time from specialized engineering teams. Therefore, there is a need for a processor design tailored to a software application that does not require a specialized engineering team and offers significantly shorter processing times.

[0011] Reducing processor design costs and times is crucial for rapid prototyping, enabling rapid testing and design iterations to reduce time to market for both generic and specialized processors.

[0012] There is also a need for processor design to be made accessible to a non-expert user.

[0013] Finally, there is a need for fine-tuning of a processor during design.

[0014] Statement of the invention

[0015] An object of the invention is a method for automatically designing a specialized processor for executing at least one program, comprising steps in which: a. the source code or a binary executable of the program is accessed, b. an analysis of the source code or the binary executable is performed to obtain processor design parameters, c. the design parameters are encoded in a design parameter data structure and the design parameter data structure is recorded, and d. a processor topology is generated from the design parameter data structure via a machine learning model.

[0016] To analyze the source code or the binary executable, steps can be performed in which: a. a binary executable disassembler is used to obtain the assembler code from a binary executable, the binary executable being directly available or obtained through a source code compilation step, b. execution metrics and execution statistics are obtained from simulation software applied to the binary executable, and c. the obtained execution metrics and / or execution statistics are converted into processor design parameters and then formatted into a data structure.

[0017] To perform the processor topology generation, steps are performed during which: a. a first topology of a processor is generated from the analysis of the source code or the binary executable and the processor design parameters and the number of design iterations of a processor topology is initialized, b. the design parameters for the processor topology generated by inference of a neural network of a machine learning module trained on a predetermined database are determined, c. whether the processor topology is not satisfactory in view of a predefined performance setpoint and whether the number of design iterations of a processor topology is less than a limit number of iterations is determined d.if so, determining additional feedback parameters via the machine learning module applied to the user's processor design parameters, saving the feedback parameters in the neural network training database of the machine learning module, retraining the neural network with the enriched database and then generating a new processor topology based on the feedback parameters and the method continuing at the step of determining design parameters for the generated new processor topology, e. if not, the processor topology is retained.

[0018] The processor can be generated through a hardware description language implementation.

[0019] When it is determined that the processor topology is not satisfactory given at least one predefined performance requirement, and the performance requirement is not a user design requirement, a design space exploration step may be performed, in which at least two processor topologies with different PPA requirements are determined, and the user then chooses a topology.

[0020] To generate a first topology of a processor, the following steps can be performed: a. accessing the data structure of the encoded design parameters b. determining values ​​of the processor design parameters by browsing the encoded design parameters, c. then converting the values ​​of the design parameters into values ​​of the input parameters of a processor generator, d. generating a routing-placement data structure of the processor from values ​​of the determined input parameters.

[0021] To generate additional feedback parameters, the following steps can be performed: a. a software development kit is generated from the processor's routing-tracing level data structure, b. an emulation data structure of an in-system programmable gate array is generated from the processor's routing-tracing level data structure and the software development kit, c. a test plan is generated based on the design parameters and at least one acceptance criterion depending on the predefined performance setpoint, d. additional feedback parameters are obtained as test results according to the test plan.

[0022] The predefined performance instruction may comprise a processing performance criterion of the specialized processor, a criterion of power consumed by the specialized processor and a criterion of surface occupied by the specialized processor, the at least one acceptance criterion corresponding to a threshold of at least one of the criteria of the performance instruction or a criterion of an algebraic combination of at least two of the criteria of the performance instruction.

[0023] The machine learning module may be of the supervised learning type. The supervised machine learning module may comprise a graph neural network, in particular a pre-pruned graph neural network.

[0024] Processor topologies can be generated through a hardware description language implementation.

[0025] Brief description of the drawings

[0026] Other aims, characteristics and advantages of the invention will appear on reading the following description, given solely by way of non-limiting example and made with reference to the appended drawings in which:

[0027] Figure [Fig. 1] illustrates the main steps of a design method according to the invention,

[0028] Figures [Fig. 2a] and [Fig. 2b] illustrate the main sub-steps of a source code or binary executable analysis in order to obtain processor design parameters.

[0029] Figure [Fig. 3] illustrates the main sub-steps of processor generation,

[0030] - Figure [Fig. 4] illustrates the main sub-steps of generating a first topology of a processor, and

[0031] Figure [Fig. 5] illustrates the main substeps associated with determining additional parameter data.

[0032] Detailed description

[0033] Systems and methods for automatically generating processor architectures from source code or binary executables allow a single user to generate an architecture and iterate on that architecture by specifying processor design parameters. Even a user who is not an expert in software architecture can use these systems and methods to develop a processor architecture.

[0034] The system automates both the design of processors and the verification of the resulting processors using both commercial and open source EDA (Electronic Design Automation) computer electronic design tools.

[0035] This automation of the use of computer-aided electronic design tools makes it possible to accomplish in a few hours or minutes what used to take considerably longer for teams of specialized engineers equipped with specialized EDA tools.

[0036] Figure [Fig. 1] illustrates the main steps of an automatic processor design method according to the invention.

[0037] In the first step 1 10, the source code or a binary executable of a program is accessed.

[0038] For example, the source code of a program can be written in a high-level language such as Python, or in low-level languages ​​such as C / C++.

[0039] The binary executable can be generated with a compiler that produces a portable executable (PE), a simple executable, an executable in extensible and linkable format (ELF), an executable in object-mach format (Mach-0), or similar executable formats.

[0040] During a second step 120, an analysis of the source code or the binary executable is carried out to obtain processor design parameters.

[0041] Source code analysis can be automated with open source tools such as Capstone or with commercial tools such as Binary Ninja. Source code or binary executable analysis is described in more detail in relation to figures [Fig. 2a] and [Fig. 2b],

[0042] During a third step 130, processor design parameters are accessed.

[0043] In a particular embodiment, the processor design parameters are design parameters of the user's processor. For example, a user can select specific design parameters from a set of design parameters (in other words, a design parameter is an instruction executable by a processor or a microarchitecture component). A specific design parameter can also be a design plan (for example, a power consumption not to be exceeded, a performance to be achieved, or a surface area occupied by the processor not to be exceeded). The selection of the parameters is carried out in particular via a graphical interface or a web interface. The design parameters are then encoded in a design parameter data structure during a step 140.Design parameters can be encoded in JSON (JavaScript Object Notation) or a similar data representation.

[0044] During a step 150, the design parameter data structure is recorded, locally on a volatile data medium, on a non-volatile data medium, or remotely on a server or in a remote, delocalized environment of the cloud type.

[0045] During a step 160, a processor topology is generated from the design parameter data structure.

[0046] Figure [Fig. 2a] illustrates a first embodiment of step 120 of analyzing source code or binary executable in order to obtain processor design parameters. In this embodiment, a source code is available as input to step 120.

[0047] Step 120 then comprises a first sub-step 210a during which the source code is compiled, using traditionally used compilation tools such as GCC or LLVM. It should be noted that this embodiment is advantageous compared to the second embodiment illustrated by the figure [Fig 2b] insofar as it is possible to extract more information on the execution context of the program or the application. It is therefore possible to generate a processor more suited to a given application (better performance / consumption / surface ratio). This information can also be used in addition to the input of sub-step 230.

[0048] Figure [Fig. 2b] illustrates a second embodiment of step 120 of analyzing source code or binary executable in order to obtain processor design parameters. In this embodiment, a binary executable is available as input to step 120.

[0049] Step 120 then comprises a first sub-step 210b during which a readable source code is obtained from a binary executable by means of a binary executable decompiler.

[0050] An example of a binary executable decompiler is the LIEF (Library to Instrument Executable Formats), which is a library for decompiling a large number of executable formats, including ELF, PE, and Mach-O. Such a library allows you to browse the binary executable and access its contents, i.e., the sections it contains.

[0051] Regardless of the embodiment, step 120 continues with the following sub-steps.

[0052] During a second sub-step 220, a binary executable disassembler is used to obtain the assembler code of the source code compiled during the first step 210a in the first embodiment or of the binary executable browsed during the first sub-step 210b in the second embodiment. The Capstone software will be cited in particular as an example of such a binary executable disassembler. Other software can of course be used.

[0053] During a third sub-step 230, execution metrics and execution statistics of simulation software applied to the binary executable are obtained.

[0054] Such simulation software can be Quick Emulator (acronym QEMU), Spike or riscvOVPsim. Other simulation software can of course be used.

[0055] The simulation software can be configured to use a deterministic software profiler (in a tracing approach) or to use a statistical profiler (in a sampling approach).

[0056] To achieve this, the Prof5 software profiler can be used.

[0057] The execution metrics and statistics may also include information from the analysis of the source code previously compiled during step 210a.

[0058] During a fourth sub-step 240, the execution metrics and / or execution statistics obtained are converted into processor design parameters and formatted in a data structure. The metrics / statistics are parameters that are encoded in the data exchange file format (JSON). They are directly used as generic parameters of the RTL code of the processor to generate its topology. This can (depending on the application classes) reduce the silicon surface area of ​​the processor, its consumption, while potentially increasing its operating frequency, and this without reducing its computational performance. It should be noted that this does not alter the semantics of the program in any way. The source code obtained in sub-step 210 makes it possible to contextualize the execution of the program.Application execution metrics / statistics typically indicate the instructions used, their frequency of occurrence, and their sequence. The program is manipulated to delete, substitute, or reorder instructions to reduce the instructions and / or memory transactions required to execute the program.

[0059] Figure [Fig. 3] illustrates the main sub-steps of processor generation step 160. A machine learning module comprising a machine learning model is notably used to iterate over an initial processor topology in order to refine its parameters according to the chosen design parameters.

[0060] The machine learning module may be of the supervised learning type. It may be a graph neural network (GNN). Such a graph neural network may be created using open source libraries such as PyTorch and / or TensorFlow. The training of the learning module is carried out using a dynamic database. This database evolves with the design parameters (possibly the user's design parameters, the processor source code or the processor routing) and the design parameters determined during successive iterations which are intended to converge towards a local or global optimum of a processor topology.

[0061] The neural network training is performed in an optimized way, i.e., from the pre-pruned GNN to reduce the density of the matrices. The reduction in the number of parameters of the GNN model is accompanied by efficient storage by compression of the graph adjacency matrices. For example, the Yale Sparse Matrix format can be used for this purpose. GNN-type neural networks use dense graphs and this optimization makes it possible to limit the impact of neural network training in terms of computational resources and memory usage. The machine learning module may require additional information such as design plan guidelines (e.g., PPA) for one or more parameters.

[0062] In a first sub-step 310, a first topology of a processor is generated from the analysis of the source code or the binary executable and the user's processor design parameters determined in steps 120 and 130. The generation of the topology of a first processor is described in more detail in relation to the figure [Fig. 4]. The number of design iterations of a processor topology is also initialized.

[0063] During a second sub-step 320, the design parameters are determined for the processor topology generated in the previous step. The determination of the design parameters is carried out by the inference of the neural network of the learning module applied to the generated processor topology.

[0064] During a third sub-step 325, it is determined whether the topology of the processor is satisfactory in view of at least one predefined performance instruction. In its most basic form, the performance instruction is defined as a specific value (threshold) to be reached or not exceeded for at least one of the criteria of the PPA. For example, the generated processor must not exceed a consumption of 100 milliwatts. It can also be a multi-criteria instruction. For example, a processor is sought having a consumption of 50 mW maximum, with a surface area of ​​0.1 mm 2 maximum and an operating frequency of at least 100 MHz.

[0065] In a particular embodiment, if this instruction is not defined by the user, a design space exploration step is carried out, i.e. a proposal of a set of processor architectures with different PPA instructions is made. For example, a low-power, high-frequency, small-area processor, etc. is proposed. The final selection of the most suitable architecture is made by the user. Processors are typically synthesized (RTL code to logic gate description), placed and routed on an FPGA and / or ASIC technology using a set of specific tools (open source or commercial). These tools do not necessarily converge to an optimum in a single iteration. After each iteration, the results of interest (metrics) are extracted from the reports generated by the tools.

[0066] The number of design iterations of a processor topology is also compared to a limiting number of iterations.

[0067] If the topology is not satisfactory in view of the at least one predefined performance setpoint and if the number of design iterations of a processor topology is less than a limit number of iterations, the method continues with a fourth sub-step 330, during which additional return parameters are determined by means of a machine learning module applied to the user's processor design parameters.

[0068] The method continues with a fifth sub-step 340 during which a new processor topology is generated based on the return parameters determined in the fourth sub-step 330. The new topology is determined in a similar manner to the determination of the first topology obtained at the end of the first sub-step 310, using the machine learning module. The database (“dataset” in English) of the machine learning module is enriched with the processor design parameters generated at each iteration. The neural network of the machine learning module is then trained again with this enriched database. The successive inferences (the data processing by the neural network) then benefit from this new training.

[0069] In a particular embodiment, the design instructions are modified when the user wishes to explore the design space to be offered compromises relating to the PPA instructions, or when the user's instructions do not allow the neural network to converge towards the desired solution. The method then resumes at the second step sub-step 320 applied to the new processor topology obtained at the end of the fifth sub-step 340.

[0070] If it has been determined, during the third sub-step 325, that the processor topology is satisfactory in view of the at least one predefined performance setpoint or if the number of design iterations of a processor topology is not less than a limit number of iterations, the sub-method ends with a sixth sub-step 350, and the processor topology is preserved.

[0071] Figure [Fig. 4] illustrates the main sub-steps included in step 310 of generating a first topology of a processor from the analysis of the source code or the binary executable and the user's processor design parameters. This generation involves the creation of an RTL (Register Transfer Level) type data structure making it possible to define the behavior of a circuit in terms of sending signals or transferring data between registers. It also makes it possible to define the logical operations performed on these signals.

[0072] During a first sub-step 410, the data structure of the encoded design parameters determined during step 140 is accessed.

[0073] During a second sub-step 420, design parameter values ​​of the processor are determined by browsing the encoded design parameters.

[0074] The values ​​of the design parameters are then converted into values ​​of the input parameters of a processor generator, during a third sub-step 430.

[0075] A data structure of the RTL source code of the processor is obtained by applying a processor generator to the values ​​of the determined input parameters, during a fourth sub-step 440. It should be noted that HCL (Hardware Control Language) software such as Migen generates RTL code without hierarchy (i.e. without visible relationship between the different components of the processor). Such RTL source code is “difficult to understand” or even uninterpretable by an engineer.

[0076] In contrast, the processor generator proposed here uses Python source code as input in a similar way to Migen but generates a processor that is fully readable and easily modifiable by a domain engineer. The generated RTL source code is also hierarchical and allows intervention during the precise positioning of processor components (the English term "floorplanning").

[0077] The processor can be generated through an HCL implementation using software such as Migen or similar. For example, generating the data structure from the RTL source code may include executing Python code to read the design parameter data structure and dynamically generate a processor design graph.

[0078] Figure [Fig. 5] illustrates the main sub-steps associated with determining additional return parameters of the fourth sub-step 330.

[0079] A software development kit (SDK) is then generated from the RTL processor, during a first sub-step 510. The SDK is generated automatically after analysis of the source code (if available), the binary executable and the RTL source code of the processor.

[0080] The software development kit may include a compiler, an assembler, at least one header file, at least one library, at least one boot loader, at least one kernel driver, a hardware abstraction layer (HAL) and / or other tools for a fully functional processing environment.

[0081] A data structure for emulating an in situ programmable gate array (FPGA) is generated from the RTL processor and the SDK software development kit in a second sub-step520.

[0082] In particular, we can use tools such as Yosys for the synthesis of the data structure of the RTL source code, Verilator for the simulation of the data structure of the RTL source code, nextpnr for place-route, OpenFPGA loader to load the FPGA binary stream onto a target FPGA.

[0083] As a reminder, placement-routing is a step in the design of a processor aimed at determining the placement of the different electronic components based on the expected logical result described in the RTL source code data.

[0084] During a third sub-step 530, a test plan is generated based on the design parameters and on at least one acceptance criterion. An acceptance criterion is defined based on at least one of the user parameters. For example, the operating frequency of the processor may be an acceptance criterion, such as the silicon surface area used, the energy consumption and / or the computing performance of the processor.

[0085] Tests based on the test plan are then performed in a fourth sub-step 540. The test results may be used as additional feedback parameters for the machine learning module in sub-step 340. These tests again make it possible to extract metrics for both design and execution of the program on the generated processor. These metrics enrich the training database (“dataset”) of the neural network used by the machine learning. For example, the test plan for the processor design may include using FPGAs.

Claims

CLAIMS 1. A method for automatically designing a specialized processor for executing at least one program, comprising steps in which: a. accessing the source code or a binary executable of the program, b. performing an analysis of the source code or the binary executable to obtain processor design parameters, c. encoding the design parameters in a design parameter data structure and saving the design parameter data structure, and d. generating a processor topology from the design parameter data structure via a machine learning model.

2. Method for automatically designing a specialized processor according to claim 1, in which, to analyze the source code or the binary executable, steps are carried out during which: a. a binary executable disassembler is used to obtain the assembler code from a binary executable, the binary executable being directly available or obtained via a source code compilation step, b. execution metrics and execution statistics are obtained from simulation software applied to the binary executable, and c. the execution metrics and / or execution statistics obtained are converted into processor design parameters.

3. Method for automatically designing a specialized processor according to claim 1 or 2, wherein, to carry out the generation of processor topology, steps are carried out during which: a. a first topology of a processor is generated from the analysis of the source code or the binary executable and the processor design parameters and a number of design iterations of a processor topology is initialized, b. design parameters are determined for the processor topology generated by inference of a neural network of a machine learning module trained on a predetermined database, c. it is determined whether the processor topology is not satisfactory in view of a predefined performance setpoint and whether the number of design iterations of a processor topology is less than a limit number of iterations d.if so, determining additional feedback parameters via the machine learning module applied to the user's processor design parameters, saving the feedback parameters in the neural network training database of the machine learning module, retraining the neural network with the enriched database and then generating a new processor topology based on the feedback parameters and the method continuing at the step of determining design parameters for the generated new processor topology, e. if not, the processor topology is retained.

4. Method for automatically designing a specialized processor according to claim 3, in which, when it is determined that the topology of the processor is not satisfactory in view of at least one predefined performance instruction and that the performance instruction is not a user design instruction, a step of exploring the design space is carried out, during which at least two processor topologies with different predefined performance instructions are determined, then the user chooses a topology.

5. Method for automatically designing a specialized processor according to claim 3 or 4, in which, to generate a first topology of a processor, the following steps are carried out: a. accessing the data structure of the encoded design parameters b. determining values ​​of design parameters of the processor by browsing the encoded design parameters, c. then converting the values ​​of the design parameters into values ​​of the input parameters of a processor generator, d. generating a routing-placement data structure of the processor from values ​​of the determined input parameters.

6. Method for automatically designing a specialized processor according to claim 5, wherein, to generate additional return parameters, the following steps are carried out: a. A software development kit is generated from the processor's routing-tracing level data structure, b. An in-situ programmable gate array emulation data structure is generated from the processor's routing-tracing level data structure and the software development kit, c. a test plan is generated based on the design parameters and at least one acceptance criterion based on the predefined performance setpoint, d. additional feedback parameters such as test results are obtained according to the test plan.

7. Method for automatically designing a specialized processor according to claim 6, in which the predefined performance instruction comprises a criterion of processing performance of the specialized processor, a criterion of power consumed by the specialized processor and a criterion of surface occupied by the specialized processor, the at least one acceptance criterion corresponding to a threshold of at least one of the criteria of the performance instruction or a criterion of an algebraic combination of at least two of the criteria of the performance instruction.

8. Method for automatically designing a specialized processor according to any one of claims 3 to 7, in which the automatic learning module is of the supervised learning type.

9. Method for automatically designing a specialized processor according to claim 8, in which the supervised machine learning module comprises a graph neural network, in particular a pre-pruned graph neural network.

10. Method for automatically designing a specialized processor according to any one of claims 1 to 9, in which the processor topologies are generated through an implementation in hardware description language.

Citation Information

Patent Citations

  • Automated Microprocessor Design

    US20220050946A1