A processor microcode automatic generation and optimization method and system based on deep reinforcement learning

CN122837845APending Publication Date: 2026-09-29ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610934207.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]然而,发明人在实践中发现上述方法存在以下局限性:第一,上述方法依赖预先构建和维护的原子原语库作为代码生成的基本单元,对于无固定指令集的新型处理器架构,原子原语库难以事先定义;第二,上述方法以大语言模型的一次性或多次生成为核心,策略改进依赖人工设计的提示词和反馈编码,缺乏通过与真实硬件执行环境交互而自主习得最优生成策略的能力;第三,上述方法的优化主要发生在任务图拓扑层面,最终微码序列仍由确定性算法后端按固定规则生成,无法在微码级别实现策略驱动的端到端优化

Benefits of technology

[0049]第一,无需人工定义原子原语库或编译器后端规则。策略网络通过与实际工具链和硬件仿真器交互,从零自主习得目标处理器的微码生成策略,大幅降低新型处理器架构的软件开发门槛。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837845A_ABST
    Figure CN122837845A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of embedded system compilation, processor design automation and deep reinforcement learning, and discloses a processor microcode automatic generation and optimization method and system based on deep reinforcement learning. The present application replaces the traditional compiler backend with a hardware-in-the-loop reinforcement learning framework, and constructs a closed-loop training system including a policy network module, a tool link interface module, a hardware-in-the-loop verification module and a reward calculation module. The policy network gradually selects microcode instructions from the action space, the tool chain provides instant syntax rewards, the hardware simulator provides function rewards, and the policy network is iteratively updated through a reinforcement learning algorithm until a processor microcode sequence that meets the acceptance criteria is generated. The present application supports multi-stage reward shaping, retrieval enhanced state representation, curriculum learning, self-correction, multi-objective Pareto optimization and cross-architecture transfer learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of embedded system compilation technology, processor design automation, and deep reinforcement learning technology, and particularly relates to a method and system for automatic generation and optimization of processor microcode based on deep reinforcement learning. This invention is especially applicable to embedded microcontrollers (MCUs), field-programmable gate arrays (FPGAs) soft cores, application-specific integrated circuits (ASICs), neural native processors without fixed instruction sets, and any application scenario requiring the automatic generation of low-level control sequences for specific target hardware. Background Technology

[0002] Traditional embedded software development relies on compilers to translate high-level language programs into assembly or machine code for the target processor. The compiler backend transforms the intermediate representation into machine code conforming to the target architecture's instruction set through three stages: instruction selection, register allocation, and code scheduling. This approach has been refined over decades on general-purpose processors, resulting in a mature toolchain system.

[0003] However, for new non-standard processor architectures, especially those that do not use traditional opcode tables and decoders, traditional compiler backends cannot be directly applied. A significant amount of manual work is required to re-customize backend rules, instruction selection modes, and scheduling strategies, resulting in extremely high development costs. Furthermore, traditional compilers generate code based on static rules, limiting their ability to optimize for specific hardware platforms and making it difficult to automatically explore efficient control paths that exceed the scope of the instruction set.

[0004] The applicant previously proposed an embedded code generation framework based on a large language model and an abstract task graph, as well as an iterative rewriting method for the task graph based on physical constraint feedback. This method uses the abstract task graph as an intermediate representation, parses user intent through a large language model, and maps the task graph to machine code through a deterministic algorithm backend, enabling advanced optimization of task logic at the semantic level.

[0005] However, the inventors found the following limitations in practice: First, the method relies on a pre-built and maintained atomic primitive library as the basic unit for code generation. For new processor architectures without a fixed instruction set, the atomic primitive library is difficult to define in advance. Second, the method focuses on the one-time or multiple generation of a large language model. Policy improvement depends on manually designed prompts and feedback codes, lacking the ability to autonomously learn the optimal generation strategy through interaction with the real hardware execution environment. Third, the optimization of the method mainly occurs at the task graph topology level. The final microcode sequence is still generated by the deterministic algorithm backend according to fixed rules, and policy-driven end-to-end optimization cannot be achieved at the microcode level.

[0006] Existing research has applied reinforcement learning to the chip physical design phase, for example, to optimize chip placement and routing. These methods run offline during design time, aiming to minimize trace length or power consumption, and are independent of microcode execution during processor runtime.

[0007] In the field of compiler optimization, existing research has used reinforcement learning for loop transformation selection or compiler optimization sequence search. However, these methods all operate under the constraints of traditional instruction sets, using manually defined compiler optimization operations as the action space, and cannot directly generate microcode sequences for bare-metal hardware.

[0008] Currently, there is no complete method or system that combines deep reinforcement learning with hardware-in-the-loop (HIL) feedback to automatically generate processor microcode sequences end-to-end. Summary of the Invention

[0009] The purpose of this invention is to provide a method and system for automatic generation and optimization of processor microcode based on deep reinforcement learning, so as to solve the above-mentioned technical problems.

[0010] To address the aforementioned technical problems, this invention provides a method and system for automatic microcode generation and optimization based on deep reinforcement learning. The method models the entire underlying execution control sequence generation process as a Markov decision process: the state is the joint encoding of the current task context, hardware constraint information, and the generated sequence; the action is selecting the next microcode instruction, assembly instruction, machine code fragment, or action vector from the set of control units supported by the target architecture; the reward is provided in real-time by a toolchain, hardware simulator, or FPGA prototyping platform. The policy network interacts with the environment through multiple iterations, autonomously learning how to gradually construct a complete underlying execution control sequence that simultaneously meets functional requirements and efficiency goals, starting from the requirements specification.

[0011] This invention designs a multi-stage reward function that organically combines the immediate grammatical reward generated by toolchain verification, the functional reward generated by hardware simulation, and the efficiency reward generated by code efficiency metrics. The grammatical reward is a dense reward, available immediately after each generation step, guiding the policy network to quickly learn grammatical compliance. The functional reward is a sparse reward, obtained after the sequence has completely passed through the toolchain, and is the core signal driving the policy network to converge towards the correct function. The efficiency reward encourages the policy network to explore more compact implementation methods. A reward shaping module normalizes, smooths, and dynamically weights these components, addressing the training instability problem caused by sparse rewards.

[0012] This invention employs a retrieval-enhanced representation to construct the state vector: it retrieves target processor register descriptions, peripheral memory mappings, calling conventions, and instruction set constraints relevant to the current task from a hardware knowledge base, and concatenates these with the encoded representation of the generated microcode sequence to form the enhanced state vector. This mechanism allows the policy network to utilize the target processor's hardware specifications without starting from scratch, significantly accelerating training convergence.

[0013] This invention employs a course-based learning strategy, progressively escalating the training task from generating single instructions to generating complete microcode sequences. The policy network is fully trained at each difficulty level before moving to the next stage, thus avoiding the problem of the policy network being unable to effectively explore in high-difficulty, sparse reward scenarios.

[0014] This invention also designs a self-correction mechanism: when the toolchain returns a syntax error, the error type, location and context are encoded as a correction state, driving the policy network to select rollback, replacement or insertion operations for autonomous correction, and the error correction capability is strengthened through special error correction rewards.

[0015] The specific technical solution is as follows:

[0016] A method for automatic generation and optimization of processor microcode based on deep reinforcement learning includes the following steps:

[0017] S1: Receive the processor's task execution requirements and parse the requirements into the initial state of the reinforcement learning training environment;

[0018] S2: Retrieve hardware constraint information that matches the target processor architecture identifier and task function description from the hardware knowledge base, and combine the hardware constraint information, the initial state, and the encoded representation of the currently generated underlying execution control sequence into an enhanced state. The enhanced state is dynamically updated as the currently generated underlying execution control sequence changes.

[0019] S3: Construct a reinforcement learning training environment, which includes at least: a state space containing the augmented state; an action space containing the set of microcode instructions, assembly instructions, machine code fragments, or action vectors supported by the target processor architecture; and a multi-stage reward function;

[0020] S4: The policy network selects the next underlying execution control unit from the action space based on the current enhancement state and adds it to the currently generated underlying execution control sequence;

[0021] S5: Submit the current underlying execution control sequence to the toolchain for verification; if a syntax error or compliance violation is returned, an immediate negative syntax reward is generated, and the error type, error location, and current sequence context are encoded as an error correction state. The policy network selects a rollback, replacement, or insertion operation to correct the current sequence. After correction, return to this step to resubmit for verification until verification is passed or the preset maximum number of corrections is reached before proceeding to the next step.

[0022] S6: When the toolchain verification is successful, the current underlying execution control sequence is loaded into the hardware emulator, cycle precision simulator or field-programmable gate array prototype platform for execution. The functional reward is calculated based on the degree of conformity between the execution output and the functional acceptance criteria, and the efficiency reward is calculated based on the number of instructions, code size or number of execution cycles.

[0023] S7: Update the network parameters based on grammar reward, function reward, efficiency reward, integrity penalty and error correction reward, and repeat S4 to S7 until convergence or the preset number of training rounds is reached, and output the underlying execution control sequence that meets the functional acceptance criteria and the efficiency optimization goal.

[0024] Furthermore, in the multi-stage reward function:

[0025] Syntax rewards are provided by the toolchain in real time after each generation step; positive rewards are given for successful verification and negative penalties are given for failure.

[0026] Functional rewards are executed by a hardware simulator or field-programmable gate array prototyping platform after the toolchain is verified and compared with the functional acceptance criteria, and rewards are given according to the degree of compliance.

[0027] Efficiency rewards are calculated based on the number of instructions generated in the sequence, code size, or number of execution cycles.

[0028] Integrity penalty is triggered when the sequence reaches the preset maximum length and still fails the functional verification.

[0029] Furthermore, in S3, the iterative training adopts a course learning strategy: the training tasks are progressively upgraded in terms of complexity, from single instruction generation and single peripheral access to complete multi-step computation sequences; when the average functional reward of the policy network in the current stage exceeds a preset threshold, it automatically enters the next stage; the reinforcement learning algorithm is selected from one or a combination of proximal policy optimization, deep Q-network, actor-critic algorithm, or soft actor-critic algorithm.

[0030] Furthermore, in S5:

[0031] The rollback operation includes deleting one or more underlying execution control units at the end of the current sequence;

[0032] The replacement operation includes replacing the underlying execution control unit at a specified location with a compliance unit of the same type;

[0033] Insertion operations include inserting pre-initialization, register configuration, or peripheral wait operations at a specified location.

[0034] Furthermore, in S4, the policy network uses a recurrent neural network, a long short-term memory network, or a Transformer structure to perform timing modeling of the underlying execution control sequence; the hardware knowledge base stores the target processor's register mapping, peripheral access timing, calling conventions, and instruction set constraints in the form of key-value pairs, structured tables, or vector indexes.

[0035] Furthermore, the method also includes a multi-objective Pareto optimization step: maintaining a Pareto front set during training, retaining effective low-level execution control sequences that are not mutually dominant in terms of code volume and execution cycle count, and providing the user with the set and corresponding code volume and execution cycle count metrics at the end of training for scenario preference selection.

[0036] Furthermore, the underlying execution control sequence includes at least one of the following: processor microcode sequence, assembly instruction sequence, machine code binary sequence, or action vector sequence used by a decoderless processor; the output underlying execution control sequence is derived as at least one of the following: assembly language text, machine code binary file, or high-dimensional action vector sequence used by a decoderless neural native processor architecture.

[0037] Furthermore, the method also includes a cross-architecture transfer learning step: transferring the policy network parameters, state encoder parameters, or hardware knowledge retrieval parameters trained on the source processor architecture to the target processor architecture, and fine-tuning the training based on the action space mapping relationship and hardware constraint differences between the source processor architecture and the target processor architecture.

[0038] This invention also discloses a processor microcode automatic generation and optimization system based on deep reinforcement learning, used to execute the method described above, the system comprising:

[0039] The requirement parsing module is used to parse the requirement specifications of the processor to execute tasks and output initial status information;

[0040] The hardware knowledge retrieval module is used to retrieve hardware constraint information of the target processor from the hardware knowledge base and form an enhanced state;

[0041] The policy network module is used to generate the underlying execution control unit based on the current enhancement state and receive reward signals to update parameters;

[0042] The tool linking module is used to call the assembler, linker, or compliance check tools to verify the generated underlying execution control sequence and return error information;

[0043] The hardware-in-the-loop verification module is used to execute verified low-level execution control sequences on instruction set simulators, periodic precision simulators, or field-programmable gate array prototyping platforms and return the execution results.

[0044] The reward calculation module is used to calculate reward values ​​by integrating toolchain verification results, hardware execution results, functional acceptance criteria, and efficiency optimization targets.

[0045] The error correction control module is used to trigger rollback, replacement, or insertion operations based on error information;

[0046] And a sequence output module, used to output the underlying execution control sequence that meets the functional acceptance criteria and efficiency optimization objectives after training convergence.

[0047] Furthermore, the system also includes a reward shaping module for normalizing, exponentially moving average smoothing, and dynamically adjusting the weights of multiple reward components; and a training trajectory caching module for storing historical training trajectories for the policy network to sample and reuse to improve sample efficiency.

[0048] The method and system for automatic generation and optimization of processor microcode based on deep reinforcement learning, as proposed in this invention, have the following advantages:

[0049] First, there is no need to manually define atomic primitive libraries or compiler backend rules. The policy network learns the microcode generation strategy for the target processor from scratch by interacting with the actual toolchain and hardware simulator, significantly reducing the software development threshold for new processor architectures.

[0050] Second, it possesses end-to-end microcode-level policy optimization capabilities. The optimization granularity of this invention is at the single microcode instruction level, and the policy network can explore efficient implementation paths that traditional compilers cannot discover at the microcode sequence level, including utilizing the instruction characteristics or hardware timing constraints unique to the target processor.

[0051] Third, it supports progressive convergence guided by multi-stage rewards. Syntax rewards provide dense feedback, functional rewards provide goal orientation, and efficiency rewards provide performance optimization incentives. The three types of rewards work together to drive the policy network to automatically optimize code efficiency while ensuring functional correctness.

[0052] Fourth, it possesses self-correction and error-healing capabilities. The policy network can identify errors returned by the toolchain and autonomously select corrective actions without manual intervention, thus improving the system's robustness and automation.

[0053] Fifth, the output format is compatible with multiple processor architectures. The generated microcode sequence can be output in the form of assembly language, machine code binary, or high-dimensional action vector sequence, which is suitable for traditional instruction set processors and decoderless neural native processor architectures, respectively.

[0054] Sixth, it supports cross-architecture transfer learning. The parameters of a policy network that converges on the first architecture can be transferred to the new architecture, significantly reducing the training samples and computational overhead on the new architecture. Attached Figure Description

[0055] Figure 1 A comparison diagram of the traditional compiler backend microcode generation process and the hardware-in-the-loop reinforcement learning microcode generation process of the present invention, provided for embodiments of the present invention.

[0056] Figure 2 This is a block diagram of the overall architecture of the processor microcode automatic generation system based on deep reinforcement learning provided in an embodiment of the present invention.

[0057] Figure 3 The flowchart for calculating the multi-stage reward function is provided for an embodiment of the present invention.

[0058] Figure 4 This is a schematic diagram of the closed loop of policy network inference, training, and update provided in an embodiment of the present invention. Detailed Implementation

[0059] To better understand the purpose, structure, and function of this invention, the following detailed description, in conjunction with the accompanying drawings, provides an automatic processor microcode generation and optimization method and system based on deep reinforcement learning.

[0060] To avoid ambiguity caused by differences in terminology between different processor architectures, the term "low-level execution control sequence" is used as a higher-level concept in this invention. It can refer to the assembly instruction sequence or machine code binary sequence on a traditional processor, the microcode sequence in a microprogrammed controller, or the high-dimensional action vector sequence that directly drives the execution base module in a decoder-less neural native processor architecture. Unless the context otherwise requires, the term "microcode sequence" in this invention can be understood as a specific implementation of the aforementioned low-level execution control sequence.

[0061] The present invention provides a method for automatic generation and optimization of processor microcode based on deep reinforcement learning, comprising the following steps, wherein the underlying execution control sequence includes at least one of the following: processor microcode sequence, assembly instruction sequence, machine code binary sequence, or action vector sequence used by a decoderless processor:

[0062] S1: Receive the requirements specification for the processor to execute the task. The requirements specification includes at least the target processor architecture identifier, task function description, function acceptance criteria and efficiency optimization target. Parse the requirements specification into the initial state of the reinforcement learning training environment.

[0063] S2: Retrieve hardware constraint information from the hardware knowledge base that matches the target processor architecture identifier and task function description; combine the hardware constraint information, the initial state, and the encoded representation of the currently generated underlying execution control sequence into an enhanced state; the enhanced state is dynamically updated as the currently generated underlying execution control sequence changes; the hardware constraint information includes at least one of register mapping, peripheral access timing, calling convention, or instruction set constraints.

[0064] S3: Construct a reinforcement learning training environment, the training environment including at least: a state space containing the augmented state; an action space containing the set of microcode instructions, assembly instructions, machine code fragments, or action vectors supported by the target processor architecture; and a multi-stage reward function, the multi-stage reward function including at least a syntax reward, a functional reward, an efficiency reward, and an integrity penalty;

[0065] In the multi-stage reward function, the syntax reward is fed back by the toolchain immediately after each generation step, with a positive reward for successful verification and a negative penalty for failure; the function reward is executed by the hardware simulator or field-programmable gate array prototype platform after the toolchain verification is successful, and is awarded according to the degree of compliance after comparison with the function acceptance criteria; the efficiency reward is calculated based on the number of instructions, code size, or number of execution cycles of the generated sequence; the integrity penalty is triggered when the sequence reaches the preset maximum length and still fails to pass the function verification.

[0066] Iterative training employs a course-based learning strategy: training tasks are progressively increased in complexity, from single instruction generation and single peripheral access to complete multi-step computation sequences; when the average functional reward of the policy network in the current stage exceeds a preset threshold, it automatically enters the next stage; the reinforcement learning algorithm is selected from one or a combination of proximal policy optimization, deep Q-network, actor-critic algorithm, or soft actor-critic algorithm.

[0067] S4: The policy network selects the next underlying execution control unit from the action space based on the current enhancement state and adds it to the currently generated underlying execution control sequence;

[0068] The policy network uses a recurrent neural network, a long short-term memory network, or a Transformer structure to perform timing modeling of the underlying execution control sequence; the hardware knowledge base stores the target processor's register mapping, peripheral access timing, calling conventions, and instruction set constraints in the form of key-value pairs, structured tables, or vector indexes.

[0069] S5: Submit the current underlying execution control sequence to the toolchain for verification; if a syntax error or compliance violation is returned, an immediate negative syntax reward is generated, and the error type, error location, and current sequence context are encoded as an error correction state. The policy network selects a rollback, replacement, or insertion operation to correct the current sequence. After correction, return to this step to resubmit for verification until verification is passed or the preset maximum number of corrections is reached before proceeding to the next step.

[0070] Rollback operations include deleting one or more underlying execution control units at the end of the current sequence; replacement operations include replacing the underlying execution control unit at a specified position with a compliant unit of the same type; and insertion operations include inserting pre-initialization, register configuration, or peripheral wait operations at a specified position.

[0071] S6: When the toolchain verification is successful, the current underlying execution control sequence is loaded into the hardware emulator, cycle precision simulator or field-programmable gate array prototype platform for execution. The functional reward is calculated based on the degree of conformity between the execution output and the functional acceptance criteria, and the efficiency reward is calculated based on the number of instructions, code size or number of execution cycles.

[0072] The output low-level execution control sequence is derived as at least one of assembly language text, machine code binary file, or high-dimensional action vector sequence used by decoderless neural native processor architecture.

[0073] S7: Update the strategy network parameters based on the syntax reward, function reward, efficiency reward, integrity penalty and error correction reward, and repeat S4 to S7 until convergence or the preset training rounds are reached, and output the underlying execution control sequence that meets the functional acceptance criteria and the efficiency optimization goal.

[0074] The method also includes a multi-objective Pareto optimization step: during training, a Pareto front set is maintained, retaining effective low-level execution control sequences that are not mutually dominant in terms of code volume and execution cycle count. At the end of training, the set and the corresponding code volume and execution cycle count metrics are provided to the user for scenario preference selection.

[0075] The method further includes a cross-architecture transfer learning step: transferring the policy network parameters, state encoder parameters, or hardware knowledge retrieval parameters trained on the source processor architecture to the target processor architecture, and fine-tuning the training based on the action space mapping relationship and hardware constraint differences between the source and target processor architectures.

[0076] The present invention provides an automatic processor microcode generation and optimization system based on deep reinforcement learning, comprising:

[0077] The requirement parsing module is used to parse the requirement specifications of the processor to execute tasks and output initial status information;

[0078] The hardware knowledge retrieval module is used to retrieve hardware constraint information of the target processor from the hardware knowledge base and form an enhanced state;

[0079] The policy network module is used to generate the underlying execution control unit based on the current enhancement state and receive reward signals to update parameters;

[0080] The tool linking module is used to call the assembler, linker, or compliance check tools to verify the generated underlying execution control sequence and return error information;

[0081] The hardware-in-the-loop verification module is used to execute verified low-level execution control sequences on instruction set simulators, periodic precision simulators, or field-programmable gate array prototyping platforms and return the execution results.

[0082] The reward calculation module is used to calculate reward values ​​by integrating toolchain verification results, hardware execution results, functional acceptance criteria, and efficiency optimization targets.

[0083] The error correction control module is used to trigger rollback, replacement, or insertion operations based on error information;

[0084] And a sequence output module, used to output the underlying execution control sequence that meets the functional acceptance criteria and efficiency optimization objectives after training convergence.

[0085] The system also includes a reward shaping module for normalizing, exponentially moving average smoothing, and dynamically adjusting the weights of multiple reward components; and a training trajectory caching module for storing historical training trajectories for the policy network to sample and reuse to improve sample efficiency.

[0086] Example 1: Overall System Architecture

[0087] like Figure 2 As shown, the processor microcode automatic generation system based on deep reinforcement learning of the present invention includes a requirement parsing module, a policy network module, a tool linking module, a hardware-in-the-loop verification module, a reward calculation module, and a microcode output module. These six modules constitute a complete hardware-in-the-loop reinforcement learning (HIL-RL) closed loop. In one embodiment, the requirement parsing module has a built-in or connected hardware knowledge retrieval module, and an error correction control module is set between the tool linking module and the policy network module. The above sub-modules can also be deployed as independent modules.

[0088] The requirements parsing module receives the task requirements specification. This specification can be natural language text (e.g., "send the string DNS__ALIVE every 500ms via UART") or a structured Hardware Requirements Description (HRD), which must at least include a target architecture identifier field (e.g., "riscv32"), a task description field, and acceptance criteria fields (e.g., the expected output string and its minimum occurrence count). The requirements parsing module parses this information into initial state information, which is used to initialize the state encoder of the policy network module.

[0089] The hardware knowledge retrieval module retrieves the target processor's register mappings, peripheral addresses, access timings, calling conventions, and instruction set constraints from the hardware knowledge base based on the target architecture identifier field and the task description field. It then combines the retrieval results with the initial state information to form enhanced state information. This enhanced state information serves as input to the policy network module, enabling the policy network to be explicitly guided by the target hardware constraints when generating the underlying execution control sequence.

[0090] The policy network module maintains the currently generated microcode sequence based on the current state and appends the current microcode instruction or action vector from the action space to the sequence based on neural network inference. When the policy network module is in training mode, it simultaneously receives the comprehensive reward signal to update the network parameters; when in inference mode, it directly generates the microcode sequence using the converged parameters.

[0091] The toolchain module receives the currently generated microcode sequence, calls the assembler and linker of the target architecture to compile and verify it, and returns the verification results and error information to the reward calculation module. If the toolchain verification fails (e.g., the assembler reports a syntax error), the hardware simulation step is skipped, and the reward calculation module directly generates a negative syntax penalty signal.

[0092] The error correction control module receives error information returned by the tool link port module, encodes the error type, error location, and current sequence context into an error correction status, and provides the policy network module with three types of error correction actions: rollback, replacement, and insertion. After the policy network module selects the appropriate correction action based on the error correction status, the tool link port module verifies the corrected sequence again, thus forming an automatic error repair closed loop.

[0093] The hardware-in-the-loop verification module receives the executable binary that has passed toolchain verification, executes it on the instruction set simulator, and returns the execution output to the reward calculation module. In one implementation, the hardware-in-the-loop verification module uses QEMU to simulate a RISC-V bare-metal environment, captures sequential output (e.g., serial port printouts), and compares it with the expected output in the acceptance criteria.

[0094] The reward calculation module integrates the toolchain verification results and simulation output results, calculates the comprehensive reward signal according to the acceptance criteria, and sends it back to the policy network module.

[0095] The microcode output module stores the microcode sequence that meets the acceptance criteria after training convergence, and provides the optimized microcode sequence to the outside world in the form of assembly language text, machine code binary, or high-dimensional action vector sequence.

[0096] Example 2: Multi-stage reward function design

[0097] like Figure 3 As shown, the multi-stage reward function calculation process includes three components: syntax reward, functional reward, and efficiency reward.

[0098] The toolchain verification node receives the currently generated microcode sequence and calls the assembler to perform syntax checks. If assembly is successful, a positive syntax reward component (e.g., +0.1) is generated; if assembly fails, a negative syntax penalty (e.g., -1.0) is generated, and the error information is encoded into the next state, triggering a self-correction process (see Example 5). The syntax reward is a dense reward, available immediately after each generation step, helping the policy network quickly learn syntax compliance and solving the problem of low exploration efficiency caused by sparse functional rewards.

[0099] Once the toolchain verification is complete, the process moves to the hardware simulation execution node. The executable binary is loaded into QEMU for execution, UART output is captured, and compared with the expected output in the acceptance criteria (e.g., the string "DNS__ALIVE" appears at least twice). The functional bonus component is calculated. A full positive functional bonus (e.g., +10.0) is generated when the acceptance criteria are fully met; an intermediate bonus is generated based on the proportion of matching output bytes when the criteria are partially met.

[0100] The code efficiency analysis node counts the total number of instructions in the generated microcode sequence or the size of the compiled code, and calculates the efficiency reward component. When the total number of instructions in the generated microcode sequence is less than the preset reference implementation baseline, provided that the function is correct, a positive efficiency reward is generated (e.g., +0.05 for each instruction less), which incentivizes the policy network to explore a more compact implementation.

[0101] The reward shaping module dynamically weights and sums the three reward components mentioned above to obtain a comprehensive reward value. In the early stages of training, the weight of the syntax reward is increased to accelerate basic compliance learning; in the mid-to-late stages, the weight of the function reward is increased to focus on function convergence; and in the final stage, the weight of the efficiency reward is increased to incentivize code optimization. Furthermore, when the sequence length exceeds a preset threshold (e.g., 100 instructions) and the function still fails, an integrity penalty (e.g., -5.0) is applied to prevent the policy network from falling into an infinite generation loop.

[0102] Example 3: Retrieval Enhancement State Representation

[0103] This embodiment describes a method for constructing policy network state vectors using retrieval-enhanced generation (RAG) technology.

[0104] The hardware knowledge base stores the target processor's peripheral register address mapping, read / write operation timing of each register, general-purpose register conventions, system call conventions, and instruction set cheat sheets in key-value pairs. For example, for the RISC-VQEMUVirt target platform, the knowledge base stores information such as the UART base address (0x10000000), transmit holding register offset (+0x00), line status register offset (+0x05), and transmit ready bit position (bit 5).

[0105] When parsing the task requirements specification, the requirements parsing module retrieves relevant entries from the hardware knowledge base based on keywords (such as "UART" and "transmit") in the task function description and encodes them into a knowledge context vector. The policy network module concatenates the knowledge context vector with the Transformer encoded representation of the currently generated microcode sequence to form an enhanced state vector, which serves as the input to the policy network.

[0106] By employing retrieval-enhanced representations, the policy network can learn the operational addresses and timing constraints of the target peripheral without starting from scratch. This allows it to focus more exploration resources on generating logically correct instruction sequences, significantly accelerating training convergence. In control experiments, the policy network using retrieval-enhanced representations required approximately 60% fewer training epochs to achieve functional convergence compared to the baseline network without this mechanism.

[0107] Example 4: Course Learning and Training Strategies

[0108] This embodiment describes a method for gradually increasing the difficulty of training tasks using a course learning strategy.

[0109] The course divides the training tasks into the following stages according to their complexity:

[0110] Phase 1 (Beginner): Single instruction generation task, such as generating the li instruction that loads a specific immediate value into a register; the acceptance criteria are that the instruction assembly passes and the expected register state is generated in the simulator.

[0111] The second stage (basic): single peripheral access task, such as generating a sequence of instructions to send a single character to the UART (usually requiring 3 to 5 instructions, including loading the base address, waiting for transmission to be ready, and writing data).

[0112] Phase 3 (Intermediate): Multi-step computation sequence tasks, such as generating instruction sequences that implement simple arithmetic operations or string processing, which may include conditional statements and simple loop structures.

[0113] Phase 4 (Advanced): Complete microcode task, such as generating a bare-metal program that meets all acceptance criteria, including complete functional modules such as initialization, main loop, delay and peripheral communication.

[0114] When the policy network's average functional reward exceeds a preset threshold (e.g., 80% of the maximum functional reward) across 100 consecutive evaluation episodes in the current stage, the system automatically advances to the next stage. This learning process avoids the problem of policy networks being unable to effectively explore challenging tasks with sparse rewards, allowing the training process to progress steadily from simple to complex.

[0115] Example 5: Self-correction mechanism

[0116] This embodiment describes the mechanism by which the policy network performs autonomous error correction when toolchain verification fails.

[0117] When the tool linker module returns an assembly error, the error message should at least include the error type (e.g., "unknown instruction", "register number out of range", "operand format mismatch") and the error line number. The requirement parsing module 201 encodes the error type and error location into independent error feature vectors, concatenates them with the current sequence encoding and the knowledge context vector to form an error correction state, and inputs it into the policy network module.

[0118] The action space of the policy network in the error correction state is expanded to include three types of correction operations: rollback operation, which deletes one or more instructions at the end of the generated sequence; replacement operation, which replaces the instruction at a specified line number with another compliant instruction of the same type; and insertion operation, which inserts a missing pre-configuration instruction at the current position (e.g., supplementing an initialization instruction before using a register).

[0119] The reward calculation module sets up special error correction rewards for error correction actions: a positive error correction reward (e.g., +2.0) is generated when the toolchain verification that was originally failed is successfully passed, and an additional penalty is added when multiple consecutive error correction failures occur to encourage the policy network to abandon the current unrepairable sequence prefix and revert to an earlier checkpoint to regenerate.

[0120] In a typical training case, the policy network incorrectly used register x0 as the target write register when generating the UART transmit sequence (x0 is a zero register in RISC-V, so writing to it is invalid). Although the toolchain passed the assembly, the simulator output was empty; the self-correction mechanism encoded "expected output not appearing" as a functional error correction state, driving the policy network to replace the relevant instructions with those using writable registers, ultimately correcting the functional error.

[0121] Example 6: Multi-objective Pareto Optimization and Offline Deployment

[0122] Multi-objective Pareto optimization:

[0123] During training, the reward calculation module records the code size (number of bytes) and execution cycle count (obtained through a cycle-accurate simulator) of each valid microcode sequence that meets the functional acceptance criteria. These two metrics are defined as a bi-objective minimization problem, maintaining a Pareto front set. The microcode sequences retained in the Pareto front set satisfy the following condition: for any two sequences in the set, there is no sequence that is superior to the other in both code size and execution cycle count.

[0124] After training, the microcode output module provides the user with all microcode sequences in the Pareto front set, along with annotations for the code size and execution cycle count of each sequence. The user can then select the solution from the Pareto front that best suits their needs, based on the flash storage limitations and real-time requirements of the target device.

[0125] Offline deployment:

[0126] The selected microcode sequence is exported as an executable binary file for the target processor architecture via the microcode output module and written to the Flash storage of the target device. The deployed device directly reads and executes the microcode sequence from the Flash, without the need for a policy network to participate in online inference. For target devices employing a decoder-less neural native processor architecture, the microcode sequence is output as a high-dimensional action vector sequence and written to the action vector storage module. The processor architecture then reads this sequence sequentially and directly drives the execution of the base module.

[0127] Example 7: Real Hardware Verification Based on FPGA-in-the-Loop

[0128] In one implementation, the hardware-in-the-loop verification module uses a field-programmable gate array (FPGA) prototyping platform instead of a software simulator to achieve real hardware-in-the-loop (HIL) training.

[0129] The RTL description of the target processor is instantiated in the FPGA prototyping platform, and the FPGA is connected to the training host via a serial port or JTAG interface. The microcode sequence generated by the toollink module is a binary file, which is loaded into the target processor's program memory through the FPGA programming interface. After execution by the FPGA, the output result is returned to the training host via the serial port. The hardware-in-the-loop verification module captures the serial port output and sends it to the reward calculation module for functional reward calculation.

[0130] Compared to software simulators, FPGA-in-the-loop training can capture real hardware timing behavior, including peripheral bus access latency, interrupt response timing, and pipelined effects, making the trained microcode policies more reliable on real hardware. Furthermore, FPGA-in-the-loop training can calculate precise efficiency rewards based on actual execution cycles, guiding the policy network to optimize microcode for real hardware timing.

[0131] Example 8: Complete Example of RISC-V Bare-Metal Microcode Generation

[0132] This embodiment describes the complete process of training and generating a cyclic UART output microcode sequence using the method of this invention on the RISC-VQEMUVirt target platform.

[0133] The requirements parsing module parses the following structured HRD file: the target architecture is RISC-VQEMUVirtUART, the task description is "continuously and repeatedly send the string 'DNS__ALIVE' with a newline via UART", and the acceptance criterion is that the string "DNS__ALIVE" appears at least twice in the simulation output. The requirements parsing module retrieves relevant entries for RISC-VQEMUVirtUART from the hardware knowledge base, obtaining information such as the base address (0x10000000), transmit register offset (+0x00), and transmit readiness detection method (polling LSR register bit 5).

[0134] The policy network module adopts the configuration of the third stage of the course. After training convergence, it generates the following core instruction sequence: load the UART base address to t0; add an outer loop label; load the ASCII code character by character to t1 and write it to offset t0+0 via the sb instruction; insert an LSR polling wait loop before each character is written; generate a delay loop after the string is output; jump back to the outer loop label. After the toolchain verification is successful, the QEMU simulation output contains the string "DNS__ALIVE" multiple times, and the functional reward reaches the preset pass threshold. The number of instructions, code size, and execution cycles of the final generated sequence are recorded by the reward calculation module and can be compared with the preset reference implementation benchmark for efficiency reward calculation.

[0135] Example 9: Cross-Architecture Transfer Learning

[0136] This embodiment describes a method for migrating a trained policy network to a new target processor architecture.

[0137] After training on the source processor architecture, the system saves the policy network parameters, state encoder parameters, and hardware knowledge retrieval unit parameters. When migrating to the target processor architecture, the requirement resolution module 201 reads the target architecture identifier and retrieves the target architecture's register mappings, peripheral addresses, calling conventions, and instruction set constraints from the hardware knowledge base. The system establishes an action mapping table based on the correspondence between the source and target architecture action spaces. For example, register loading, storage, conditional jumps, and peripheral write operations in the source architecture are mapped to functionally equivalent or nearly equivalent low-level execution control units in the target architecture.

[0138] For the state encoding layer, task semantic encoding layer, and reward shaping parameters shared by the source and target architectures, the system directly loads the training parameters from the source architecture. For the action output layer and the embedding layer, which is strongly correlated with the target hardware constraints, the system re-initializes or partially initializes the action space of the target architecture. Subsequently, fine-tuning training is performed on the toolchain and hardware-in-the-loop verification module of the target architecture. During the fine-tuning phase, the closed-loop feedback of syntactic rewards, functional rewards, and efficiency rewards is preserved, but the exploration amplitude is reduced and the stability constraints on the passed sequences are increased.

[0139] Through the aforementioned transfer learning mechanism, the policy network can utilize the task decomposition, sequence organization, and error correction policies already learned in the source architecture without having to be trained from scratch on each new processor architecture, thereby reducing the training sample requirements and hardware-in-the-loop verification overhead on the target architecture.

[0140] This invention can be used for the automatic software development of embedded microcontrollers, field-programmable gate array soft cores, application-specific integrated circuits and edge intelligent chips, and is suitable for the automatic generation of bare-metal programs in scenarios such as IoT sensor nodes, motor controllers, real-time control systems, industrial edge computing devices and low-power wearable devices.

[0141] For novel processors employing non-standard instruction sets or decoder-less neural native processor architectures, this invention provides a general method for automatically generating target code without requiring manual compilation of the compiler backend, offering broad industrial applicability. The reinforcement learning algorithm, neural network architecture, toolchain calling interface, and hardware simulator used in this invention are all existing technologies or can be implemented using existing engineering methods, thus possessing engineering feasibility.

[0142] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A method for automatic generation and optimization of processor microcode based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Receive the processor's task execution requirements and parse the requirements into the initial state of the reinforcement learning training environment; S2: Retrieve hardware constraint information that matches the target processor architecture identifier and task function description from the hardware knowledge base, and combine the hardware constraint information, the initial state, and the encoded representation of the currently generated underlying execution control sequence into an enhanced state. The enhanced state is dynamically updated as the currently generated underlying execution control sequence changes. S3: Construct a reinforcement learning training environment, the training environment including at least: a state space containing the reinforcement states; The action space includes the set of microcode instructions, assembly instructions, machine code fragments, or action vectors supported by the target processor architecture; and a multi-stage reward function. S4: The policy network selects the next underlying execution control unit from the action space based on the current enhancement state and adds it to the currently generated underlying execution control sequence; S5: Submit the current underlying execution control sequence to the toolchain for verification; if a syntax error or compliance violation is returned, an immediate negative syntax reward is generated, and the error type, error location, and current sequence context are encoded as an error correction state. The policy network selects a rollback, replacement, or insertion operation to correct the current sequence. After correction, return to this step to resubmit for verification until verification is passed or the preset maximum number of corrections is reached before proceeding to the next step. S6: When the toolchain verification is successful, the current underlying execution control sequence is loaded into the hardware emulator, cycle precision simulator or field-programmable gate array prototype platform for execution. The functional reward is calculated based on the degree of conformity between the execution output and the functional acceptance criteria, and the efficiency reward is calculated based on the number of instructions, code size or number of execution cycles. S7: Update the network parameters based on grammar reward, function reward, efficiency reward, integrity penalty and error correction reward, and repeat S4 to S7 until convergence or the preset number of training rounds is reached, and output the underlying execution control sequence that meets the functional acceptance criteria and the efficiency optimization goal.

2. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, In the multi-stage reward function: Syntax rewards are provided by the toolchain in real time after each generation step; positive rewards are given for successful verification and negative penalties are given for failure. Functional rewards are executed by a hardware simulator or field-programmable gate array prototyping platform after the toolchain is verified and compared with the functional acceptance criteria, and rewards are given according to the degree of compliance. Efficiency rewards are calculated based on the number of instructions generated in the sequence, code size, or number of execution cycles. Integrity penalty is triggered when the sequence reaches the preset maximum length and still fails the functional verification.

3. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, In S3, iterative training adopts a course learning strategy: the training task is progressively upgraded in terms of complexity, from single instruction generation and single peripheral access to a complete multi-step computation sequence; when the average functional reward of the policy network in the current stage exceeds a preset threshold, it automatically enters the next stage; the reinforcement learning algorithm is selected from one or a combination of proximal policy optimization, deep Q-network, actor-critic algorithm, or soft actor-critic algorithm.

4. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, In S5: The rollback operation includes deleting one or more underlying execution control units at the end of the current sequence; The replacement operation includes replacing the underlying execution control unit at a specified location with a compliance unit of the same type; Insertion operations include inserting pre-initialization, register configuration, or peripheral wait operations at a specified location.

5. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the policy network uses a recurrent neural network, a long short-term memory network, or a Transformer structure to perform timing modeling of the underlying execution control sequence; the hardware knowledge base stores the target processor's register mapping, peripheral access timing, calling conventions, and instruction set constraints in the form of key-value pairs, structured tables, or vector indexes.

6. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, The method also includes a multi-objective Pareto optimization step: during training, a Pareto front set is maintained, retaining effective low-level execution control sequences that are not mutually dominant in terms of code volume and execution cycle count. At the end of training, the set and the corresponding code volume and execution cycle count metrics are provided to the user for scenario preference selection.

7. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, The underlying execution control sequence includes at least one of the following: processor microcode sequence, assembly instruction sequence, machine code binary sequence, or action vector sequence used by a decoderless processor; the output underlying execution control sequence is derived as at least one of the following: assembly language text, machine code binary file, or high-dimensional action vector sequence used by a decoderless neural native processor architecture.

8. The method for automatic generation and optimization of processor microcode based on deep reinforcement learning according to claim 1, characterized in that, The method further includes a cross-architecture transfer learning step: transferring the policy network parameters, state encoder parameters, or hardware knowledge retrieval parameters trained on the source processor architecture to the target processor architecture, and fine-tuning the training based on the action space mapping relationship and hardware constraint differences between the source and target processor architectures.

9. A processor microcode automatic generation and optimization system based on deep reinforcement learning, used to execute the method as described in any one of claims 1-8, characterized in that, The system includes: The requirement parsing module is used to parse the requirement specifications of the processor to execute tasks and output initial status information; The hardware knowledge retrieval module is used to retrieve hardware constraint information of the target processor from the hardware knowledge base and form an enhanced state; The policy network module is used to generate the underlying execution control unit based on the current enhancement state and receive reward signals to update parameters; The tool linking module is used to call the assembler, linker, or compliance check tools to verify the generated underlying execution control sequence and return error information; The hardware-in-the-loop verification module is used to execute verified low-level execution control sequences on instruction set simulators, periodic precision simulators, or field-programmable gate array prototyping platforms and return the execution results. The reward calculation module is used to calculate reward values ​​by integrating toolchain verification results, hardware execution results, functional acceptance criteria, and efficiency optimization targets. The error correction control module is used to trigger rollback, replacement, or insertion operations based on error information; And a sequence output module, used to output the underlying execution control sequence that meets the functional acceptance criteria and efficiency optimization objectives after training convergence.

10. The processor microcode automatic generation and optimization system based on deep reinforcement learning according to claim 9, characterized in that, The system also includes a reward shaping module, which is used to normalize, exponentially move average smooth, and dynamically adjust the weights of multiple reward components. It also includes a training trajectory caching module, which stores historical training trajectories for the policy network to sample and reuse, thereby improving sample efficiency.