Instruction processing system and method based on dynamic configuration register and processor
Through dynamic configuration of the register system, the problems of multi-precision calculation and energy efficiency optimization in edge computing are solved, and efficient and flexible computing resource management is achieved to adapt to the low-power consumption needs of edge devices.
Patent Information
- Application Number
- CN202510387881.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-15
AI Technical Summary
The existing processor architectures have problems in the edge computing with insufficient multi-precision computing support, poor energy efficiency optimization, and poor computing resource flexibility, which is difficult to meet the needs of hybrid precision computing and multi-protocol data format compatible processing.
The dynamic configuration register system is adopted, including a register array, a logic control unit, an independent power control unit and a dynamic voltage regulation unit. By dynamically combining low-bit physical registers and real-time voltage regulation, computing resource configuration and energy consumption management are optimized.
It realizes efficient and flexible computing capabilities in multi-precision computing, reduces power consumption, adapts to the miniaturization needs of edge computing devices, and improves energy efficiency ratio.
Smart Images

Figure CN120492030A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer systems, and in particular to an instruction processing system, method, and processor based on dynamic configuration registers. Background Art
[0002] With the widespread application of edge computing scenarios, current edge computing devices and chips have put forward new requirements for specific application scenarios such as mixed-precision computing and multi-protocol data format compatible processing: on the one hand, flexible configuration of computing resources is required, and on the other hand, there is an urgent need to reduce computing power consumption.
[0003] When faced with these demands, edge devices based on existing processor architectures and optimization technologies have numerous shortcomings. For example, computing devices based on ARM and X86 architectures lack support for multi-precision computing and energy efficiency optimization. RISC-V, as an open-source instruction set architecture, is highly scalable, but its support for low-order physical register design and dynamic combination is limited. DSPs are optimized for specific computing tasks, such as FFT and filter implementation, but their versatility is poor and they struggle to meet diverse application requirements. GPU architectures support large-scale parallel computing, but their high power consumption and large chip area make them unsuitable for resource-constrained scenarios like edge computing. Another type of software optimization technology virtually splits high-order registers into multiple low-order physical registers based on computing needs during the computational process to increase computing resource flexibility. However, these high-order registers always operate at full power, making it difficult to meet the low-power requirements of edge devices. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an instruction processing system, method, processor, chip and storage medium based on dynamic configuration registers.
[0005] In a first aspect, an embodiment of the present invention provides an instruction processing system based on a dynamic configuration register, the system comprising:
[0006] A register array comprising a plurality of low-order physical registers dynamically connected via a programmable interconnect medium;
[0007] The logic control unit includes a control module equipped with a one-to-one correspondence with the multiple low-order physical registers, and is used to dynamically configure the connection mode, data flow direction and operation mode between the registers according to the configuration signal received and processed from the central controller.
[0008] An independent power control unit is provided for each low-order physical register, so that the power supply voltage of each low-order physical register can be set independently;
[0009] The dynamic voltage regulation unit is used to control the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the operation in real time according to the operation accuracy requirements corresponding to the currently received instruction.
[0010] Optionally, the system further comprises:
[0011] The carry and overflow detection circuit is used to monitor and process the carry and overflow status information between multiple low-order physical registers currently participating in the operation in real time, and to feed back the carry and overflow status information to the corresponding low-order physical registers. It is also used to update the value in the flag register after each instruction cycle.
[0012] Optionally, the system further comprises:
[0013] The asynchronous timing control module is used to control different register units to participate in instruction processing in different clock cycles.
[0014] Optionally, the system further comprises:
[0015] The multi-layer clock domain control module determines the corresponding clock domain to be used according to the calculation accuracy requirements corresponding to the currently received instruction.
[0016] Optionally, determining a corresponding clock domain to be used according to a computational accuracy requirement corresponding to a currently received instruction specifically includes:
[0017] When the calculation precision corresponding to the currently received instruction is low precision, the high-precision clock domain is closed.
[0018] In a second aspect, an embodiment of the present invention provides an instruction processing method based on the instruction processing system of the first aspect, characterized in that the method includes:
[0019] According to the current instruction, the central controller sends a configuration signal to each low-order physical register unit to determine the number and connection method of the low-order physical registers required to run the current instruction;
[0020] Adjust the data flow direction and operation mode of each low-order physical register involved in running the current instruction, and establish corresponding connections with adjacent low-order physical registers;
[0021] According to the calculation accuracy requirements of the currently received instruction, the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the calculation is controlled in real time;
[0022] The corresponding clock domain to be used is determined according to the calculation accuracy requirement corresponding to the currently received instruction, and all low-order physical registers involved in instruction processing are controlled to work collaboratively within the clock domain.
[0023] Optionally, the method further includes:
[0024] After each instruction cycle, the carry and overflow detection circuits update the values in the flag registers, while the dynamic voltage regulation unit adjusts the power supply strategy of each low-order physical register according to the requirements of the next instruction processing.
[0025] In a third aspect, an embodiment of the present invention provides a processor, characterized in that the processor includes the instruction processing system described in the first aspect.
[0026] In a fourth aspect, an embodiment of the present invention provides a chip, characterized in that the chip includes the instruction processing system described in the first aspect.
[0027] In a fifth aspect, an embodiment of the present invention provides a storage medium on which computer program instructions are stored, characterized in that when the computer program instructions are executed, the instruction processing method described in the second aspect is implemented.
[0028] The instruction processing system, method, processor, chip and storage medium based on dynamic configuration registers provided by the embodiments of the present invention demonstrate their excellent capabilities in multi-precision calculations throughout the entire process of receiving instructions, responding to requirements and completing processing through dynamic combination control of all low-order physical registers. The compact design of low-order physical registers reduces the physical implementation area, which is conducive to the integration of miniaturized devices. With the application of edge scenarios as the goal, dynamic voltage regulation and energy consumption management technology are adopted to enable it to maintain excellent energy efficiency in different computing tasks, providing strong technical support for intelligent data acquisition, edge computing equipment and high-performance computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application.
[0030] Figure 1 A schematic diagram of the structure of an instruction processing system based on a dynamically configurable register provided by an embodiment of the present invention;
[0031] Figure 2 A flowchart of an instruction processing method provided by an embodiment of the present invention;
[0032] Figure 3 A flowchart illustrating an example process of a method for processing a 32-bit addition instruction provided by an embodiment of the present invention;
[0033] Figure 4 A schematic diagram of the structure of a chip provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0035] Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0036] Edge computing is a distributed computing architecture that pushes computing tasks and data storage from centralized data centers to the edge of the network, close to where the data is generated.
[0037] With the rapid development of the Internet of Things (IoT), autonomous driving, smart manufacturing, telemedicine, and other fields, edge computing is becoming increasingly widespread. These scenarios demand fast response, low latency, high performance, and data privacy protection. The diverse nature of edge computing applications places higher demands on the performance, flexibility, and compatibility of edge computing devices and chips. These devices and chips must adapt to diverse application scenarios and process a variety of data types while maintaining low power consumption and high efficiency.
[0038] With the widespread application of edge computing scenarios, current edge computing devices and chips have put forward new requirements for specific application scenarios such as mixed-precision computing and multi-protocol data format compatible processing: on the one hand, flexible configuration of computing resources is required, and on the other hand, there is an urgent need to reduce computing power consumption.
[0039] Mixed-precision computing refers to the use of data representations of different precisions (such as 32-bit floating point, 16-bit floating point, and 8-bit integers) within the same computing task. In edge computing, mixed-precision computing can improve computing efficiency and reduce energy consumption while maintaining sufficient computing accuracy to meet application requirements. For example, in autonomous driving, the high demand for real-time performance may require certain computing tasks to use lower precision to speed up processing, while safety-critical tasks use higher precision to ensure accuracy.
[0040] Furthermore, edge computing devices need to process data from various devices and sensors, which may use different communication protocols and data formats. Therefore, edge computing devices and chips must be able to process data in multiple protocols and formats, parsing, converting, and storing data in various formats. This helps achieve seamless data integration and efficient utilization, and promotes interoperability between different devices and systems.
[0041] When faced with these demands, edge devices based on existing processor architectures and optimization technologies have numerous shortcomings. For example, computing devices based on ARM and X86 architectures lack support for multi-precision computing and energy efficiency optimization. RISC-V, as an open-source instruction set architecture, is highly scalable, but its support for low-order physical register design and dynamic combination is limited. DSPs are optimized for specific computing tasks, such as FFT and filter implementation, but their versatility is poor and they struggle to meet diverse application requirements. GPU architectures support large-scale parallel computing, but their high power consumption and large chip area make them unsuitable for resource-constrained scenarios like edge computing. Another type of software optimization technology virtually splits high-order registers into multiple low-order physical registers based on computing needs during the computational process to increase computing resource flexibility. However, these high-order registers always operate at full power, making it difficult to meet the low-power requirements of edge devices.
[0042] Based on this, the embodiment of the present invention proposes an instruction processing system based on dynamic configuration registers, as shown in the attached Figure 1 As shown, the system specifically includes:
[0043] The register array 110 includes a plurality of low-order physical registers dynamically connected via a programmable interconnect medium.
[0044] The instruction processing system in the embodiments of the present invention is also referred to as HEX4. HEX4 utilizes multiple modular low-order physical registers as the basic register group, which, depending on the type, specifically includes a low-order floating-point register group, a low-order integer register group, a low-order vector register group, and a low-order flag register group. The "low-order physical registers" in the embodiments of the present invention refer to both native 4-bit, 8-bit, and 16-bit physical registers in the system, as well as physical registers of other bit widths in a custom system, such as 3-bit, 7-bit, or any other custom bit width. The "low-order physical registers" in the "low-order physical registers" are relative to high-order bit width registers such as 256-bit and 512-bit, and are used to illustrate the concept that high-order bit width instruction operations can be achieved through the dynamic combination of low-order physical registers in the embodiments of the present invention. Therefore, the "low-order physical registers" in the embodiments of the present invention can be understood as the basic physical registers involved in instruction operations in the system, whose bit width is completely customized by the system to a specific bit width value, and all low-order physical registers in the system have this bit width value. Optionally, it can be considered that the bit width of the low-order physical registers is generally limited to no more than 128 bits or 256 bits.
[0045] These low-order physical registers can be dynamically connected via programmable interconnect media, such as spanning wires or network media, providing extreme flexibility. This feature allows HEX4 to flexibly adjust its internal structure to optimize performance based on different computing requirements.
[0046] Specifically, each low-order physical register unit in HEX4 is an independent computing unit capable of storing and processing low-order commonly used data types. These units do not exist in isolation, but can be connected through a series of programmable interconnect networks. In practical applications, when higher-precision calculations are required, HEX4 can dynamically program multiple low-order physical register units together to form a larger register group. In this way, HEX4 can support a wide range of needs, from simple low-order operations to complex high-precision operations.
[0047] For example, when the system performs 8-bit operations and the low-order physical register is 4 bits, HEX4 can combine two 4-bit physical registers into an 8-bit register; when performing 16-bit operations, four low-order physical registers can be combined; and so on, until the required operation accuracy is achieved. This flexible combination method not only improves the computing power of HEX4, but also enables it to more efficiently utilize hardware resources, reducing power consumption and costs. Unlike existing technologies, the registers in HEX4 are all low-order physical registers to meet the flexible requirements of low-bit width in edge computing, rather than traditional architectures such as x86, ARM, and RISC-V, which are usually fixed at 128, 256, or 512 bits. In low-bit width operation scenarios for edge computing scenarios, low-order physical registers need to be split from high-order registers to participate in the operation, resulting in a large amount of power waste and unable to meet the low-power requirements of edge devices. The compact design of low-order physical registers reduces the physical implementation area, which is conducive to the integration of miniaturized devices.
[0048] Furthermore, HEX4's programmable interconnect network allows users to customize the connections between register units and data transmission paths based on specific application requirements. This customization capability enables HEX4 to better adapt to various complex computing scenarios, providing users with more flexible and efficient computing solutions.
[0049] The logic control unit 120 includes control modules corresponding to the multiple low-order physical registers, and is used to dynamically configure the connection mode, data flow direction and operation mode between the registers according to the configuration signal received and processed from the central controller.
[0050] The logic control unit is the core component of HEX4 that manages and configures the lower physical registers. These lower physical registers are the fundamental units for data processing and storage, and the logic control unit ensures that these registers can work together efficiently and flexibly through a series of carefully designed control modules.
[0051] Specifically, each control module in the logical control unit corresponds one-to-one with a low-level physical register. This one-to-one correspondence ensures precise and personalized control of each register. The primary responsibility of these control modules is to dynamically configure the various parameters and behaviors of the registers based on configuration signals received and processed from the central controller. The central controller, also commonly referred to as the controller portion of the central processing unit, primarily performs instruction control, data control, and timing control.
[0052] Configuration signals usually contain specific instructions about how registers are connected, the direction of data flow, and the operation mode.
[0053] The connection method refers to how registers are connected or organized to form specific data paths. This determines how data flows between registers and how they interact with each other. The logic control unit interprets these instructions and adjusts the register connection topology as needed. For example, it can configure registers in series, parallel, or more complex network structures to accommodate different data processing requirements.
[0054] Data flow direction refers to the path and direction of data flow within the register network. It determines how data is transferred from one register to another and how this data is processed. Regarding data flow direction, the logic control unit ensures that data flows along a predetermined path between registers. This means that data can be transmitted unidirectionally, circulate bidirectionally, or traverse the register network according to specific algorithms. This flexible data flow management is crucial for implementing complex computing tasks and data processing workflows.
[0055] The logic control unit also configures the register operation mode. The operation mode refers to the type of data operation performed within the register network, determining how the registers process input data and produce the desired output. The operation mode determines how the registers behave when performing computational tasks, such as addition, subtraction, logical operations, or other specific algorithms. By dynamically adjusting the operation mode, the logic control unit can adapt the registers to a variety of different computational scenarios, thereby improving the overall performance and flexibility of the system.
[0056] It is worth noting that these configuration capabilities of the logic control unit are dynamic, meaning they can be adjusted at any time as needed during system operation. This dynamic configuration feature enables the system to quickly adapt to changing needs, optimize resource utilization, and improve processing efficiency.
[0057] The independent power control unit 130 is provided for each low-order physical register so that the power supply voltage of each low-order physical register can be set independently.
[0058] The independent power control unit in the embodiments of the present invention is a dedicated hardware component responsible for providing power control for the connected physical registers. "Independent" means that each unit is autonomous and can independently control the supply voltage of the registers to which it is connected. This design allows for fine-grained management of the power supply to each register, which may be used to optimize performance, save energy, or achieve specific voltage requirements. Instead of all registers sharing the same voltage control strategy, it may be used to meet the operating voltage requirements of different registers, optimize power consumption, or improve overall system performance.
[0059] It's understandable that by setting different voltages for different registers, their operating speed and power consumption can be optimized, thereby improving overall system performance. When full-speed operation isn't required, power consumption can be reduced by lowering the supply voltage for certain registers. Certain registers may require a specific voltage range to function properly. Independent power control units allow each register to be supplied with the required voltage, ensuring system stability.
[0060] The dynamic voltage adjustment unit 140 is used to control the supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the operation in real time according to the operation accuracy requirement corresponding to the currently received instruction.
[0061] The dynamic voltage adjustment unit in this embodiment of the present invention is a hardware unit that adjusts the supply voltage in real time based on the current computational precision requirements. It connects to the independent power control unit corresponding to each low-order physical register and regulates the register's supply voltage by controlling these units. The core function of the dynamic voltage adjustment unit is to dynamically adjust the supply voltage of the registers involved in the computation in real time based on the computational precision requirements of the instructions.
[0062] When the processor receives an instruction, the dynamic voltage regulator analyzes the instruction's precision requirements. Based on these requirements, it calculates the required supply voltage for each low-order physical register involved in the operation. It then sends control signals to the independent power control units corresponding to these registers to adjust their supply voltages. The dynamic voltage regulator is able to adjust the supply voltage in real time based on the instruction's precision requirements. This means that as the processor executes instructions, the supply voltage can be dynamically adjusted as needed to optimize performance, reduce power consumption, or meet specific voltage requirements.
[0063] Based on the above embodiments, Figure 1 As shown, the instruction processing system in the embodiment of the present invention further includes:
[0064] The carry and overflow detection circuit 105 is used to monitor and process the carry and overflow status information between multiple low-order physical registers currently participating in the operation in real time, and to feed back the carry and overflow status information to the corresponding low-order physical registers. It is also used to update the value in the flag register after each instruction cycle.
[0065] The carry and overflow detection circuit in the embodiment of the present invention is responsible for monitoring and processing the carry (i.e., the result exceeds the maximum value that the register can represent, and a unit needs to be added to the higher bit) and overflow (for signed number operations, the result exceeds the maximum or minimum value range that the data type can represent) that may occur during the operation. These situations are two common special cases in digital computing and are crucial to ensuring the accuracy of the calculation results and the stability of the system. The circuit continuously monitors multiple low-order physical registers currently involved in the operation in the computer. This means that whenever the data in these registers changes or participates in the operation, the carry and overflow detection circuit will immediately notice it. During the operation, if a carry or overflow occurs, the detection circuit will identify these states.
[0066] Once a carry or overflow is detected, the detection circuitry feeds this status information back to the relevant physical registers, allowing the registers to know whether the result of the operation is completely accurate or whether special measures (such as triggering an exception handler) are required to handle the carry or overflow situation.
[0067] HEX4 typically has one or more flag registers to store various status information, such as the zero flag, sign flag, carry flag, and overflow flag. Carry and overflow detection circuits update these flags after each instruction is executed, based on whether a carry or overflow occurred during the operation. This is crucial for subsequent instruction execution and operations such as conditional jumps, as they may determine the next step based on the status of these flags.
[0068] The asynchronous timing control module 106 is used to control different register units to participate in instruction processing in different clock cycles.
[0069] The main task of the asynchronous timing control module in the embodiment of the present invention is to control different register units to participate in instruction processing in different clock cycles. This means that the module can coordinate the working rhythm of each register unit to ensure that they execute corresponding operations at the correct time point.
[0070] Specifically, the asynchronous timing control module in the embodiments of the present invention can employ partially asynchronous logic, meaning that some parts of the system are asynchronous while others maintain a synchronous design. This design approach combines the advantages of both synchronous and asynchronous circuits, maintaining system stability while improving flexibility.
[0071] In asynchronous timing control modules, partially asynchronous logic allows different register units to communicate and operate without global clock signal synchronization. This reduces the problems caused by clock skew and clock delay, improving system response speed and efficiency.
[0072] Accordingly, to accommodate the asynchronous logic or partially asynchronous logic in the asynchronous timing control module, embodiments of the present invention may also employ loosely coupled data channels. This refers to data exchange between different modules or components being conducted via relatively independent channels with weak dependencies between these channels. Within the asynchronous timing control module, loosely coupled data channels enable different register units to independently transmit and process data within different clock cycles. This reduces the degree of coupling between modules, improving the system's scalability and fault tolerance. Because the asynchronous timing control module allows different register units to flexibly operate within different clock cycles, system resources can be more efficiently utilized. This reduces latency and idle cycles, improving overall system throughput. Partially asynchronous logic and loosely coupled data channels enable the system to respond more flexibly to failures. For example, if a register unit fails, other units can continue to operate without causing a complete system crash. Furthermore, the asynchronous design reduces interference and errors caused by clock signals, further improving system stability.
[0073] The multi-layer clock domain control module 107 determines the corresponding clock domain to be used according to the calculation accuracy requirement corresponding to the currently received instruction.
[0074] The multi-layer clock domain control module, a design strategy in embodiments of the present invention, is used to dynamically manage clock signals for different components of a complex digital system based on current task requirements. Clock signals are reference signals used to synchronize operations in digital circuits, controlling data reading, writing, and processing. This module identifies the required computational precision of the currently received instruction and determines which clock domains should be activated or deactivated based on this requirement.
[0075] A clock domain refers to a logical section controlled by the same clock signal. Using independent clock domains allows operations of different precision levels to proceed independently without interfering with each other. This design allows the system to shut down the clock to high-precision components during low-precision tasks, thereby reducing unnecessary power consumption. It also allows the system to more flexibly respond to diverse computational requirements.
[0076] When the system performs low-precision tasks, the power consumption of these parts can be significantly reduced by turning off the clock signals of the high-precision parts. This is because the clock signal itself consumes energy, and at the clock edge (that is, the rising or falling edge of the clock signal), the transistors in the digital circuit will perform switching operations, thereby consuming additional energy. When performing higher-precision calculations, the system needs to ensure that all related register units can operate in a coordinated manner. This usually means that these register units need to be activated synchronously to ensure that they receive and process data at the same point in time. Synchronous activation ensures that the outputs of different parts can be correctly correlated with each other, thereby avoiding data inconsistencies or errors. This is the key to achieving high-precision calculations.
[0077] Based on any of the above embodiments, an embodiment of the present invention provides an instruction processing method, as shown in the attached Figure 2 As shown, the instruction processing method specifically includes the following steps:
[0078] Step S210: According to the current instruction, a configuration signal is sent to each low-order physical register unit through the central controller to determine the number and connection mode of the low-order physical registers required to run the current instruction;
[0079] Step S220, adjusting the data flow direction and operation mode of each low-order physical register involved in running the current instruction, and establishing corresponding connections with adjacent low-order physical registers;
[0080] Step S230, controlling in real time the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the operation according to the operation accuracy requirement corresponding to the currently received instruction;
[0081] Step S240 , determining a corresponding clock domain to be used according to the computational accuracy requirement corresponding to the currently received instruction, and controlling all low-order physical registers involved in instruction processing to work collaboratively within the clock domain.
[0082] Step S250: After each instruction cycle ends, the carry and overflow detection circuit updates the value in the flag register, and the dynamic voltage adjustment unit adjusts the power supply strategy of each low-order physical register according to the requirements of the next instruction processing.
[0083] The core of the instruction processing method provided by the embodiment of the present invention is to respond to the current instruction requirements in real time through the central controller and flexibly configure the various parameters of the low-order physical registers. First, the system will accurately determine the number and connection method of the required registers according to the instruction requirements, and then adjust the data flow direction and operation mode to adapt to different processing requirements. On this basis, the system will also adjust the power supply voltage of each register in real time according to the calculation accuracy to balance performance and power consumption. Finally, within the selected clock domain, all registers involved in instruction processing will be activated collaboratively to ensure data consistency and efficient execution of instructions.
[0084] The instruction processing method provided by the embodiments of the present invention not only significantly improves system flexibility and response speed, but also effectively reduces power consumption and improves energy efficiency through sophisticated power management and clock domain control strategies. This dynamically configurable design enables the system to meet complex instruction processing requirements while maintaining efficient energy utilization.
[0085] Based on any of the above embodiments, as shown in the attached Figure 3 As shown, an embodiment of the present invention provides an example process of a 32-bit addition instruction processing method, and the specific process is as follows.
[0086] HEX4 demonstrates its outstanding capabilities in multi-precision computing through innovative register dynamic combination control throughout the entire process of receiving instructions, responding to needs and completing processing. HEX4 focuses on edge scenarios. Through this design, it not only optimizes energy efficiency, but also provides reliable support for expanding high-performance requirements such as artificial intelligence and graphics processing. In the HEX4 instruction set architecture, register dynamic combination control is a core function that enables it to efficiently support multi-precision computing. The following is the specific response and processing flow when the system receives an instruction that requires a 32-bit addition operation and the low-order physical register is a 4-bit physical register:
[0087] Step 310: Receive instruction
[0088] Instruction fetch: HEX4 reads an instruction from memory. Assume that the instruction is addition Nb, Nd, Ns1, Ns2, which is used to perform multi-precision addition operations.
[0089] Instruction decoding: The instruction decoding unit recognizes that this is a 32-bit addition operation that requires the use of eight 4-bit physical registers for combination.
[0090] Step 320, respond to the instruction
[0091] Flag register configuration: Determines whether carry handling and overflow detection are enabled based on the settings in Nb. The calculation precision level is configured to 32 bits, requiring the dynamic combination of eight 4-bit physical registers.
[0092] Step 330: Dynamic combination of registers
[0093] Register combination: The control logic unit sends a signal to combine the target register Nd with 7 adjacent 4-bit physical registers to form a 32-bit operation unit.
[0094] Connection configuration: Set up the data bus and control bus so that the 4-bit physical registers can work together. Configure the carry propagation path to ensure the correct transmission of low-order bits to high-order bits.
[0095] Step 340, load operand
[0096] Source operand load: Loads 32 bits of source data from addresses Ns1 and Ns2, storing each 4-bit physical register segment into the corresponding register. This is accomplished using data load instructions, ensuring that each 4-bit physical register segment is correctly read into the register.
[0097] Step 350, perform addition operation
[0098] Execute in 4-bit parts: Starting from the least significant bit (LSB), perform the addition operation on each 4-bit physical register. After each calculation, check whether a carry occurs and transfer it to the next higher register.
[0099] Cross-register carry handling: A dedicated carry management network is used to pass carry information between each 4-bit physical register to ensure calculation accuracy. If an overflow occurs in the highest bit, the overflow flag is triggered.
[0100] Step 360: Result combination and storage
[0101] Result reassembly: Combines the calculation results of all eight 4-bit physical registers into a complete 32-bit result. Checks for overflow and, if overflow occurs, determines how to handle it (such as automatic truncation or triggering an exception) based on the settings in Nb.
[0102] Result storage: Use the HESAVE instruction (data storage instruction in HEX4) to write the final 32-bit result back to the target memory address.
[0103] Step 370: Update the flag register
[0104] Carry flag: If a carry occurs in the highest bit, the corresponding bit in Nb is set to 1.
[0105] Overflow flag: If the result of the operation exceeds the 32-bit range, the overflow flag is set to 1.
[0106] Other status updates: Based on the calculation results, update the zero flag and sign flag for use by subsequent instructions.
[0107] Step 380, exception handling (if necessary)
[0108] Exception detection: If an overflow or other predefined exception occurs during the operation, the hardware will trigger an exception interrupt.
[0109] Exception response: The system suspends execution of the current instruction and executes the exception handling service routine instead. Based on the configuration in Nb, it decides whether to automatically adjust the calculation method (such as switching to scientific notation) or execute other preset error handling steps.
[0110] Step 390, restore the context and continue
[0111] Restore register combination: After completing the processing of the current instruction, the control logic unit releases the previously combined 32-bit register configuration and restores each 4-bit physical register to an independent state.
[0112] Continue executing subsequent instructions: The program counter increments and the system is ready to receive and execute the next instruction.
[0113] Through the above steps, HEX4 achieves efficient processing of 32-bit addition operations while maintaining the advantages of low power consumption and small chip area. This dynamic register combination mechanism enables it to demonstrate extremely high flexibility and performance in multi-precision calculations.
[0114] Based on any of the above embodiments, an embodiment of the present invention provides a processor, which includes the instruction processing system described in the embodiment of the present invention.
[0115] Based on any of the above embodiments, an embodiment of the present invention provides a chip, which includes the instruction processing system described in the embodiment of the present invention.
[0116] Based on any of the above embodiments, Figure 4 The schematic diagram of the physical structure of a chip provided by an embodiment of the present invention is shown. The electronic device may include: a processor (processor) 410, a communication interface (Communications Interface) 420, a memory (memory) 430 and a communication bus 440. The processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the following method:
[0117] According to the current instruction, the central controller sends a configuration signal to each low-order physical register unit to determine the number and connection method of the low-order physical registers required to run the current instruction;
[0118] Adjust the data flow direction and operation mode of each low-order physical register involved in running the current instruction, and establish corresponding connections with adjacent low-order physical registers;
[0119] According to the calculation accuracy requirements of the currently received instruction, the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the calculation is controlled in real time;
[0120] The corresponding clock domain to be used is determined according to the calculation accuracy requirement corresponding to the currently received instruction, and all low-order physical registers involved in instruction processing are controlled to work collaboratively within the clock domain.
[0121] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0122] On the other hand, an embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:
[0123] According to the current instruction, the central controller sends a configuration signal to each low-order physical register unit to determine the number and connection method of the low-order physical registers required to run the current instruction;
[0124] Adjust the data flow direction and operation mode of each low-order physical register involved in running the current instruction, and establish corresponding connections with adjacent low-order physical registers;
[0125] According to the calculation accuracy requirements of the currently received instruction, the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the calculation is controlled in real time;
[0126] The corresponding clock domain to be used is determined according to the calculation accuracy requirement corresponding to the currently received instruction, and all low-order physical registers involved in instruction processing are controlled to work collaboratively within the clock domain.
[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An instruction processing system based on a dynamically configured register, characterized in that: The system comprises: A register array comprising a plurality of low-order physical registers dynamically connected via a programmable interconnect medium; a logic control unit, comprising control modules corresponding one to one with the plurality of low-order physical registers, and configured to dynamically configure the connection mode, data flow direction, and operation mode between the registers according to the configuration signal received and processed from the central controller; An independent power control unit is provided for each low-order physical register, so that the power supply voltage of each low-order physical register can be set independently; The dynamic voltage regulation unit is used to control the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the operation in real time according to the operation accuracy requirements corresponding to the currently received instruction.
2. The instruction processing system according to claim 1, wherein: The system further comprises: The carry and overflow detection circuit is used to monitor and process the carry and overflow status information between multiple low-order physical registers currently participating in the operation in real time, and to feed back the carry and overflow status information to the corresponding low-order physical registers. It is also used to update the value in the flag register after each instruction cycle.
3. The instruction processing system according to claim 1, wherein: The system further comprises: The asynchronous timing control module is used to control different register units to participate in instruction processing in different clock cycles.
4. The instruction processing system according to claim 1, wherein: The system further comprises: The multi-layer clock domain control module determines the corresponding clock domain to be used according to the calculation accuracy requirements corresponding to the currently received instruction.
5. The instruction processing system according to claim 4, characterized in that: The determining of the corresponding clock domain to be used according to the computational accuracy requirement corresponding to the currently received instruction specifically includes: When the calculation precision corresponding to the currently received instruction is low precision, the high-precision clock domain is closed.
6. An instruction processing method based on the instruction processing system according to any one of claims 1 to 5, characterized in that: The method comprises: According to the current instruction, the central controller sends a configuration signal to each low-order physical register unit to determine the number and connection method of the low-order physical registers required to run the current instruction; Adjust the data flow direction and operation mode of each low-order physical register involved in running the current instruction, and establish corresponding connections with adjacent low-order physical registers; According to the calculation accuracy requirements of the currently received instruction, the power supply voltage provided by the independent power control unit corresponding to each low-order physical register involved in the calculation is controlled in real time; The corresponding clock domain to be used is determined according to the calculation accuracy requirement corresponding to the currently received instruction, and all low-order physical registers involved in instruction processing are controlled to work collaboratively within the clock domain.
7. The instruction processing method according to claim 6, characterized in that: The method further comprises: After each instruction cycle, the carry and overflow detection circuits update the values in the flag registers, while the dynamic voltage regulation unit adjusts the power supply strategy of each low-order physical register according to the requirements of the next instruction processing.
8. A processor, characterized in that: The processor comprises the instruction processing system according to any one of claims 1 to 5.
9. A chip, characterized in that: The chip includes the instruction processing system according to any one of claims 1 to 5.
10. A storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed, the instruction processing method according to any one of claims 6 to 7 is implemented.