Overflow optimization method and related device for very long instruction word architecture
By dynamically calculating the overflow priority of virtual registers in the ultra-long instruction font architecture and performing overflow optimization operations, the data overflow problem caused by insufficient registers is solved, and system performance and stability are improved.
Patent Information
- Application Number
- CN202510082161.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
When performing complex operations, the ultra-long instruction font architecture causes data to overflow into memory due to insufficient registers, extends data access time, increases memory burden, and seriously restricts system performance.
In the instruction preprocessing stage, the instruction set is traversed, the target instruction is parsed, the virtual register is identified, its running status data is obtained, the overflow priority is calculated dynamically, and the overflow optimization operation is performed based on the priority and the mapping relationship between the virtual register and the physical register, and the data is moved from the physical register to memory.
By dynamically configuring the overflow priority of virtual registers, optimizing register resource utilization, reducing non-essential data overflow, reducing memory management pressure, and improving system performance and stability.
Smart Images

Figure CN119536809B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to an overflow optimization method and related devices for a very long instruction word architecture. Background Art
[0002] With the rapid development of computer technology, especially in the fields of high-performance computing, embedded technology and data center core, unprecedented high standards and requirements are put forward for the computing performance and efficiency of processors. As an outstanding model of parallel computing technology, the Very Long Instruction Word (VLIW) architecture has greatly accelerated the data processing speed with its unique advantage of parallel processing within the instruction set, opening up a new direction for improving computing performance. However, while the very long instruction set architecture brings a high degree of parallelism, it also brings higher register pressure. This limitation restricts the further release of the potential of the very long instruction word architecture.
[0003] In the related technology, when executing complex operations, very long instruction words often need to process multiple operands. When the register is not large enough to accommodate them, data overflow to the memory becomes inevitable, which not only prolongs the data access time, but also increases the memory burden, seriously restricting the system performance. Traditional schedulers use simple strategies such as fixed priority or random selection to manage overflows, which are rigid in dealing with changing computing needs, and frequent unnecessary overflows become performance bottlenecks.
[0004] Therefore, it is urgent to design a technical solution to solve at least one of the above technical problems. Summary of the invention
[0005] In response to the technical problems existing in the prior art, the present application provides an overflow optimization method and related devices for a very long instruction word architecture, so as to realize register resource optimization of the very long instruction word architecture, improve resource utilization efficiency, meet the computing requirements in the very long instruction word architecture, and improve the system performance of the very long instruction word architecture.
[0006] In a first aspect, an embodiment of the present application provides an overflow optimization method for a very long instruction word architecture, the method comprising:
[0007] In the instruction preprocessing stage, the instruction set is traversed and the target instruction is parsed to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register;
[0008] Acquire the operation status data of each virtual register; wherein the operation status data at least includes: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction;
[0009] Dynamically acquiring the overflow priority of each virtual register according to the running status data; wherein the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register;
[0010] According to the overflow priority and the mapping relationship between the virtual register and the physical register, an overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0011] In a second aspect, an embodiment of the present application provides an overflow optimization device for a very long instruction word architecture, wherein
[0012] The parsing unit is configured to traverse the instruction set and parse the target instruction in the instruction preprocessing stage to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register;
[0013] An acquisition unit is configured to acquire operation status data of each virtual register; wherein the operation status data at least includes: a data item life cycle of each virtual register, a reuse frequency in the target instruction, and a predicted use time in the target instruction;
[0014] a priority unit configured to dynamically obtain the overflow priority of each virtual register according to the running status data; wherein the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register;
[0015] The execution unit is configured to perform the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, so as to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising:
[0017] at least one processor, memory, and input-output unit;
[0018] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the overflow optimization method for a very long instruction word architecture of the first aspect.
[0019] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions. When the instructions are executed on a computer, the computer executes the overflow optimization method for a very long instruction word architecture of the first aspect.
[0020] The beneficial effect of the present application is: an overflow optimization method and related device for a very long instruction word architecture are provided. In the technical scheme, first, in the instruction preprocessing stage, the instruction set is traversed and the target instruction is parsed to identify the virtual registers called by the target instruction. Among them, each virtual register includes at least: a read register and a storage register. Then, the operating status data of each virtual register is obtained. Among them, the operating status data at least includes: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction. Then, the overflow priority of each virtual register is dynamically obtained according to the operating status data. Among them, the farther the predicted use time is from the current moment, the higher the overflow priority corresponding to the virtual register. Finally, according to the overflow priority and the mapping relationship between the virtual register and the physical register, the overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0021] The technical solution of the present application dynamically configures the overflow priority of each virtual register, executes the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, and moves the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage, thereby realizing the register resource optimization of the very long instruction word architecture, improving the resource utilization efficiency, meeting the computing requirements in the very long instruction word architecture, and improving the system performance of the very long instruction word architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow chart of an overflow optimization method for a very long instruction word architecture according to an embodiment of the present application;
[0023] Figure 2 It is a structural schematic diagram of an overflow optimization method device for a very long instruction word architecture according to an embodiment of the present application;
[0024] Figure 3 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application;
[0025] Figure 4 It is a structural schematic diagram of a medium device according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0027] In the related technology, when executing complex operations, very long instruction words often need to process multiple operands. When the register is not large enough to accommodate them, data overflow to the memory becomes inevitable, which not only prolongs the data access time, but also increases the memory burden, seriously restricting the system performance. Traditional schedulers use simple strategies such as fixed priority or random selection to manage overflows, which are rigid in dealing with changing computing needs, and frequent unnecessary overflows become performance bottlenecks.
[0028] Specifically, in the related technologies, traditional schedulers rely on fixed priorities or random strategies to manage overflows, which are particularly clumsy when faced with complex and changeable computing tasks. They are unable to dynamically adjust overflow management strategies in real time according to computing needs and resource status, resulting in limited performance.
[0029] In addition, in related technologies, due to the lack of intelligent pre-judgment mechanisms, it is difficult to accurately distinguish the urgency of data, resulting in a large amount of unnecessary data overflowing into the memory, which not only increases the pressure on memory management, but also significantly slows down the overall performance of the system. In particular, when data overflows into the memory due to insufficient register capacity, the memory access speed is much lower than the register access speed, resulting in a significant increase in the time to access this data, which in turn prolongs the execution time of computing tasks, affecting the system's response speed and throughput.
[0030] Therefore, it is urgent to design a technical solution to solve at least one of the above technical problems.
[0031] In order to solve at least one technical problem in the related art, an embodiment of the present application provides an overflow optimization method and related devices for a very long instruction word architecture.
[0032] In the technical solution provided by the present application, first, in the instruction preprocessing stage, the instruction set is traversed and the target instruction is parsed to identify the various virtual registers called by the target instruction. Among them, each virtual register includes at least: a read register and a storage register. Then, the operating status data of each virtual register is obtained. Among them, the operating status data at least includes: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction. Then, the overflow priority of each virtual register is dynamically obtained according to the operating status data. Among them, the farther the predicted use time is from the current moment, the higher the overflow priority corresponding to the virtual register. Finally, according to the overflow priority and the mapping relationship between the virtual register and the physical register, the overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0033] Specifically, firstly, the technical solution of this application can more accurately determine which data can be temporarily moved out of the register and which data should be retained in the register by accurately analyzing the operating status data of the virtual register, including the data item life cycle, reuse frequency and predicted usage time. This can avoid unnecessary data from being frequently moved between registers and memory, make more efficient use of register resources, give full play to the advantages of high-speed reading and writing of registers, and improve the overall data processing efficiency of the system.
[0034] Secondly, the technical solution of this application dynamically obtains overflow priority based on the running status data, and preferentially overflows the target data corresponding to the virtual registers with a longer predicted usage time to the memory, effectively reducing the overflow of unnecessary data. This means that the memory will not be occupied by a large amount of data that is not needed in the short term, reducing the burden of memory management, reducing the additional overhead caused by the transmission of unnecessary data back and forth between the memory and the registers, and improving system performance and stability.
[0035] Third, since unnecessary overflows are reduced, the technical solution of this application reduces the time to wait for data to be read back from the memory to the register when executing instructions, and reduces data access latency. This enables instructions to obtain required data faster, speeds up the execution of instructions, and significantly improves the response speed and throughput of the system, especially when processing complex computing tasks, enabling smoother operation and improved user experience.
[0036] Fourthly, the technical solution of this application can dynamically adjust the overflow strategy according to the requirements of different instructions in the instruction set to adapt to the changing computing environment and task requirements. Whether it is a data-intensive or computing-intensive task, the system performance can be ensured to be unaffected by flexibly managing register resources and overflow operations, which enhances the adaptability and flexibility of the system in different application scenarios and provides strong support for complex computing tasks.
[0037] The technical solution of the present application dynamically configures the overflow priority of each virtual register, executes the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, and moves the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage, thereby realizing the register resource optimization of the very long instruction word architecture, improving the resource utilization efficiency, meeting the computing requirements in the very long instruction word architecture, and improving the system performance of the very long instruction word architecture.
[0038] The overflow optimization method scheme for the very long instruction word architecture provided in the embodiment of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a special device (such as a special terminal device with an overflow optimization method system for the very long instruction word architecture, etc.). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing the overflow optimization method scheme for the very long instruction word architecture.
[0039] Figure 1 A flow chart of an overflow optimization method for a very long instruction word architecture provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method comprises the following steps:
[0040] 101, in the instruction preprocessing stage, traverse the instruction set and parse the target instruction to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register;
[0041] 102, obtaining operation status data of each virtual register; wherein the operation status data at least includes: a data item life cycle of each virtual register, a reuse frequency in the target instruction, and a predicted use time in the target instruction;
[0042] 103, dynamically acquiring the overflow priority of each virtual register according to the running status data; wherein, the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register;
[0043] 104. Execute an overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, so as to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0044] In the embodiment of the present application, a very long instruction word (VLIW) architecture processor is an architecture for parallel processing of instruction sets. In this architecture, each instruction word contains multiple instructions that are executed in parallel, and these instructions share limited physical register resources to temporarily store operands and intermediate results. Due to the limited physical register resources, when the number of instructions increases or the data processing is complex, it is easy to have a shortage of register resources, and at this time, overflow operations are required to manage register resources.
[0045] Overflow operation means that when the physical register is full or about to be full, the data of the register being used is temporarily saved to the memory space, and then the vacated register is allocated to other calculation operations. When the data stored in the memory space needs to be used, the data is read from the memory space to the register to participate in the calculation. Its main purpose is to use the memory space to store the register data that is not used temporarily, so that the entire calculation has more available registers, which can realize more levels of program pipeline and improve the running efficiency of the program.
[0046] Traditional schedulers often adopt fixed priority or random selection strategies when facing overflow operations. These strategies have obvious defects. For example, the fixed priority strategy cannot be flexibly adjusted according to the actual computing tasks and data usage, which may cause important data to overflow, while some temporarily less important data occupies registers. The random selection strategy lacks consideration of the importance and timing of data use, which can easily cause unnecessary overflow, that is, overflowing the data to be used to the memory, resulting in frequent reading of data from the memory, which greatly increases the data access time. Because the access speed of the memory is much lower than the access speed of the register, this will slow down the overall performance of the system and affect the response speed and throughput of the system.
[0047] In principle, in the instruction preprocessing stage, the called virtual registers (including read registers and store registers) are identified by traversing the instruction set and parsing the target instructions. Virtual registers are an abstraction of physical registers, which makes it easier to analyze from a logical level. Then the operating status data of each virtual register is obtained, such as the data item life cycle, the frequency of reuse in the target instruction, and the predicted usage time. The data item life cycle can help determine the active period of the data in the entire computing process; the data of virtual registers with high reuse frequency is usually the key data in the computing process and needs to be retained in the register; the predicted usage time is an important basis for judging the urgency of the data.
[0048] Next, the overflow priority of each virtual register is dynamically calculated based on the acquired running status data. This is mainly based on the predicted usage time, that is, the farther the predicted usage time is from the current moment, the higher the overflow priority of the virtual register. At the same time, other factors such as reuse frequency will also be considered comprehensively. In this way, the virtual registers corresponding to data that is unlikely to be used in the short term can be overflowed first, avoiding unnecessary overflows and enabling register resources to be allocated to more urgent computing operations first.
[0049] Finally, according to the calculated overflow priority and the mapping relationship between virtual registers and physical registers, overflow optimization operations are performed. The target data associated with the target instruction is moved from the current physical register to the corresponding memory space for temporary storage. In this process, factors such as memory write performance and storage layout are also considered to ensure that data can be stored in memory and subsequently read efficiently. Through reasonable overflow operations, memory space is effectively utilized to alleviate the shortage of register resources, while reducing performance degradation caused by unnecessary overflow and frequent memory access, thereby improving the overall operation efficiency of the system.
[0050] The following is a detailed introduction to the principles and effects of each step:
[0051] In a very long instruction word architecture processor, instruction operations usually involve multiple operands, which need to be stored in registers during instruction execution. In 101, by traversing the instruction set and parsing each target instruction, the read register (used to provide operand input) and the storage register (used to store the operation result) used in the instruction can be determined. Due to the limited number of physical registers, in order to more effectively manage and track the usage of these registers, the concept of virtual registers is introduced to abstract the registers actually used in the instruction into virtual registers, so as to perform unified analysis and processing in subsequent steps, without being limited to the actual hardware limitations of physical registers and complex hardware mapping relationships.
[0052] Here, a clear mapping between instructions and virtual registers is established, laying the foundation for the subsequent precise analysis of the demand and usage patterns of instructions for register resources. This enables the use of registers to be sorted out at a logical level without directly considering the hardware details of the physical registers, improving the flexibility and scalability of the analysis, and facilitating the system to more efficiently plan and manage register resources to cope with complex and changing instruction sequences and computing tasks.
[0053] In 102, the life cycle of a data item is first introduced. The life cycle of a virtual register is determined by recording the time period from when it is first defined (or assigned) to when it is last used. During the instruction parsing process, when a virtual register is assigned or used for the first time, its start time is marked. Each time the virtual register is referenced again, its active status record is updated. When it is not referenced again within a certain time range, its life cycle is marked as over. This helps to understand the survival time and active period of the data during the entire instruction execution process, so as to judge its occupation and necessity of register resources at different stages.
[0054] Secondly, the frequency of reuse. During the instruction analysis process, the number of times each virtual register is reused in the target instruction set is counted. A virtual register with a high frequency of reuse means that its data is used multiple times in a short period of time, which plays an important role in the system's computing process. It is necessary to give priority to keeping it in the register to reduce data reading latency; while a virtual register with a low frequency of reuse is relatively less likely to be used again in a short period of time, and it can be considered to overflow it to the memory when resources are tight.
[0055] Next, predict the usage time. According to the execution logic and data dependencies of the instructions, combined with the existing instruction sequence patterns and historical execution data (if any), estimate the time point when each virtual register may be used next. For example, if the execution result of an instruction is usually used again as an operand after several subsequent instructions, then the predicted usage time of the virtual register can be roughly determined. This requires an in-depth understanding and analysis of the overall structure and data flow of the instruction set, and a more accurate prediction can be achieved by establishing a model of instruction execution or based on empirical rules.
[0056] In this way, these operating status data provide detailed information for a comprehensive understanding of the usage characteristics of virtual registers, allowing the system to evaluate the importance and urgency of each virtual register to the computing process from multiple dimensions. Based on this data, more intelligent and reasonable resource allocation decisions can be made later, such as determining which virtual register data should be kept in high-speed registers first and which can be temporarily moved out to relatively slow memory, thereby optimizing system performance to the greatest extent while ensuring the correctness of calculations, reducing unnecessary data transmission and storage overhead, and improving the efficiency and speed of instruction execution.
[0057] In 103, the predicted usage time is used as the main basis, combined with other operating status data (such as reuse frequency, etc.) for comprehensive consideration to determine the overflow priority of the virtual register. Since the farther the predicted usage time is from the current moment, the lower the possibility that the data of the virtual register will be used again in the short term, then the impact of overflowing its corresponding target data to the memory on the current computing process is relatively small. Therefore, such a virtual register is given a higher overflow priority. At the same time, if the predicted usage time of two virtual registers is similar, their reuse frequency can be further referenced. The virtual register with a low reuse frequency is relatively more likely to overflow, that is, its overflow priority is higher. This dynamic priority determination mechanism based on multiple factors can be flexibly adjusted according to the actual situation during the execution of instructions, adapt to different computing tasks and data dependencies, and ensure that the most appropriate overflow decision can be made at any time.
[0058] Here, a flexible and intelligent overflow decision strategy is implemented, which enables the system to accurately determine which data can be overflowed first according to the real-time status of the current instruction set and the specific situation of the virtual register, avoiding the blindness and inefficiency of traditional fixed priority or random overflow strategies. By giving priority to overflowing data that has less impact on the current calculation to the memory, the data stored in the register can be effectively kept to be the most urgently needed for the current calculation, reducing the performance overhead caused by frequent switching of data between registers and memory, improving the overall operating efficiency and response speed of the system, and enabling the processor to more efficiently process complex instruction sequences and large-scale data computing tasks.
[0059] In 104, according to the overflow priority determined in the previous step and the pre-established mapping relationship between the virtual register and the physical register, the physical register corresponding to the target virtual register that needs to overflow is found. Then, the target data in the physical register is moved to the corresponding memory space for temporary storage through a memory write operation. In this process, it is necessary to consider the memory write performance and data storage layout, and select a suitable memory area for storage to ensure that the data can be stored in the memory and subsequently read efficiently. At the same time, it is also necessary to update relevant system status information, such as register mapping tables, memory occupancy records, etc., so that the storage location and status of the data can be accurately obtained and managed during the subsequent instruction execution process.
[0060] Here, by purposefully overflowing temporarily unnecessary data to the memory, valuable physical register resources are released, allowing the system to process more instructions and data with limited physical register resources, improving the system's concurrent processing capabilities and resource utilization. Moreover, a reasonable memory storage strategy and system status update mechanism ensure that temporary storage of data in memory is orderly and efficient. When subsequent instructions need to use the overflowed data again, they can quickly read it from the memory and restore it to the register, reducing the additional delay caused by data overflow and recovery operations, further improving the overall performance and stability of the system, and ensuring that the entire computing process can be carried out smoothly and efficiently.
[0061] As an optional embodiment, in 101, the target instruction is parsed to obtain the calling information of the target instruction; the registers called by the target instruction, the register calling order, and the data dependency are obtained from the calling information; and corresponding numbers are configured for the registers based on the register calling order and the data dependency to obtain each virtual register.
[0062] Exemplarily, the target instruction is first parsed to obtain its call information. For example, for the instruction "r1=addr2, r3; r4 = mul r1, r5", it can be clearly parsed that it calls registers r1, r2, r3, r4, and r5, and the calling order is to first use r2 and r3 for addition to get the result and store it in r1, and then use r1 and r5 for multiplication to store it in r4. At the same time, the data dependency is analyzed, such as the value of r1 depends on the addition result of r2 and r3, and the value of r4 depends on the multiplication result of r1 and r5. Based on this register call order and data dependency, registers are configured and numbered. Assuming that they are numbered from 0 at the first appearance without dependency, for the above instructions, r2 and r3 have no dependency and can be numbered 0 and 1 respectively at the first appearance. r1 depends on r2 and r3, and its number is associated with r2 and r3 that participated in the previous operation. Since r2 is numbered 0 and r3 is numbered 1, r1 can be numbered 2 (indicating that it depends on the operation results of registers numbered 0 and 1). r5 has no dependency and can be numbered 3 at the first appearance. r4 depends on r1 and r5, and its number is associated with r1 (numbered 2) and r5 (numbered 3), and can be numbered 4. In this way, each virtual register is obtained.
[0063] In this way, through this numbering method based on the calling order and data dependency, the flow and dependency context of the data in the instruction can be clearly reflected, and the usage and importance of the data can be more accurately grasped when the running status data of the virtual register is subsequently analyzed (such as the life cycle of the data item, the frequency of reuse, the predicted usage time, etc.). For example, when determining the overflow priority, for the virtual register that is on the key data dependency path and has a high frequency of reuse (such as r1 in the above example, if it is used multiple times in subsequent instructions), it can be quickly identified according to its number and given a lower overflow priority, and its data is preferentially guaranteed to be retained in the register to avoid affecting the computing efficiency due to unreasonable overflow operations, thereby improving the resource management efficiency and data processing performance of the entire very long instruction word architecture processor when executing instructions, reducing the performance bottleneck and delay caused by improper register resource allocation, and enhancing the system's processing ability and adaptability to complex instruction sets, so that the system can run various computing tasks more efficiently and stably. Whether it is a data-intensive or computing-intensive task, better execution results can be obtained through this optimized register management method.
[0064] Further optionally, in 101, corresponding numbers are respectively configured for registers based on the register calling order and the data dependency to obtain each virtual register, which can be implemented as follows:
[0065] Determine whether the register to be numbered appears for the first time. If the register to be numbered appears for the first time and there is no data dependency between the register to be numbered and other registers, then the number corresponding to the register to be numbered is set incrementally based on the number corresponding to the previous register that appeared for the first time. If the register to be numbered does not appear for the first time, or there is a data dependency between the register to be numbered and other registers, then the number corresponding to the register to be numbered is set to the number corresponding to the dependent register.
[0066] In the overflow optimization related process of the very long instruction word architecture processor, this method focuses on the register number configuration link. By carefully considering the register calling order and data dependency, a corresponding virtual register number is assigned to each register, so as to better manage and analyze register resources.
[0067] In the specific implementation, each register to be numbered will be judged. First, it is judged whether it is the first appearance. This step is the basic distinction point of the entire numbering rule.
[0068] When it is determined that the register to be numbered appears for the first time and there is no data dependency between it and other registers that have appeared, a sequentially increasing numbering method will be used. That is, based on the number corresponding to the previous register that appeared for the first time, the number of the register to be numbered is set in ascending order. For example, assuming that the number of the first register with no dependency is set to 0, the number of the next register that also appears for the first time and has no dependency will be set to 1, and the next one will be 2, and so on. This method can clearly and orderly assign unique numbers to those registers that appear independently and have no data association, making it easier to distinguish and track them in the entire instruction analysis framework.
[0069] If the register to be numbered does not appear for the first time, or there is a data dependency relationship between it and other registers, it will no longer be numbered according to the ascending rule. Instead, the number corresponding to the register to be numbered is set to the number corresponding to the register it depends on. For example, there is an instruction that first operates on registers A and B to obtain the result and stores it in register C. Subsequently, there is another instruction that uses register C and another register D for operation. For register C, it depends on the operation results of A and B, so its number will be consistent with the number corresponding to the operation results of A and B. The purpose of this is to intuitively reflect the data dependency and transmission path from the numbering level, so that when analyzing the status of the virtual register and performing overflow priority judgment and other operations in the future, it is easy to trace the source of the data and the dependency chain according to the number, and accurately grasp the relevance and importance of the data.
[0070] Through this numbering configuration method, data dependencies can be presented very clearly at the virtual register level of the entire instruction set. Just like building a data flow diagram, the number of each virtual register contains information about its relationship with other registers. This is extremely helpful for subsequent in-depth analysis of the execution logic of instructions and understanding the data transfer process between different instructions. For example, when debugging the instruction implementation of a complex algorithm or optimizing program performance, by checking the virtual register number, you can quickly understand the ins and outs of the data, which helps to accurately identify possible problem points or links that can be further optimized.
[0071] The core goal of overflow optimization is register resource management. When determining which virtual registers can be overflowed to memory first, the importance and urgency of each virtual register can be more accurately judged based on the dependency relationship reflected by the number and the order of occurrence reflected by the increasing order of the number. Virtual registers whose numbers reflect that they are in the critical dependency path and are frequently used (through subsequent analysis of state data such as reuse frequency) can be identified and given a lower overflow priority, so that their data can be retained in the register for as long as possible, avoiding unnecessary overflows, thereby improving the overall performance of the system. For virtual registers whose numbers are relatively less critical and less likely to be used again in the short term, overflow operations can be reasonably arranged as needed, and memory space can be fully utilized to alleviate the tight situation of register resources, achieve efficient resource allocation and overflow optimization, and ensure that the processor with a very long instruction word architecture can use limited register resources in a better way when executing various complex instruction sets, improve computing efficiency, reduce performance bottlenecks and delays caused by unreasonable resource allocation, and enhance the system's adaptability and processing capabilities for different types of computing tasks.
[0072] As an optional embodiment, in 102, the operation status data of each virtual register is obtained, including:
[0073] Obtain the parsing result of the target instruction; extract the operation timestamp corresponding to each virtual register from the parsing result; determine the first appearance time, the last appearance time, and the time point of each reference of each virtual register based on the operation timestamp; calculate the active time period and the inactive time period of each virtual register according to the first appearance time, the last appearance time, and the time point of each reference, so as to obtain the data item life cycle of each virtual register; based on the time point of each reference of each virtual register, count the reuse frequency of each virtual register in the target instruction; based on the time point of each reference of each virtual register, determine the predicted usage time of each virtual register in the target instruction.
[0074] In this optional embodiment, the operation status data of the virtual register is obtained through a series of operations based on the parsing results of the target instruction, which has many important effects.
[0075] First, by extracting the operation timestamp and further determining the first appearance time, the last appearance time and the time point of each reference of the virtual register, the active time period and the inactive time period are calculated, and then the data item life cycle is obtained. In this way, it is possible to clearly know the specific time periods when the data carried by each virtual register is in an active and available state and an idle state during the entire instruction execution process. This plays a key role in the subsequent judgment of the timeliness and importance of the data. For example, it can be clear which data is only effective in a short period of time and which data will run through the entire calculation process for a long time. When making register resource allocation and overflow decisions, this information can be used to prioritize those data that are active for a long time or are needed in the critical stage to remain in the register, avoiding the arbitrary overflow of important data to the memory, thereby ensuring the efficiency and consistency of the calculation, and reducing the performance loss caused by the subsequent frequent reading of data from the memory due to incorrect overflow operations.
[0076] By counting the reuse frequency based on the time point each time a virtual register is referenced, it is possible to intuitively present how frequently the data corresponding to each virtual register is reused during instruction execution. A high reuse frequency means that the data plays a more core role in the calculation and needs to be kept in the register with higher priority to reduce the time overhead of repeatedly obtaining data, because the access speed of the register is much faster than that of the memory. Through this indicator, when resources are tight and overflow operations are required, it is easy to filter out those data that have a greater impact on computing efficiency and need to be retained, reasonably arrange register resources, and improve the overall data processing speed of the system, so that the system can more smoothly cope with complex computing tasks.
[0077] Then, the predicted usage time of the virtual register in the target instruction is determined through these referenced time points, and the next time the data of each virtual register may be used can be predicted in advance. In the overflow optimization process, this information is crucial. Combined with other status data, it can accurately determine which virtual register data will not be used in the short term, thereby determining its overflow priority, and giving priority to overflowing the virtual registers corresponding to data that is still far from the next use time, freeing up register resources for data that will be used soon, effectively avoiding unnecessary data switching between registers and memory, improving the utilization efficiency of register resources, and reducing data access latency, ensuring that the system can quickly obtain the required data when executing instructions, improving the operating performance and response speed of the entire very long instruction word architecture processor, so that it can better adapt to computing tasks of different types and sizes, and ensure the efficiency and stability of complex instruction execution.
[0078] Exemplarily, the target instruction is first parsed in detail. During this process, when the operation of the virtual register in the instruction (such as read and write operations) is identified, the system time at this time is recorded as the operation timestamp. These operation timestamps are like time stamps for virtual register activities, which can accurately record the moment when each operation occurs. For example, for the instruction "r1 = add r2, r3", when the write operation of r1 (storing the addition result) is processed, the timestamp of this moment is recorded. As the instructions are traversed and parsed, each virtual register will accumulate a series of operation timestamps, which constitute the time series of the virtual register during the entire instruction execution process. These time series provide rich raw data for the subsequent determination of register usage.
[0079] In the timestamp sequence of a virtual register, the earliest timestamp corresponds to the first appearance time of the virtual register. This time point marks the beginning of the data represented by the virtual register entering the calculation process. For example, if the timestamp sequence of a virtual register is [10ms, 15ms, 20ms], then 10ms is its first appearance time. On the contrary, the latest timestamp in the timestamp sequence corresponds to the last appearance time. This represents the moment when the data carried by the virtual register last participated in the calculation, and may no longer have a direct effect on the current calculation task. For example, in the above timestamp sequence, 20ms is the last appearance time. Each timestamp in the timestamp sequence (except the first appearance time, if there are no other special circumstances) represents the time point when the virtual register is referenced (read or written). These time points reflect the active moments of the virtual register during the execution of instructions. By counting the number and distribution of these time points, we can understand the usage frequency and usage period of the virtual register.
[0080] The active time period of a virtual register is the time interval from the first appearance to the last appearance. If the virtual register is still active after the last appearance (for example, in a loop structure, it may continue to be used after a cycle ends), then the active time period will extend to the current time or the end time specified by the instruction set. For example, if the first appearance time is 10ms and the last appearance time is 20ms, then the active time period is from 10ms to 20ms. This time period reflects the actual length of time that the data carried by the virtual register plays a role in the calculation process, which is very helpful for judging the timeliness and importance of the data.
[0081] The inactive time period is the time interval between two adjacent active time periods. By analyzing the inactive time period, we can understand the time periods when the virtual register is idle, which helps to arrange register resources reasonably. For example, if a virtual register has no operation record between 30ms and 40ms, then this period is its inactive time period. The data item life cycle is composed of the active time period and the inactive time period, which fully describes the existence status of the data represented by the virtual register in the entire computing process.
[0082] By counting the number of time points each virtual register is referenced (the first time point is counted only once), the reuse frequency of the virtual register in the target instruction can be obtained. For example, if a virtual register has 5 reference time points (excluding the first time point), then its reuse frequency is 5. Virtual registers with high reuse frequency are used multiple times during the calculation process and are usually critical data, which need to be given special consideration when allocating resources.
[0083] Based on the distribution of the reference time points of each virtual register, its next usage time can be predicted. If the reference time points of the virtual register show a certain regularity (such as in a periodic instruction sequence), the next usage time can be predicted based on this regularity. If there is no obvious regularity, it can be estimated based on the intervals between the most recent reference time points. For example, if a virtual register is referenced at 10ms, 20ms, and 30ms, then it can be predicted that the next usage time may be around 40ms. This predicted usage time is very important for determining the overflow priority, and can help the system plan the use of register resources in advance to avoid unnecessary data overflows.
[0084] Further optionally, in 102, based on the time point at which each virtual register is referenced each time, determining the predicted usage time of each virtual register in the target instruction may be implemented as follows:
[0085] In the absence of register pressure, the instruction set is executed simulated; based on the simulated execution data of the instruction set and the time point when each virtual register is referenced each time, the instruction position where each virtual register will be referenced next is predicted; based on the predicted instruction position, the predicted usage time of each virtual register in the target instruction is determined.
[0086] In this way, the instruction set is completely scheduled once without register pressure. By simulating the execution order and logic of each instruction, the instruction position of each virtual register that may be used again can be accurately located, so that the timing of data use can be fully considered in the subsequent calculation of overflow priority, thereby improving the rationality and accuracy of overflow decisions and further optimizing system performance.
[0087] In addition, in some other embodiments, the usage of registers can be determined by statically analyzing the syntax and semantics of the instruction set. The opcode and operand of the instruction are analyzed to identify which instructions read or write specific registers. For example, in assembly language, instructions such as "MOV" (move data) and "ADD" (addition) involve register operations. By scanning the entire instruction set, a mapping table of register usage is constructed to record which instructions each register is used as a source operand (read operation) or a destination operand (write operation).
[0088] This method can clearly see the static usage pattern of registers in the instruction sequence. For example, in a simple arithmetic calculation program, instruction stream analysis can find that some registers are specifically used to store intermediate calculation results, and it can be determined in which subsequent instructions these results will be reused, thereby judging their reuse frequency and life cycle.
[0089] Alternatively, in some other embodiments, the control flow structure in the program is considered, such as branch statements (such as "IF - ELSE"), loop statements (such as "FOR", "WHILE"), etc. For branch statements, the usage of registers under different branch paths is analyzed. For example, in a conditional judgment branch, some registers may be used only in branches that meet specific conditions, and remain unchanged or not used in other branches.
[0090] For loop statements, determine the difference in register usage inside and outside the loop body. For example, a register may be repeatedly read and written inside the loop body to store intermediate results for each loop iteration, while outside the loop body, it may only be initialized before the loop starts and used to store the final result after the loop ends. Through control flow analysis, you can understand the scope and frequency of register usage under different control structures, which helps predict register usage.
[0091] Processors are usually equipped with hardware performance counters that can record various hardware-related events, such as the number of register reads, writes, cache hit / miss ratios, etc. Using these hardware performance counters, the actual usage frequency and pattern of registers can be obtained.
[0092] For example, by reading the register read count counter, we can directly know how many times a register has been read within a period of time, thereby determining its reuse frequency. At the same time, combined with other indicators such as cache miss rate, we can infer whether the data in the register is frequently accessed and the flow of this data between memory and registers, thereby understanding the usage of the register.
[0093] Alternatively, you can set hardware breakpoints on the processor to trigger breakpoints when specific registers are accessed (read or written). In this way, you can track the time and context of register usage. For example, when debugging a complex program or performing a detailed analysis of the usage of a specific register, set a hardware breakpoint to pause program execution when a key register is accessed, and then view the instruction address, program status, and other information at this time to determine the usage and purpose of the register.
[0094] Furthermore, the data dependency between instructions is analyzed and a data dependency graph is constructed. In the graph, nodes represent instructions or registers, and edges represent the dependency of data flowing from one instruction or register to another. For example, if the result of instruction A is used as the input of instruction B, then there is an edge from the node corresponding to instruction A (or the node corresponding to its output register) to the node corresponding to instruction B (or the node corresponding to its input register).
[0095] By traversing this data dependency graph, the data source and data flow direction of each register can be determined. For a register, if it has multiple data sources or its data flows to multiple key instructions, then it may be in a relatively important position in the data transfer process, and its usage frequency may be high. The distribution and weight of the edges in the data dependency graph (if any) can be used to further determine its usage probability and importance in different computing stages.
[0096] Based on the data dependency graph, analyze the lifetime of the data in each register. From the moment the data is written to the register, track the path and number of times it is used in subsequent instructions until the data is no longer needed (no subsequent instructions depend on the data in the register). In this way, the lifetime and reuse frequency of the data items in the register can be determined. For example, if the data in a register is depended on by subsequent instructions in a long instruction sequence, then its data item lifetime is long and the reuse frequency is also high, indicating that this register plays a relatively important role in the calculation process.
[0097] As an optional embodiment, in 102, after predicting the instruction position of each virtual register to be referenced next time based on the simulated execution data of the instruction set and the time point when each virtual register is referenced each time, the probability distribution of each virtual register under different instruction execution paths can be obtained based on the predicted instruction position; and the instruction position weight of each virtual register is configured based on the probability distribution.
[0098] In the overflow optimization scenario of the very long instruction word architecture processor, it is of great significance to further obtain the corresponding probability distribution under different instruction execution paths after predicting the next instruction location referenced by the simulated execution data based on the instruction set and the time point when each virtual register is referenced each time.
[0099] In actual program execution, there are often multiple execution paths for instructions. For example, in a program containing conditional judgments, branch statements, or loop structures, the direction of instructions under different conditions is different. By analyzing the probability distribution of virtual registers being referenced under these different execution paths, we can grasp the usage patterns of data from a more macro and comprehensive perspective. For example, for a certain virtual register, after simulation and analysis, it is found that it has a higher probability of being referenced when the program execution enters a certain loop structure; while in another branch path that executes error handling, the probability of being referenced is very low. This probability distribution is like a "data usage map", which clearly shows the difference in importance of virtual register data in various possible execution scenarios, and provides a strong basis for more accurate resource allocation and overflow decisions in the future.
[0100] By configuring instruction position weights for each virtual register based on the above-mentioned probability distribution, the system can become more intelligent and efficient when processing register resource management and overflow operations.
[0101] The weight configuration can intuitively reflect the relative importance of different virtual registers in the entire instruction execution process. For those virtual registers that are in high-probability execution paths and are likely to be referenced, a higher instruction position weight is given, which means that when making resource allocation decisions such as overflow priority judgment, the system will be more inclined to keep their data in the register. This is because their data plays a key role in the smooth progress of the current computing task. Keeping them in the register can reduce the frequency of reading data from the memory, and use the advantages of high-speed reading and writing of registers to accelerate instruction execution, avoiding slowing down the performance of the entire system due to frequent memory access (if these key data are unreasonably overflowed to the memory, it will cause frequent reading).
[0102] On the contrary, for those virtual registers with low probability of being referenced in low-probability execution paths, their weights are relatively low, and their data can be given priority to overflow to memory when resources are tight. This not only makes reasonable use of limited register resources, but also does not have a significant negative impact on the overall computing efficiency, because these data themselves are not easy to be used in the current main computing scenarios.
[0103] In summary, by configuring instruction position weights in this way, the system can dynamically and accurately adjust the register resource allocation strategy based on the uncertainty of the instruction execution path and the probability of data use in different paths, effectively avoiding problems such as resource waste or unreasonable overflow of key data that may be caused by traditional fixed strategies. It significantly improves the performance and adaptability of very long instruction word architecture processors when processing complex and diverse computing tasks, making the entire instruction execution process smoother and more efficient, and ensuring that the system can make reasonable resource management decisions under different program logic and data dependency situations, thereby optimizing the overall operating efficiency and response speed of the system.
[0104] Further optionally, in 102, when assigning a weight value to each possible usage position according to the probability distribution of instruction execution (i.e., the instruction position weight introduced above), a machine learning algorithm is used to train the historical instruction execution data to predict the usage probability distribution of each virtual register under different instruction combinations, so as to more accurately determine the weight value, so that the system can adapt to the execution modes of different types of programs, optimize overflow decisions, improve the system's adaptability and processing capabilities to diversified computing tasks, and further improve the overall performance.
[0105] Specifically, instruction set-related data is collected from the historical execution records of the system, including instruction sequences of different types of programs, usage of virtual registers (such as which instruction positions are referenced, the number of references, etc.), and types of programs (such as data processing programs, graphics rendering programs, etc.). These data form the basis of the training data set. Feature extraction is performed on the collected data, for example, the instruction sequence is converted into a vector form that can be processed by the machine learning algorithm. It can be encoded according to information such as the opcode and operand type of the instruction. For the usage of virtual registers, its position in the instruction sequence, its relationship with other registers, and other features are recorded. At the same time, the program type is classified and encoded so that the model can learn the characteristics of different types of programs. Noise data in the data set is cleaned, such as error records or incomplete instruction information. The data is normalized, for example, numerical features such as instruction positions are mapped to specific intervals so that different features have the same scale, which is convenient for machine learning algorithms to train.
[0106] According to the characteristics of the problem and the nature of the data, select a suitable machine learning model, such as a neural network, decision tree, support vector machine, etc. For example, the neural network model can handle complex nonlinear relationships well and is suitable for learning complex mappings between instruction combinations and virtual register usage probabilities. Divide the preprocessed historical instruction execution data into a training set and a validation set. Use the training set to train the selected model, taking the instruction combination features as input and the virtual register usage probability distribution as output. By adjusting the parameters of the model (such as the weights and biases of the neural network), the model can learn the usage rules of virtual registers under different instruction combinations. During the training process, use the validation set to evaluate the performance of the model to avoid overfitting. For example, through techniques such as cross-validation, ensure that the model has good predictive ability on unseen data.
[0107] After the model training is completed, the features of the new instruction combination (from the instruction set currently to be processed) are input into the model, and the model will output the usage probability distribution of each virtual register at different instruction positions. This probability distribution reflects the possibility of the virtual register being used at each position under a given instruction combination. According to the predicted usage probability distribution, a weight value is assigned to each possible usage position of each virtual register. For example, the usage probability can be used as a weight, or the probability can be converted into a weight through a certain mapping function (such as a logarithmic function, an exponential function, etc.), so that the weight can better reflect the importance of the virtual register in instruction execution. The higher the weight value, the greater the possibility that the virtual register will be used at this position, and a higher priority should be given in resource management processes such as overflow decisions.
[0108] As the system processes different types of programs, the model continuously learns the instruction combinations and virtual register usage patterns of various programs. When encountering new program types or instruction combinations, the model can make adaptive adjustments based on existing knowledge and new data to optimize the prediction of the probability distribution of virtual register usage. During instruction execution, overflow decisions are optimized based on dynamically determined instruction location weights. For example, for virtual registers with higher weights, try to avoid overflowing their data to memory to ensure that their data can be accessed in a timely manner, thereby improving the system's processing capabilities and overall performance for diversified computing tasks. At the same time, new instruction execution data is collected regularly, returned to the data collection and preprocessing stage, and the model is updated and optimized to ensure that the model can always adapt to dynamic changes during program execution.
[0109] Using this method, machine learning algorithms can be used to train based on historical instruction execution data, and the association between different instruction combinations and the probability of virtual register usage can be deeply explored, making the predicted usage probability distribution more realistic. After accurately determining the weight value, when making overflow decisions, resources can be more reasonably allocated based on the actual importance of each virtual register, avoiding premature overflow of critical data, reducing unnecessary memory reading and writing, and improving data processing efficiency. For example, for different types of programs, their execution modes can be adaptive, and overflow strategies can be flexibly adjusted for both data processing and computationally intensive programs. This enhances the system's adaptability to diverse computing tasks, allowing it to still run efficiently in the face of complex and changing instruction scenarios, ensuring steady improvement in overall performance, and giving very long instruction word architecture processors more advantages in resource management and task execution.
[0110] Further optionally, in 102, when using a machine learning algorithm to predict the probability distribution of virtual register usage, an online learning method is adopted, which can update and optimize the model in real time according to new instruction execution data, so that the system can quickly adapt to dynamic changes in the program execution process, such as modifications to code logic, changes in data input, etc., continuously improve the accuracy and adaptability of overflow decisions, and ensure that the system always maintains high performance in long-term operation and complex and changeable computing scenarios.
[0111] Specifically, during the operation of the processor, a data collection module is specially configured to capture new instruction execution data in real time. This module can record the detailed information of each executed instruction, including the instruction's opcode, operands (especially the part involving virtual registers), the execution order of the instructions, and the execution time. In order to avoid excessive impact of frequent data updates on the model training process, a data buffer is set. The newly collected instruction execution data is first temporarily stored in the buffer, and the data in the buffer is batched for model update according to a certain strategy (such as reaching a certain amount of data or after a certain time interval). Set clear trigger conditions to start the model update process. For example, when the amount of new instruction execution data in the buffer reaches a preset threshold (such as 100 new instruction records), or when the system detects major changes that may affect the use of virtual registers during program execution (such as modification of key branches of code logic, a large amount of new data input, etc.), the model update is triggered. The frequency of model updates is dynamically adjusted according to the system load and the dynamic degree of instruction execution. When the system load is low and the instruction execution is relatively stable, the update frequency can be appropriately reduced; when the load is high or the program changes frequently, the update frequency can be accelerated to ensure that the model can adapt to changes in time.
[0112] Considering real-time performance and computational efficiency, choose a machine learning algorithm suitable for online learning. For example, incremental learning algorithms (such as incremental decision trees, online support vector machines, etc.) or online learning algorithms based on gradient descent (such as Adagrad, Adadelta, etc.). These algorithms can effectively update model parameters when new data arrives without retraining the entire model.
[0113] When the update is triggered, the new instruction execution data in the buffer is integrated with the previous training data (if necessary), and the model is updated according to the rules of the selected online learning algorithm. For algorithms based on gradient descent, the gradient of the new data relative to the model parameters is calculated, and the parameters are updated according to the gradient; for incremental learning algorithms, the decision rules of the model are gradually adjusted according to the characteristics of the new data and the existing model structure. A set of real-time performance evaluation indicators are established to measure the effect of the model update. These indicators can include the accuracy of overflow decisions (such as the correct prediction of the proportion of virtual registers that need to overflow), the response time of the system (due to the improved accuracy of overflow decisions, the system's data access delay is reduced), the efficiency of instruction execution (by reducing unnecessary register-memory data exchanges, the speed of instruction execution is improved), etc. According to the results of performance evaluation, a feedback adjustment mechanism is established. If the performance of the updated model is not improved or has decreased, analyze the possible reasons (such as data quality issues, improper learning rate settings, etc.), and adjust the model update strategy (such as data preprocessing methods, learning rates, update frequency, etc.) to ensure that the model can be continuously optimized and the system can always maintain high performance in long-term operation and complex and changing computing scenarios.
[0114] Therefore, on the one hand, the data can be updated and optimized according to the new instruction execution data in real time, so that the prediction of the probability distribution of virtual register usage is more in line with the current actual situation. When dynamic changes such as code logic modification and data input changes occur during program execution, the system can adjust quickly to avoid overflow decision errors caused by the incompatibility of the old model. On the other hand, continuously improving the accuracy and adaptability of overflow decisions means more reasonable arrangement of register resources, giving priority to the retention of key data, reducing unnecessary data switching between registers and memory, and improving overall operating efficiency. In long-term operation and complex and changeable computing scenarios, the system relies on this real-time optimization capability to always maintain high-efficiency performance, better cope with various complex tasks, and enhance the stability and reliability of the very long instruction word architecture processor to cope with different working conditions.
[0115] Further, from the perspective of data scale, if the instruction execution data is large and high-dimensional (including multiple instruction types, complex virtual register relationships, etc.), algorithms such as stochastic gradient descent (SGD) and its variants (Adagrad, Adadelta, Adam, etc.) are good choices. These algorithms can effectively update model parameters when processing large-scale data, and can adaptively adjust the learning rate to converge to better results in complex data environments. For example, in a system containing a large number of different types of program instructions, the Adam algorithm can quickly adjust the model according to the dynamic changes of the data to adapt to the prediction of the probability of virtual register usage under different instruction combinations.
[0116] From the perspective of data distribution, when the distribution of instruction execution data changes significantly over time (such as frequent changes in program logic and changing data input patterns), it is necessary to select an algorithm that can quickly adapt to changes in data distribution. Incremental learning algorithms (such as incremental decision trees) are more suitable. This algorithm can directly update the existing model when new data arrives, without the need to retrain the entire model. For example, for a software system that frequently updates its functions, the distribution of its instruction execution data will continue to change. The incremental decision tree can update the model structure in real time based on new data to adapt to changes in the probability of virtual register usage.
[0117] From the perspective of data relationships, if there is a linear relationship between instructions and the probability of virtual register usage, a simple linear online learning algorithm (such as an online learning version of the perceptron) may be sufficient. But if the relationship is complex and nonlinear, more powerful nonlinear models such as neural networks (such as long short-term memory networks - LSTM for processing sequence data) or kernel methods (such as online support vector machines) are more appropriate. For example, when dealing with instruction sets with complex nested loops and conditional judgments, the relationship between instructions and the probability of virtual register usage is often nonlinear. At this time, LSTM can effectively capture this complex relationship, thereby more accurately predicting the probability of virtual register usage.
[0118] In resource-constrained environments (such as embedded systems), algorithms with lower computational complexity need to be selected. Simple Bayesian online learning algorithms or rule-based incremental learning algorithms may be better choices. These algorithms usually do not require a lot of computing resources to update the model and can run under limited hardware conditions. For example, in a resource-constrained IoT device processor, a simple Bayesian online learning algorithm can update the virtual register usage probability model based on new instruction execution data in real time without placing an excessive burden on the device's performance.
[0119] For systems with extremely high real-time requirements, the training and update speed of the algorithm is crucial. Some fast-converging algorithms (such as variants of online gradient descent) or algorithms that can be trained in parallel (such as some distributed online learning algorithms based on the Map-Reduce framework) are more suitable. For example, in the scenario where a high-performance server processes a large number of concurrent instructions, overflow decisions need to be made quickly. An online gradient descent algorithm that can converge quickly can update the model in time to predict the probability of virtual register usage, thereby ensuring the efficient operation of the system.
[0120] If you need to explain the prediction results of the virtual register usage probability (for example, when debugging a complex system or showing the decision basis to the user), choose an algorithm with good interpretability. Decision trees and their incremental versions are usually more interpretable because their model structure is intuitive and can clearly show the impact of different instruction features on the virtual register usage probability. For example, during the development phase, developers can understand how instruction changes affect the prediction of virtual register usage probability by looking at the structure of the incremental decision tree.
[0121] When there is noise or outliers in the data, it is necessary to select an algorithm with strong stability. For example, the Median Stochastic Gradient Descent (SGD) algorithm with good robustness can resist the interference of outliers in the data to a certain extent, thereby updating the model more stably and ensuring that the prediction of the probability of virtual register usage will not fluctuate significantly due to abnormal data. This stability is very important for the system to continuously and accurately make overflow decisions in complex and changing computing scenarios.
[0122] As an optional embodiment, in 103, dynamically obtaining the overflow priority of each virtual register according to the running status data can be implemented as follows:
[0123] Starting from the current moment, multiple unit duration intervals are set according to preset strategies; according to the data item life cycle of each virtual register, the active time period and inactive time period of each virtual register starting from the current moment are determined; according to the active time period and inactive time period, the reuse frequency and the predicted usage time of each virtual register starting from the current moment, the predicted usage of each virtual register in each unit duration interval is determined; based on the predicted usage, an overflow priority list is established; the overflow priority list is used to indicate the priority scheduling order of the physical registers corresponding to each virtual register in each unit duration interval.
[0124] It is understandable that by setting multiple unit time intervals starting from the current moment, this is to discretize the future time, so as to facilitate the analysis and comparison of the usage of virtual registers at different time stages. It is like dividing the future timeline into small time periods, each of which can be used as an independent observation window to evaluate the status of virtual registers. The active and inactive time periods starting from the current moment are determined according to the life cycle of the data items of the virtual register, which is based on the actual usage rules of the data. The active time period represents the time range when the data is being used or is likely to be used soon, and the inactive time period represents the time when the data will not be used temporarily. In this way, the availability of each virtual register at different times can be clearly defined.
[0125] Combined with active and inactive time periods, reuse frequency, and predicted usage time, the predicted usage of virtual registers in each unit time interval is determined. The reuse frequency reflects how frequently data is reused, and the predicted usage time provides clues about the next time the data may be used. Combining these factors, it is possible to more comprehensively estimate whether a virtual register will be used and the likelihood of use in each small time interval. Based on the above predicted usage, an overflow priority list is constructed, which is essentially a time-priority scheduling guide. For each unit time interval, the priority order of the physical registers corresponding to each virtual register is clearly defined. In this order, the data of the virtual register corresponding to the physical register with a high priority is considered more important in the current interval and should be retained in the register, while the data of the virtual register with a low priority can be considered to overflow to the memory when resources are tight.
[0126] This method realizes fine management of virtual register usage in the time dimension. By dividing the unit time interval, the allocation of register resources can be flexibly adjusted according to the needs of different time stages. For example, within a unit time interval, if the predicted usage of a virtual register shows that its data is unlikely to be used, while the data of other virtual registers urgently need register resources, the data of the virtual register can be overflowed according to the priority list, thereby optimizing the utilization of register resources in each time period. The priority is determined by comprehensively considering the active time period, inactive time period, reuse frequency and predicted usage time, making the overflow decision more scientific and accurate. It avoids making potentially unreasonable decisions based on a single factor (such as only considering the predicted usage time). For example, although a virtual register is predicted to be used soon, if its reuse frequency is very low and it is currently in an inactive time period, then when resources are tight, its priority can be appropriately lowered to allocate resources to more critical virtual registers.
[0127] In the embodiment of the present application, the overflow priority list provides a flexible priority scheduling order that can adapt to dynamic changes during instruction execution. As instructions are executed, the state of the virtual register may change, and the new predicted usage can update the priority list in a timely manner. In this way, at different stages of instruction execution, the system can reasonably arrange register resources according to the real-time priority order to ensure that the system can run efficiently in various complex computing scenarios, reducing performance degradation caused by unreasonable resource allocation, such as frequently overflowing important data into memory and increasing data access delays.
[0128] Further optionally, in 103, after dynamically acquiring the overflow priority of each virtual register according to the operating status data, the real-time operating status of each virtual register can also be monitored; each time a change in the real-time operating status of the virtual register is detected, the real-time pressure of the register is predicted based on the changed real-time operating status to obtain a real-time pressure prediction result; based on the real-time pressure prediction result, a binary heap data structure is used to sort the overflow priorities to update the overflow priorities of each virtual register in real time.
[0129] Specifically, an optional implementation method of using a binary heap data structure to achieve efficient overflow priority sorting is that when building a real-time updated priority list for overflow priority calculation, the binary heap can be initialized, and the currently active physical registers and their corresponding virtual register information can be inserted into the heap as node elements. During the insertion process, the initial priority value of each node is determined based on the expected next use position of the virtual register (as the main sorting basis) and other relevant factors (such as data reuse frequency, criticality of instructions, etc., and the weight distribution of specific factors can be determined based on system performance testing and optimization experience), and an initial priority binary heap structure is constructed to lay the foundation for subsequent fast priority query and update operations. Data structure.
[0130] Furthermore, whenever the state of a virtual register changes and the priority list needs to be updated, first, the corresponding node is quickly located through a fast positioning algorithm in a binary heap (such as a search algorithm based on node identification or index), and its priority value is recalculated based on the updated virtual register information (such as the change in the expected next use position, the update of the number of references, etc.), to ensure that the priority value can accurately reflect the virtual register.
[0131] Further optionally, after the overflow priority is sorted using the binary heap data structure based on the real-time pressure prediction result in 103, the number of virtual registers in the current operating environment and the instruction update frequency can also be obtained. Further, based on the number of virtual registers in the active state and the instruction update frequency, at least one parameter of the number of nodes and the branching factor of the binary heap is adjusted to balance the time complexity and space complexity of the sorting operation.
[0132] Further optionally, in 102, when using timestamps to accurately calculate the active time periods and inactive time periods of virtual registers, a time window mechanism is introduced to divide the instruction execution process into multiple continuous time windows, and the activity of virtual registers is counted for each time window, and the analysis granularity of the virtual register life cycle is further refined, so that the idleness and demand for computing resources can be accurately judged at different time scales, so that in a dynamically changing computing environment, the system can more flexibly adjust the overflow strategy to improve resource utilization efficiency and system performance.
[0133] Further optionally, in 103, when the time window mechanism is used to refine the granularity of the virtual register lifecycle analysis, different weight coefficients are set for the activity in different time windows, and a higher weight is given to the recent time window to better reflect the real-time needs of the current computing task, so that when determining the overflow priority, it is possible to better balance the usage of recent and long-term data, optimize resource allocation, and improve the system's response speed and processing efficiency for computing tasks with higher real-time requirements.
[0134] Specifically, the instruction execution process is divided into multiple time windows, and each time window is like a small cycle for observing the activity of virtual registers. In each time window, the activity-related information such as the first appearance time, the last appearance time, and the number of references of the virtual register are counted. This division method can capture the data usage pattern of virtual registers in different time periods in a more detailed manner.
[0135] Set different weight coefficients for different time windows. A higher weight is given to the recent time window because the recent instruction execution can better reflect the real-time needs of the current computing task. As the time window moves farther away from the current moment, its weight gradually decreases. The reason for this is that the recent data usage has a more direct impact on the current calculation, while the long-term data usage, although also of reference value, is relatively less important.
[0136] When determining the overflow priority, the activity of the virtual registers in each time window and the corresponding weight coefficients are comprehensively considered. In this way, the recent and distant data usage can be balanced. For example, a virtual register is very active in the recent time window, but is rarely used in the distant time window. Due to the high recent weight, its overall importance will still be considered high; conversely, a virtual register that is active in the distant time window but rarely used recently will have a relatively low overall importance. This balance helps to arrange register resources more reasonably according to the actual computing task requirements.
[0137] By better balancing the use of recent and distant data, it is possible to more accurately determine which virtual registers' data is more important in the current computing task, thereby making more reasonable decisions when overflowing. For the physical registers corresponding to those virtual registers that are recently active and have higher weights, overflow will be avoided as much as possible to ensure that the current computing task can proceed smoothly; and for virtual registers that are highly active in the distant future but less important in the near future, overflow can be given priority when resources are tight, freeing up register resources for more urgent computing operations, thus achieving optimal resource allocation.
[0138] For computing tasks with high real-time requirements, this method enables the system to respond quickly. Because the system focuses more on recent instruction execution needs when allocating resources, it can ensure that the data that is currently needed is retained in the register, reducing the time to read data from the memory. For example, in a real-time data processing system, based on the high weight of the recent time window, the system can promptly provide sufficient register resources for the data being processed, improving the speed and efficiency of data processing, so that the system can better cope with tasks with high real-time requirements, such as real-time video processing, real-time financial data calculation, etc.
[0139] As an optional embodiment, in 104, the overflow optimization operation of the target instruction is performed according to the overflow priority and the mapping relationship between the virtual register and the physical register, which can be implemented as follows:
[0140] According to the overflow priority, the physical register corresponding to the virtual register whose predicted usage time is farthest from the current time is selected from each virtual register as the target physical register to be executed. If there are multiple virtual registers whose predicted usage time is farthest from the current time, the physical register corresponding to the virtual register with the lowest reuse frequency is selected as the target physical register to be executed. Finally, the overflow operation is performed on the target physical register to reduce the register resource pressure in the current operating environment.
[0141] Specifically, in the overflow optimization operation, the overflow priority is first considered to select the target physical register. The overflow priority is mainly determined by the predicted usage time of the virtual register, because the farther the predicted usage time is from the current time, the lower the possibility that the data corresponding to the virtual register will be used in the short term. Therefore, the physical registers corresponding to these data are preferentially selected as the overflow objects, so that the register resources can be freed up to the maximum extent without affecting the current computing task.
[0142] When the predicted usage time of multiple virtual registers is the farthest from the current time, the reuse frequency factor is further considered. The data carried by virtual registers with low reuse frequency is reused less frequently during the entire calculation process and is relatively less important. Therefore, in these cases, the physical register corresponding to the virtual register with the lowest reuse frequency is selected as the target physical register for the overflow operation to ensure that the overflowed data has the least impact on the current calculation process.
[0143] By performing an overflow operation on the selected target physical register, the data in it is temporarily stored in the memory, thereby releasing the physical register resources and reducing the pressure on the register resources in the current operating environment. This is because in a very long instruction word architecture processor, the number of physical registers is limited. When register resources are tight, a reasonable overflow operation can avoid calculation delays or errors caused by insufficient resources, ensuring that the system can continue to efficiently execute subsequent instructions.
[0144] This overflow optimization operation can accurately select the physical register that best suits the overflowed data, so that register resources are fully and reasonably utilized. By giving priority to overflowing data that is unlikely to be used in the short term, register space is left for more urgent computing operations, which improves register utilization efficiency and avoids resource waste. For example, in a complex computing task, multiple virtual registers compete for limited register resources at the same time. In this way, data that is critical to the current computing step can be kept in the register, speeding up the computing process.
[0145] Since the physical registers corresponding to the virtual registers that are predicted to be used in the long term and have a low reuse frequency are selected for overflow, the impact on the current computing performance is minimized. This can reduce the frequent reading of data from the memory due to unreasonable overflow operations (such as overflowing data that is about to be used). Because the memory access speed is much lower than the register access speed, frequent memory access will greatly reduce the overall performance of the system. By reducing the pressure on register resources, the system can execute instructions more smoothly, reduce computing delays, and improve response speed and throughput.
[0146] This method of selecting the target physical register based on overflow priority and multiple factors enables the system to adapt to different computing tasks and data usage patterns. Whether it is a data-intensive or computing-intensive task, the overflow strategy can be dynamically adjusted according to the actual instruction execution situation, ensuring that the system always maintains good performance and stability in various complex and changing computing scenarios, and effectively responds to the register resource management requirements under different program logics.
[0147] As an optional embodiment, after executing the overflow optimization operation of the target instruction in 104 according to the overflow priority and the mapping relationship between the virtual register and the physical register, the temporary storage duration and the corresponding return position of the target data can also be set based on the overflow priority; the return position is used to indicate the physical register corresponding to when the target data is returned from the memory space. If the temporary storage duration is reached, the target data is moved from the memory space to the physical register corresponding to the return position. Alternatively, when a call instruction for the target data is received, the target data is moved from the memory space to the physical register indicated by the call instruction.
[0148] It should be noted here that after the overflow optimization operation is performed, the temporary storage time and return location of the target data are further arranged according to the previously determined overflow priority. The overflow priority reflects the importance and urgency of the data corresponding to the virtual register in subsequent calculations. The higher the priority of the data, the shorter its temporary storage time will be. The physical register corresponding to the return location is also the place where the data is more urgently needed in subsequent calculations. In this way, the data can be managed in an orderly manner according to its importance and order of use.
[0149] The purpose of setting the temporary storage time is to reasonably control the storage time of data in the memory, to avoid data being idle in the memory for a long time and affecting the overall efficiency, and to prevent data from being returned to the register too early or too late. For example, based on the estimation of the instruction execution process and the frequency of data use, a longer temporary storage time is given to some data with a slightly lower priority, while a shorter temporary storage time is set for data that is about to be used (high priority), to ensure that the data can be returned to the register at the appropriate time to participate in the operation.
[0150] The physical register corresponding to the return position is determined based on the data dependency between instructions and the data requirements for subsequent calculations. When data is returned from memory, it accurately enters the physical register needed for subsequent calculations, ensuring that the data can seamlessly enter the calculation process and reducing the extra data handling or waiting time caused by improper data return location.
[0151] On the one hand, when the set temporary storage time is reached, the system automatically moves the target data from the memory space to the corresponding physical register, and ensures that the data is available in a timely manner according to the pre-planned rhythm. On the other hand, when a call instruction for the target data is received, regardless of whether the temporary storage time is reached, the data is immediately returned to the physical register indicated by the call instruction. This reflects the flexibility of responding to real-time data needs and ensures that the instructions being executed can smoothly obtain the required data for calculation.
[0152] In the above steps, by setting the temporary storage time and return position based on the overflow priority, the data flow between the register and the memory is made more orderly and accurate. The blindness of data return is avoided, so that each data can return to the most appropriate physical register at the most appropriate time to participate in subsequent calculations, optimize the entire data usage process, and improve the consistency and efficiency of data processing. The data can be returned in time when the call instruction is received, which effectively responds to the real-time demand for data during the instruction execution process, reduces the instruction execution delay caused by waiting for data to be returned from the memory, and improves the response speed of the system, especially for those computing tasks with high real-time requirements, such as real-time control systems, real-time image rendering and other scenarios, which can better ensure the smooth operation of the system. In addition, the reasonable setting of the temporary storage time avoids unnecessary long-term retention of data in the memory or premature return to occupy register resources, so that both register and memory resources can be used more reasonably. At the same time, the orderly data return mechanism reduces the additional overhead caused by poor data management, further improves the overall performance of the system, and makes the very long instruction word architecture processor more efficient and stable when processing complex instruction sets and diversified computing tasks.
[0153] As an optional embodiment, in 104, when performing an overflow operation to temporarily save the data in the selected physical register to the memory space, an optimized memory write strategy is adopted, that is, according to the distribution of free blocks in the memory and the expected usage time of the data, a memory area with a faster write speed and lower future read latency is selected for storage, so as to reduce the storage and reading time of the data in the memory, further improve the overall data processing efficiency of the system, and enhance the advantages of the method of the present invention in practical applications.
[0154] Here, when performing overflow operations, the free block distribution of the memory itself is considered, aiming to find free memory areas where data can be written quickly and reduce the write waiting time. At the same time, the expected use time of the data is combined to make decisions, and data that is expected to be read soon is preferentially stored in areas with low future read latency, so that when the data is needed later, it can be quickly obtained from the memory, optimizing the storage and reading links of data in the memory from the source and improving the overall efficiency. By regularly monitoring and evaluating the write and read performance of the memory, a memory performance distribution model is constructed. Since the state of the memory hardware will change over time, such as the aging of the memory causing the read and write performance of some areas to decrease, and the distribution of free blocks changes after defragmentation, the data storage location selection strategy is dynamically adjusted according to the model, which can flexibly select the optimal storage area according to the actual state of the memory hardware in real time, ensuring that data can always be stored and read in the memory with high efficiency.
[0155] Further optionally, when selecting a memory area with a faster write speed and lower future read latency for storage, the write and read performance of the memory is regularly monitored and evaluated, a memory performance distribution model is established, and the data storage location selection strategy is dynamically adjusted according to the model to adapt to changes in the memory hardware status (such as memory aging, fragmentation, etc.), always maintaining the storage and reading efficiency of data in the memory at a high level, and further enhancing the adaptability and practicality of the method of the present invention in different hardware environments.
[0156] Further optionally, when establishing the memory performance distribution model, not only the characteristics of the memory hardware itself are considered, but also the memory competition situation is comprehensively modeled in combination with other processes running in the current system. By real-time monitoring the memory usage patterns and resource requirements of other processes, the storage location selection strategy of the data in the memory is dynamically adjusted to avoid the data storage and reading efficiency of the present invention being reduced due to interference from other processes, thereby further enhancing the adaptability and stability of the system in a complex multi-tasking environment.
[0157] Further optionally, when comprehensively modeling the memory usage of other processes, a memory resource negotiation mechanism is established between processes. When the method of the present invention detects an urgent need for memory resources from other processes, it can actively release a portion of the memory space occupied by relatively non-urgent data, and reclaim the corresponding memory resources after the resource pressure on other processes is relieved. Through this collaborative approach, the memory resource utilization and stability of the entire system are improved, ensuring that the present invention can run stably and efficiently in a multi-tasking environment.
[0158] Specifically, when establishing the memory performance distribution model, we not only focus on the characteristics of the memory hardware itself, but also take into account the competition for memory by other processes currently running in the system. The memory usage patterns and resource requirements of other processes are monitored in real time, because the memory occupancy and read and write operations of different processes will affect each other. By integrating these factors into the model, we can more accurately grasp the dynamic changes of the entire memory usage environment, thereby formulating a storage location selection strategy that is more in line with the actual situation, and avoiding the loss of data storage and reading efficiency involved in the present invention due to interference from other processes.
[0159] In order to cope with the situation of tight memory resources in a multi-tasking environment, a memory resource negotiation mechanism between processes is established. When it is detected that other processes have urgent needs for memory resources, the method of the present invention can actively determine which of the data stored in itself is relatively less urgent, release the memory space occupied by these data to other processes in urgent need, and then reclaim the corresponding memory resources after the resource pressure of other processes is relieved. This mechanism is based on the judgment of the urgency of the memory demand of each process, and realizes the flexible allocation of memory resources between different processes.
[0160] Through the above steps, the optimized memory write strategy and the dynamic adjustment of the storage location selection strategy are adopted to reduce the writing and reading time of data in the memory. For example, data that may have taken a long time to wait to be written to the memory can now be quickly stored in the appropriate area and can be obtained faster during subsequent reading. This directly speeds up the data processing rhythm of the entire system and improves processing efficiency, allowing the very long instruction word architecture processor to operate more smoothly when executing complex instruction sets. By selecting a memory area with lower future read latency to store data, the delay caused by memory read waiting is effectively reduced, making it more timely to obtain data during instruction execution, further enhancing the system response speed, especially for computing tasks with higher real-time requirements. This optimization effect is more significant and can improve the overall performance of the system.
[0161] By dynamically adjusting the data storage location selection strategy to adapt to changes in the memory hardware state, whether the memory has aging problems due to long-term use, or the free block distribution has changed after operations such as defragmentation, the method of the present invention can ensure that the storage and reading efficiency of data in the memory is maintained at a high level, so that it can function stably in systems with different hardware conditions, and improve the applicability of the method in various hardware environments. Comprehensively consider the memory usage of other processes and establish a memory resource negotiation mechanism, so that the system can better coordinate the needs of each process for memory resources in a complex multi-tasking environment. On the one hand, it avoids the decrease in data storage and reading efficiency caused by interference from other processes. On the other hand, through flexible memory resource allocation, the stability of the entire system during multi-tasking operation is enhanced, ensuring that the method of the present invention can still operate efficiently in complex scenarios where multiple processes compete for memory resources, and improving the overall stability and practicality of the system.
[0162] Through the memory resource negotiation mechanism, memory space is reasonably allocated between different processes to achieve full utilization of memory resources. When other processes have urgent needs, the memory occupied by relatively non-urgent data is released in time, and then recovered after the situation is alleviated, avoiding idle waste of memory resources, so that the memory resources of the entire system can be dynamically allocated according to the actual needs of each process, improving the overall utilization of memory resources in a multi-tasking environment, and ensuring the efficient and stable operation of the system in complex multi-tasking scenarios.
[0163] In another optional embodiment, in 104, when reading data from the memory space back to the register for calculation, a pre-read buffer is set. When the system detects that an instruction is about to use data that has overflowed into the memory, the data is read from the memory to the pre-read buffer in advance. Once the instruction is executed to the stage where the data is needed, the data can be quickly obtained directly from the pre-read buffer and loaded into the register, thereby significantly shortening the instruction waiting time caused by memory reading, further improving the response speed and throughput of the system, and giving full play to the innovation and practicality of the present invention in optimizing data access processes.
[0164] Further optionally, the pre-read buffer adopts a multi-level cache structure, and the data pre-read from the memory is stored in different levels of cache according to the expected usage time and priority of the data, and the data to be used is preferentially stored in the fastest access cache level. At the same time, a cache elimination algorithm, such as the least recently used (LRU) algorithm, is adopted to promptly clean up data that has not been used for a long time, so as to improve the space utilization and data access speed of the pre-read buffer, further improve the response speed and throughput of the system, and give full play to the role of the pre-read buffer in optimizing the data access process.
[0165] Further optionally, when using the least recently used (LRU) algorithm to clean up data in the pre-read buffer that has not been used for a long time, optimization is performed in combination with the heat value of the data. The heat value is calculated based on the historical access frequency of the data and the recent access time interval. Data with lower heat values are cleaned up first. At the same time, for data with higher heat values, appropriate cache retention strategies are adopted. Even if it has not been accessed recently, it is retained in the pre-read buffer for a certain period of time to prevent frequent cleaning of data that may be used again due to misjudgment, thereby improving the hit rate and data access efficiency of the pre-read buffer and improving the overall performance of the system.
[0166] Further optionally, when combined with the data heat value optimization algorithm, an adaptive heat value decay function is used to dynamically adjust the decay speed of the heat value according to the operating status of the system (such as idle time, load level, etc.), so that the heat value can more accurately reflect the actual use value of the data, further optimize the cache management strategy of the pre-read buffer, improve data access efficiency and system performance, especially when the system operating status changes frequently, and can better adapt to different workload requirements.
[0167] In an embodiment of the present application, the overflow priority of each virtual register is dynamically configured, and according to the overflow priority and the mapping relationship between the virtual register and the physical register, the overflow optimization operation of the target instruction is performed, and the target data associated with the target instruction is moved from the current physical register to the corresponding memory space for temporary storage, thereby realizing the register resource optimization of the very long instruction word architecture, improving the resource utilization efficiency, meeting the computing requirements in the very long instruction word architecture, and improving the system performance of the very long instruction word architecture.
[0168] In another embodiment of the present application, an overflow optimization device for a very long instruction word architecture is also provided. Figure 2 The device comprises the following units:
[0169] The parsing unit is configured to traverse the instruction set and parse the target instruction in the instruction preprocessing stage to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register;
[0170] An acquisition unit is configured to acquire operation status data of each virtual register; wherein the operation status data at least includes: a data item life cycle of each virtual register, a reuse frequency in the target instruction, and a predicted use time in the target instruction;
[0171] a priority unit configured to dynamically obtain the overflow priority of each virtual register according to the running status data; wherein the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register;
[0172] The execution unit is configured to perform the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, so as to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0173] Further optionally, after the execution unit performs the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, the execution unit is further configured to:
[0174] Based on the overflow priority, setting the temporary storage duration of the target data and the corresponding return position; the return position is used to indicate the physical register corresponding to when the target data is returned from the memory space;
[0175] If the temporary storage time is reached, the target data is moved from the memory space to the physical register corresponding to the return position; or,
[0176] When a call instruction for the target data is received, the target data is moved from the memory space where the target data is located to the physical register indicated by the call instruction.
[0177] Further optionally, the parsing unit parses the target instruction to identify each virtual register called by the target instruction, and is configured to:
[0178] Parsing the target instruction to obtain calling information of the target instruction;
[0179] Acquire the registers called by the target instruction, the register calling order, and the data dependency from the calling information;
[0180] Based on the register calling sequence and the data dependency, corresponding numbers are configured for the registers respectively to obtain various virtual registers.
[0181] Further optionally, the parsing unit configures corresponding numbers for registers based on the register calling order and the data dependency relationship to obtain each virtual register, which is configured as follows:
[0182] Determine whether the register to be numbered appears for the first time;
[0183] If the register to be numbered appears for the first time, and there is no data dependency between the register to be numbered and other registers, then the number corresponding to the register to be numbered is set incrementally based on the number corresponding to the previous register that appears for the first time;
[0184] If the register to be numbered does not appear for the first time, or there is a data dependency relationship between the register to be numbered and other registers, the number corresponding to the register to be numbered is set to the number corresponding to the dependent register.
[0185] Further optionally, the acquiring unit acquires the running status data of each virtual register and is configured to:
[0186] Obtaining the parsing result of the target instruction;
[0187] Extracting the operation timestamp corresponding to each virtual register from the parsing result;
[0188] Determine the first appearance time, the last appearance time, and the time point of each reference of each virtual register based on the operation timestamp;
[0189] According to the first appearance time, the last appearance time, and the time point of each reference of each virtual register, the active time period and the inactive time period of each virtual register are calculated to obtain the data item life cycle of each virtual register;
[0190] Based on the time point at which each virtual register is referenced each time, counting the reuse frequency of each virtual register in the target instruction;
[0191] Based on a time point each time each virtual register is referenced, a predicted use time of each virtual register in the target instruction is determined.
[0192] Further optionally, the acquisition unit determines the predicted usage time of each virtual register in the target instruction based on the time point each time each virtual register is referenced, and is configured to:
[0193] Simulating execution of the instruction set without register pressure;
[0194] Based on the simulated execution data of the instruction set and the time point when each virtual register is referenced each time, predict the instruction position at which each virtual register will be referenced next time;
[0195] Based on the predicted instruction position, a predicted usage time of each virtual register in the target instruction is determined.
[0196] Further optionally, the acquisition unit, after predicting the instruction position at which each virtual register is referenced next time based on the simulated execution data of the instruction set and the time point at which each virtual register is referenced each time, is further configured to:
[0197] Based on the predicted instruction positions, the probability distribution of each virtual register under different instruction execution paths is obtained;
[0198] The instruction position weight of each virtual register is configured based on the probability distribution.
[0199] Further optionally, the priority unit dynamically obtains the overflow priority of each virtual register according to the running status data, and is configured as follows:
[0200] Starting from the current moment, multiple unit duration intervals are set according to the preset strategy;
[0201] According to the life cycle of the data items of each virtual register, determine the active time period and the inactive time period of each virtual register starting from the current moment;
[0202] Determine the predicted usage of each virtual register in each unit time interval according to the active time period and the inactive time period, the reuse frequency and the predicted usage time of each virtual register from the current moment;
[0203] An overflow priority list is established based on the predicted usage; the overflow priority list is used to indicate the priority scheduling order of the physical registers corresponding to each virtual register within each unit time interval.
[0204] Further optionally, after dynamically acquiring the overflow priority of each virtual register according to the running status data, the priority unit is further configured to:
[0205] Monitor the real-time operating status of each virtual register;
[0206] Each time a change in the real-time operating state of the virtual register is detected, the real-time pressure of the register is predicted based on the changed real-time operating state to obtain a real-time pressure prediction result;
[0207] Based on the real-time pressure prediction result, a binary heap data structure is used to perform overflow priority sorting to update the overflow priority of each virtual register in real time.
[0208] Further optionally, the priority unit, after performing overflow priority sorting using a binary heap data structure based on the real-time pressure prediction result, is further configured as follows:
[0209] Get the number of active virtual registers and instruction update frequency in the current running environment;
[0210] At least one parameter of the number of nodes and the branching factor of the binary heap is adjusted based on the number of virtual registers in an active state and the instruction update frequency to balance the time complexity and space complexity of the sorting operation.
[0211] Further optionally, the execution unit performs the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, and is configured to:
[0212] According to the overflow priority, selecting from each virtual register a physical register corresponding to a virtual register whose predicted usage time is farthest from the current time as a target physical register to be executed;
[0213] If there are multiple virtual registers whose predicted usage time is farthest from the current time, the physical register corresponding to the virtual register with the lowest reuse frequency is selected as the target physical register to be executed;
[0214] A spill operation is performed on a target physical register to reduce register resource pressure in the current operating environment.
[0215] The device can implement various steps in the above method embodiment, which will not be expanded here.
[0216] In an embodiment of the present application, an overflow optimization device for a very long instruction word architecture is used to optimize register resources of the very long instruction word architecture, improve resource utilization efficiency, meet computing requirements in the very long instruction word architecture, and improve system performance of the very long instruction word architecture.
[0217] See also Figure 3 , Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present application. Figure 3As shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the following steps are implemented: in the instruction preprocessing stage, the instruction set is traversed, and the target instruction is parsed to identify the virtual registers called by the target instruction; wherein each virtual register includes at least: a read register and a store register; the running status data of each virtual register is obtained; wherein the running status data includes at least: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction; the overflow priority of each virtual register is dynamically obtained according to the running status data; wherein, the farther the predicted use time is from the current moment, the higher the overflow priority corresponding to the virtual register; according to the overflow priority and the mapping relationship between the virtual register and the physical register, the overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0218] See also Figure 4 , Figure 4 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present application. Figure 4 As shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the following steps are implemented: in the instruction preprocessing stage, the instruction set is traversed, and the target instruction is parsed to identify the virtual registers called by the target instruction; wherein each virtual register includes at least: a read register and a store register; the running status data of each virtual register is obtained; wherein the running status data includes at least: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction; the overflow priority of each virtual register is dynamically obtained according to the running status data; wherein, the farther the predicted use time is from the current moment, the higher the overflow priority corresponding to the virtual register; according to the overflow priority and the mapping relationship between the virtual register and the physical register, the overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage.
[0219] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0220] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0221] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0222] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An overflow optimization method for a very long instruction word architecture, characterized in that: The method comprises: In the instruction preprocessing stage, the instruction set is traversed and the target instruction is parsed to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register; Acquire the operation status data of each virtual register; wherein the operation status data at least includes: the data item life cycle of each virtual register, the reuse frequency in the target instruction, and the predicted use time in the target instruction; Dynamically acquiring the overflow priority of each virtual register according to the running status data; wherein the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register; According to the overflow priority and the mapping relationship between the virtual register and the physical register, an overflow optimization operation of the target instruction is performed to move the target data associated with the target instruction from the current physical register to the corresponding memory space for temporary storage; After performing the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, the method further includes: Based on the overflow priority, setting the temporary storage duration of the target data and the corresponding return position; the return position is used to indicate the physical register corresponding to when the target data is returned from the memory space; If the temporary storage time is reached, the target data is moved from the memory space where it is located to the physical register corresponding to the return position; or, when a call instruction for the target data is received, the target data is moved from the memory space where it is located to the physical register indicated by the call instruction.
2. The overflow optimization method according to claim 1, characterized in that: The parsing of the target instruction to identify each virtual register called by the target instruction includes: Parsing the target instruction to obtain calling information of the target instruction; Acquire the registers called by the target instruction, the register calling order, and the data dependency from the calling information; Based on the register calling sequence and the data dependency, corresponding numbers are configured for the registers respectively to obtain various virtual registers.
3. The overflow optimization method according to claim 2, characterized in that: The configuring corresponding numbers for registers based on the register calling order and the data dependency to obtain each virtual register includes: Determine whether the register to be numbered appears for the first time; If the register to be numbered appears for the first time, and there is no data dependency between the register to be numbered and other registers, then the number corresponding to the register to be numbered is set incrementally based on the number corresponding to the previous register that appears for the first time; If the register to be numbered does not appear for the first time, or there is a data dependency relationship between the register to be numbered and other registers, the number corresponding to the register to be numbered is set to the number corresponding to the dependent register.
4. The overflow optimization method according to claim 1, characterized in that: The step of obtaining the operation status data of each virtual register includes: Obtaining the parsing result of the target instruction; Extracting the operation timestamp corresponding to each virtual register from the parsing result; Determine the first appearance time, the last appearance time, and the time point of each reference of each virtual register based on the operation timestamp; According to the first appearance time, the last appearance time, and the time point of each reference of each virtual register, the active time period and the inactive time period of each virtual register are calculated to obtain the data item life cycle of each virtual register; Based on the time point at which each virtual register is referenced each time, counting the reuse frequency of each virtual register in the target instruction; Based on a time point each time each virtual register is referenced, a predicted use time of each virtual register in the target instruction is determined.
5. The overflow optimization method according to claim 4, characterized in that: The step of determining the predicted usage time of each virtual register in the target instruction based on the time point at which each virtual register is referenced each time includes: Simulating execution of the instruction set without register pressure; Based on the simulated execution data of the instruction set and the time point when each virtual register is referenced each time, predict the instruction position at which each virtual register will be referenced next time; Based on the predicted instruction position, a predicted usage time of each virtual register in the target instruction is determined.
6. The overflow optimization method according to claim 5, characterized in that: After predicting the instruction position of each virtual register to be referenced next time based on the simulated execution data of the instruction set and the time point when each virtual register is referenced each time, the method further includes: Based on the predicted instruction positions, the probability distribution of each virtual register under different instruction execution paths is obtained; The instruction position weight of each virtual register is configured based on the probability distribution.
7. The overflow optimization method according to claim 1, characterized in that: The dynamically acquiring the overflow priority of each virtual register according to the running status data includes: Starting from the current moment, multiple unit duration intervals are set according to the preset strategy; According to the life cycle of the data items of each virtual register, determine the active time period and the inactive time period of each virtual register starting from the current moment; Determine the predicted usage of each virtual register in each unit time interval according to the active time period and the inactive time period, the reuse frequency and the predicted usage time of each virtual register from the current moment; An overflow priority list is established based on the predicted usage; the overflow priority list is used to indicate the priority scheduling order of the physical registers corresponding to each virtual register within each unit time interval.
8. The overflow optimization method according to claim 7, characterized in that: After dynamically acquiring the overflow priority of each virtual register according to the running status data, the method further includes: Monitor the real-time operating status of each virtual register; Each time a change in the real-time operating state of the virtual register is detected, the real-time pressure of the register is predicted based on the changed real-time operating state to obtain a real-time pressure prediction result; Based on the real-time pressure prediction result, a binary heap data structure is used to perform overflow priority sorting to update the overflow priority of each virtual register in real time.
9. The overflow optimization method according to claim 8, characterized in that: After the overflow priority is sorted using a binary heap data structure based on the real-time pressure prediction result, the method further includes: Get the number of active virtual registers and instruction update frequency in the current running environment; At least one parameter of the number of nodes and the branching factor of the binary heap is adjusted based on the number of virtual registers in an active state and the instruction update frequency to balance the time complexity and space complexity of the sorting operation.
10. The overflow optimization method according to claim 1, characterized in that: The performing the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register includes: According to the overflow priority, selecting from each virtual register a physical register corresponding to a virtual register whose predicted usage time is farthest from the current time as a target physical register to be executed; If there are multiple virtual registers whose predicted usage time is farthest from the current time, the physical register corresponding to the virtual register with the lowest reuse frequency is selected as the target physical register to be executed; A spill operation is performed on a target physical register to reduce register resource pressure in the current operating environment.
11. An overflow optimization device for a very long instruction word architecture, characterized in that: The device comprises the following units, wherein: The parsing unit is configured to traverse the instruction set and parse the target instruction in the instruction preprocessing stage to identify each virtual register called by the target instruction; wherein each virtual register includes at least: a read register and a store register; An acquisition unit is configured to acquire operation status data of each virtual register; wherein the operation status data at least includes: a data item life cycle of each virtual register, a reuse frequency in the target instruction, and a predicted use time in the target instruction; a priority unit configured to dynamically obtain the overflow priority of each virtual register according to the running status data; wherein the farther the predicted usage time is from the current time, the higher the overflow priority corresponding to the virtual register; an execution unit, configured to perform an overflow optimization operation of the target instruction according to the overflow priority and a mapping relationship between the virtual register and the physical register, so as to move the target data associated with the target instruction from the current physical register to a corresponding memory space for temporary storage; The execution unit is further configured to: after executing the overflow optimization operation of the target instruction according to the overflow priority and the mapping relationship between the virtual register and the physical register, set the temporary storage time and the corresponding return position of the target data based on the overflow priority; the return position is used to indicate the physical register corresponding to when the target data is returned from the memory space where it is located; if the temporary storage time is reached, move the target data from the memory space where it is located to the physical register corresponding to the return position; or, when a call instruction for the target data is received, move the target data from the memory space where it is located to the physical register indicated by the call instruction.
12. An electronic device, characterized in that: include: Memory for storing computer software programs; A processor is used to read and execute the computer software program, thereby implementing the overflow optimization method described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that: The storage medium stores a computer software program, and when the computer software program is executed by the processor, the overflow optimization method according to any one of claims 1 to 10 is implemented.
14. A chip, characterized in that: The chip is loaded with a computer software program and / or a hardware unit, and the computer software program and / or the hardware unit are used to implement the overflow optimization method as described in any one of claims 1-10.
Citation Information
Patent Citations
Instruction scheduling and register allocation method on optimized clustered VLIW (Very Long Instruction Word) processor
CN104484160A