Vector instruction ordering queue source operand multi-level maintenance method and system
Patent Information
- Application Number
- CN202611154814.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-31
AI Technical Summary
[0004]本申请提供了一种向量指令定序队列源操作数多级维护方法及系统,旨在解决当生产者指令的数据通路宽度小于消费者指令的数据通路宽度时,生产者需多个时钟周期才能完成全部位宽的数据写入,消费者若按常规机制提前就绪,会导致取数时数据不完整,引发执行错误的问题
[0015]本申请通过兼顾正确性与执行效率:针对窄数据通路生产者到宽数据通路消费者的场景,通过多级移位链延迟唤醒,确保消费者在数据完整写入后才标记就绪,避免取数错误;针对等宽指令交互场景,保持快旁路直接就绪,无额外延迟,适配不同交互场景的最优时序。
Smart Images

Figure CN122653692B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vector processor instruction scheduling technology, and in particular to a method and system for multi-level maintenance of vector instruction ordering queue source operands. Background Technology
[0002] In the instruction issue pipeline of a vector processor, a sequence queue is used to buffer vector instructions to be issued and maintain the readiness status of each instruction's source operands. The instruction can only be issued and executed after all source operands are ready. Existing vector instructions generally adopt a register broadcast wake-up mechanism with a fixed lead time: when the producer instruction starts execution, it broadcasts the destination physical register number, and the source operands of the consumer instruction are directly marked as ready after hitting this number.
[0003] This mechanism has significant drawbacks: when the data path width of the producer instruction is smaller than that of the consumer instruction, the producer needs multiple clock cycles to complete the writing of the full bit width of data. If the consumer is ready in advance according to the conventional mechanism, it will lead to incomplete data retrieval and cause execution errors. If a conservative long-delay wake-up strategy is uniformly adopted, redundant waiting cycles will be introduced in equal-width instruction interaction scenarios, significantly reducing instruction issuance efficiency. At the same time, the existing mechanism cannot quickly respond to branch flushing events, and incomplete delayed wake-up states cannot be synchronously canceled in a timely manner, which can easily cause pipeline state chaos. Summary of the Invention
[0004] This application provides a method and system for multi-level maintenance of source operands in a vector instruction ordering queue. It aims to solve the problem that when the data path width of a producer instruction is smaller than the data path width of a consumer instruction, the producer needs multiple clock cycles to complete the writing of the full bit width of data. If the consumer is ready in advance according to the conventional mechanism, it will lead to incomplete data when fetching data, causing execution errors.
[0005] In a first aspect, embodiments of this application provide a method for multi-level maintenance of source operands in a vector instruction ordering queue, the method comprising: During the decoding stage, the data path width between the source operand and the destination operand of each vector instruction is determined based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions, and the corresponding splitting number is calculated; the vector instructions include producer instructions and consumer instructions; Assign a physical register number to the destination operand of the producer instruction. Broadcast the physical register number corresponding to the destination operand during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, compare the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction. If the data path width of the producer instruction is less than that of the consumer instruction, inject the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all flags in the corresponding shift chain are cleared and the ready state of the corresponding source operand is canceled.
[0006] In some embodiments, determining the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions includes: identifying the category to which the vector instruction belongs, calculating the basic data path width of the source operand by combining the element bit width and the vector register group multiple; if the vector instruction is of the widening type, setting the data path width of the destination operand to twice the basic data path width of the source operand; if the vector instruction is of the narrowing type, setting the data path width of the source operand to twice the basic data path width of the destination operand.
[0007] In some embodiments, the calculation of the corresponding number of splits includes: obtaining the reference width corresponding to the full-width data path, calculating the ratio of the current data path width to the reference width, and determining the number of splits for a single instruction through the data path based on the ratio; the smaller the ratio value, the more splits are required, and the number of splits corresponding to the reference width is one.
[0008] In some embodiments, the process of assigning physical register numbers to the destination operand of a producer instruction and broadcasting the physical register number corresponding to the destination operand during the first cycle of producer instruction execution includes: retrieving an idle physical register number from the renaming free table and assigning it to the destination operand of the producer instruction; and sending the corresponding physical register number to the all source operand monitoring unit when the producer instruction enters the first cycle of the emit pipeline.
[0009] In some embodiments, if the source operand of the consumer instruction hits the broadcast physical register number, comparing the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction includes: pre-storing the physical register numbers that each source operand of the consumer instruction is waiting for, comparing the broadcast physical register number with each pre-stored number one by one; if the numbers match, it is determined to be a hit, and the data path width of the destination operand of the producer instruction is retrieved and compared with the data path width of the current source operand.
[0010] In some embodiments, if the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, injecting the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag shifts with the shift chain step by step, includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick.
[0011] In some embodiments, when the ready marker is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set, including: real-time detection of the marker state at the lowest level of the shift chain; when the ready marker moves to the lowest level of the shift chain and is moved out, a ready valid signal is generated, and the state of the corresponding source operand is updated to ready.
[0012] In some embodiments, the step of setting the ready signal directly after the source operand hits the broadcast physical register number if the producer instruction data path width is greater than or equal to the consumer instruction data path width includes: when it is determined that the producer data path width is greater than or equal to the consumer data path width, skipping the shift chain delay process, and directly updating the state of the corresponding source operand to ready in the same clock cycle when the physical register number is hit.
[0013] In some embodiments, when a branch flush occurs during the shift process, clearing all markers in the corresponding shift chain and canceling the ready state of the corresponding source operand includes: when a branch flush signal is received, comparing the instruction identifiers associated with each marker in the shift chain; if the identifiers match successfully, clearing all ready markers at all levels of the corresponding shift chain, and simultaneously canceling the ready state of the corresponding source operand, restoring it to the unready state.
[0014] Secondly, this application provides a multi-level maintenance system for source operands of a vector instruction ordering queue, the system comprising: The clock count calculation unit is used to determine the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction type, element bit width, and vector register group multiple corresponding to multiple vector instructions, and to calculate the corresponding clock count; the vector instructions include producer instructions and consumer instructions; The numbering allocation unit is used to allocate physical register numbers to the destination operands of producer instructions. The physical register number corresponding to the destination operand is broadcast during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, the data path widths of the destination operand of the producer instruction and the source operand of the consumer instruction are compared. If the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, the ready flag of the corresponding source operand is injected into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. The status cancellation unit is used to set the ready signal of the corresponding source operand when the ready mark is moved out of the lowest level of the shift chain; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all marks in the corresponding shift chain are cleared and the ready state of the corresponding source operand is cancelled.
[0015] This application balances correctness and execution efficiency: for narrow data path producers to wide data path consumers, it uses a multi-level shift chain to delay wake-up, ensuring that consumers are marked as ready only after the data is completely written, thus avoiding data retrieval errors; for equal-width instruction interaction scenarios, it maintains fast bypass direct readiness without additional delay, adapting to the optimal timing for different interaction scenarios.
[0016] By matching different width difference scenarios at different entry levels of the shift chain, it can also accommodate the wake-up requirements of long-delay execution units such as division and reduction, ensuring precise alignment between the ready time and the actual data availability time. When a branch flush event occurs, all ready flags in the corresponding shift chain can be cleared directly, quickly canceling the ready state of the source operand, ensuring the consistency of the pipeline state, and improving the operational stability of the ordered queue.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic flowchart illustrating the steps of a multi-level maintenance method for source operands in a vector instruction ordering queue, provided in an embodiment of this application. Figure 2This is a schematic diagram illustrating the principle of a multi-level maintenance method for source operands of a vector instruction ordering queue provided in an embodiment of this application; Figure 3 This is a schematic block diagram of a multi-level maintenance system for source operands of a vector instruction ordering queue provided in one embodiment of this application; Figure 4 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0023] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0024] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0026] In the instruction issue pipeline of a vector processor, a sequence queue is used to buffer vector instructions to be issued and maintain the readiness status of each instruction's source operands. The instruction can only be issued and executed after all source operands are ready. Existing vector instructions generally adopt a register broadcast wake-up mechanism with a fixed lead time: when the producer instruction starts execution, it broadcasts the destination physical register number, and the source operands of the consumer instruction are directly marked as ready after hitting this number.
[0027] This mechanism has significant drawbacks: when the data path width of the producer instruction is smaller than that of the consumer instruction, the producer needs multiple clock cycles to complete the writing of the full bit width of data. If the consumer is ready in advance according to the conventional mechanism, it will lead to incomplete data retrieval and cause execution errors. If a conservative long-delay wake-up strategy is uniformly adopted, redundant waiting cycles will be introduced in equal-width instruction interaction scenarios, significantly reducing instruction issuance efficiency. At the same time, the existing mechanism cannot quickly respond to branch flushing events, and incomplete delayed wake-up states cannot be synchronously canceled in a timely manner, which can easily cause pipeline state chaos.
[0028] Please refer to Figure 1 and Figure 2 This application provides a method for multi-level maintenance of source operands in a vector instruction ordering queue, applicable to computer devices. The computer device can be deployed on a single server or a server cluster. It can also be deployed on handheld terminals, laptops, wearable devices, or robots, etc. It should be noted that all information involved in the method provided in this application is extracted with the authorization of the relevant user and in accordance with relevant regulations, and will not infringe on user privacy.
[0029] This embodiment provides a multi-level maintenance method for source operands in a vector instruction ordering queue, applicable to the instruction issue scheduling system of a vector processor. The system can be integrated into a general-purpose processor, digital signal processor, application-specific integrated circuit (ASIC), or programmable logic device (PLC). The system includes at least a decoding module, a renaming module, an ordering queue module, a source operand monitoring unit, a delayed wake-up shift chain module, an execution unit, and a branch processing module. The method differentiates instruction interaction scenarios with different data path widths, employing two strategies: delayed wake-up via shift chain and direct wake-up via fast bypass. It also supports rapid state revocation under branch flushing, maximizing instruction issue efficiency while ensuring data read correctness.
[0030] The provided method for multi-level maintenance of source operands in a vector instruction ordering queue includes steps S101 to S103. Details are as follows: Step S101. In the decoding stage, the data path width between the source operand and the destination operand of each vector instruction is determined according to the vector instruction category, element bit width and vector register group multiple corresponding to multiple vector instructions, and the corresponding splitting number is calculated; the vector instructions include producer instructions and consumer instructions.
[0031] Specifically, this step is executed during the decoding stage of the vector instructions and is completed collaboratively by the decoding module and the clock count calculation unit.
[0032] When a vector instruction enters the decoding pipeline, the instruction's category identifier, element bit width configuration information, and vector register set multiplier configuration information are extracted first. The vector instruction category includes various types such as regular operations, widened operations, narrowed operations, long-latency operations, and memory access operations. The element bit width indicates the binary bit length of a single vector element and can be configured to 8 bits, 16 bits, 32 bits, or 64 bits. The vector register set multiplier indicates the number of vector registers occupied by a single instruction and can be configured to fractional multipliers, 1x, 2x, 4x, or 8x.
[0033] The clock cycle calculation unit first calculates the basic data bit width of a single operand based on the element bit width and the multiple of the vector register group. Then, it determines the data path width level corresponding to the source operand and the destination operand based on the instruction type. The data path width level is used to characterize the amount of data that can pass through the data path within a single clock cycle. From high to low, it includes five levels: full bit width, half bit width, quarter bit width, eighth bit width, and one-sixteenth bit width, corresponding to the amount of data that can be transmitted in a single clock cycle: all memory bank data, two memory bank data, one memory bank data, half memory bank data, and one-quarter memory bank data, respectively.
[0034] After determining the data path width level, the time-splitting calculation unit calculates the number of time-splitting steps for a single instruction. The calculation rule for the number of time-splitting steps is as follows: taking one time-splitting step corresponding to a full-width data path as the baseline, the number of time-splitting steps doubles for each level decrease in data path width. The number of time-splitting steps corresponding to a full-width data path is one, meaning the instruction can pass directly through the data path in a single time; half-width corresponds to two time-splitting steps, quarter-width corresponds to four time-splitting steps, and so on. The number of time-splitting steps is used to control the dwell time of the instruction at the transmit port. Each time-splitting step processes a data slice of the corresponding width until all data processing is completed.
[0035] The vector instructions processed in this step can be divided into producer instructions and consumer instructions according to the data flow. Producer instructions are those that produce results and write them to the vector register after execution, while consumer instructions are those that need to read existing data from the vector register as operands. The same instruction can simultaneously serve as a consumer of the preceding instruction and a producer of the following instruction.
[0036] Step S102. Assign a physical register number to the destination operand of the producer instruction. Broadcast the physical register number corresponding to the destination operand during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, compare the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction. If the data path width corresponding to the producer instruction is less than that corresponding to the consumer instruction, inject the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain.
[0037] Specifically, this step is completed collaboratively by the renaming module, the source operand monitoring unit, and the shift chain control unit.
[0038] First, the renaming module assigns an independent physical register number to the destination operand of each producer instruction. These physical register numbers are retrieved sequentially from the free physical register table. After allocation, the physical register number is bound and stored with the renaming identifier of the corresponding instruction. For grouped instructions with a vector register group multiple greater than 1, the renaming module allocates multiple consecutive physical register numbers per group, and all physical register numbers in the grouped instructions share the same renaming identifier.
[0039] When a producer instruction is initiated and enters the first stage of the emit pipeline, the execution path corresponding to the producer instruction broadcasts the physical register number of the instruction's destination operand. The broadcast signal is sent to all source operand monitoring units in the sequence queue. The broadcast signal also carries the data path width level, split stage number, pipeline stage identifier, and forward path identifier for the corresponding destination operand.
[0040] Each source operand monitoring unit internally stores the physical register numbers that the source operands of its respective consumer instruction are waiting for. When the broadcast physical register number matches the pre-stored physical register number of a source operand, it is determined that the corresponding source operand has hit the broadcast.
[0041] After the hit determination is completed, the width comparison unit immediately retrieves the data path width level of the producer instruction destination operand and compares it with the data path width level of the current consumer instruction source operand.
[0042] If the comparison result shows that the data path width of the producer instruction's destination operand is less than the data path width of the consumer instruction's source operand, a width upgrade scenario is determined. In this case, the shift chain control unit generates a ready flag for the corresponding source operand and determines the entry level of the delayed wake-up shift chain based on the width difference, injecting the ready flag into the corresponding entry level. The larger the width difference, the higher the corresponding entry level, and the more shift cycles the ready flag needs to traverse. After injection, the ready flag moves to a higher level of the shift chain with each clock cycle.
[0043] If the comparison result shows that the data path widths of the two are equal, the shift chain delay process will not be entered, and the fast bypass ready path will be triggered directly.
[0044] Step S103. When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all flags in the corresponding shift chain are cleared and the ready state of the corresponding source operand is canceled.
[0045] Specifically, this step is completed collaboratively by the shift chain output unit, the ready arbitration unit, and the branch processing unit.
[0046] The shift chain output unit monitors the status of the marker at the lowest level of the shift chain in real time. When the ready marker reaches and leaves the highest level of the shift chain after being shifted step by step, the shift chain output unit generates a ready valid signal and sends it to the status register of the corresponding source operand, updating the status of the corresponding source operand to ready.
[0047] For fast bypass paths in equal-width scenarios, the ready arbitration unit directly generates a ready valid signal in the same clock tick when the physical register number is hit, updating the state of the corresponding source operand to ready, without needing to go through the shift chain delay.
[0048] Once all source operands of a consumer instruction are updated to the ready state, the launch arbitration module arbitrates according to the instruction age and selects the oldest ready instruction to launch to the execution unit.
[0049] Within any clock cycle of the ready marker shift, if the branch processing module generates a branch flush signal, it immediately compares the rename identifier carried by the flush signal with the rename identifiers bound to all ready markers in the shift chain. If the identifiers match successfully, the ready markers in all levels of the corresponding shift chain are immediately cleared, and the ready state of the corresponding source operand is canceled, restoring it to the unready state, awaiting subsequent wake-up. For split instructions that have entered the transmit pipeline, the branch flush signal is synchronously sent to the corresponding pipeline stage, canceling the transmitted split clock cycles to ensure consistent pipeline states.
[0050] In some embodiments, determining the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions includes: identifying the category to which the vector instruction belongs, calculating the basic data path width of the source operand by combining the element bit width and the vector register group multiple; if the vector instruction is of the widening type, setting the data path width of the destination operand to twice the basic data path width of the source operand; if the vector instruction is of the narrowing type, setting the data path width of the source operand to twice the basic data path width of the destination operand.
[0051] The decoding module first identifies the opcode of the vector instruction to determine the operation category to which the instruction belongs. Simultaneously, it reads the element bit width parameter and vector register set multiplier parameter configured in the instruction to calculate the basic data path width of the source operand. The basic data path width is the path width required to read all the data of the source operand within a single cycle, and is obtained by multiplying the element bit width by the number of elements processed within a single cycle, and then multiplying by the vector register set multiplier.
[0052] When the instruction belongs to the widened operation class, the bit width of the instruction's operation result is twice the bit width of the source operand. In this case, the data path width of the destination operand is set to twice the basic data path width of the source operand, and the source operand is still read according to the basic data path width.
[0053] When the instruction belongs to the narrowing operation class, the bit width of the instruction's operation result is half the bit width of the source operand. In this case, the data path width of the source operand is set to twice the basic data path width of the destination operand, and the destination operand is written to the register according to the basic data path width.
[0054] This implementation method can accurately adapt to the source and destination bit width differences of special vector instructions such as widening and narrowing, providing accurate input parameters for subsequent width comparison and wake-up delay calculation.
[0055] In some embodiments, the calculation of the corresponding number of splits includes: obtaining the reference width corresponding to the full-width data path, calculating the ratio of the current data path width to the reference width, and determining the number of splits for a single instruction through the data path based on the ratio; the smaller the ratio value, the more splits are required, and the number of splits corresponding to the reference width is one.
[0056] The reference width value corresponding to the full-width data path is pre-stored. The reference width is the maximum single-shot transmission bit width that the vector processor data path can support, corresponding to the state where all memory cells are simultaneously turned on.
[0057] Retrieve the data path width value corresponding to the current command, and calculate the ratio of the current data path width to the reference width. The ratio value is a positive value obtained by dividing the current width by the reference width, and the ratio value includes multiple levels such as one-half, one-quarter, one-eighth, and one-sixteenth.
[0058] The number of data slices a single instruction passes through is determined by a ratio value. The smaller the ratio value, the more slices are required. When the ratio value is one, the number of slices is one, and the instruction can pass through the entire data path in a single slice. When the ratio value is one-half, the number of slices is two, and the instruction passes through two data slices in two slices. When the ratio value is one-quarter, the number of slices is four, and the instruction passes through four data slices in four slices, and so on.
[0059] This implementation method can map instructions of different bit widths into a unified integer-step splitting process, which facilitates timing control and resource scheduling of the transmit port.
[0060] In some embodiments, the process of assigning physical register numbers to the destination operand of a producer instruction and broadcasting the physical register number corresponding to the destination operand during the first cycle of producer instruction execution includes: retrieving an idle physical register number from the renaming free table and assigning it to the destination operand of the producer instruction; and sending the corresponding physical register number to the all source operand monitoring unit when the producer instruction enters the first cycle of the emit pipeline.
[0061] The renaming module internally maintains a free physical register table, which stores the numbers of all currently unused physical registers. When a producer instruction enters the renaming phase, the renaming module retrieves a free physical register number from the head of the free physical register table and assigns it to the destination operand of the corresponding instruction; at the same time, it removes the corresponding number from the free table, marks it as allocated, and binds it to the renaming identifier corresponding to the instruction.
[0062] When a producer instruction enters the first stage of the launch pipeline through the launch arbitration, the broadcast module of the corresponding execution path packages information such as the physical register number of the target operand of the corresponding instruction, the data path width level, the number of split stages, the current pipeline stage, and the forward path identifier into a broadcast signal and sends it to all source operand monitoring units in the given sequence queue.
[0063] For grouped vector instructions, the broadcast signal carries all physical register numbers within the group. When the source operand of a consumer instruction hits any number within the group, the corresponding wake-up logic is triggered.
[0064] This implementation method enables unified management of physical registers and global synchronization of wake-up signals, ensuring that all consumer instructions can capture the producer's execution status in a timely manner.
[0065] In some embodiments, if the source operand of the consumer instruction hits the broadcast physical register number, comparing the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction includes: pre-storing the physical register numbers that each source operand of the consumer instruction is waiting for, comparing the broadcast physical register number with each pre-stored number one by one; if the numbers match, it is determined to be a hit, and the data path width of the destination operand of the producer instruction is retrieved and compared with the data path width of the current source operand.
[0066] When each consumer instruction enters the ordering queue, its corresponding source operand monitoring unit will pre-store the physical register numbers that all source operands of the corresponding instruction are waiting for, and each source operand corresponds to an independent comparator.
[0067] When a broadcast signal arrives, the comparator corresponding to each source operand compares the physical register number carried in the broadcast with its own pre-stored physical register numbers. If the two numbers match exactly, the corresponding source operand is determined to have matched the broadcast signal.
[0068] In the same clock cycle after the hit determination is completed, the width comparison unit extracts the data path width level of the producer instruction destination operand from the broadcast signal, and at the same time extracts the data path width level of the current source operand from the configuration register of the consumer instruction, and compares the values of the two levels.
[0069] This implementation method can quickly complete hit detection and width difference recognition, providing a basis for subsequent selection of wake-up path. The entire process can be completed within a single frame without introducing additional timing overhead.
[0070] In some embodiments, if the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, injecting the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag shifts with the shift chain step by step, includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick.
[0071] The shift chain control unit pre-stores a mapping table between width differences and entry levels. When a width upgrade scenario is detected, the shift chain control unit calculates the difference level between the producer data path width and the consumer data path width. The difference level is in width increments, with each increment corresponding to one level of difference.
[0072] The corresponding shift chain entry level is retrieved from the mapping table based on the width difference level. The larger the width difference, the higher the corresponding entry level, and the more frames the ready marker stays in the shift chain. For example, a width difference of one increment corresponds to the first entry level, a difference of two increments corresponds to the second entry level, and so on.
[0073] Once the ready flag enters the shift chain from the corresponding entry level, it automatically moves one level to a higher level in the shift chain every clock cycle until it reaches and moves out of the highest level.
[0074] This implementation method can flexibly adapt to different width upgrade scenarios, ensuring that the ready time is precisely aligned with the actual data write-back time, avoiding data errors caused by waking up too early or performance losses caused by waking up too late.
[0075] In some embodiments, when the ready marker is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set, including: real-time detection of the marker state at the lowest level of the shift chain; when the ready marker moves to the lowest level of the shift chain and is moved out, a ready valid signal is generated, and the state of the corresponding source operand is updated to ready.
[0076] The lowest level output of the shift chain is connected to a readiness detection circuit, which samples the flag status of the lowest level of the shift chain in real time.
[0077] When the ready marker moves to the lowest level of the shift chain with the clock tick, and moves out of the lowest level on the current clock edge, the ready detection circuit captures the transition edge of the corresponding marker and generates a pulse-shaped ready valid signal.
[0078] A ready signal is sent to the ready status register of the corresponding source operand, and the status bit in the register is toggled from low to high, marking that the corresponding source operand is ready.
[0079] After the status update is completed, the transmit arbitration circuit immediately samples the status bits of all source operands of the corresponding instruction. If all source operands are high, the corresponding instruction is added to the transmit candidate list and waits for the next arbitration transmit.
[0080] This implementation method allows for precise control over the generation time of the ready signal, ensuring that the source operand's ready time is perfectly aligned with the time when the data is fully available, thus guaranteeing the correctness of data reading.
[0081] In some embodiments, the step of setting the ready signal directly after the source operand hits the broadcast physical register number if the producer instruction data path width is greater than or equal to the consumer instruction data path width includes: when it is determined that the producer data path width is greater than or equal to the consumer data path width, skipping the shift chain delay process, and directly updating the state of the corresponding source operand to ready in the same clock cycle when the physical register number is hit.
[0082] When the width comparison unit determines that the data path width of the producer instruction destination operand and the consumer instruction source operand are exactly equal, it generates a fast bypass enable signal to block the shift chain injection path.
[0083] In the same clock cycle when the physical register number hit is valid, the fast bypass path directly generates a ready valid signal and sends it to the ready status register of the corresponding source operand, updating the status bit to ready.
[0084] The entire process does not go through any level of the shift chain and does not introduce any additional clock delay, achieving a fast bypass effect that is ready as soon as it hits.
[0085] This implementation method can maintain optimal wake-up efficiency in normal scenarios with no width difference, avoid the performance loss caused by a uniform latency strategy, and balance correctness and execution efficiency.
[0086] In some embodiments, when a branch flush occurs during the shift process, clearing all markers in the corresponding shift chain and canceling the ready state of the corresponding source operand includes: when a branch flush signal is received, comparing the instruction identifiers associated with each marker in the shift chain; if the identifiers match successfully, clearing all ready markers at all levels of the corresponding shift chain, and simultaneously canceling the ready state of the corresponding source operand, restoring it to the unready state.
[0087] Each ready flag in the shift chain is bound to a rename identifier for the corresponding instruction. The rename identifier is a unique identifier assigned to the instruction during the renaming phase, and all source operands of the same instruction share the same rename identifier.
[0088] When a branch prediction fails, the branch processing module generates a branch flush signal, which carries a renaming identifier for all instructions on the failed prediction path.
[0089] After the flushing signal arrives at the shift chain control unit, it compares the renaming identifier bound to all ready markers in the shift chain with the renaming identifier carried by the flushing signal. If the two match, a zeroing operation is immediately triggered, clearing all ready markers in all levels of the corresponding shift chain to an invalid state.
[0090] Simultaneously, the zeroing operation triggers a ready state cancellation signal, restoring the ready state register of the corresponding source operand to a low level, i.e., a not-ready state. For instructions that have entered the emit pipeline, a flush signal is simultaneously sent to the corresponding stage of the pipeline, canceling the currently executing split-step operation and ensuring the consistency of the entire pipeline state.
[0091] This implementation method enables rapid response to branch prediction failure events, timely cleanup of invalid ready states, prevention of erroneous instructions from continuing to occupy launch and shift chain resources, and improvement of pipeline operational stability.
[0092] In some embodiments, such as Figure 2 As shown, the timing horizontal axis of this embodiment represents continuous clock beats, arranged sequentially from the first beat to the eighth beat; the vertical axis represents different operation levels and signal states, including nine levels: producer instruction, consumer instruction, physical register number broadcast and source operand hit judgment, data path width comparison, shift chain state, ready flag, ready signal, instruction issuance, and branch flush.
[0093] The first scenario is a delayed wake-up process from narrow producer to wide consumer, corresponding to the upper half of the timing diagram. The execution process is as follows: In the first phase, the producer instruction enters the first phase of the transmit pipeline and starts execution; at the same time, the broadcast physical register number signal takes effect, and the source operand hit judgment module starts comparison; the source operand of the consumer instruction completes the hit judgment within this phase, and the data path width comparison module starts width comparison synchronously.
[0094] In the second step, the width comparison is completed, and it is determined that the width of the producer's data path is less than the width of the consumer's data path, indicating a need for width upgrade. The ready marker is generated and injected into the first entry level of the shift chain, and the state of the first level of the shift chain becomes valid.
[0095] In the third step, the ready flag is shifted from the first level to the second level, the second level of the shift chain becomes valid, and the first level is restored to invalid.
[0096] In the fourth step, the ready flag is shifted from the second level to the third level, the third level of the shift chain becomes valid, and the second level is restored to invalid.
[0097] In the fifth step, the ready flag is shifted from the third level to the fourth level, the fourth level of the shift chain becomes valid, and the third level is restored to invalid.
[0098] In the sixth step, the ready flag is moved out of the lowest level of the shift chain, the ready signal is set, and the corresponding source operand of the consumer instruction is updated to the ready state; if all source operands of the consumer instruction are ready at this time, it enters the list of candidates to be issued.
[0099] On the seventh beat, the arbitration is completed, the consumer instruction is transmitted from the sequence queue to the execution unit, and the instruction transmission signal takes effect.
[0100] The second scenario is the equal-width instruction fast bypass ready process, corresponding to the middle part of the timing diagram. The execution process is as follows: In the first cycle, the producer instruction broadcasts the physical register number, and the source operand of the consumer instruction completes the hit determination within this cycle.
[0101] In the second step, the width comparison is completed, and it is determined that the widths of the two data paths are equal. This triggers the fast bypass path, directly sets the ready signal, and updates the source operand to the ready state. At the same time, the instruction can directly enter the transmit arbitration without going through the shift chain delay.
[0102] In the third step, the command completes arbitration and launches, significantly shortening the waiting period compared to the width upgrade scenario.
[0103] The third part is the processing sequence of branch flushing, and the execution process is as follows: At any beat in the shifting process, such as the fourth beat, the branch processing module generates a branch flush signal, which carries the renaming identifier of the corresponding instruction.
[0104] The shift chain control unit completes the identification comparison within this cycle. Once a match is successful, it immediately clears the ready flags of all levels in the shift chain and cancels the ready signal, restoring the source operand to the unready state. The instruction cycles that have entered the pipeline are synchronously canceled, and the entire shift chain returns to its initial idle state.
[0105] This implementation method can fully present the timing performance of the multi-level maintenance method in different scenarios, and clearly demonstrate the collaborative working logic of the three mechanisms of delayed wake-up, fast bypass and flushing cancellation.
[0106] Please see Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of the vector instruction ordering queue source operand multi-level maintenance system 200 provided in this application embodiment. The vector instruction ordering queue source operand multi-level maintenance system 200 is used to execute the steps of the vector instruction ordering queue source operand multi-level maintenance method shown in the above embodiments. The vector instruction ordering queue source operand multi-level maintenance system 200 can be a single server or a server cluster, or it can be a terminal, such as a handheld terminal, laptop computer, wearable device, or robot.
[0107] like Figure 3 As shown, the vector instruction ordering queue source operand multi-level maintenance system 200 includes: The time-of-decomposition calculation unit 201 is used to determine the data path width between the source operand and the destination operand of each vector instruction based on the vector instruction category, element bit width and vector register group multiple corresponding to multiple vector instructions during the decoding stage, and to calculate the corresponding time-of-decomposition; the vector instructions include producer instructions and consumer instructions; The numbering allocation unit 202 is used to allocate physical register numbers to the destination operand of the producer instruction. The physical register number corresponding to the destination operand is broadcast during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, the data path width of the destination operand of the producer instruction is compared with that of the source operand of the consumer instruction. If the data path width corresponding to the producer instruction is less than that corresponding to the consumer instruction, the ready flag of the corresponding source operand is injected into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. The status cancellation unit 203 is used to set the ready signal of the corresponding source operand when the ready mark is moved out of the lowest level of the shift chain; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all marks in the corresponding shift chain are cleared and the ready state of the corresponding source operand is cancelled.
[0108] During the decoding stage, the data path width between the source operand and the destination operand of each vector instruction is determined based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions, and the corresponding splitting number is calculated; the vector instructions include producer instructions and consumer instructions; Assign a physical register number to the destination operand of the producer instruction. Broadcast the physical register number corresponding to the destination operand during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, compare the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction. If the data path width of the producer instruction is less than that of the consumer instruction, inject the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all flags in the corresponding shift chain are cleared and the ready state of the corresponding source operand is canceled.
[0109] In some embodiments, determining the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions includes: identifying the category to which the vector instruction belongs, calculating the basic data path width of the source operand by combining the element bit width and the vector register group multiple; if the vector instruction is of the widening type, setting the data path width of the destination operand to twice the basic data path width of the source operand; if the vector instruction is of the narrowing type, setting the data path width of the source operand to twice the basic data path width of the destination operand.
[0110] In some embodiments, the calculation of the corresponding number of splits includes: obtaining the reference width corresponding to the full-width data path, calculating the ratio of the current data path width to the reference width, and determining the number of splits for a single instruction through the data path based on the ratio; the smaller the ratio value, the more splits are required, and the number of splits corresponding to the reference width is one.
[0111] In some embodiments, the process of assigning physical register numbers to the destination operand of a producer instruction and broadcasting the physical register number corresponding to the destination operand during the first cycle of producer instruction execution includes: retrieving an idle physical register number from the renaming free table and assigning it to the destination operand of the producer instruction; and sending the corresponding physical register number to the all source operand monitoring unit when the producer instruction enters the first cycle of the emit pipeline.
[0112] In some embodiments, if the source operand of the consumer instruction hits the broadcast physical register number, comparing the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction includes: pre-storing the physical register numbers that each source operand of the consumer instruction is waiting for, comparing the broadcast physical register number with each pre-stored number one by one; if the numbers match, it is determined to be a hit, and the data path width of the destination operand of the producer instruction is retrieved and compared with the data path width of the current source operand.
[0113] In some embodiments, if the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, injecting the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag shifts with the shift chain step by step, includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick.
[0114] In some embodiments, when the ready marker is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set, including: real-time detection of the marker state at the lowest level of the shift chain; when the ready marker moves to the lowest level of the shift chain and is moved out, a ready valid signal is generated, and the state of the corresponding source operand is updated to ready.
[0115] In some embodiments, the step of setting the ready signal directly after the source operand hits the broadcast physical register number if the producer instruction data path width is greater than or equal to the consumer instruction data path width includes: when it is determined that the producer data path width is greater than or equal to the consumer data path width, skipping the shift chain delay process, and directly updating the state of the corresponding source operand to ready in the same clock cycle when the physical register number is hit.
[0116] In some embodiments, when a branch flush occurs during the shift process, clearing all markers in the corresponding shift chain and canceling the ready state of the corresponding source operand includes: when a branch flush signal is received, comparing the instruction identifiers associated with each marker in the shift chain; if the identifiers match successfully, clearing all ready markers at all levels of the corresponding shift chain, and simultaneously canceling the ready state of the corresponding source operand, restoring it to the unready state.
[0117] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the vector instruction ordering queue source operand multi-level maintenance system and its modules described above can be referred to the corresponding content in the various embodiments of the vector instruction ordering queue source operand multi-level maintenance method, and will not be repeated here.
[0118] The aforementioned method for multi-level maintenance of source operands in a vector instruction ordering queue can be implemented as a computer program, which can be used in various ways, such as... Figure 3 It runs on the system shown.
[0119] Please see Figure 4 , Figure 4 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application. The computer device includes a processor, a memory, and a network interface connected via a device bus, wherein the memory may include a storage medium and internal memory.
[0120] The storage medium can store operating devices and computer programs. The computer program includes program instructions that, when executed, cause the processor to perform any vector instruction ordering queue source operand multilevel maintenance method.
[0121] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0122] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any vector instruction ordering queue source operand multilevel maintenance method.
[0123] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the terminal to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0124] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0125] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: During the decoding stage, the data path width between the source operand and the destination operand of each vector instruction is determined based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions, and the corresponding splitting number is calculated; the vector instructions include producer instructions and consumer instructions; Assign a physical register number to the destination operand of the producer instruction. Broadcast the physical register number corresponding to the destination operand during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, compare the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction. If the data path width of the producer instruction is less than that of the consumer instruction, inject the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all flags in the corresponding shift chain are cleared and the ready state of the corresponding source operand is canceled.
[0126] In some embodiments, determining the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions includes: identifying the category to which the vector instruction belongs, calculating the basic data path width of the source operand by combining the element bit width and the vector register group multiple; if the vector instruction is of the widening type, setting the data path width of the destination operand to twice the basic data path width of the source operand; if the vector instruction is of the narrowing type, setting the data path width of the source operand to twice the basic data path width of the destination operand.
[0127] In some embodiments, the calculation of the corresponding number of splits includes: obtaining the reference width corresponding to the full-width data path, calculating the ratio of the current data path width to the reference width, and determining the number of splits for a single instruction through the data path based on the ratio; the smaller the ratio value, the more splits are required, and the number of splits corresponding to the reference width is one.
[0128] In some embodiments, the process of assigning physical register numbers to the destination operand of a producer instruction and broadcasting the physical register number corresponding to the destination operand during the first cycle of producer instruction execution includes: retrieving an idle physical register number from the renaming free table and assigning it to the destination operand of the producer instruction; and sending the corresponding physical register number to the all source operand monitoring unit when the producer instruction enters the first cycle of the emit pipeline.
[0129] In some embodiments, if the source operand of the consumer instruction hits the broadcast physical register number, comparing the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction includes: pre-storing the physical register numbers that each source operand of the consumer instruction is waiting for, comparing the broadcast physical register number with each pre-stored number one by one; if the numbers match, it is determined to be a hit, and the data path width of the destination operand of the producer instruction is retrieved and compared with the data path width of the current source operand.
[0130] In some embodiments, if the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, injecting the ready flag of the corresponding source operand into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag shifts with the shift chain step by step, includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick.
[0131] In some embodiments, when the ready marker is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set, including: real-time detection of the marker state at the lowest level of the shift chain; when the ready marker moves to the lowest level of the shift chain and is moved out, a ready valid signal is generated, and the state of the corresponding source operand is updated to ready.
[0132] In some embodiments, the step of setting the ready signal directly after the source operand hits the broadcast physical register number if the producer instruction data path width is greater than or equal to the consumer instruction data path width includes: when it is determined that the producer data path width is greater than or equal to the consumer data path width, skipping the shift chain delay process, and directly updating the state of the corresponding source operand to ready in the same clock cycle when the physical register number is hit.
[0133] In some embodiments, when a branch flush occurs during the shift process, clearing all markers in the corresponding shift chain and canceling the ready state of the corresponding source operand includes: when a branch flush signal is received, comparing the instruction identifiers associated with each marker in the shift chain; if the identifiers match successfully, clearing all ready markers at all levels of the corresponding shift chain, and simultaneously canceling the ready state of the corresponding source operand, restoring it to the unready state.
[0134] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the steps of the vector instruction ordering queue source operand multilevel maintenance method provided in any embodiment of this application.
[0135] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for multi-level maintenance of source operands in a vector instruction ordering queue, characterized in that, include: During the decoding stage, the data path width between the source operand and the destination operand of each vector instruction is determined based on the vector instruction category, element bit width, and vector register group multiple corresponding to multiple vector instructions, and the corresponding splitting number is calculated; the vector instructions include producer instructions and consumer instructions; Assign a physical register number to the destination operand of the producer instruction. Broadcast the physical register number corresponding to the destination operand during the first execution of the producer instruction. If the source operand of the consumer instruction matches the broadcast physical register number, compare the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction. If the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, the ready flag of the corresponding source operand is injected into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. This includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick. When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all flags in the corresponding shift chain are cleared and the ready state of the corresponding source operand is canceled.
2. The method according to claim 1, characterized in that, The process of determining the data path width between the source and destination operands of each vector instruction during the decoding stage, based on the vector instruction type, element bit width, and vector register set multiples corresponding to multiple vector instructions, includes: Identify the category of the vector instruction and calculate the basic data path width of the source operand by combining the element bit width and the multiple of the vector register set; if the vector instruction is of the widening type, set the data path width of the destination operand to twice the basic data path width of the source operand; if the vector instruction is of the narrowing type, set the data path width of the source operand to twice the basic data path width of the destination operand.
3. The method according to claim 2, characterized in that, The calculation of the corresponding number of split beats includes: Obtain the reference width corresponding to the full-width data path, calculate the ratio of the current data path width to the reference width, and determine the number of times a single instruction is split through the data path based on the ratio; the smaller the ratio value, the more splitting steps there are, and the number of splitting steps corresponding to the reference width is one.
4. The method according to claim 1, characterized in that, The process of assigning physical register numbers to the destination operand of a producer instruction, and broadcasting the physical register number corresponding to the destination operand during the first execution of the producer instruction, includes: Retrieve an available physical register number from the rename free table and assign it to the destination operand of the producer instruction; when the producer instruction enters the first step of the emit pipeline, send the corresponding physical register number to the full source operand monitoring unit.
5. The method according to claim 1, characterized in that, If the source operand of the consumer instruction matches the broadcast physical register number, comparing the data path width of the destination operand of the producer instruction with that of the source operand of the consumer instruction includes: The system pre-stores the physical register numbers that the source operands of each consumer instruction are waiting for. It then compares the broadcast physical register numbers with each of the pre-stored numbers one by one. If the numbers match, it is considered a hit. The system then retrieves the data path width of the target operand of the producer instruction and compares it with the data path width of the current source operand.
6. The method according to claim 1, characterized in that, When the ready flag is moved out of the lowest level of the shift chain, the ready signal of the corresponding source operand is set, including: The status of the marker at the lowest level of the shift chain is monitored in real time. When the ready marker moves to the lowest level of the shift chain and is removed, a ready valid signal is generated, and the status of the corresponding source operand is updated to ready.
7. The method according to claim 1, characterized in that, If the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number, including: When it is determined that the producer's data path width is greater than or equal to that of the consumer, the shift chain delay process is skipped, and the status of the corresponding source operand is directly updated to ready in the same clock cycle when the physical register number is hit.
8. The method according to claim 1, characterized in that, When a branch flush occurs during the shift process, all markers in the corresponding shift chain are cleared, and the ready state of the corresponding source operand is revoked, including: When a branch flushing signal is received, the instruction identifier associated with each marker in the shift chain is compared. If the identifier matches successfully, the ready markers of all levels of the corresponding shift chain are cleared, and the ready state of the corresponding source operand is canceled and restored to the unready state.
9. A multi-level maintenance system for source operands of a vector instruction ordering queue, used to implement the method as described in any one of claims 1-8, characterized in that, include: The clock count calculation unit is used to determine the data path width between the source operand and the destination operand of each vector instruction during the decoding stage based on the vector instruction type, element bit width, and vector register group multiple corresponding to multiple vector instructions, and to calculate the corresponding clock count; the vector instructions include producer instructions and consumer instructions; The numbering allocation unit is used to allocate physical register numbers to the destination operand of the producer instruction. The first clock cycle of the producer instruction execution broadcasts the physical register number corresponding to the destination operand. If the source operand of the consumer instruction matches the broadcast physical register number, the data path width of the destination operand of the producer instruction and the source operand of the consumer instruction are compared. If the data path width corresponding to the producer instruction is less than the data path width corresponding to the consumer instruction, the ready flag of the corresponding source operand is injected into the corresponding entry level of the delayed wake-up shift chain, so that the ready flag is shifted step by step with the shift chain. This includes: calculating the width difference between the producer and consumer data paths, and determining the entry level of the delayed wake-up shift chain based on the width difference; the larger the width difference, the higher the corresponding entry level, and the ready flag enters the shift chain from the corresponding entry level, moving step by step to a lower level with the clock tick. The status cancellation unit is used to set the ready signal of the corresponding source operand when the ready mark is moved out of the lowest level of the shift chain; if the data path width of the producer instruction is greater than or equal to the data path width of the consumer instruction, the ready signal is set directly after the source operand hits the broadcast physical register number; when a branch flush occurs during the shift process, all marks in the corresponding shift chain are cleared and the ready state of the corresponding source operand is cancelled.
Citation Information
Patent Citations
Configurable data processor with multi-length instruction set architecture
CN1625731A
Mobile FLOW readout and mobile FLOW sequencer features
US20100164753A1