Cyclic shift register, data cyclic shift method and storage medium
By constructing a multi-level processing network using a hybrid approach of 2x2 switching units and 2x1 multiplexers, the critical path delay and hardware complexity of the cyclic shifter are optimized, achieving efficient and low-latency cyclic shift operations suitable for modern communication systems.
Patent Information
- Application Number
- CN202511393892.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Traditional barrel shifters suffer from excessively long critical paths and tight timing due to multi-stage cascading and global interconnection, making it difficult to meet the needs of high-throughput applications.
A multi-level processing network is constructed by combining 2x2 switching units and 2x1 multiplexers. Cyclic shift operations are achieved through shift control word generation stage control signals. Data paths are optimized by combining Omega and Banyan network topologies.
It reduces critical path latency and dynamic power consumption, improves data processing efficiency, simplifies hardware implementation complexity, and is suitable for high performance and high integrability requirements.
Smart Images

Figure CN120915747A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication transmission, in particular to a cyclic shift register, a data cyclic shift method and a storage medium. BACKGROUND
[0002] In the fifth generation (5G) and future communication systems, in order to guarantee the reliability of data transmission, channel coding technology is generally used to combat noise and interference in the transmission process. The quasi-cyclic low-density parity-check (QC-LDPC) code has become one of the core coding schemes of 5G and future communication standards due to its excellent error correction performance and structured design for easy hardware implementation.
[0003] The hardware decoder of QC-LDPC code is a key component for implementing this coding scheme, and its performance directly determines the throughput and delay of the communication system. Among the many operations of the decoder, cyclic shift (Cyclic Shift) is a basic operation that is frequently executed and is the core of the decoder. The scale of cyclic shift processing is determined by the expansion factor Z in the codeword structure. Traditional hardware implementation generally uses a barrel shifter, which is composed of log2(Z) levels of large-scale multiplexers (MUX) cascaded. However, when the value of Z is large, the selection signal of each level of multiplexer is heavily loaded, and the global connection is too long, resulting in a significant increase in signal transmission path delay, which in turn causes the critical path delay to be too large. At the same time, the excessively long critical path not only makes the timing difficult to converge, limiting the further improvement of data processing throughput, but also increases power consumption and layout complexity.
[0004] Therefore, the existing implementation scheme based on barrel shifter is difficult to meet the application requirements of high throughput while maintaining excellent timing characteristics, and there is an urgent need for a low-delay and high-efficiency cyclic shift hardware architecture. SUMMARY
[0005] The present application provides a cyclic shift register, a data cyclic shift method and a storage medium, which can solve the technical problems of long critical path and timing tension caused by multi-level cascading and global connection of traditional barrel shifter.
[0006] In a first aspect, the embodiments of the present application provide a cyclic shift register, comprising: a data input port for receiving an input vector with a bit width of Z; a data output port for outputting a result vector after cyclic shift with a bit width of Z; a control port for receiving a shift control word with a bit width of K, wherein K=log2(Z); A multi-stage processing network comprises K stages of processing units, each stage of processing units comprising a plurality of basic units; the plurality of basic units of each stage of processing units are independently controlled by a common stage control signal generated from one or more bits of the shift control word; the multi-stage processing network is configured to, under the control of the stage control signal, cause the input vector to achieve a result of cyclically shifting K bits after being selected by the plurality of basic units of each stage of processing units. In some embodiments, the basic units comprise two types of 2x2 permutation units and 2x1 multiplexers, the plurality of basic units of each stage of processing units are of the same type, and at least two stages of processing units employ different types of basic units.
[0007] In some embodiments, among the K stages of processing units, the first K-1 stages of processing units employ an Omega network topology, each stage comprising Z / 2 2x2 permutation units, and the Kth stage of processing units comprises Z 2x1 multiplexers. The 0th bit to the (K-2)th bit of the shift control word are respectively used as the stage control signals of the first K-1 stages of processing units to independently control the permutation units in the first K-1 stages of processing units; and the (K-1)th bit of the shift control word is used as the stage control signal of the Kth stage of processing units to control the multiplexers in the Kth stage of processing units.
[0008] In some embodiments, among the first K-1 stages of processing units, the interconnection connection relationship between the permutation units of the kth stage and the permutation units of the (k+1)th stage is defined by a perfect shuffle permutation function; where k = 1, 2, …, K-1.
[0009] In some embodiments, among the K stages of processing units, the first stage of processing units comprises Z 2x1 multiplexers, and the last K-1 stages of processing units employ a Banyan network topology, each stage comprising Z / 2 2x2 permutation units. The 0th bit of the shift control word is used as the stage control signal of the first stage of processing units to control the multiplexers in the first stage of processing units; and the 1st bit to the (K-1)th bit of the shift control word are respectively used as the stage control signals of the last K-1 stages of processing units to independently control the permutation units in the last K-1 stages of processing units.
[0010] In some embodiments, among the last K-1 stages of processing units, the interconnection connection relationship between the permutation units of the kth stage and the permutation units of the (k+1)th stage is defined by a preset rule; where k = 2, 3, …, K. The preset rule comprises a butterfly permutation, a shuffle permutation, or a self-defined permutation rule.
[0011] In some embodiments, each of the 2x2 switching units is composed of two 2x1 multiplexers, and is configured to be in pass-through mode or cross mode according to the stage control signal.
[0012] In some embodiments, the control port is further configured to receive a direction control signal for representing a direction of cyclic shift; and the multi-stage processing network is further configured to select a left cyclic shift operation or a right cyclic shift operation according to the direction control signal.
[0013] In a second aspect, the embodiments of the present application provide a data cyclic shift method, comprising: obtaining an input vector with a bit width of Z and a shift control word with a bit width of K; wherein K=log2(Z); inputting the input vector into a pre-constructed multi-stage processing network, and controlling a basic unit in a current stage processing unit to perform a shift operation according to a stage control signal corresponding to the current stage of the multi-stage processing network; generating and outputting a result vector with a bit width of Z after K-bit cyclic shift; The multi-stage processing network comprises K-stage cascaded processing units, each stage comprising a plurality of basic units; the plurality of basic units of each stage processing unit are independently controlled by a common stage control signal, and the stage control signal is generated by one or more bits of the shift control word; the basic units comprise two types of 2x2 switching units and 2x1 multiplexers, and the plurality of basic units of each stage processing unit are of the same type.
[0014] In some embodiments, when the basic unit in the current stage processing unit is a 2x2 switching unit, the control of the basic unit in the current stage processing unit to perform a shift operation comprises: when the stage control signal is at a first level, outputting a pair of input data directly; when the stage control signal is at a second level, outputting a pair of input data in cross mode; In some embodiments, when the basic unit in the current stage processing unit is a 2x1 multiplexer, the control of the basic unit in the current stage processing unit to perform a shift operation comprises: selecting one of two input data sources to output according to the logic level of the stage control signal.
[0015] In a third aspect, the embodiments of the present application provide a data cyclic shift device, comprising at least a memory and a processor; the memory is configured to store a computer-executed program or instruction, and the processor is configured to execute the computer-executed program or instruction to realize the data cyclic shift method according to any of the embodiments of the second aspect.
[0016] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, the data cyclic shift method according to any of the embodiments of the second aspect is implemented.
[0017] The cyclic shift register and the data cyclic shift method provided by the embodiments of the present application can greatly improve the data processing efficiency and quickly realize cyclic shift K bits by using a multi-stage processing network constructed by a 2x2 switching unit and a 2x1 multiplexer and by generating a level control signal from a shift control word. Meanwhile, compared with a traditional pure multiplexer structure, the 2x2 switching unit greatly reduces the fan-out load of the control signal and the global connection pressure, thereby reducing the critical path delay and dynamic power consumption and making the timing more easily convergent. The multiplexer ensures the accuracy and flexibility of the shift function, reduces the hardware implementation complexity as a whole, is conducive to chip integration, and effectively controls the cost.
[0018] In addition, the present application also provides a computer readable storage medium, which has the same beneficial effects as the verification method of the randomization effect. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate one embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.
[0020] Figure 1 A structural schematic diagram of a cyclic shift register is provided for an embodiment of the present application.
[0021] Figure 2 A structural schematic diagram of a 2x2 switching unit is provided for an embodiment of the present application.
[0022] Figure 3 A structural schematic diagram of a cyclic shift register is provided for another embodiment of the present application.
[0023] Figure 4 A structural schematic diagram of a cyclic shift register is provided for another embodiment of the present application.
[0024] Figure 5 A flowchart of a data cyclic shift method is provided for an embodiment of the present application.
[0025] Figure 6 A structural schematic diagram of a data cyclic shift device is provided for an embodiment of the present application.
[0026] The specific embodiments of the application have been shown by way of example in the above figures, and will be described in more detail hereafter. These drawings and the associated description are not meant to limit the scope of the inventive concept in any way but merely to illustrate the inventive concept by way of specific embodiments. DETAILED DESCRIPTION
[0027] The application will be described in further detail with reference to specific embodiments in conjunction with the attached drawings. Like elements in the drawings are denoted by like reference numerals. In the following description, numerous specific details are described to provide a thorough understanding of the application. However, it will be apparent to one skilled in the art that the application can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order to avoid obscuring the application. In the following description, the terms "couple" and "coupled" refer to an operational coupling, whether mechanical, electrical, or magnetic, or some combination thereof. The term "directly coupled" means that two elements are directly in contact with each other. The term "connected" or "coupled" can also mean that two or more elements are either no longer in contact with each other or are not direct ly in contact with each other, but nonetheless still co-operate.
[0028] In addition, features, operations, or steps described in the specification can be combined in any suitable manner without departing from the scope of the application. Similarly, the various steps or acts in a method can be combined, reordered, or omitted without departing from the scope of the application. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
[0029] The terms "first", "second", and the like, in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the embodiments of the application described herein are capable of operation in other sequences than described or illustrated herein. The terms "couple" and "coupled" and the like should be interpreted broadly to mean directly coupled, indirectly coupled through one or more intermediary coupled devices, and the like, as well as any combination thereof. Other definitions and terminology will be apparent to those of skill in the art in light of the teachings herein.
[0030] The technical solutions of the application and how the technical solutions solve the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the application will be described below with reference to the drawings.
[0031] Figure 1 The structure diagram of a cyclic shift register is provided for an embodiment of the present application. As shown in the figure, the cyclic shift register provided by the embodiment comprises a data input port 10, a data output port 20, a control port 30, and a multi-stage processing network 40 connected between the data input port 10 and the data output port 20. Figure 1
[0032] In the embodiment, the data input port 10 is used to receive an input vector with a bit width of Z.
[0033] The data input port 10 is an interface for the cyclic shift register to interact with an external data source, and its main function is to receive an input vector with a bit width of Z, which is a data set to be subjected to a cyclic shift operation. The bit width Z defines the maximum amount of data that can be processed at a time by the system, for example, Z=32, which means that the register can process a 32-bit vector at a time. In an actual circuit, the data input port 10 is not a simple interface, and it is usually followed by an input buffer register (Buffer) for temporarily storing input data under the control of a clock, so as to ensure that the input data remains stable during the shift operation and avoid calculation errors caused by input changes.
[0034] The data output port 20 is used to output a result vector after cyclic shift with a bit width of Z.
[0035] The data output port 20 is responsible for outputting the result vector after the cyclic shift processing to an external device or system. The bit width of the output data remains consistent with the input, which is still Z bits, but the content of the output data has been rearranged according to the control signal. In a communication system, the data after the cyclic shift processing can be used for signal modulation, coding, and other links. The data output port 20 can timely and accurately send the processed data to meet the real-time requirements of the system and ensure the normal operation of the entire system. In an actual circuit, similar to the input port, the output port is also usually followed by an output buffer for latching the final result, so that it can be stably read by the downstream circuit in the next clock cycle.
[0036] The control port 30 is used to receive a shift control word with a bit width of K, where K=log2(Z).
[0037] The process of cyclic shift operation is to cyclically shift the input Z data bits to achieve the expected order of output data, and the shift control word is the key parameter to control the cyclic shift operation, which determines the number of bits that the input vector needs to be cyclically shifted. For Z data bits, there are Z possible shift bit numbers (from 0 to Z-1 bits), and K-bit binary number can represent 2^K states, so K = log2(Z) is the minimum and sufficient control bit width. The shift control word is a binary number, whose value directly represents the number of bits requested to be cyclically shifted. By changing the value of the shift control word, the degree of cyclic shift can be flexibly controlled to meet the shifting needs of different data. For example, for a Z = 8 (K = 3) system, the control word "101" (decimal 5) means to request a cyclic right shift (or left shift) of 5 bits.
[0038] The multi-stage processing network 40 is used to control the input vector to achieve the result of cyclic shift K bits under the control of the stage control signal after selecting the multiple basic units of each stage processing unit. The multi-stage processing network 40 includes K stages of cascaded processing units, each stage processing unit including multiple basic units; the multiple basic units of each stage processing unit are independently controlled by a common stage control signal, and the stage control signal is generated by one or more bits in the shift control word.
[0039] The multi-stage processing network 40 is the core execution component of the cyclic shift register, which undertakes the specific task of implementing the cyclic shift of the input vector. In this embodiment, the structure of the multi-stage processing network 40 is a pipeline network composed of K stages of processing units, each stage processing unit is responsible for completing a specific "granularity" of shift operation, such as the i-th stage is responsible for shifting or not shifting 2^i bits, all K stages are connected in series through inter-stage interconnection, and cooperate to complete any shift amount from 0 to (Z-1) bits. This multi-stage cascaded structure is the hallmark of the barrel shifter, which sacrifices a small amount of gate delay to achieve extremely high parallel processing speed. Compared with the traditional serial shift register, it has achieved a great leap in performance.
[0040] Further, each stage processing unit includes multiple basic units, and the multiple basic units are of the same type and work in parallel. The basic units in each stage processing unit share the same stage control signal, i.e., are independently controlled by a common stage control signal, and the stage control signal is generated by one or more bits in the shift control word, so that the accurate control of each stage processing unit according to the shift control word can be realized. For example, in the simplest design, the control signal of the i-th stage (i starts from 0) is directly equal to the i-th bit of the control word. If the bit is '1', all units in the stage perform shift operation; if it is '0', it remains straight through.
[0041] In this embodiment, the basic unit includes two types of 2x2 switching unit SW and 2x1 multiplexer MUX, the basic units of each stage processing unit are of the same type, and at least two stages of processing units use basic units of different types.
[0042] The basic unit is the most basic switching switch that constitutes the network. In this embodiment, the basic unit can be selected from two types of 2x2 switching unit SW and 2x1 multiplexer MUX. The 2x2 switching unit SW is a unit with high symmetry, having two inputs and two outputs. According to the control signal, it can be switched between the pass-through state (input A→output A, input B→output B) and the cross state (input A→output B, input B→output A), that is, it can perform the exchange operation on the input data. The 2x1 multiplexer MUX has two data inputs, one data output, and one control signal. According to the control signal, it selects one of the two inputs to be connected to the output. Through the combination and cooperative work of these basic units, partial shift operations are performed on the input vector in each stage processing unit. After continuous processing of K stages of processing units, the final result of the input vector circularly shifted by K bits is obtained, and the function of the entire circular shift is completed.
[0043] In this embodiment, the basic unit in the multi-stage processing network 40 is not necessarily of the same type in all stages, but can be flexibly configured according to the design target. At least two stages of processing units use basic units of different types. That is, in the entire multi-stage processing network 40, at least two stages of processing units use basic units of different types. Through the combination and collocation of basic units of different types in different stages, the circular shift operation of the input vector can be more flexibly and efficiently realized, and the diversified data processing requirements can be met.
[0044] The 2x2 switching unit SW and the 2x1 multiplexer MUX are mixed to form the multi-stage processing network 40, which is mainly used to alleviate the signal load and wiring bottleneck problem faced by the traditional pure MUX structure. Specifically, the 2x2 switching unit SW with symmetrical interconnection can be used in the front stage (high-bit control signal stage). The key advantage is that the stage control signal only needs to drive a simple switching switch, rather than a multi-input multiplexer MUX, which greatly reduces the fan-out load of the control signal, significantly reduces the transmission delay and dynamic power consumption of the signal at this stage. In the rear stage (low-bit control signal stage), the 2x1 multiplexer MUX is still used because it has a more direct structure when implementing fine shifting. This mixed scheme breaks the deadlock that each stage of the pure MUX structure faces high load delay. By optimizing the first few stages with the heaviest key path, the overall key path delay is effectively shortened, and the timing is more easily converged.
[0045] In summary, the cyclic shift register provided by the embodiment shortens the critical path, ensures higher operating clock frequency in timing performance, greatly optimizes the wiring complexity and delay problem of super large scale integrated circuit in physical implementation, and is fully compatible with the functional characteristics of the traditional barrel shifter, seamlessly replaces the shift module in the existing architecture without changing the upper architecture, and is an optimized solution with high performance, high integrability and high compatibility.
[0046] In some embodiments, each 2x2 switching unit SW is composed of two 2x1 multiplexers MUX, and is configured in pass-through mode or cross mode according to the stage control signal.
[0047] Figure 2 The structure diagram of the 2x2 switching unit SW provided by an embodiment of the application is shown in FIG. 2. Figure 2 As shown in FIG. 2, the two MUXs share the same stage control signal, but the input connection modes are opposite, one MUX selects the input X0 when the control signal S is 0 and selects the input X1 when the control signal S is 1; the other MUX is configured to select the input X1 when the control signal S is 0 and select the input X0 when the control signal S is 1. The symmetrical connection mode enables the entire composite unit to seamlessly switch between the pass-through mode (X0→Y0, X1→Y1) and the cross mode (X0→Y1, X1→Y0) under the action of the control signal.
[0048] In some embodiments, the 2x2 switching unit SW can also adopt a transmission gate structure, or use NAND gate or NOR gate to build a pure combinational logic structure based on logic gates, and the functions of the implementation are controllable interconnection (pass-through or cross) of two data paths, which will not be described in detail here.
[0049] In some embodiments, the control port 30 is further configured to receive a direction control signal representing a direction of the cyclic shift; and the multi-stage processing network 40 is further configured to select a cyclic left shift operation or a cyclic right shift operation according to the direction control signal.
[0050] In the embodiment, the function of the control port 30 of the cyclic shift register is expanded, and in addition to receiving a shift control word with a bit width of K (K = log2(Z)) to determine the number of cyclic shifts, the direction control signal representing the direction of the cyclic shift is also received, the direction control function is added to the cyclic shift register, and the cyclic shift register is upgraded from a single-direction shifter to a general-purpose bidirectional cyclic shifter.
[0051] For example, the direction control signal is a global signal, which "reinterprets" or "modulates" the meaning of the original stage control signal through logical operation. Specifically, the system will logically combine the received direction control signal (e.g., 0 represents right shift, 1 represents left shift) with each bit in the shift control word (i.e., the original stage control signal) to actually control the 2x2 switching unit SW or multiplexer MUX of each stage. Under the control of the direction signal, the same set of stage control signals produces opposite responses: the operation of right shifting 2^i bits is "converted" to left shifting 2^i bits when the direction signal is valid. Such a design, without duplicating or adding extra data path hardware, can achieve powerful bidirectional shift function with only a little extra control logic, greatly improving hardware utilization and flexibility of function, and meeting the needs of modern processor instruction sets for left shift, right shift and other operations.
[0052] In some embodiments, in a cyclic shift register, K-stage cascaded processing units, the first K-1 stage processing units adopt Omega network topology, each stage is composed of Z / 2 2x2 switching units SW, and the Kth stage processing unit is composed of Z 2x1 multiplexers MUX. The 0th to K-2th bits of the shift control word are respectively used as the stage control signals of the first K-1 stage processing units, independently controlling the switching units SW in the first K-1 stage processing units; the K-1th bit of the shift control word is used as the stage control signal of the Kth stage processing unit, controlling the multiplexer MUX in the Kth stage processing unit.
[0053] Omega network is a classic multi-stage interconnection network topology, which is composed of multiple stages of switching units, each stage containing a full-mix topology line (i.e., complete mixing connection) and a column of controllable four-function switching units (supporting four operations of direct connection, exchange, upcast, and downcast). Its core feature is to adopt a complete mixing mode for inter-stage connection, i.e., the input data will be mixed by a binary bit cyclic left shift of one bit before entering the next stage, ensuring uniform distribution of data. The network operates through unit control mode, and each stage of switching units independently receives control signals, automatically sets the state according to the binary bit value of the target address, and realizes the path selection of data from the source node to the target node.
[0054] In this embodiment, the efficient and regular interconnection characteristics of Omega network are used to complete most of the coarse-grained shift, and then a final multiplexer MUX stage is used to correct the inherent limitations of Omega network and achieve accurate final output selection. Specifically, the first K-1 stages form a standard Omega network, each stage having Z / 2 2x2 switching units SW, which are independently controlled by the low bits (0 to K-2 bits) of the shift control word. These stages work together to efficiently move any input bit to the vicinity of the target area. Meanwhile, considering that the standard Omega network may have a fixed deviation in the implementation of the cyclic shift or cannot directly implement the accurate selection of the last step, the Kth stage (the last stage) is designed as an array of Z 2x1 multiplexers MUX, which is uniformly controlled by the highest bit (K-1 bit) of the control word. The role of this stage is like a fine tuner, each MUX selects one of the two possible candidate results from the Omega network output as the final output, thereby perfectly solving the boundary problem of cyclic shift and ensuring that each bit can be accurately cyclically moved to the specified final position.
[0055] In some embodiments, in the first K-1 stages of processing units, the interconnection connection relationship between the switching unit SW of the kth stage and the switching unit SW of the k+1th stage is defined by a perfect shuffle permutation function; wherein k=1, 2, …, K-1.
[0056] In this embodiment, after all the switching units SW of the kth stage have processed the data, the output results will not be directly sent to the corresponding port of the next stage, but will undergo a fixed rearrangement process called "perfect shuffle" before being used as the input of the switching unit SW of the k+1th stage. That is, for the switching units SW of the kth stage and the k+1th stage, the connection relationship between them is precisely defined by the perfect shuffle permutation function. The function definition of perfect shuffle is: for a Z-bit output vector, it is regarded as a deck of cards. First, evenly divide it into two piles (Z / 2 bits each) from the middle, and then take one bit from each pile alternately to recombine, just like shuffling. In short, the perfect shuffle permutation function is like a "data allocation guide". It accurately allocates the output data of the kth stage switching unit SW to the input port of the k+1th stage switching unit SW according to the complete shuffling method, so as to ensure that the data can flow and transmit in order and efficiently between multiple processing units, and thus ensure that the entire cyclic shift register can implement the cyclic shift operation of the data according to the design requirements.
[0057] Figure 3 The structure diagram of the cyclic shift register provided by another embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, the cyclic shift register includes K stages of processing units, and each stage of processing units includes Z / 2 switching units SW. Figure 3As shown, the cyclic shift register provided in the embodiment can process the cyclic shift of 8-bit data (bit width Z = 8), and has 3 stages of cascaded processing units, wherein the first 2 stages of processing units adopt an Omega network topology, each stage is composed of Z / 2 = 4 2x2 switching units SW, and the 3rd stage of processing units is composed of 8 2x1 multiplexers MUX. The 0th to 1st bits of the shift control word are respectively used as the stage control signals of the first 2 stages of processing units, and independently control the switching units SW in the first 2 stages of processing units; and the 2nd bit of the shift control word is used as the stage control signal of the 3rd stage of processing units, and controls the multiplexers MUX in the 3rd stage of processing units.
[0058] As a specific example, when the control word received by the control port 30 is in binary form (S2, S1, S0) = (1, 0, 1), which is 5 in decimal form, that is, the operation of cyclically right shifting the input vector by 5 bits is required, and the input vector is [D0, D1,..., D7], the register will operate the input vector according to the preset shift logic, and the final output vector after the cyclic right shift by 5 bits is [D3, D4, D5, D6, D7, D0, D1, D2].
[0059] In some embodiments, in the K-stage cascaded processing units in the cyclic shift register, the 1st stage of processing units is composed of Z 2x1 multiplexers MUX, and the last K-1 stages of processing units adopt a Banyan network topology, each stage is composed of Z / 2 2x2 switching units SW. The 0th bit of the shift control word is used as the stage control signal of the 1st stage of processing units, and controls the multiplexers MUX in the 1st stage of processing units; and the 1st to K-1th bits of the shift control word are respectively used as the stage control signals of the last K-1 stages of processing units, and independently control the switching units SW in the last K-1 stages of processing units.
[0060] The Banyan network is a multi-stage space division switching network based on tree structure, which is composed of 2x2 switching units in a butterfly interconnection manner, and has the characteristics of single path, self-routing and scalability. Its core mechanism is to control the switching unit state level by level through the binary encoding of the target address, to realize the automatic path finding of data from the input port to the output port, and the inter-stage connection follows the fixed rules of butterfly permutation. Each switching unit can independently make a routing decision (straight or cross) according to a bit of the target output address (usually the bit corresponding to the stage it is in), so as to accurately route any input to the specified output. This distributed self-routing control mechanism avoids complex global scheduling, making the hardware implementation very efficient.
[0061] In contrast to the previous embodiment, fine shifting is prioritized over large-scale routing. Specifically, the first stage processing unit is composed of a large Z-select-1 multiplexer (MUX) array, which is controlled by the least significant bit (bit 0) of the shift control word. This stage is responsible for implementing the minimum granularity (1-bit) cyclic shift, which can move the entire input vector by 0 or 1 bit in one cycle, thus solving the path complexity problem of pure switching networks when implementing single-bit shifting; the following K-1 stages are composed of 2x2 switching element (SW) arrays based on the Banyan network topology, which are controlled by the higher bits (bits 1 to K-1) of the control word. This network is responsible for completing the remaining coarse-grained shifts with a step size of 2 raised to the power (2, 4, 8...). This "fine first, coarse later" cascade order greatly simplifies the input pattern of the MUX array in the first stage, optimizes the data path, and thus can achieve better timing performance and lower wiring complexity than the "Omega first, MUX later" or other structures.
[0062] In some embodiments, in the K-1 stages of processing units, the interconnection relationship between the switching elements (SWs) of the kth stage and the switching elements (SWs) of the k+1th stage is defined by a preset rule; where k = 2, 3,..., K.
[0063] In this embodiment, after all the switching elements (SWs) of the kth stage have processed the data, the connection relationship between each of its output ports and the input ports of the specific switching elements (SWs) of the k+1th stage is strictly defined by a certain, pre-designed mathematical mapping rule (or permutation function). Common such rules include the butterfly permutation, shuffle permutation, or other custom permutation rules, etc. For example, in the butterfly permutation, the connection relationship is determined by a specific bit operation (such as XOR operation) on the binary address of the output port. This pre-defined, regular interconnection pattern is the basis for the high-efficiency self-routing of the Banyan network structure.
[0064] Figure 4 The structure diagram of the cyclic shift register provided by another embodiment of the present application is shown in FIG. 6. As shown in FIG. 6, the cyclic shift register is composed of a large Z-select-1 multiplexer (MUX) array, which is controlled by the least significant bit (bit 0) of the shift control word. This stage is responsible for implementing the minimum granularity (1-bit) cyclic shift, which can move the entire input vector by 0 or 1 bit in one cycle, thus solving the path complexity problem of pure switching networks when implementing single-bit shifting; the following K-1 stages are composed of 2x2 switching element (SW) arrays based on the Banyan network topology, which are controlled by the higher bits (bits 1 to K-1) of the control word. This network is responsible for completing the remaining coarse-grained shifts with a step size of 2 raised to the power (2, 4, 8...). This "fine first, coarse later" cascade order greatly simplifies the input pattern of the MUX array in the first stage, optimizes the data path, and thus can achieve better timing performance and lower wiring complexity than the "Omega first, MUX later" or other structures. Figure 4As shown, the circular shift register provided in this embodiment can handle the circular shift of 8 bits of data (bit width Z=8). This circular shift register has three cascaded processing units. The first-stage processing unit consists of eight 2x1 multiplexers (MUX). The latter two stages employ a Banyan network topology, with each stage consisting of Z / 2=4 2x2 switching units (SW). The 0th bit of the shift control word serves as the stage control signal for the first-stage processing unit, controlling the multiplexers (MUX) within it. The 1st and 2nd bits of the shift control word serve as the stage control signals for the latter two stages, independently controlling the switching units (SW) within each stage.
[0065] It is worth noting that both of the aforementioned multi-stage processing networks 40 with different architectures possess strong scalability. For larger bit widths Z (such as 16-bit, 32-bit, 64-bit, and even 1024-bit), expansion can be achieved by strictly adhering to the same design principles, namely, expanding the number of stages K to log2(Z). The number of 2x2 switching units SW in stage K-1 is then expanded to Z / 2, ensuring sufficient data exchange between stages. The number of 2x1 multiplexers MUX in another stage is expanded to Z, thus fully constructing a circular shift register system capable of processing larger bit widths. This expansion method ensures that the growth of hardware resources is linearly related to the bit width Z, while the latency of the critical path only increases logarithmically. This allows the design to maintain excellent timing characteristics, controllable power consumption and area, and orderly layout and routing even when dealing with high-bit-width, high-performance computing requirements.
[0066] Figure 5 This is a flowchart illustrating a data cyclic shifting method provided in one embodiment of this application. Figure 5 As shown, the data cyclic shifting method provided in this embodiment can be applied to the processor or controller in the cyclic shift register described in any of the above embodiments, and the multi-level processing network 40 is controlled by a program or instruction to perform the data cyclic shifting process. As mentioned above, the multi-level processing network 40 includes K cascaded processing units, each level containing multiple basic units; the multiple basic units of each level processing unit are independently controlled by a common level control signal, which is generated by one or more bits in the shift control word; the basic units include two types: 2x2 switching units (SW) and 2x1 multiplexers (MUX), the multiple basic units of each level processing unit are of the same type, and at least two levels of processing units use different types of basic units.
[0067] This data cyclic shifting method specifically includes the following steps: Step S510: Obtain an input vector with a bit width of Z and a shift control word with a bit width of K; where K = log2(Z).
[0068] Step S520, input the input vector into the pre-constructed multi-stage processing network, and control the basic units in the current stage processing unit to perform the shift operation according to the stage control signal corresponding to the current stage of the multi-stage processing network.
[0069] Step S530, generate and output the result vector with a bit width of Z which has completed K-bit cyclic shift.
[0070] In the embodiment, when performing data cyclic shift, the system first receives the input data vector with a bit width of Z and the shift control word with a bit width of K to be processed from an external data source or an internal storage unit, completes data loading and instruction analysis, wherein the bit width of the shift control word is associated with the bit width of the data to be processed, and the value of K is accurately calculated by the formula K = log2(Z); then the obtained input vector is input into the pre-constructed multi-stage processing network 40, each stage processing unit of the network receives a stage control signal generated by a specific bit of the total shift control word, and the multi-stage processing network 40 independently and in parallel performs the corresponding shift operation (such as straight-through or exchange) of the basic units (such as switching units SW and multi-way selectors MUX) in the current stage processing unit according to the stage control signal corresponding to each stage, each stage completes a 2 power granularity shift, and the multi-stage cascade completes complex cyclic shift; finally, the data processed by the last stage is latched and output, forming a result vector with a bit width of Z which has completed the specified K-bit cyclic shift, completing single-cycle, low-delay, high-throughput shift operation, and the further output result vector can be further used or processed by subsequent circuit modules or processing units.
[0071] In some embodiments, if the basic unit in the current stage processing unit is a 2x2 switching unit, the basic unit in the current stage processing unit is controlled to perform the shift operation, which specifically includes: When the stage control signal is at the first level, a pair of input data is output straight through; when the stage control signal is at the second level, a pair of input data is output crosswise.
[0072] It can be understood that the basic unit of the current processing unit is a 2x2 switching unit, based on its working principle, i.e. a pair of input data is processed by the 2x2 switching unit, the level state of the stage control signal received by the switching unit can be used to switch between two determined operation modes. Specifically, when the stage control signal is at a first level (e.g. logic '0'), the unit works in a pass-through mode, and the input data X0 and X1 are directly sent to the corresponding output port (X0→Y0, X1→Y1), i.e. the pass-through transmission of data is realized, and the corresponding position relationship between the input port and the output port is not changed; when the stage control signal is at a second level (e.g. logic '1'), the unit switches to a cross mode, and the paths of the input data are exchanged before output (X0→Y1, X1→Y0), i.e. the cross processing of data is realized, so that the data originally at the first input port is output to the second output port, and the data originally at the second input port is output to the first output port. The single switching unit completes the data exchange between two local bits, and the cooperative work of all Z / 2 switching units in the whole stage under the same control signal realizes the global cyclic shift operation through the pass-through or cross output mode controlled by the level.
[0073] If the basic unit in the current processing unit is a 2x1 multiplexer, the basic unit in the current processing unit is controlled to perform a shift operation, specifically including: According to the logic level of the stage control signal, one of the two input data sources is selected for output.
[0074] It can be understood that when the basic unit of the current processing unit is a 2x1 multiplexer MUX, based on its working principle, i.e. each 2x1 multiplexer MUX has two input data sources (usually from two different outputs of the previous stage network) and a stage control signal (generated by a bit of the total shift control word), when working, it does not modify or operate the data, but selects one of the two input data sources according to the logic level of the stage control signal and directly transmits the selected data to the output port. Specifically, when the stage control signal is at a low level (0), it selects and outputs the first input data source; when the signal is at a high level (1), it selects and outputs the second input data source. In the whole stage processing unit, all Z (or Z / 2) MUXs perform this selection operation in parallel and independently, but are commanded by the same stage control signal. By carefully configuring the connection mode of the two input sources of each MUX, for example, one is connected to the corresponding bit data, and the other is connected to the cyclic data separated by a certain number of bits, the cooperative selection operation of the MUXs in the whole stage can realize the cyclic shift function of a specific step size specified by the stage.
[0075] Figure 6The structural schematic diagram of the data cyclic shift device provided by an embodiment of the present application is shown in the figure. Figure 6 As shown in the figure, the data cyclic shift device provided by the embodiment includes at least a memory 610 and a processor 620, which can be connected through a bus.
[0076] In the embodiment, the memory 610 is configured to store computer-executed instructions or programs.
[0077] The memory 610 can include a volatile memory or a non-volatile memory, or the memory 610 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), or a Direct Rambus RAM (DRRAM).
[0078] The memory 610 includes but is not limited to these and any other suitable type of memory.
[0079] In the embodiment, the processor 620 is configured to execute computer-executed programs or instructions to implement the processes of any of the embodiments of the data cyclic shift method, and the same technical effects can be achieved. To avoid repetition, details are not described herein.
[0080] The processor 620 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0081] The embodiment of the present application further provides a readable storage medium, which stores programs or instructions, and the programs or instructions are executed by a processor to realize the processes of any embodiment of the data cyclic shift method and achieve the same technical effects. To avoid repetition, details are not described herein.
[0082] The processor can be a central processing unit (CPU) or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0083] The readable storage medium can be a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0084] Those skilled in the art can understand that all or part of the functions of the various methods in the above embodiments can be implemented in the form of hardware or in the form of a computer program. When all or part of the functions in the above embodiments are implemented in the form of a computer program, the program can be stored in a computer readable storage medium, which can include a read only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, a hard disk, etc. The above functions are implemented by executing the program by a computer. For example, the program is stored in a memory of a device, and when the program in the memory is executed by a processor, all or part of the above functions are implemented. In addition, when all or part of the functions in the above embodiments are implemented in the form of a computer program, the program can also be stored in a server, another computer, a disk, an optical disk, a flash disk or a mobile hard disk, etc. The program is downloaded or copied and saved in the memory of a local device, or the system of the local device is updated, and when the program in the memory is executed by a processor, all or part of the functions in the above embodiments are implemented.
[0085] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above specific embodiments, and the above specific embodiments are only illustrative and not restrictive. Those skilled in the art can make some simple deductions, modifications or replacements according to the idea of the present application without departing from the scope of the present application and the protection scope of the claims.
Claims
1. A cyclic shift register, characterized by, include: The data input port is used to receive an input vector with a bit width of Z; The data output port is used to output the result vector after cyclic shifting with a bit width of Z; The control port is used to receive shift control words with a bit width of K, where K = log2(Z); A multi-level processing network includes K cascaded processing units, each of which contains multiple basic units. The multiple basic units of each processing unit are independently controlled by a common level control signal, which is generated by one or more bits of the shift control word. The multi-level processing network is used to ensure that, under the control of the level control signal, the input vector achieves a cyclic shift of K bits after being selected by the multiple basic units of each processing unit. The basic unit includes two types: a 2x2 switching unit and a 2x1 multiplexer. The basic units of each processing unit are of the same type, and at least two processing units use basic units of different types.
2. The cyclic shift register of claim 1, wherein, In the K-level cascaded processing units, the first K-1 levels of processing units adopt the Omega network topology, each level consists of Z / 2 of the 2x2 switching units, and the K-level processing unit consists of Z of the 2x1 multiplexers. Bits 0 to K-2 of the shift control word serve as level control signals for the first K-1 level processing units, independently controlling the switching units in the first K-1 level processing units; bit K-1 of the shift control word serves as the level control signal for the K level processing unit, controlling the multiplexer in the K level processing unit.
3. The cyclic shift register of claim 2, wherein, In the first K-1 level processing units, the interconnection relationship between the k-th level switching unit and the k+1-th level switching unit is defined by the perfect shuffle permutation function; where k=1,2,…,K-1.
4. The cyclic shift register of claim 1, wherein, In the K-level cascaded processing units, the first-level processing unit consists of Z 2x1 multiplexers, and the subsequent K-1 level processing units adopt a Banyan network topology, with each level consisting of Z / 2 2x2 switching units. The 0th bit of the shift control word serves as the level control signal for the first-level processing unit, controlling the multiplexer in the first-level processing unit; the 1st to K-1th bits of the shift control word serve as the level control signals for the next K-1 level processing units, independently controlling the switching units in the next K-1 level processing units.
5. The cyclic shift register of claim 4, wherein, In the K-1 level processing unit, the interconnection relationship between the k-th level switching unit and the k+1-th level switching unit is defined by a preset rule; where k = 2, 3, ..., K. The preset rules include butterfly displacement, mixed shuffling displacement, or custom displacement rules.
6. The cyclic shift register of claim 1, wherein, Each of the 2x2 switching units consists of two 2x1 multiplexers, configured in either cut-through or crossover mode according to the stage control signal.
7. The cyclic shift register of claim 1, wherein, The control port is also used to receive a direction control signal that characterizes the direction of the cyclic shift; the multi-level processing network is also used to select to perform a cyclic left shift or a cyclic right shift operation according to the direction control signal.
8. A method of cyclically shifting data, the method comprising: include: Obtain an input vector with a bit width of Z and a shift control word with a bit width of K; where K = log2(Z); inputting the input vector into a pre-constructed multi-stage processing network, and controlling a basic unit in the current stage processing unit to perform a shift operation according to a stage control signal of a corresponding stage of each stage of the multi-stage processing network; generating and outputting a result vector with a bit width of Z and which has completed K-bit cyclic shift; wherein the multi-stage processing network comprises K stages of cascaded processing units, each stage comprising a plurality of basic units; the plurality of basic units of each stage of processing units are independently controlled by a common stage control signal, the stage control signal being generated by one or more bits in the shift control word; the basic units comprise two types of 2x2 cross-over units and 2x1 multiplexers, and the plurality of basic units of each stage of processing units are of the same type.
9. The data cyclic shift method of claim 8, wherein, when the basic unit in the current stage processing unit is a 2x2 cross-over unit, the control of the basic unit in the current stage processing unit to perform a shift operation comprises: when the stage control signal is at a first level, outputting a pair of input data directly; when the stage control signal is at a second level, outputting a pair of input data in a crossed manner; when the basic unit in the current stage processing unit is a 2x1 multiplexer, the control of the basic unit in the current stage processing unit to perform a shift operation comprises: selecting one of two input data sources to output according to a logic level of the stage control signal.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are used to implement the data cyclic shift method according to any one of claims 8 or 9 when executed by the processor.
Citation Information
Patent Citations
Method and device for reducing number of cascade stages of multi-stage cyclic shift network
CN109687877A
Code rate compatible LDPC encoder based on quasi-cyclic generation matrix
CN112039535A
Data cyclic shift device and method, chip, computer equipment and storage medium
CN114265625A
NR LDPC coding and decoding cyclic shift implementation device
CN117081608A
Reconfigurable barrel shifter and rotator
US8713399B1