Cyclic redundancy check hardware acceleration module integrated on processor and check acceleration method
By integrating a cyclic redundancy check (CRC) hardware acceleration module into the processor, the problems of insufficient flexibility and high latency in CRC calculation are solved, enabling efficient and flexible CRC calculation, optimizing the data processing flow, reducing CPU load, and ensuring data integrity and accuracy.
Patent Information
- Application Number
- CN202510937424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, CRC calculation suffers from insufficient flexibility, difficult configuration, low integration with processors, and large processing latency, resulting in resource-intensive and time-consuming software implementation that fails to meet the demands of high-speed data processing.
Design a hardware acceleration module for cyclic redundancy check integrated into the processor, including an instruction extension unit, a programmable register set, a pipelined computing unit, an address cache queue, a data cache queue, a checksum cache queue, and an arbitration unit. Flexible configuration is achieved through custom instructions and a programmable register set, and pipelined computing optimizes the data processing flow and reduces the CPU load.
It significantly improves CRC calculation efficiency, enhances flexibility and configurability, optimizes data processing flow, reduces CPU load, provides reliable error detection support, and ensures the integrity and accuracy of data transmission.
Smart Images

Figure CN120849178A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of integrated circuits and computers, and specifically relates to a cyclic redundancy check hardware acceleration module and a verification acceleration method integrated into a processor. Background Technology
[0002] In modern computing systems, data integrity and accuracy are paramount. Cyclic Redundancy Check (CRC), as a mechanism for detecting errors in data transmission or storage, has been widely applied in various systems such as IoT devices, wireless communications, embedded security applications, and network device configurations to provide error detection functionality. The CRC algorithm generates a short, fixed-length check value (or CRC code) by applying a specific polynomial to a data block, thereby achieving error detection. Currently, commonly used CRC verification methods in communication systems mainly include software and hardware implementations. Software implementation, implemented through software programming, involves bit manipulation and polynomial division. While relatively flexible and customizable to specific needs, it can become a performance bottleneck as data volume increases and real-time processing requirements rise. Hardware implementation involves developing dedicated hardware circuits to accelerate the CRC verification process. Hardware implementation offers higher performance and efficiency, making it suitable for applications with high speed requirements. The development of hardware acceleration modules has become crucial for improving CRC calculation efficiency. Especially in embedded systems and network devices, these modules can significantly improve data processing speed and reduce system latency. In the existing technology, there are some hardware-accelerated CRC solutions, but these solutions usually face problems such as insufficient flexibility, difficult configuration, low integration with the processor, and large processing latency.
[0003] Based on this, the present invention proposes a hardware acceleration module and a verification acceleration method for cyclic redundancy check integrated into a processor. Summary of the Invention
[0004] To address the aforementioned problems in the prior art, namely insufficient flexibility, difficult configuration, low integration with the processor, and high processing latency, this invention provides a cyclic redundancy check hardware acceleration module and a check acceleration method integrated into the processor.
[0005] A first aspect of the present invention provides a cyclic redundancy check hardware acceleration module integrated into a processor, the module comprising:
[0006] The instruction extension unit implements multiple sets of custom instructions through the coprocessor interface. The format of the custom instructions includes a function code field that controls the reading and writing of processor architecture registers.
[0007] The programmable register set includes a configuration register, a polynomial register, an initial value register, and an XOR value register. The configuration register supports dynamically setting the module enable state, automatic check code extraction function, polynomial bit width selection, input / output data inversion enable, and the amount of data corresponding to a single check code.
[0008] The pipelined computing unit includes a checksum calculation module with a multi-level processing structure. It performs finite field division operations according to the programmable register group configuration and controls the calculation process through a counting unit.
[0009] An address cache queue stores memory addresses and automatically triggers write operations after calculations are completed.
[0010] A data buffer queue that receives a stream of data to be verified from memory or registers;
[0011] A checksum cache queue stores reference checksums to be compared.
[0012] The arbitration unit coordinates the priority of memory access requests from different instructions;
[0013] The instruction response unit generates execution feedback based on the instruction type and module status.
[0014] Furthermore, the custom instructions include control register read / write instructions, memory data loading instructions, register data loading instructions, result storage instructions, and checksum comparison instructions;
[0015] When the result storage instruction is executed, the address cache queue immediately stores the target address and returns a response signal, while the calculation result is asynchronously written to memory.
[0016] Furthermore, each register in the programmable register group is configured as follows:
[0017] The polynomial register is used to store programmable polynomial coefficients;
[0018] The initial value register is used to define the starting value for calculation;
[0019] The XOR value register is used to configure the result mask value;
[0020] The data quantity field of the configuration register defines the number of data bytes corresponding to a single checksum.
[0021] Furthermore, the address cache queue operates as follows:
[0022] When receiving a storage instruction, the target address is cached and an execution signal is fed back immediately.
[0023] When the pipeline calculation unit completes the signal and the queue is not empty, a memory write request is automatically initiated.
[0024] Furthermore, the data cache queue operates as follows:
[0025] Receive verification data from different sources through the data selection channel;
[0026] When the queue is not full, a signal indicating that the instruction execution is complete should be sent.
[0027] A continuous data stream is provided to the pipeline computing unit through a ready signal handshake mechanism.
[0028] Furthermore, the working method of the checksum cache queue in the automatic extraction enabled state is as follows:
[0029] The reference check code is extracted from the end of the data stream using the counting unit;
[0030] When the output of the calculation unit is valid, a checksum comparison is performed automatically.
[0031] Furthermore, the priority method for the arbitration unit is as follows:
[0032] When the data cache queue is full, storage requests from the address cache queue are processed first.
[0033] In other scenarios, memory accesses that require data loading are prioritized.
[0034] Furthermore, the pipeline calculation unit includes:
[0035] The data selector selects the initial value or intermediate result input based on the start signal;
[0036] Multi-level arithmetic units perform polynomial division and pass intermediate results level by level;
[0037] The result processing module performs bit order adjustment and masking operations on the final result according to the configuration.
[0038] The checksum extraction module extracts the end segment of the data stream based on the counting signal as a reference checksum.
[0039] Furthermore, the processing logic of the instruction response unit for the result write-back instruction is as follows:
[0040] If the calculation is not completed and the data cache queue is empty, an error response is returned;
[0041] If the calculation is complete, wait for the operation channel handshake to succeed before writing to the register.
[0042] In another aspect, the present invention proposes a verification acceleration method based on a cyclic redundancy check (CRC) hardware acceleration module integrated into a processor. The method includes:
[0043] Configure polynomial parameters, initial values, and data volume parameters using control commands;
[0044] The data stream is written to the data buffer queue using a load instruction;
[0045] The pipeline calculation unit automatically extracts data and performs pipeline calculations.
[0046] The target address is pre-stored by a storage instruction, and the calculation result is automatically written into memory when it is ready.
[0047] By comparing the pre-stored reference check code with the comparison command, the system automatically performs a data integrity check when the calculation result is ready.
[0048] The beneficial effects of this invention are:
[0049] Significantly improve CRC calculation efficiency: Through the high integration of hardware acceleration modules and processors, CRC calculation efficiency will be greatly improved, meeting the needs of high-speed data processing and accelerating data transmission and processing.
[0050] Enhanced flexibility and configurability of CRC calculation: The configurable CRC module will support CRC calculation requirements under multiple communication protocols, making the system more flexible and adaptable, and providing customized CRC calculation solutions for different application scenarios.
[0051] Optimize data processing flow: The design of custom instruction sets, pipeline structure and built-in counters will optimize the data processing flow, realize efficient processing of continuous data blocks and automated CRC result generation, and improve the overall data processing efficiency of the system.
[0052] Reduce CPU load: Reducing the number of times the CPU is involved in CRC calculations will reduce the CPU load, freeing up processor cores for other important tasks and improving system performance and response speed.
[0053] Enhancing the flexibility of result output: Introducing an address FIFO queue will improve the flexibility and efficiency of result output, enabling the pre-configuration of result storage addresses and ensuring that data processing results can be output to the specified location in a timely and accurate manner.
[0054] Provides reliable error detection support: Through efficient CRC calculation and flexible configuration options, this technology will provide reliable error detection support for data communication and storage systems, ensuring the integrity and accuracy of data transmission. Attached Figure Description
[0055] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0056] Figure 1 This is an architecture diagram of the first embodiment of the present invention;
[0057] Figure 2 This is a schematic diagram showing the positional relationship between the coprocessor and the processor in the first embodiment of the present invention;
[0058] Figure 3 This is a schematic diagram of the programmable register group in the first embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the arbitration unit in the first embodiment of the present invention;
[0060] Figure 5 This is a schematic diagram of the flow calculation unit in the first embodiment of the present invention;
[0061] Figure 6 This is a schematic diagram of the structure of the instruction response unit in the first embodiment of the present invention;
[0062] Figure 7 This is a schematic diagram of the structure of the built-in decoding module in the instruction extension unit in the first embodiment of the present invention. Detailed Implementation
[0063] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0064] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0065] This invention provides a hardware acceleration module for cyclic redundancy check integrated into a processor, the module comprising:
[0066] The instruction extension unit implements multiple sets of custom instructions through the coprocessor interface. The format of the custom instructions includes a function code field that controls the reading and writing of processor architecture registers.
[0067] The programmable register set includes a configuration register, a polynomial register, an initial value register, and an XOR value register. The configuration register supports dynamically setting the module enable state, automatic check code extraction function, polynomial bit width selection, input / output data inversion enable, and the amount of data corresponding to a single check code.
[0068] The pipelined computing unit includes a checksum calculation module with a multi-level processing structure. It performs finite field division operations according to the programmable register group configuration and controls the calculation process through a counting unit.
[0069] An address cache queue stores memory addresses and automatically triggers write operations after calculations are completed.
[0070] A data buffer queue that receives a stream of data to be verified from memory or registers;
[0071] A checksum cache queue stores reference checksums to be compared.
[0072] The arbitration unit coordinates the priority of memory access requests from different instructions;
[0073] The instruction response unit generates execution feedback based on the instruction type and module status.
[0074] To more clearly illustrate the cyclic redundancy check (CRC) hardware acceleration module integrated into a processor according to the present invention, this embodiment specifically describes a CRC hardware acceleration module integrated into a processor, which will be discussed in conjunction with the following description. Figure 1 The various units in the embodiments of the present invention are described in detail below:
[0075] The instruction extension unit implements multiple sets of custom instructions through the coprocessor interface. The format of the custom instructions includes a function code field that controls the reading and writing of processor architecture registers.
[0076] In this embodiment, the instruction extension unit implements eight RISC-V custom instructions (crc_csrw, crc_csrr, crc_load, crc_store, crc_rload, crc_rstore, crc_cmp_r, crc_cmp_m) through the coprocessor interface. The instruction encoding adopts the R-type format, and the read and write of the processor architecture registers are controlled by the xd / xs1 / xs2 bits of the func3 field.
[0077] like Figure 7 As shown, the instruction extension unit has a built-in decoding module that parses the binary encoding field of the custom instruction in real time and generates multiple control signals according to the preset instruction mapping rules.
[0078] The decoding module decodes different control signals according to different instructions. The signals are as follows: Figure 7 As shown, when different instructions arrive, the decoder pulls different control lines high, thereby selecting different signal paths.
[0079] In this embodiment, the custom instructions include control register read / write instructions, memory data loading instructions, register data loading instructions, result storage instructions, and checksum comparison instructions; when the result storage instruction is executed, the address cache queue immediately stores the target address and returns a response signal, and the calculation result is asynchronously written to memory.
[0080] Specifically, one of the significant advantages of RISC_V is its scalability. Users can define custom instructions to implement different hardware functions. In the existing RISC_V architecture, there are already four sets of predefined instruction types, namely custom-0, custom-1, custom-2, and custom-4 in Table 1. Therefore, hardware accelerators can be designed using custom instructions.
[0081] Table 1: opcode(inst[6:0]) custom-0 0001011 custom-1 0101011 custom-2 1011011 custom-3 1111011
[0082] Table 1 shows the 32-bit user-defined instruction opcodes in the RISC_V architecture.
[0083] The main functions of the CRC hardware system include, as the sender, performing CRC calculations on the data to obtain the CRC checksum and sending it to the receiver; and as the receiver, receiving the data and checksum and performing verification. The instructions in Table 2 are used to perform the CRC verification function. The encoding format of the instructions is the R-type instruction in the risc_v architecture. Specifically, the xd, xs1, and xs2 bits in the func3 field indicate whether the CRC coprocessor module needs to read or write the corresponding architecture register in the processor. xd indicates whether the result needs to be written back to the architecture register pointed to by the rd field; 1 indicates writing is required, and 0 indicates not writing. The CRC coprocessor module will output a corresponding valid signal to the processor for the processor to determine whether to write. xs1 or xs2 indicates whether the architecture register pointed to by the rs1 or rs2 field needs to be read. The CRC coprocessor module will use whether this field is 0 to select whether to mask the value of the register. For rs1 and rs2, after receiving the instructions defined by the coprocessor, the processor can choose to pass the values of the rs1 and rs2 registers to the coprocessor. The coprocessor will mask unnecessary registers based on the different values of the func3 field in the instruction. After the instruction is completed, the processor can determine whether the rd value has been written based on the valid output signal of rd. Alternatively, the processor can directly send the values of the registers required by the coprocessor based on the value of the func3 field of the instruction. For rd, the processor can also determine whether to write the result based on the valid signals related to rd in the coprocessor, or determine whether to write based on the value of the func3 field of the instruction. Alternatively, since the instruction number rd, which does not need to write to the rd register, is 0, the processor can do nothing and write all the results by default, thus allowing for flexible processor design. The processing logic of rs1, rs2, and rd mentioned above can also be implemented in hardware without using the func3 field of the instruction; this selection can be directly fixed based on the function implemented by the instruction. Since the rs2 source register is not used in the design of the CRC coprocessor, the value of func7 is kept consistent, while the encoding space of rs2 is used to distinguish instructions. Execution of computational instructions requires a longer combinational logic path than data transmission instructions. Considering that all data entering this module needs CRC verification, no dedicated CRC calculation instruction was designed. The specific behavior of the instruction is as follows:
[0084] (1) crc_csrw: Used to write to the control registers of the CRC processor module. There are 4 control registers, numbered 1-4, occupying 2 bits of the rd field (see the control register design module for details on control registers). This instruction reads the value of the architecture register, numbered rs1, and writes it to the register of the CRC module, numbered the lower two bits of the rd field.
[0085] (2) crc_csrr: Used to read the control registers of the CRC processor module. This instruction reads the value of the register of the CRC module, the lower two bits of the rd field, and writes it into the architecture register, the rs1 field.
[0086] (3) crc_load: Used to load data from memory into the crc module. This instruction reads the value in the rs1 architecture register and uses it as the address to access memory.
[0087] (4) crc_store: Used to store the result of the CRC calculation back into memory. This instruction will read the value in the rs1 architecture register and use it as the address to access memory.
[0088] (5) c_rload: Used to load the value of the processor's rs1 architecture register into the rc module for rc calculation.
[0089] (6) crc_rstore: Used to write the result of the CRC calculation to the processor's rs1 architecture register.
[0090] (7) crc_cmp_r: is used to receive the check code from the architecture register and compare it with the calculated check code to determine whether the data is abnormal. The received check code is passed into the module through the rs1 register and compared with the check code calculated in the module to determine whether the received data is abnormal.
[0091] (8) crc_cmp_m: Used to receive the check code from memory and compare it with the calculated check code to determine whether the data is abnormal. The received check code is passed into the module through the rs1 register and compared with the check code calculated in the module to determine whether the received data is abnormal.
[0092] Table 2:
[0093] Table 2 is the CRC coprocessor instruction table.
[0094] The architecture diagram of the CRC coprocessor is as follows: Figure 1 As shown, the coprocessor interface has two channels: the instruction channel sends instructions to the coprocessor, which processes them and returns the results; the LSU channel interacts with memory by controlling the main processor's LSU. During integration with the main processor, it can be broken down into more detailed channels for interconnection. The specific design of each module is as follows:
[0095] The programmable register set includes a configuration register, a polynomial register, an initial value register, and an XOR value register. The configuration register supports dynamically setting the module enable state, automatic check code extraction function, polynomial bit width selection, input / output data inversion enable, and the amount of data corresponding to a single check code.
[0096] The registers in the programmable register group are configured as follows:
[0097] The polynomial register is used to store programmable polynomial coefficients;
[0098] The initial value register is used to define the starting value for calculation;
[0099] The XOR value register is used to configure the result mask value;
[0100] The data quantity field of the configuration register defines the number of data bytes corresponding to a single checksum.
[0101] In this embodiment, the configuration register CRC_CFG supports dynamic setting of module enable status, automatic check code extraction function, polynomial bit width selection, input / output data inversion enable, and the amount of data corresponding to a single check code. This register group is read and written through dedicated instructions: when the processor executes the control register write instruction, the decoding module connects the data path and writes the value in the architecture register into the control register with the specified number; when the control register read instruction is executed, the output path is selected to transmit the value of the target register back to the architecture register.
[0102] The polynomial register CRC_POLY stores programmable polynomial coefficients, and its bit width configuration must be consistent with the polynomial bit width selection field in the configuration register, supporting multiple polynomial specifications such as 4-bit, 8-bit, 16-bit, and 32-bit; the initial value register CRC_INIT defines the calculation start value and is loaded into the calculation unit when the pipeline starts; the XOR value register CRC_XOR configures the result mask value and is used to perform XOR operation on the final check code.
[0103] The configuration register adopts a multi-field composite structure: the least significant bit is the module enable bit, which controls the opening and closing of the hardware acceleration function; the second least significant bit is the automatic checksum extraction enable bit, which automatically extracts the checksum of the last bit of the data stream in receive mode; the second and third bits are the polynomial bit width selection field, which specifies the effective bit length of the polynomial coefficients in binary encoding; the fourth bit controls the input data bit order reversal function, and the fifth bit controls the output data bit order reversal function, which are adapted to the bit order specifications of different communication protocols respectively.
[0104] The high 26 bits of the configuration register constitute the data volume field, which defines the number of data bytes corresponding to a single check code. When the automatic check code extraction function is enabled, the value of this field controls the check code truncation position through a counter, thereby realizing the automatic separation of the data stream and the check code.
[0105] Specifically, the control registers are mainly used to control the enabling / disabling of C and configure some parameters in the CRC calculation process. This allows the CRC module to be applicable to different communication protocols, greatly enhancing its versatility and flexibility. As shown in Table 3, there are four control registers, all 32 bits wide. The crc_cfg register contains different fields, each with a different meaning. Bit 0 is the coprocessor enable bit, which can disable the module to reduce power consumption when the coprocessor is idle; bit 1 enables automatic checksum stripping, used in the receiving state, with a polynomial width of 32, to achieve automatic checksum stripping and automatic verification; bits 2 and 3 represent the polynomial width, i.e., the number of bits of the highest power of the given polynomial. The CRC coprocessor supports all polynomials with a highest power of 4, 8, 16, and 32 bits, and the polynomial width should be consistent with the value of the polynomial in the CRC_POLY register. Bits 4 and 5 indicate whether the input and output data in the CRC calculation are reversed, which is also part of some communication protocol specifications. Bits 6-31 indicate how many bytes of data are placed in one checksum. CRC_POLY, COC_INIT, and CRC_XOR are parameters required for CRC calculation, namely the value of the CRC polynomial, the initial value used in the CRC calculation, and the value to be XORed at the end.
[0106] The programmable register set, also known as the crc_control_reg module, has the following structure: Figure 3 As shown, after the decoding module recognizes the crc_csw or crc_csr instruction, it will enable the write or read operation of that module. If the instruction is a write instruction, demux will activate the input data path of the control register module, and the value of rs1 will be written to the corresponding numbered register. If the instruction is a read instruction, mux will activate the output path of the control register module, sending the data to the processor for reception via the rd interface.
[0107] Specifically, the hardware acceleration module accesses the programmable register set through custom instructions crc_csrw (write control register) and crc_csrr (read control register).
[0108] Write operation (crc_csrw) implementation: When the instruction decoding unit, that is, the decode module, recognizes the crc_csrw instruction:
[0109] The decoding unit generates the corresponding control signal ("write enable signal").
[0110] This control signal selects the demux multiplexer, connecting the data value from the processor architecture register (specified by the rs1 field of the instruction) to the input path of the target control register (specified by the lower two bits of the rd field of the instruction).
[0111] The target control register samples and stores the input data value on the active edge of the clock.
[0112] The instruction response unit, namely the instr_respond module, can return a response signal indicating that the instruction execution is complete in the next clock cycle, indicating that the write operation is completed immediately.
[0113] When the instruction decoding unit recognizes the read operation crc_csrr instruction:
[0114] The decoding unit generates a corresponding control signal, namely the "read enable" signal.
[0115] This control signal selects the multiplexer mux, which outputs the current value stored in the target control register (specified by the lower two bits of the instruction's rd field).
[0116] The output register value is transmitted back to the processor through the module's output interface, namely the rd interface.
[0117] The processor writes this value to the target architecture register specified by the instruction rs1 field.
[0118] The instruction response unit can return a response signal indicating that the instruction execution is complete in the next clock cycle, indicating that the read operation is completed immediately.
[0119] Connections with other modules:
[0120] The configuration signals (crc_width, in_reversion, out_reversion, autopeel_en, data_width) output by the programmable register group are directly connected to modules such as crc_calculate (pipeline calculation unit) and CRC_FIFO (checksum buffer queue) to dynamically control their calculation behavior, data processing flow and automatic checksum extraction / comparison logic.
[0121] The values output by the CRC_POLY, CRC_INIT, and CRC_XOR registers are directly used as input parameters for the CRC calculation by the crc_calculate module.
[0122] Table 3:
[0123] Table 3 is the description table of the control registers.
[0124] The pipelined computing unit, namely the crc_calculate module, contains a checksum calculation module with a multi-level processing structure. It performs finite field division operations according to the programmable register group configuration and controls the calculation process through the counting unit.
[0125] The flow calculation unit includes:
[0126] The data selector selects the initial value or intermediate result input based on the start signal;
[0127] Multi-level arithmetic units perform polynomial division and pass intermediate results level by level;
[0128] The result processing module performs bit order adjustment and masking operations on the final result according to the configuration.
[0129] The checksum extraction module extracts the end segment of the data stream based on the counting signal as a reference checksum.
[0130] In this embodiment, as Figure 5 As shown, the pipeline computing unit adopts a multi-level processing structure and includes the following key sub-modules:
[0131] Data selector: Receives the data stream and initial value (crc_init) from the upper-level data buffer queue (D_FIFO).
[0132] When the data packet header arrives (determined by the counter signal), the start signal is set, and the selector outputs the initial value as the calculation input.
[0133] When processing data in a non-first round, the selector outputs the value obtained by XORing the intermediate result (data_temp) of the previous round of calculation with the current input data.
[0134] Multi-level arithmetic unit (corresponding to LSRF_calculate): Performs polynomial division operations on the Galois field (GF(2)), with the core operation being modulo-2 division with the polynomial register (CRC_POLY) in the programmable register set. A pipelined architecture is adopted, with data passing through multiple levels of logic gates in a time-sharing manner, each level generating a partial result and passing it to the next level.
[0135] Result processing module: Receives the final output of the multi-level arithmetic unit.
[0136] The result is bit-reversed based on the out_reversion bit in the configuration register (CRC_CFG).
[0137] Perform a mask operation (XOR operation) with the value of the XOR register (CRC_XOR).
[0138] Checksum extraction module (integrated into counter logic): When autopeel_en is activated, a fixed-length bit segment is extracted from the end of the data stream based on the configured data_width field as a reference checksum and written to the checksum buffer queue (CRC_FIFO).
[0139] The computational flow control uses a counter-driven mechanism: a built-in counter (cnt) continuously monitors the amount of data processed. The header (first round of calculation) and tail (last round of calculation) of the data packet are identified based on the counter value.
[0140] The header data triggers the start signal, initiating the calculation process.
[0141] The tail data triggers the finish signal, indicating that the current data packet calculation is complete.
[0142] Pipeline-level coordination: After the last round of data enters the pipeline, the valid flag (valid_temp) is raised step by step with the finish signal it transmits, and finally the result valid signal (valid_out) is generated.
[0143] The lower-level module uses the ready signal to apply back pressure: if ready = 0, the result processing module (inversion and masking operations) is paused; however, multi-level operation units that do not handle the last round of data can still continue processing (to avoid global blocking).
[0144] The behavior of the unit is controlled in real time by the programmable register set: the polynomial width (crc_width) determines the bit width of the operation (4 / 8 / 16 / 32 bits). The initial value (CRC_INIT) provides the starting point for calculation. The input / output inversion enable (in_reversion / out_reversion) controls the data bit order adjustment. The data width field (data_width) defines the checksum truncation position (used when autopeel_en is enabled).
[0145] Interfaces and Collaboration:
[0146] Input: Streaming data access is achieved via a handshake signal (valid / ready) with the data buffer queue (D_FIFO). Real-time parameter updates to the configuration register group are received.
[0147] Output: The calculation result is passed to the storage or comparison module via the valid_out / ready handshake. The automatically extracted reference checksum is written to the checksum buffer queue (CRC_FIFO). The finish signal triggers a memory write to the address buffer queue (A_FIFO).
[0148] An address cache queue stores memory addresses and automatically triggers write operations after calculations are completed.
[0149] The address cache queue works as follows:
[0150] When receiving a storage instruction, the target address is cached and an execution signal is fed back immediately.
[0151] When the pipeline calculation unit completes the signal and the queue is not empty, a memory write request is automatically initiated.
[0152] Specifically, in this embodiment, the address cache queue is an address cache FIFO, used to store the addresses loaded into the module in the crc_store instruction. The execution of the crc_store instruction is divided into two parts. When the module receives the crc_store instruction, it stores the address data of rs1 into A_FIFO, and then immediately feeds back the instruction execution result through the instruction channel in the next cycle. At this time, the next instruction can continue to be executed. Subsequent memory access behavior will be executed automatically. That is, when there is address data in A_FIFO (i.e., the FIFO is not empty), when the crc_calculate calculation module issues a finish signal, the LSU channel issues a write response, and the calculation result of crc_calculate and the address in A_FIFO are simultaneously sent into the channel for memory writing. By separating the execution of instructions, the instruction processing changes from blocking to non-blocking, which will greatly improve the instruction processing efficiency.
[0153] A data buffer queue that receives a stream of data to be verified from memory or registers;
[0154] The data cache queue operates as follows:
[0155] Receive verification data from different sources through the data selection channel;
[0156] When the queue is not full, a signal indicating that the instruction execution is complete should be sent.
[0157] A continuous data stream is provided to the pipeline computing unit through a ready signal handshake mechanism.
[0158] In this embodiment, the data buffer FIFO has two data sources: data from the crc_rload instruction and data from the crc_load instruction. A data selector selects one data source to store in the FIFO. After receiving the crc_load or crc_rload instruction, the decoder identifies the instruction and selects either rs1 or rdata. When the FIFO is not full, the instruction execution result is immediately fed back to the instruction channel after data is stored. When the FIFO is full, the data waits until the FIFO is not full before being stored, and then the instruction execution result is fed back to the instruction channel. Subsequent CRC calculation is performed automatically. When there is data in D_FIFO, the valid signal will go high. After the crc_calculate module's ready signal goes high, the handshake is successful, and the FIFO write enable goes high. The data in D_FIFO enters the crc_calculate module for CRC calculation. When the FIFO is empty, the valid signal is low, and lower-level modules cannot obtain data.
[0159] A checksum cache queue stores reference checksums to be compared.
[0160] The working method of the check code cache queue in the automatic extraction enabled state is as follows:
[0161] The reference check code is extracted from the end of the data stream using the counting unit;
[0162] When the output of the calculation unit is valid, a checksum comparison is performed automatically.
[0163] In this embodiment, the CRC_FIFO is used to store the CRC checksum to be verified. When `autopeel_en` of the CRC_CFG is disabled, it is written by the `crc_cmp_r` / `crc_cmp_m` instructions, and then automatically compared. When the decode module detects the `crc_cmp_r` / `crc_cmp_m` instruction, the write enable of the CRC_FIFO goes high, the mux selects the data path, and when the FIFO is not full, the write operation is performed while simultaneously returning the instruction execution result. When the CRC_FIFO is not empty, and the valid signal of the output data from the `crc_calculate` module goes high, the CRC checksum is read and compared with the CRC checksum output by the `crc_calculate` module. If `autopeel_en` of the CRC_CFG is enabled, the CRC checksum will be automatically extracted from the data stream. This can be determined by the counter `cnt` module. The CRC checksum data is then sent into the CRC_FIFO. When the calculation is complete and the `valid_out` of the CRC calculation module is valid, the calculated checksum and the checksum in the CRC_FIFO are sent together into the check module for comparison.
[0164] The arbitration unit coordinates the priority of memory access requests from different instructions;
[0165] The priority method for the arbitration unit is as follows:
[0166] When the data cache queue is full, storage requests from the address cache queue are processed first.
[0167] In other scenarios, memory accesses that require data loading are prioritized.
[0168] The arbitration unit dynamically schedules memory access requests by monitoring system status signals in real time. When a crc_load instruction request exists, the module checks the empty / full status of the address cache queue (A_FIFO), the completion signal (finish) of the computation unit, and the capacity status of the data cache queue (D_FIFO), and executes the following arbitration logic:
[0169] If A_FIFO is empty, the memory read request of crc_load will be processed immediately;
[0170] If A_FIFO is not empty and the calculation completion signal is valid, the data is split according to the state of D_FIFO. When D_FIFO is not full, the crc_load instruction is executed first to ensure data supply. When D_FIFO is full, the data request is switched to A_FIFO to prevent buffer overflow.
[0171] If A_FIFO is not empty but computation is not complete (finish = 0), the crc_load instruction is always executed first to maintain the continuous operation of the computation unit. This arbitration mechanism is implemented through a priority decoder, which continuously collects the empty / not_empty signal and finish signal of A_FIFO and the full / not_full signal of D_FIFO, and generates a channel selection signal to drive the multiplexer according to the above rules. See details. Figure 4 The selector output is connected to the memory read channel of crc_load and the write channel of A_FIFO result, respectively.
[0172] When a load request is approved, memory reading is immediately initiated and data is written to D_FIFO; when a storage request is approved, the target address output by A_FIFO is bound to the verification result generated by the pipeline calculation unit and an asynchronous write operation is initiated.
[0173] The instruction response unit generates execution feedback based on the instruction type and module status.
[0174] The processing logic of the instruction response unit for the result write-back instruction is as follows:
[0175] If the calculation is not completed and the data cache queue is empty, an error response is returned;
[0176] If the calculation is complete, wait for the operation channel handshake to succeed before writing to the register.
[0177] The instruction response unit via Figure 6 The structure shown implements dynamic feedback of instruction execution results, and its behavior logic is strictly related to the instruction type and module state:
[0178] (1) When the decoding unit indicates a control register read / write instruction (crc_csrw / crc_csrr), since the register operation can be completed instantaneously, the module immediately feeds back the normal execution result to the processor in the next clock cycle (marking the end of the instruction), without waiting for additional conditions.
[0179] (2) When the decoding unit indicates a memory data load instruction (crc_load), the arbitration unit selection signal and memory handshake status need to be monitored synchronously: if the arbitration unit selects the instruction and the data buffer queue (D_FIFO) is not full (LSU channel handshake successful), the data is stored in D_FIFO and a normal result is immediately fed back; if D_FIFO is full, it continues to wait until the conditions are met.
[0180] (3) When the decoding unit indicates a storage instruction (crc_store) or a memory check code comparison instruction (crc_cmp_m), check the status of the address cache queue (A_FIFO): if A_FIFO is not empty (address write successful), return a normal response immediately; otherwise, block and wait.
[0181] (4) When the decoding unit indicates a register data loading instruction (crc_rload), the D_FIFO status is directly checked: if the D_FIFO is not fully loaded (data is successfully written), the normal execution result is immediately fed back.
[0182] (5) When the decoding unit indicates the result write-back instruction (crc_rstore), perform two-level judgment: if D_FIFO is empty and the calculation completion signal (finish) is invalid, immediately feed back an error response; if finish is valid, wait for the handshake to succeed and then perform data write-back and return the normal result.
[0183] (6) When the decoding unit indicates a register check code comparison instruction (crc_cmp_r), check the capacity of the check code buffer queue (CRC_FIFO): if the queue is not full, write the reference check code and immediately provide a normal response.
[0184] It should be noted that the cyclic redundancy check hardware acceleration module integrated into the processor provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.
[0185] A second embodiment of the present invention proposes a verification acceleration method based on a cyclic redundancy check (CRC) hardware acceleration module integrated into a processor. The method, characterized by the following features:
[0186] Configure polynomial parameters, initial values, and data volume parameters using control commands;
[0187] The data stream is written to the data buffer queue using a load instruction;
[0188] The pipeline calculation unit automatically extracts data and performs pipeline calculations.
[0189] The target address is pre-stored by a storage instruction, and the calculation result is automatically written into memory when it is ready.
[0190] By comparing the pre-stored reference check code with the comparison command, the system automatically performs a data integrity check when the calculation result is ready.
[0191] Although the steps in the above embodiments are described in the above order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not need to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple variations are all within the protection scope of this invention.
[0192] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related explanations of the methods described above can be found in the corresponding processes in the aforementioned module embodiments, and will not be repeated here.
[0193] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0194] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0195] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A hardware acceleration module for cyclic redundancy check integrated into a processor, characterized in that, This module includes: The instruction extension unit implements multiple sets of custom instructions through the coprocessor interface. The format of the custom instructions includes a function code field that controls the reading and writing of processor architecture registers. The programmable register set includes a configuration register, a polynomial register, an initial value register, and an XOR value register. The configuration register supports dynamically setting the module enable state, automatic check code extraction function, polynomial bit width selection, input / output data inversion enable, and the amount of data corresponding to a single check code. The pipelined computing unit includes a checksum calculation module with a multi-level processing structure. It performs finite field division operations according to the programmable register group configuration and controls the calculation process through a counting unit. An address cache queue stores memory addresses and automatically triggers write operations after calculations are completed. A data buffer queue that receives a stream of data to be verified from memory or registers; A checksum cache queue stores reference checksums to be compared. The arbitration unit coordinates the priority of memory access requests from different instructions; The instruction response unit generates execution feedback based on the instruction type and module status.
2. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The custom instructions include control register read / write instructions, memory data loading instructions, register data loading instructions, result storage instructions, and checksum comparison instructions; When the result storage instruction is executed, the address cache queue immediately stores the target address and returns a response signal, while the calculation result is asynchronously written to memory.
3. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, Each register in the programmable register group is configured as follows: The polynomial register is used to store programmable polynomial coefficients; The initial value register is used to define the starting value for calculation; The XOR value register is used to configure the result mask value; The data quantity field of the configuration register defines the number of data bytes corresponding to a single checksum.
4. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The address cache queue works as follows: When receiving a storage instruction, the target address is cached and an execution signal is fed back immediately. When the pipeline calculation unit completes the signal and the queue is not empty, a memory write request is automatically initiated.
5. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The data cache queue operates as follows: Receive verification data from different sources through the data selection channel; When the queue is not full, a signal indicating that the instruction execution is complete should be sent. A continuous data stream is provided to the pipeline computing unit through a ready signal handshake mechanism.
6. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The working method of the check code cache queue in the automatic extraction enabled state is as follows: The reference check code is extracted from the end of the data stream using the counting unit; When the output of the calculation unit is valid, a checksum comparison is performed automatically.
7. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The priority method for the arbitration unit is as follows: When the data cache queue is full, storage requests from the address cache queue are processed first. In other scenarios, memory accesses that require data loading are prioritized.
8. The cyclic redundancy check hardware acceleration module integrated into a processor according to claim 1, characterized in that, The flow calculation unit includes: The data selector selects the initial value or intermediate result input based on the start signal; Multi-level arithmetic units perform polynomial division and pass intermediate results level by level; The result processing module performs bit order adjustment and masking operations on the final result according to the configuration. The checksum extraction module extracts the end segment of the data stream based on the counting signal as a reference checksum.
9. A hardware acceleration module for cyclic redundancy check integrated into a processor according to claim 1, characterized in that, The processing logic of the instruction response unit for the result write-back instruction is as follows: If the calculation is not completed and the data cache queue is empty, an error response is returned; If the calculation is complete, wait for the operation channel handshake to succeed before writing to the register.
10. A method for accelerating cyclic redundancy check (CRC) hardware acceleration modules integrated into a processor, based on the CRC hardware acceleration module integrated into a processor as described in any one of claims 1-9, characterized in that, The method includes: Configure polynomial parameters, initial values, and data volume parameters using control commands; The data stream is written to the data buffer queue using a load instruction; The pipeline calculation unit automatically extracts data and performs pipeline calculations. The target address is pre-stored by a storage instruction, and the calculation result is automatically written into memory when it is ready. By comparing the pre-stored reference check code with the comparison command, the system automatically performs a data integrity check when the calculation result is ready.
Citation Information
Cited By
A data detection system and method
CN122493919A