5GNR air interface protocol stack data plane acceleration method
Patent Information
- Application Number
- CN202280102174.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-08-29
AI Technical Summary
In existing technologies, the data plane processing latency of the 5G NR air interface protocol stack is relatively long, which affects the correctness and performance of user plane service processing.
By caching data after physical layer processing and using the Data Path Accelerator (DPHA) for interruption notification, pre-parses header fields of each layer (MAC, RLC, PDCP, and SDAP), and configures ROHC and SEC accelerators at the PDCP layer, the data processing flow is optimized and latency is reduced.
It significantly reduces downlink user plane latency of 5G baseband processors, enhances processor reliability and performance, and ensures the correctness of data processing.
Smart Images

Figure CN120569906A_ABST
Abstract
Description
A 5GNR air interface protocol stack data plane acceleration method Technical Field
[0001] The present invention relates to the field of hardware accelerators, and in particular to a 5GNR air interface protocol stack data plane acceleration method. Background Art
[0002] How can we ensure that the involvement of hardware accelerators does not have side effects on normal uplink and downlink user plane service processing, that is, how can we ensure the correctness of service processes? Traditional software practices involve first demodulating the PDSCH channel data at the physical layer, performing a CRC check on the LDPC code block, and then caching the MAC TB in a designated DDR shared memory. The MAC downlink task then initiates the demultiplexing process. The MAC Entity extracts the MACsubPDUs on all logical channels. The MAC layer processes all MACCEs, and all RLC data (DCCH PDUs) are delivered to the RLC Entities. The RLC entities decode all downlink data, perform window boundary checks, and check for abnormal and redundant data. Complete SDUs that meet the requirements are then delivered to the PDCP layer for further processing. The RLC temporarily caches segmented data in a local buffer. Once all segmented data is collected, it is delivered to the PDCP for sorting and aggregation. In traditional protocol stack software, the uplink MAC layer appends a subPDU header, and the RLC layer appends an RLC header only when the UE receives an uplink authorization from the base station. This inevitably increases user plane processing latency and reduces processing performance.
[0003] Summary of the Invention
[0004] The purpose of the present invention is to provide a 5GNR air interface protocol stack data plane acceleration method, which can reduce the latency of existing products, enhance reliability, improve performance, and optimize the accelerator through the 5GNR air interface protocol stack data plane design, while ensuring the correctness of data processing and making corresponding adjustments to the software architecture.
[0005] The embodiment of the present invention is achieved as follows:
[0006] The embodiment of the present application provides a 5GNR air interface protocol stack data plane acceleration method, which includes the following steps: after the physical layer processes the downlink shared channel PDSCH data, the L1C controls the data flow to be cached in the MAC / PHY shared memory and notifies the PS core through an interrupt;
[0007] S1-0. The CPU initializes the Data Path Accelerator (DPHA) and configures the following parameters: a. TB address and size, b. packet address after de-ciphering, c. L2 header fields buffer address, and d. de-ciphering algorithm. The CPU then starts the Data Path Accelerator (DPHA).
[0008] S1-1, DPHA first parses the MAC, RLC, PDCP and SDAP header fields from the shared memory ShMem and stores them in PS SRAM;
[0009] S1-2. The DPHA mac parser parses the MAC header, extracts a MAC subPDU, and caches the MAC header fields. If it is a MAC CE, it outputs the MAC CE Data and ends the PDU processing. If it is a MAC SDU, it triggers the RLC header parser.
[0010] S1-3. DPHA RLC header parser parses the RLC header and first caches the RLC header fields: if an RLC segment is received, DPHA outputs the segment address information and ends the PDU processing; if a control PDU is received, DPHA outputs the control PDU field and ends the PDU processing; if a complete RLC PDU is received, DPHA triggers the PDCP header parser;
[0011] S1-4. The DPHA PDCP header parser parses the PDCP header and first caches the PDCP header fields. If it is a PDCP control PDU, the DPHA outputs the PDU address and length information and ends the PDU processing. If it is a PDCP data PDU, the SEC accelerator context is configured and the SEC is started. The SEC accelerator reads the payload to be decrypted from the ShMem via DMA, performs the decryption operation, and writes it to the destination address configured by the PS, then triggers the SDAP processing.
[0012] S1-5. DPHA SDAP header parser parses the SDAP header and caches the SDAP header fields. If it is an SDAP control PDU, it parses the corresponding fields and ends the PDU processing. If it is an SDAP data PDU, it outputs the data length and address information, ends the processing of this MAC subPDU, jumps to step 2, and starts processing the next MAC subPDU.
[0013] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: the data parsed by DPHA is configured with multiple HFD data formats according to different logical channels to form MAC_CE_GROUP and MAC_SDU_GROUP, the former containing all MAC CEs in a TB block, and the latter containing header fileds of all MAC SDUs in a TB.
[0014] In some embodiments of the present invention, the HFD data formats include HFD_MAC_CE, HFD_AM_CTRL, HFD_AM_PDCP_CTRL, HFD_AM_PDCP_DATA, HFD_AM_PDCP_SDAP_DATA, HFD_AM_SEG_DATA, HFD_UM_PDCP_DATA, HFD_UM_PDCP_SDAP_DATA, HFD_UM_SEG_DATA, and HFD_UM_PDCP_CTRL.
[0015] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: for the hierarchical management of TB blocks, the TB block memory is allocated by L1C and released by PS; the protocol stack and PHY jointly maintain an idle TB block linked list, and before starting DPHA each time, the protocol stack configures the address of the TB block received in this TTI to the register of DPHA.
[0016] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: the backup solution of the TB block memory is allocated by the PS and notified to the PHY through the mac-phy Mailbox mechanism interaction message.
[0017] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: DPHA includes configuration registers for multiple logical channels, and the multiple configuration registers are used to indicate whether the DPHA is in idle state or running state, whether the HFD buffer overflows, MAC CE / SDU HFD buffer base address register, and HFD buffer configuration register.
[0018] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps, and the DPHA accelerator is designed to include: DHPA_LCx_CFG_REG, DHPA_LCx_PDCPRXDELIV_REG, DHPA_CTL_REDHPA_ITR_CTDHPA_STATUS_RE, DHPA_CE_HFD_BUF_BASE_ADDR_RE, DHPA_MAC_SDU_HFD_BUF_BASE_ADDR_REG, DHPA_MAC_SDU_HFD_CFG_REG, DHPA_SEC_CTX_BUF_ADDR_REG, DHPA_SEC_CTX_BUF_CFG_REG and DHPA_TB_BUFFER_ADDR_REG multiple registers.
[0019] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: the shared memory allocation principles include static allocation, dynamic allocation by PHY, memory recovery, hierarchical management and Link-list read head and write tail; each memory block corresponds to a Desc descriptor, and all Desc descriptors are stored in SRAM in the form of a Desc Chain. During the system initialization phase, the protocol stack creates the Desc Chain; the allocation and release of memory are achieved by maintaining the two registers DHPAFreeDescAddrReg and DHPATailDescAddrReg.
[0020] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps, and the memory allocation process includes: Allocation 1: The DLTB block pool is initialized by the PS, and the block pool is managed in the form of a TB DESC chain; Allocation 2: After the initialization is completed, the PS notifies the PHY through a MAC-PHY event message, which contains at least the first address and the total number of blocks of the DL TB block pool; Allocation 3: Before a HARQ entity of the PHY starts to receive TB data, it reads an idle TB address from the FreeDescAddrReg field of the TB DESC chain, and updates the FreeDescAddrReg to the address of the next idle TB; Allocation 4: When the PHY correctly receives the TB data of the HARQ entity, it notifies the PS through a PHY-MAC message; Allocation 5: If the PHY cannot correctly receive the TB of the HARQ entity, the PHY notifies the protocol stack through a PHY-MAC message, and the message carries the TB where the TB is located. DESC address; Allocation 6: If the PS receives a PHY message, it updates the statistical information and directly updates TailDescAddrReg to reclaim the TB; Allocation 7: If Allocation 5 occurs, the PS receives a PHY message and performs DL data processing. After all MAC SDUs in the TB are delivered to the AP, TailDescAddrReg is updated to reclaim the TB.
[0021] In some embodiments of the present invention, when DPHA parses the TB, for each processed MAC subPDU: if it is a MAC CE, the MAC CE HFD is output to the SRAM; if it is an RLC segment or a PDCP status report, the MAC SDU HFD is output to the SRAM, and the HFD contains the absolute address of the RLC segment and the PDCP payload in the source TB; if it is a complete RLC data packet, and the logical channel to which the data packet belongs is configured with decryption or security, a SEC context is configured for it, and the SEC context is stored in the SRAM space, and the decrypted address is assigned to the packet.
[0022] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0023] The present application provides a 5GNR air interface protocol stack data plane acceleration method, which implements integrity verification and decryption activation based on header decompression of downlink data, and the PDCP layer is responsible for configuring and starting the downlink ROHC, and the SEC accelerator executes them respectively. The ROHC and SEC accelerators feed back the corresponding header decompression and integrity verification results to the PDCP. The PDCP layer submits all successfully verified data to the AP application processor and discards all data that fails the test. After adding the DPHA accelerator, the above scheme changes the downlink processing method of traditional software to a certain extent, reduces the data plane delay of existing products, and the physical layer demodulates the PDSCH channel. After the LDPC CRC code block check is completed, the MAC TB is cached in the specified DDR shared memory, and the DPHA accelerator is activated to perform downlink header parse. At the same time, the timing of attaching the MAC subPDU header and RLCPDU header is advanced, and there is no need to wait for the uplink authorization on the base station side. After each new PDCP data arrives, the MAC subPDU and RLCPDU header can be attached according to the complete SDU. An optimization technical solution for pre-grouping headers is proposed to optimize the uplink user plane time overhead, and the software architecture and code flow are modified and adjusted to a certain extent to enhance reliability and improve performance. The present invention reduces the latency of existing products, enhances reliability, improves performance, and optimizes the accelerator through the design of the protocol stack data path. At the same time, it ensures the correctness of data processing and makes corresponding adjustments and optimizations to the software architecture. The DPHA mac parser parses the MAC header, extracts the MAC subPDU for processing, and saves the processing results in the corresponding HFD data format. The PS subsequently performs protocol specification processing on various HFD data. The DPHA accelerator significantly reduces the downlink user plane latency of the existing 5G baseband processor, enhances the reliability of the baseband processor, and improves its performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only represent certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0025] FIG1 is a schematic diagram of a method for accelerating the data plane of a 5GNR air interface protocol stack according to an embodiment of the present invention;
[0026] FIG2 is a flow chart of a method for accelerating the data plane of a 5GNR air interface protocol stack according to an embodiment of the present invention;
[0027] FIG3 is a flow chart of DPHA implementing NR L2 protocol header parsing at each layer according to an embodiment of the present invention;
[0028] FIG4 is a schematic diagram of a memory initialization process according to an embodiment of the present invention;
[0029] FIG5 is a schematic diagram of a specific memory allocation process according to an embodiment of the present invention;
[0030] FIG6 is a schematic diagram of a specific memory release process according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0032] Example
[0033] As shown in Figures 1 and 2, the embodiment of the present application provides a 5GNR air interface protocol stack data plane acceleration method, which includes the following steps: after the physical layer processes the downlink shared channel PDSCH data, the L1C controls the data flow to be cached in the MAC / PHY shared memory and notifies the PS core through an interrupt; S1-0, the CPU initializes the Data Path accelerator and configures the following parameters: a. TB address and size, b. data packet address after de-ciphering, c. L2 header fields buffer address, d. de-ciphering algorithm used; then starts the Data Path accelerator (DPHA); S1-1, DPHA first parses the MAC, RLC, PDCP and SDAP layer header fields from the shared memory ShMem and stores them in the PS SRAM; S1-2, DPHA mac parser parses the MAC header and takes out a MAC subPDU, caches the MAC header fields: if it is a MAC CE, outputs MAC CE Data and ends the PDU processing; if it is a MAC SDU, triggers the RLC header parser; S1-3, DPHA RLC header The parser parses the RLC header and caches the RLC header fileds first: if an RLCsegment is received, the DPHA outputs the segment address information and ends the PDU processing; if a control PDU is received, the DPHA outputs the control PDU field and ends the PDU processing; if a complete RLC PDU is received, the PDCP header parser is triggered; S1-4, the DPHA PDCP header parser parses the PDCP header and caches the PDCP header fileds first: if it is a PDCP control PDU, the DPHA outputs the PDU address and length information and ends the PDU processing; if it is a PDCP data PDU, the SEC accelerator context is configured and the SEC is started. The SEC accelerator reads the Payload to be decrypted from the ShMem through DMA, performs the decryption operation and writes it to the destination address configured by the PS, and then triggers the SDAP processing; S1-5, the DPHA SDAP header parser parses the SDAP header and caches the SDAP header fileds: if it is an SDAP control PDU, parses the corresponding field and ends the PDU processing; if it is an SDAP data PDU, outputs the data length and address information and ends this MAC SubPDU processing, jump to step 2, and start processing the next MAC subPDU.
[0034] In some embodiments of the present invention, the above-mentioned 5G air interface protocol stack data plane acceleration method further includes the following steps: the data parsed by DHPA is configured with multiple HFD data formats according to different logical channels to form MAC_CE_GROUP and MAC_SDU_GROUP, the former contains all MAC CEs in a TB block, and the latter contains header fileds of all MAC SDUs in a TB.
[0035] In some embodiments of the present invention, the HFD data formats include HFD_MAC_CE, HFD_AM_CTRL, HFD_AM_PDCP_CTRL, HFD_AM_PDCP_DATA, HFD_AM_PDCP_SDAP_DATA, HFD_AM_SEG_DATA, HFD_UM_PDCP_DATA, HFD_UM_PDCP_SDAP_DATA, HFD_UM_SEG_DATA, and HFD_UM_PDCP_CTRL.
[0036] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: for the hierarchical management of TB blocks, the TB block memory is allocated by L1C and released by PS; the protocol stack and PHY jointly maintain an idle TB block linked list, and before starting DPHA each time, the protocol stack configures the address of the TB block received in this TTI to the register of DPHA.
[0037] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: the backup solution of the TB block memory is allocated by the PS and notified to the PHY through the mac-phy Mailbox mechanism interaction message.
[0038] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: DPHA includes configuration registers for multiple logical channels, and the multiple configuration registers are used to indicate whether the DPHA is in idle state or running state, whether the HFD buffer overflows, MAC CE / SDU HFD buffer base address register, and HFD buffer configuration register.
[0039] In some embodiments of the present invention, the above-mentioned 5G air interface protocol stack data plane acceleration method further includes the following steps, and the DPHA accelerator is designed to include: DHPA_LCx_CFG_REG, DHPA_LCx_PDCPRXDELIV_REG, DHPA_CTL_REDHPA_ITR_CTDHPA_STATUS_RE, DHPA_CE_HFD_BUF_BASE_ADDR_RE, DHPA_MAC_SDU_HFD_BUF_BASE_ADDR_REG, DHPA_MAC_SDU_HFD_CFG_REG, DHPA_SEC_CTX_BUF_ADDR_REG, DHPA_SEC_CTX_BUF_CFG_REG and DHPA_TB_BUFFER_ADDR_REG multiple registers.
[0040] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps: the shared memory allocation principles include static allocation, dynamic allocation by PHY, memory recovery, hierarchical management and Link-list read head and write tail; each memory block corresponds to a Desc descriptor, and all Desc descriptors are stored in SRAM in the form of a Desc Chain. During the system initialization phase, the protocol stack creates the Desc Chain; the allocation and release of memory are achieved by maintaining the two registers DHPAFreeDescAddrReg and DHPATailDescAddrReg.
[0041] In some embodiments of the present invention, the above-mentioned 5GNR air interface protocol stack data plane acceleration method further includes the following steps, and the memory allocation process includes: Allocation 1: The DLTB block pool is initialized by the PS, and the block pool is managed in the form of a TB DESC chain; Allocation 2: After the initialization is completed, the PS notifies the PHY through a MAC-PHY event message, which contains at least the first address and the total number of blocks of the DL TB block pool; Allocation 3: Before a HARQ entity of the PHY starts to receive TB data, it reads an idle TB address from the FreeDescAddrReg field of the TB DESC chain, and updates the FreeDescAddrReg to the address of the next idle TB; Allocation 4: When the PHY correctly receives the TB data of the HARQ entity, it notifies the PS through a PHY-MAC message; Allocation 5: If the PHY cannot correctly receive the TB of the HARQ entity, the PHY notifies the protocol stack through a PHY-MAC message, and the message carries the TB where the TB is located. DESC address; Allocation 6: If the PS receives a PHY message, it updates the statistical information and directly updates TailDescAddrReg to reclaim the TB; Allocation 7: If Allocation 5 occurs, the PS receives the PHY message and performs DL data processing. After all MAC SDUs in the TB are delivered to the AP, TailDescAddrReg is updated to reclaim the TB.
[0042] In some embodiments of the present invention, when DPHA parses the TB, for each processed MAC subPDU: if it is a MAC CE, the MAC CE HFD is output to the SRAM; if it is an RLC segment or a PDCP status report, the MAC SDU HFD is output to the SRAM, and the HFD contains the absolute address of the RLC segment and the PDCP payload in the source TB; if it is a complete RLC data packet, and the logical channel to which the data packet belongs is configured with decryption or security, a SEC context is configured for it, and the SEC context is stored in the SRAM space, and the decrypted address is assigned to the packet.
[0043] The user plane of the protocol stack consists of four layers: SDAP, PDCP, RLC, and MAC. The SDAP layer allocates the corresponding service channel (for DRBs) based on QoS mapping, ultimately mapping it to the logical channel. The PDCP layer performs header compression and decompression, encryption and decryption, and sorting. The RLC layer performs necessary segmentation based on the size indicated by the UL grant. The MAC concatenates and multiplexes all logical channel data RLC PDUs, providing TB blocks of the specified size to the physical layer.
[0044] As shown in Figure 3, the Downlink Packet High-Speed Accelerator (DHPA) implements NR L2 protocol header parsing for each layer. In eMBB scenarios, user plane data throughput reaches up to 2 Gbps, which consumes significant CPU resources. The current SoC architecture design leverages the advantages of AISC by offloading the L2 header parsing function, a relatively fixed process with simple processing logic, into the AISC. This reduces CPU load and power consumption while maintaining system throughput. Specifically, the Downlink Packet High-Speed Accelerator (DPHA) centrally parses the headers of each sub-PDU within the air interface transport block (TB) and stores the results in an internally defined data structure. This avoids the CPU from performing extensive bit-level unpacking operations. Furthermore, the MAC CE is located at the front of the TB. After completing MAC CE parsing, the DPHA immediately notifies the MAC task, enabling it to start earlier and reducing the latency of its response to the MAC CE. After parsing the RLC / PDCP / SDAP headers, the DPHA immediately notifies the protocol stack (RLC task) to perform logical processing of the header fields at each layer and simultaneously initiates the SEC for data decryption. In addition, for scenarios with fewer or even no MAC CEs, DPHA can also be configured to notify the PS once after all headers at each layer are parsed.
[0045] Data communication between the HW and SW uses shared content and registers, while message notifications are transmitted via interrupts. The protocol stack starts the DPHA by setting the corresponding bits in the configuration register. Before starting, the protocol stack configures the DPHA configuration register. The detailed process after DPHA startup is described in the next section. After the DPHA completes header parsing, it triggers the corresponding interrupt and sets the corresponding status register.
[0046] In step 2, MAC header parser: DPHA mac parser parses the MAC header to extract a MAC subPDU and caches the MAC header fileds. If it is a MAC CE, MAC CE Data is output and the PDU processing is ended. After all MAC CE parsing is completed, if MAC CE is completed (MAC_CE_INFORM), the PS MAC task is interrupted and notified. If it is a MAC SDU, the RLC header parser is triggered. In step 3, RLCheader parser: DPHA RLC header parser parses the RLC header and caches the RLC header fileds first. In step 4, PDCP header parser: DPHA PDCP header parser parses the PDCP header and caches the PDCP header fileds first. Among them, if it is a PDCP data PDU, the SEC accelerator context is configured and SEC is started. The SEC accelerator reads the Payload to be decrypted from ShMem through DMA, performs the decryption operation and writes it to the destination address configured by PS (DDR or SRAM), and then triggers SDAP processing. Step 5, SDAP header parser: DPHA SDAP header parser parses the SDAP header and caches the SDAP header fields.
[0047] In DHPA_LCx_CFG_REG, DHPA contains 32 logical channel configuration registers, which store the configuration information of logical channels 1 to 32. The 32 logical channel configuration registers are specifically represented as DHPA_LC1_CFG_REG, DHPA_LC2_CFG_REG…DHPA_LC32_CFG_REG, as shown in the following table:
[0048]
[0049]
[0050]
[0051] In DHPA_LCx_PDCPRXDELIV_REG, DHPA contains 32 PDCPRXDELIV registers, specifically DHPA_LC1_PDCPRXDELIV_REG, DHPA_LC2_PDCPRXDELIV_REG…DHPA_LC32_PDCPRXDELIV_REG. The 32 PDCPRXDELIV registers are used to store the PDCPRXDELIV values of each LC. PDCPRXDELIV is the COUNT value of the next data packet to be delivered to each logical PDCP entity. For example, if PDCP has delivered packets with a COUNT value of N, PDCPRXDELIV should be configured as N+1. Before starting DHPA, the protocol stack ensures that the value of the PDCPRXDELIV register group has been updated with the latest PDCPRXDELIV value in each PDCP logical channel context. The update timing is as follows: after PDCP delivers the data packet to SDAP, the value of this register group is updated synchronously when the RB (LCID) context is updated.
[0052] DHPA_CTL_REG is represented as:
[0053] FieldR / WBitsDescriptionNoteReserve 31-2 DPA_STOPW1PS set this bit to stop DPA. DPA_STARTW0PS set this bit to start DPA.
[0054] DHPA_ITR_CTL is represented as:
[0055]
[0056]
[0057] DHPA_STATUS_REG is represented as:
[0058]
[0059] In the DHPA_CE_HFD_BUF_BASE_ADDR_REG, the DHPA generates a header field descriptor (HFD) when it parses a MAC subPDU from a TB. The HFD output storage area is planned based on the maximum length of each HFD. The protocol stack configures a starting SRAM buffer address (HFD buffer) for each MAC CE HFD and MAC SDU HFD. The DHPA stores the generated HFDs in the HFD buffers sequentially when parsing the TB. The parsing process includes the MAC CE HFD output buffer register, the MAC CE HFD config register, the MAC SDU HFD output buffer register, and the MAC SDU HFD config register.
[0060] In the DHPA_MAC_SDU_HFD_BUF_BASE_ADDR_REG, the MAC SDU HFD buffer is managed similarly to the MAC CE HFD. The CPU is also responsible for initialization and the MAC SDU HFD buffer is reused every TTI. The number of HFDs in the MAC SDU class depends on the payload length of each MAC SDU. The HFD buffer length is estimated based on the average SDU size.
[0061] 160kbytes / 400 (bytes)*20≈8kbytes, where 160kbytes is the maximum TB size (bytes) per TTI when the system operates at a 2Gbps downlink rate. Small packets require a larger HFD buffer. Considering that control PDUs at each layer may vary in length, assuming each HFD is 32 bytes long and using a 64-byte packet as an example for evaluation, the HFD buffer size is calculated as follows: the number of packets per TB is 160kbytes / 64 (bytes) = 2560 packets.
[0062] HFD buffer size: 2560 * 20 bytes = 51200 = 50k. Considering that the packet length is usually large at maximum throughput, 16kBytes of this address space is reserved.
[0063] DHPA_MAC_SDU_HFD_BUF_BASE_ADDR_REG is represented as:
[0064]
[0065] DHPA_MAC_SDU_HFD_CFG_REG is represented as:
[0066]
[0067]
[0068] DHPA_SEC_CTX_BUF_ADDR_REG is represented as:
[0069]
[0070] DHPA_SEC_CTX_BUF_CFG_REG is represented as:
[0071]
[0072] DHPA_TB_BUFFER_ADDR_REG, this register points to the address space of the TB received by the PS in the current TTI. The PS must configure this register before starting DHPA. DHPA reads data from the address space indicated by this register to perform header parsing of MAC CE and MAC SDU, which is expressed as:
[0073]
[0074] The memory initialization process is shown in Figure 4. In the DL Memory Desc defination allocation principle, static allocation: the protocol stack statically initializes the memory space before DHPA is started. The memory space takes into account the maximum TB size of each slot and the maximum RTT that the protocol stack can cache. Memory recycling: For the data of the SRB logical channel, after the data is delivered by SDAP to RRC, the RRC notifies to mark the recycling status of the SDU. For DRB data, after SDAP successfully delivers the data to the AP, the recycling status of the SDU is marked. Each time the recycling status of an SDU is marked, it is only necessary to reduce the remain sdu num in the TB mem block (after decryption) where it is located by one. When it reaches 0, it indicates that the TB memblock can be recycled. Hierarchical management: To improve memory allocation efficiency and reduce bus access, space is allocated in SRAM for mapping actual DDR addresses. Link-list read and write: When allocating memory, the head pointer register, FreeDescAddrReg, is read to obtain a free block desc. When releasing memory, the block desc pointed to by TailDescAddrReg is updated and the next field of the block desc is set to 0. Each memory block corresponds to a desc descriptor, and all desc descriptors are stored in SRAM as a desc chain. During system initialization, the protocol stack (CPU) creates the desc chain.
[0075] Memory allocation and release are implemented by maintaining two registers, DHPAFreeDescAddrReg and DHPATailDescAddrReg. The memory allocation process is shown in Figure 5, and the memory release process is shown in Figure 6. In the DL MAC-PHY TB Memory Scheme allocation process, the PS is responsible for initializing the DL TB block pool, which is managed as a TB DESC chain. The TB DESC chain is mapped to an SRAM area shared by the PS and PHY. After initialization, the PS notifies the PHY via a MAC-PHY message. This message includes at least the starting address and total number of DL TB block pool blocks. The specific format will be jointly determined by the PS and PHY. The starting address of the DL TB block pool is represented by FreeDescAddrReg. Before a PHY HARQ entity begins receiving TB data, it reads a free TB address from the FreeDescAddrReg field in the TB DESC chain and updates FreeDescAddrReg with the address of the next free TB. When the PHY correctly receives (CRC succeeds) the TB data of the HARQ entity, it notifies the PS through a PHY-MAC message, which carries the TB DESC address of the TB. This message reuses the currently defined PHY_MAC_NR_PDSCH_IND message. The difference is that the data address information needs to be modified to the address of TB DESC. It can also be compatible with the previous type and add an additional TB DESC address field. If the PHY cannot correctly receive the TB process of the HARQ entity, for example, all RV versions cannot be correctly parsed after being merged, the PHY notifies the protocol stack through a PHY-MAC message, which carries the TB DESC address of the TB. The specific format of the message is to be determined jointly by PS / phy. After receiving the PHY message, the PS updates the statistical information and directly updates TailDescAddrReg to reclaim the TB. If the PHY correctly receives the TB data of the HARQ entity, it notifies the PS through a PHY-MAC message, and then performs the DL data processing process (DHPA-SEC-rlc / pdcp / sdap). After delivering all MAC SDUs (including segments) in the TB to the AP, the TailDescAddrReg is updated to reclaim the TB. In principle, this memory management mechanism can be used for all types of data on the PDSCH channel.
[0076] When DHPA parses TB, for each processed MAC subPDU: if it is a MAC CE, it outputs the MAC CE HFD to SRAM. The starting address configuration is shown in the following table DHPA_CTL_REG:
[0077] FieldR / WBitsDescriptionReserve 31-2 DPA STOPW1PS set this bit to stop DPA.DPA STARTW0PS set this bit to start DPA.
[0078] If it is an RLC segment or PDCP status report, the MAC SDU HFD is output to SRAM. The HFD contains the absolute address of the RLC segment and PDCP payload in the source TB. If it is a complete RLC data packet, and the logical channel to which the data packet belongs is configured for decryption or full protection, the SEC context (MCP) is configured for it. The SEC context is stored in the SRAM space. The configuration register is shown in DHPA_CE_HFD_BUF_BASE_ADDR_REG. The starting address of the HFD buffer memory area of the MAC CE is initialized by the CPU and reused for each TTI. Each time the DPHA is started, it starts from the starting address and saves the parsed MAC CE HFD in the HFD buffer in sequence, as shown in the following table:
[0079]
[0080] , and at the same time assign the decrypted address to the packet (this address is the original address of the data, that is, decryption in situ).
[0081] TB block memory can be allocated by L1C and released by PS (the backup solution is that PS allocates it and notifies PHY through the mac-phy Mailbox mechanism). The protocol stack and PHY jointly maintain a free TB block list. Before starting DHPA each time, the protocol stack configures the address of the TB block received in the current TTI to the DHPA register DHPA_TB_BUFFER_ADDR_REG.
[0082] The terms and abbreviations used in this document are shown in the following table:
[0083] English abbreviation English full name PS Protocol stack L2 Layer 2 (sdap\pdcp\rlc\mac) DMA Direct Memory Access DPA Data path accelerator DPHA Downlink Packet High-speed accelerator
[0084] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0085] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0086] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0087] In summary, the embodiment of the present application provides a 5GNR air interface protocol stack data plane acceleration method, which implements integrity verification and decryption activation based on header decompression of downlink data, and the PDCP layer is responsible for configuring and starting the downlink ROHC, and the SEC accelerator executes them respectively. The ROHC and SEC accelerators feed back the corresponding header decompression and integrity verification results to the PDCP. The PDCP layer submits all successfully verified data to the AP application processor and discards all data that fails the verification. After adding the DPHA accelerator, the above scheme changes the downlink processing method of traditional software to a certain extent, reduces the data plane delay of existing products, and the physical layer demodulates the PDSCH channel data. After the LDPC CRC code block check is completed, the MAC TB is cached in the specified DDR shared memory, and the DPHA accelerator is activated to perform downlink header parse. At the same time, the timing of attaching the MAC subPDU header and RLC PDU header is advanced, and there is no need to wait for the uplink authorization on the base station side. After each new PDCP data arrives, the MAC subPDU and RLC PDU header can be attached according to the complete SDU. An optimization technical solution for pre-grouping the header is proposed to optimize the uplink user plane time overhead, and the software architecture and code flow are modified and adjusted to a certain extent to enhance reliability and improve performance. The present invention reduces the latency of existing products, enhances reliability, improves performance, and optimizes the accelerator through the design of the protocol stack data path. At the same time, it ensures the correctness of data processing and makes corresponding adjustments and optimizations to the software architecture. The DPHA mac parser parses the MAC header, extracts the MAC subPDU for processing, and saves the processing results in the corresponding HFD data format. The PS subsequently performs protocol specification processing on various HFD data. The DPHA accelerator significantly reduces the downlink user plane latency of the existing 5G baseband processor, enhances the reliability of the baseband processor, and improves its performance.
[0088] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A 5GNR air interface protocol stack data plane acceleration method, characterized in that: The steps include: After the physical layer processes the downlink shared channel PDSCH data, the L1C controls the data flow and caches it in the MAC / PHY shared memory, and notifies the PS core through an interrupt. S1-0. The CPU initializes the Data Path Accelerator (DPHA) and configures the following parameters: a. TB address and size, b. packet address after de-ciphering, c. L2 header fields buffer address, and d. de-ciphering algorithm. The CPU then starts the Data Path Accelerator (DPHA). S1-1, DPHA first parses the MAC, RLC, PDCP and SDAP header fields from the shared memory ShMem and stores them in PS SRAM; S1-2. The DPHA mac parser parses the MAC header, extracts a MAC subPDU, and caches the MAC header fields. If it is a MAC CE, it outputs the MAC CE Data and ends the PDU processing. If it is a MAC SDU, it triggers the RLC header parser. S1-3, DPHA RLC header parser parses the RLC header and first caches the RLC header fields: if an RLC-segment is received, DPHA outputs the segment address information and ends the PDU processing; If a control PDU is received, DHPA outputs the control PDU field and ends the PDU processing; if a complete RLC PDU is received, the PDCP header parser is triggered; S1-4. The DPHA PDCP header parser parses the PDCP header and first caches the PDCP header fields. If it is a PDCP control PDU, the DPHA outputs the PDU address and length information and ends the PDU processing. If it is a PDCP data PDU, the SEC accelerator context is configured and the SEC is started. The SEC accelerator reads the payload data to be decrypted from the ShMem via DMA, performs the decryption operation, and writes it to the destination address configured by the PS, then triggers the SDAP processing. S1-5, DPHA SDAP header parser parses the SDAP header and caches the SDAP header fields: If it is an SDAP control PDU, parse the corresponding fields and end the PDU processing; If it is an SDAP data PDU, output the data length and address information, end the processing of this MAC subPDU, jump to step 2, and start processing the next MAC subPDU.
2. A 5GNR air interface protocol stack data plane acceleration method according to claim 1, characterized in that: The following steps are also included: the data parsed by DPHA, including multiple HFD data formats, are composed of MAC_CE_GROUP and MAC_SDU_GROUP according to different logical channel configurations. The former contains all MAC CEs in a TB block, and the latter contains the header fileds of all MAC SDUs in a TB.
3. A 5GNR air interface protocol stack data plane acceleration method according to claim 2, characterized in that: The method further includes the following steps, wherein the HFD data format includes HFD_MAC_CE, HFD_AM_CTRL, HFD_AM_PDCP_CTRL, HFD_AM_PDCP_DATA, HFD_AM_PDCP_SDAP_DATA, HFD_AM_SEG_DATA, HFD_UM_PDCP_DATA, HFD_UM_PDCP_SDAP_DATA, HFD_UM_SEG_DATA and HFD_UM_PDCP_CTRL.
4. A 5GNR air interface protocol stack data plane acceleration method according to claim 2, characterized in that: The following steps are also included: for hierarchical management of TB blocks, the L1C is responsible for allocating TB block memory, and the PS is responsible for releasing it; the protocol stack and PHY jointly maintain an idle TB block list. Before starting DPHA each time, the protocol stack configures the address of the TB block received in this TTI to the DPHA register.
5. A 5GNR air interface protocol stack data plane acceleration method according to claim 4, characterized in that: The following steps are also included: the backup solution of TB block memory is allocated by PS and notified to PHY through the mac-phy mailbox mechanism interaction message.
6. A 5GNR air interface protocol stack data plane acceleration method according to claim 1, characterized in that: There are two logical channel configuration methods: DPHA includes multiple logical channel configuration registers, which are used to indicate whether DPHA is in idle or running state, whether HFD buffer overflows, MAC CE / SDU HFD buffer base address register, and HFD buffer configuration register.
7. A 5GNR air interface protocol stack data plane acceleration method according to claim 6, characterized in that: The DPHA accelerator is designed with multiple registers including: DHPA_LCx_CFG_REG, DHPA_LCx_PDCPRXDELIV_REG, DHPA_CTL_REDHPA_ITR_CTDHPA_STATUS_RE, DHPA_SEC_CTX_BUF_ADDR_REG, DHPA_CE_HFD_BUF_BASE_ADDR_RE, DHPA_MAC_SDU_HFD_BUF_BASE_ADDR_REG, DHPA_MAC_SDU_HFD_CFG_REG, DHPA_SEC_CTX_BUF_CFG_REG and DHPA_TB_BUFFER_ADDR_REG.
8. A 5GNR air interface protocol stack data plane acceleration method according to claim 1, characterized in that: The shared memory allocation principles include static allocation, dynamic allocation by PHY, memory recycling, hierarchical management, and Link-list head and tail reads. Each memory block corresponds to a Desc descriptor, and all Desc descriptors are stored in SRAM as a Desc Chain. During the system initialization phase, the protocol stack creates the Desc Chain. Memory allocation and release are achieved by maintaining the two registers DHPAFreeDescAddrReg and DHPATailDescAddrReg.
9. A 5GNR air interface protocol stack data plane acceleration method according to claim 8, characterized in that: The memory allocation process includes: Allocation 1: The DLTB block pool is initialized by the PS and managed in the form of a TB DESC chain; Allocation 2: After initialization is complete, the PS notifies the PHY via a MAC-PHY event message, which contains at least the first address and total number of DL TB block pools; Allocation 3: Before a HARQ entity of the PHY starts to receive TB data, it reads an idle TB address from the FreeDescAddrReg field of the TB DESC chain and updates FreeDescAddrReg to the address of the next idle TB; Allocation 4: After the PHY correctly receives the TB data of the HARQ entity, it notifies the PS through the PHY-MAC message; Allocation 5: If the PHY cannot correctly receive the TB of the HARQ entity, the PHY notifies the protocol stack through a PHY-MAC message, which carries the TB DESC address where the TB is located; Allocation 6: If the PS receives a PHY message, it updates the statistics and directly updates TailDescAddrReg to reclaim the TB; Allocation 7: If allocation 5 occurs, the PS receives the PHY message and performs the DL data processing process. After delivering all MAC SDUs in the TB to the AP, it updates TailDescAddrReg to reclaim the TB.
10. A 5GNR air interface protocol stack data plane acceleration method according to claim 8, characterized in that: When DPHA parses the TB, for each processed MAC subPDU: if it is a MAC CE, it outputs the MAC CE HFD to SRAM; If it is an RLC segment or PDCP status report, the MAC SDU HFD is output to SRAM. The HFD contains the absolute address of the RLC segment and PDCP payload in the source TB. If it is a complete RLC data packet and the logical channel to which the data packet belongs is configured for decryption or protection, the SEC context is configured for it. The SEC context is stored in the SRAM space and the decrypted address is assigned to the packet.