A communication system physical layer signal processing method, a communication system and an SDR system

CN117693029BActive Publication Date: 2026-09-29INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311564548.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2026-09-29
Estimated Expiration
2043-11-22

AI Technical Summary

Technical Problem

然而,由于基于通用处理器的SDR平台处理流程和算法的灵活性以及计算核心的异构性导致子任务间负载均衡困难,以及软件控制的子任务处理启动会引入时延,且增加了实现难度,上述的流水线处理模式不适用于基于通用处理器的物理层信号处理

Benefits of technology

[0021]与现有技术相比,本发明的优点在于:本发明方案中通过构建异构多核处理器上满足实时性处理要求的隔离核心数量确定方法、隔离核心组划分方法以及与线程池绑定方法、线程池与时隙绑定方法、异构多核通用处理器上数据块并行处理结构,能够确定处理一个数据块所需的最少核心数,保证数据块处理时间小于数据块到达时间,满足实时性处理需求;并能够确定每个隔离的计算核心所用的核类型,以及线程池使用的隔离核心组;以及确定每帧每个待处理的时隙使用的线程池;同时根据分配的线程池,数据块可并行处理,完成后交付数据至上一层协议。满足实时性处理要求的同时避免了需要设计复杂的负载均衡机制以及处理启动机制,易于部署与实现。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117693029B_ABST
    Figure CN117693029B_ABST
Patent Text Reader

Abstract

The application provides a communication system physical layer signal processing method, the physical layer signal comprises multiple frames of data, each frame of data contains multiple continuous time slots, the method comprises continuously numbering the multiple continuous time slots in each frame of data in time sequence and performing the following steps: S1, based on the number of time slots of each frame of data in the physical layer signal and the transmission time of each frame of data, the communication system is preconfigured to construct multiple mutually isolated thread pools, and the thread pools are continuously numbered according to the time slot numbering rule; S2, each time slot in each frame of data is bound to a thread pool, and the time slots with the same number in different frames are bound to the same thread pool; S3, based on the binding relationship between the time slots and the thread pools in the step S2, the physical layer signal is processed. The scheme of the application meets the real-time processing requirement, avoids the need to design a complex load balancing mechanism and a processing starting mechanism, and is easy to deploy and implement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, specifically to physical layer signal simulation processing technology in the field of wireless communication, and more specifically, to a physical layer signal processing method for a communication system, a communication system, and an SDR system. Background Technology

[0002] With the increasing demand for broadband transmission from services such as ultra-high-definition video and virtual reality, communication systems have evolved to 5G. Data transmission rate is one of the key performance indicators of wireless communication systems. For example, 5G NR, based on Orthogonal Frequency Division Multiplexing (OFDM), uses high-bandwidth, high-spectral-efficiency modulation and coding schemes, massive MIMO antennas, and Low Density Parity Check Code (LDPC) encoding and decoding technologies, with a designed downlink theoretical peak transmission rate of 20Gbps. Before deploying a 5G communication system, it needs to be tested to verify its physical layer algorithm performance and check its peak rate. In existing technologies, to verify the physical layer algorithm performance and check key indicators such as peak rate, a Software Defined Radio (SDR) platform is often used for link-level testing.

[0003] SDR employs a general-purpose, open hardware platform for its radio frequency (RF) section. The host computer uses a general-purpose Intel x86 architecture processor to implement physical layer baseband processing and the protocol stack. It also supports the integration of devices such as channel simulators and signal analyzers into the link to simulate real-world communication environments, providing significant flexibility and ease of use for communication system development and testing. However, communication systems have stringent real-time requirements for physical layer signal processing. All processing must be completed within a specified time; that is, the data processing rate must exceed the data arrival rate. This is especially crucial when testing peak transmission rates, where the amount of data processed per unit time is large. If the expected data throughput cannot be met, the anticipated transmission rate cannot be achieved. Therefore, SDR must be able to fully simulate the physical layer signal processing mode of the communication system for better testing. In practical applications, to meet real-time requirements, commercial communication equipment often uses dedicated multi-processor system-on-chip (MPSoC, hereinafter referred to as a communication-specific processor) hardware to implement physical layer signal streaming processing. The pipelined processing model employed by dedicated communication processors divides the physical layer processing flow into several independent subtasks (such as channel estimation and demodulation), each processed by a dedicated computing core. Data is processed by one subtask before being passed to the next, and data exchange is completed directly on the hardware. SDR typically uses general-purpose processors based on the Intel x86 architecture as the primary computing resource for baseband signal processing. It is developed using C / C++ to facilitate the deployment of various communication protocols and algorithms, significantly reducing development difficulty and cost. However, this flexibility and simplicity make it difficult to implement the pipelined processing model used in dedicated communication processors on SDR-based communication systems. This means that testing of this type of physical layer processing model is impossible on SDR systems. The main reasons for this are twofold: First, pipelined processing requires finely dividing subtasks and rationally designing the computing resources for each subtask and performing load balancing to avoid blocking and waiting problems caused by large differences in computation time between subtasks. However, since physical layer processing modules on the SDR platform can be flexibly added or removed as needed, and the algorithms used by each module can be flexibly selected, changes in the number and processing time of subtasks in the pipelined processing model require re-load balancing, which greatly increases the workload and complexity of load balancing between subtasks. On the other hand, in the pipelined processing mode, subtasks can transfer data through Direct Memory Access (DMA). When the DMA transfer data length reaches the pre-configured transfer length, the DMA issues a data fragmentation interrupt to notify the subsequent subtask to start data processing and continue to receive subsequent data, thus realizing the control of data streaming processing at the hardware level.However, if the above pipelined processing mode is adopted on SDR, the startup of each subtask and the interruption of data fragmentation must be implemented in software using thread communication methods, which will introduce additional processing latency and implementation difficulty.

[0004] In summary, existing physical layer processing architectures fall into two categories: serial processing and pipelined processing. Serial processing processes each time slot sequentially, employing Single Instruction Multiple Data (SIMD) instruction sets and parallel programming acceleration techniques to ensure that the data processing time for each time slot does not exceed the duration of a single time slot, guaranteeing continuous data reception and processing. However, with the increase in bandwidth and the spectral efficiency of the modulation and coding schemes used, the computational load per time slot increases, easily exceeding the time of a single time slot. The aforementioned methods fail to achieve the required throughput, causing buffer overflows and preventing normal data transmission and reception. Furthermore, the serial architecture introduces significant context overhead, leading to low utilization of computational resources. Offloading some computational tasks (such as LDPC encoding / decoding) to GPUs or FPGAs for hardware acceleration is a primary method to alleviate this problem. However, the time required to transfer the data to be processed to the GPU or FPGA and to read the computation results still increases with the computational load, and the time slot processing time remains limited to the duration of a single time slot, still failing to meet real-time processing requirements. Pipeline processing architectures for communication-specific processors divide the physical layer signal processing process into sub-modules, each processed by an independent computing core. Data is processed in each computational subtask according to a First-In-First-Out (FIFO) rule, and DMA technology is used to control data exchange between subtasks. Because modules can process in parallel, throughput can be significantly improved, and real-time requirements can be met if the data processing time in the pipeline is less than the data arrival interval. However, the flexibility of the processing flow and algorithms of SDR platforms based on general-purpose processors, as well as the heterogeneity of computing cores, makes load balancing between subtasks difficult. Furthermore, software-controlled subtask startup introduces latency and increases implementation complexity. Therefore, the aforementioned pipelined processing model is not suitable for physical layer signal processing based on general-purpose processors. Thus, a new physical layer signal processing method is needed that can meet real-time requirements without adding extra latency or implementation complexity. Summary of the Invention

[0005] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a new physical layer signal processing method, communication system, and SDR system for communication systems.

[0006] According to a first aspect of the present invention, a physical layer signal processing method for a communication system is provided. The physical layer signal includes multiple frames of data, each frame of data containing multiple consecutive time slots. The method includes sequentially numbering the multiple consecutive time slots in each frame of data according to time order and performing the following steps: S1, pre-configuring the communication system based on the number of time slots in each frame of data in the physical layer signal and the transmission time of each frame of data to construct multiple mutually isolated thread pools, and sequentially numbering the thread pools according to the time slot numbering rules; S2, binding each time slot in each frame of data to a thread pool, and binding time slots with the same number in different frames to the same thread pool; S3, processing the physical layer signal based on the binding relationship between the time slots and the thread pools in step S2.

[0007] Preferably, the communication system includes multiple processing cores, with multiple of each type of processing core, and in step S1, each thread pool contains only one type of processing core.

[0008] Preferably, in step S1, each thread pool is bound to one or more processing cores, and the number of processing cores bound to each thread pool is determined by: obtaining the configuration parameters of the cores in the communication system to determine the data processing capability of each core; obtaining the transmission time of each frame of data; configuring the number of cores bound to each thread pool so that the time for each thread pool to process a time slot is less than or equal to the difference between the transmission time of each frame of data and the preset reserved time.

[0009] Preferably, S2 includes: S21, calculating the number of time slots that each thread pool can handle in the following manner:

[0010]

[0011] Among them, ICG j This represents the j-th thread pool. T represents the number of time slots that the j-th thread pool can handle. frame Indicates the transmission time of each frame of data. This represents the time it takes for the j-th thread pool to process one time slot, with the symbol... Indicates rounding down;

[0012] S22. Bind time slots to thread pools as follows:

[0013] P(i) = i mod N

[0014] Where P(i) represents the thread pool bound to the i-th time slot, and N represents the total number of thread pools.

[0015] Preferably, in step S1, the number of threads in the thread pool is the same as the number of time slots in each frame of data.

[0016] Preferably, in step S2, each time slot in each frame of data is bound to a thread pool with the same number.

[0017] Preferably, step S3 includes: S31, the communication system reads data from each time slot according to the time slot interval, and saves the read time slot data to the corresponding thread pool memory every frame; S32, when each thread pool receives a new time slot task, it determines whether the previous time slot task has been completed. If it has been completed, the data from the previous time slot task is retrieved and the new time slot task is added to the thread pool. If it has not been completed, the previous time slot task is abandoned and the new time slot task is processed.

[0018] Preferably, step S3 further includes: S33, if the previous time slot task was not completed in step S32, the reserved time is adjusted so that the time for each thread pool to process the time slot task meets the real-time requirement, wherein the real-time requirement means that the time for the thread pool to process the time slot task is less than or equal to the difference between the transmission time of each frame of data and the preset reserved time.

[0019] According to a second aspect of the present invention, a communication system is provided, the communication system being configured to process physical layer signals as described in the first aspect of the present invention.

[0020] According to a third aspect of the present invention, an SDR system is provided, the SDR system being configured to perform analog processing on physical layer signals of a communication system using the method described in the first aspect of the present invention.

[0021] Compared with existing technologies, the advantages of this invention are as follows: This invention, through constructing a method for determining the number of isolated cores on a heterogeneous multi-core processor to meet real-time processing requirements, a method for dividing isolated core groups, a method for binding them to thread pools, a method for binding thread pools to time slots, and a parallel data block processing structure on a heterogeneous multi-core general-purpose processor, can determine the minimum number of cores required to process a data block, ensuring that the data block processing time is less than the data block arrival time, thus meeting real-time processing requirements. It can also determine the core type used by each isolated computing core, the isolated core group used by the thread pool, and the thread pool used for each time slot to be processed in each frame. Simultaneously, based on the allocated thread pool, data blocks can be processed in parallel, and after completion, the data is delivered to the upper-layer protocol. While meeting real-time processing requirements, it avoids the need to design complex load balancing and processing startup mechanisms, making it easy to deploy and implement. Attached Figure Description

[0022] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0023] Figure 1 This is a schematic diagram of the downlink transceiver process of the physical layer in a 5G communication system based on SDR.

[0024] Figure 2 This is a schematic diagram of the serial processing flow;

[0025] Figure 3 This is a schematic diagram of the assembly line processing flow.

[0026] Figure 4 This is a schematic diagram of a processing structure based on data block computing resource isolation according to an embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram of a processing structure based on data block computing resource isolation according to an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0029] As described in the background section, existing physical layer processing structures cannot be implemented on general-purpose processors while meeting real-time requirements. This invention addresses the shortcomings of pipelined processing modes used in dedicated communication processors on SDR platforms employing general-purpose processors. It proposes a physical layer processing structure based on data block and computational resource isolation. This structure processes physical layer signals in units of data blocks and places these blocks into isolated thread pools for parallel processing, thus achieving a physical layer processing structure within a general-purpose processor. In this invention, the processing between data blocks is independent (e.g., subframes in LTE, time slots in NR), with each data block occupying its allocated computational resources for parallel processing. This invention avoids the complexity of software implementation of pipelined processing modes, eliminating the need for load balancing within time slots and communication between threads, thus offering the advantage of easy deployment. Furthermore, this invention considers the heterogeneity of general-purpose processor cores and proposes a method for computational resource isolation based on real-time processing requirements.

[0030] In summary, the present invention provides a physical layer signal processing method for a communication system. The physical layer signal includes multiple frames of data, each frame containing multiple consecutive time slots. The method includes sequentially numbering the multiple consecutive time slots in each frame and performing the following steps: S1, pre-configuring the communication system based on the number of time slots in each frame and the transmission time of each frame to construct multiple isolated thread pools, and sequentially numbering the thread pools according to the time slot numbering rules; S2, binding each time slot in each frame to a thread pool, with time slots with the same number in different frames bound to the same thread pool; S3, processing the physical layer signal based on the binding relationship between time slots and thread pools in step S2. This invention is applicable not only to homogeneous processors but also to heterogeneous processors, with the same resource isolation method. For ease of understanding, this embodiment focuses on heterogeneous processors. It should be noted that the core idea of ​​data segmentation and resource isolation proposed in this invention is as follows: For data blocks that can be processed independently, data blocks (in this embodiment, time slots are used as the unit data blocks) are processed in parallel using isolated computing cores without interfering with each other. Based on the real-time processing requirements and the computing power of heterogeneous cores, an appropriate number of computing cores are allocated to the processing of each data block (in heterogeneous processors, the number of different computing cores required for the same data block is different; a computing core with strong computing power requires fewer cores to process a data block than a computing core with weak computing power requires fewer cores to process a data block). Compared with existing methods, the processing structure of this invention processes data in blocks, which meets the requirements of real-time processing while avoiding the design of complex subtask load balancing mechanisms and processing startup mechanisms for each computing task, making it easier to deploy and implement. Furthermore, this invention isolates computing resources. Addressing the issues that serial processing structures cannot meet the real-time requirements and that the pipelined processing mode of communication processors is not suitable for communication systems based on general-purpose processors, the computing resource isolation scheme considers both the computing power of heterogeneous cores and the real-time processing requirements when allocating computing resources. This allows for the allocation of computing resources while meeting processing speed requirements, avoiding the need to design complex load balancing and processing startup mechanisms, making it easy to deploy and implement.

[0031] To better understand this invention, the following detailed description, in conjunction with the accompanying drawings and embodiments, covers several aspects, including an example of data transmission in a 5G NR communication system using SDR, the physical layer processing structure in the prior art, and the physical layer processing structure of this invention based on data block computing resource isolation.

[0032] I. Example of Data Transmission in a 5G NR Communication System Based on SDR

[0033] This invention uses downlink data transmission in a 5G NR communication system based on SDR as an example for illustration. Figure 1 As shown, this demonstrates the downlink transceiver process of the 5G NR physical layer based on SDR. The transmitting end is the base station gNB, and the receiving end is the user terminal UE. Taking the UE side as an example, the hardware part is a general-purpose SDR hardware platform, including modules such as RF processing and AD / DA conversion, which can convert the wireless signals received by the antenna into baseband IQ data. The baseband IQ data is transmitted to the host computer through a high-speed interface (such as PCI-E). The host computer uses a multi-core general-purpose processor based on the Intel x86 architecture as a computing resource for physical layer processing, etc. The multi-core general-purpose processor can adopt homogeneous or heterogeneous cores, such as the i9-10900X containing 20 homogeneous logical cores, and the i9-12900K containing 16 logical performance cores (P-Cores) and 8 logical efficiency cores (E-Cores). The P-Core's clock speed is higher than that of the E-Core. The physical layer baseband processing program reads the baseband IQ data from the SDR hardware through the driver interface and performs baseband processing. 5G NR uses time slots as the scheduling granularity, with each time slot being an independent data block. For each downlink time slot to be processed, the physical layer processing mainly includes OFDM demodulation, de-mapping, channel estimation, QAM demodulation, and LDPC decoding. After decoding, the 0 and 1 bits of data are delivered to the MAC layer protocol for processing. According to the 5G NR physical layer protocol design, the time length of one frame is set to T. frame Each time slot has a duration of T. slot The number of sampling points in each time slot is N. sample The time slot can be configured with a transmission bandwidth of W. W Resource parameters such as modulation and coding scheme (MCS).

[0034] Under the above-mentioned transmission and reception process and physical layer protocol settings, the following uses the UE side as an example to illustrate the working mode of the serial processing structure, pipelined processing structure and data block computing resource isolation processing structure.

[0035] II. Physical layer processing structure under existing technology

[0036] Figure 2 This demonstrates a serial processing structure, in which every T... slot Read a time slot from the SDR. Each time slot has N. sample There are 10 sampling points. Each time slot to be processed is added to the thread pool P as a computation task. The thread pool P processes only one computation task at a time, but can use all computing resources. Therefore, under this processing structure, the data arrival time interval is T. slot When the data processing time T of thread pool P proc Less than T slot If the enqueue rate is less than the dequeue rate, continuous reception and processing are possible; otherwise, data overflow will occur.

[0037] Figure 3 This demonstrates a pipelined processing architecture. Building upon a serial processing architecture, it divides the physical layer processing flow into M subtasks. Each subtask creates a thread pool and occupies independent CPU computing resources. The CPU group exclusively used by a thread pool is denoted as an ICG (Isolated Core Group). Each thread pool processes only one subtask at a time. Each time slot is considered complete only after all subtasks have been processed sequentially. Different subtasks in different time slots can be processed in parallel; that is, when time slot i is in subtask m, time slot i+1 can be processed simultaneously in subtask m-1. Under this processing architecture, the data processing time T for one time slot is... proc This refers to the time after M subtasks. The execution time of each subtask (after load balancing) is less than one time slot time T. slot When, i.e., T proc <T slot This means that the data enqueue rate is less than the dequeue rate. This structure can achieve continuous data reception and processing; otherwise, overflow will occur and the data cannot be processed normally.

[0038] III. Physical Layer Processing Structure Based on Data Block Computation Resource Isolation

[0039] In a processing architecture based on data block-based computational resource isolation, similar to a pipelined architecture, each thread pool occupies independent CPU computing resources, and the CPU group exclusively occupied by a thread pool is denoted as an ICG. However, in the processing architecture based on data block (i.e., time slot) computational resource isolation of this invention, each time slot is processed by a specific thread pool (data blocks are bound to thread pools), and each thread pool occupies an independent computing core. Once a time slot is processed in a thread pool, it is completed without going through other threads. When a time slot i is being processed, time slot i+1 can be processed in parallel in its corresponding thread pool. Under this processing architecture, when each ICG processes 1 time slot per frame and the data arrival time is T... frame If the data processing time T proc Less than T frame If the enqueue rate is less than the dequeue rate, this structure can continuously receive and process data; otherwise, data overflow will occur and the data cannot be processed normally.

[0040] In the processing structure based on data block computing resource isolation of the present invention, there is no need to load balance the data processed in parallel, and no need to design a new structure when facing changes in the process.

[0041] First, for pipelined architectures, to prevent blocking issues caused by excessive differences in the execution time of various subtasks, pipelined processing architectures need to perform load balancing to ensure that the computation time of parallel processing subtasks is consistent. However, in the architecture proposed in this invention, all physical layer processing in each time slot uses a dedicated CPU core. After processing is completed, the complete physical layer data block of this time slot can be directly delivered to the upper layer protocol. Therefore, the processing progress of other time slots does not need to be considered. Thus, the architecture proposed in this invention does not require load balancing for parallel processing time slots.

[0042] Secondly, when a physical layer processing subprocess is added or the algorithm of a subprocess is adjusted, the computation time of this subprocess changes, and the original load balancing result of the pipeline structure is no longer effective. The pipeline structure needs to be redesigned (such as splitting the algorithm of a subprocess) and load balancing needs to be performed. However, the processing structure proposed in this invention does not divide computing resources according to subprocesses or subtasks. Even if the implementation of a subprocess is adjusted, it will only affect the computation time of a time slot on a specific computing core and will not affect other time slots that are processed in parallel, so no adjustment is required.

[0043] Furthermore, adjacent thread pools in the pipelined processing structure need to communicate to determine whether the predecessor task has finished before the processing of the successor subtask can begin. However, the processing structure proposed in this invention does not require communication between thread pools, thus avoiding the latency and implementation difficulty introduced in this aspect.

[0044] According to an embodiment of the present invention, the specific implementation steps of the communication physical layer processing method based on data block (i.e., time slot) computational resource isolation are as follows:

[0045] T1. First, perform pre-configuration:

[0046] T11: Configure core isolation. Isolate all cores of the general-purpose processor that can be used for physical layer processing, i.e., configure the CPU cores that need to be isolated.

[0047] T12: Determine the minimum number of computational cores required for a thread pool to exclusively utilize an ICG. Under the time slot resource configuration that reaches peak rate (where computational load is at its maximum), test the time T required for the UE to process one downlink time slot using different ICG configurations. proc This refers to the time from the start of OFDM demodulation to the end of PDSCH channel decoding. To avoid delays introduced by the P-Core waiting for the E-Core, homogeneous cores are used within the ICG; therefore, ICGs are divided into two categories: ICGs containing P-cores (P-ICG) and ICGs containing E-cores (E-ICG). The time T varies depending on the number of P-cores used. proc The minimum number of cores k within the P-ICG can be determined. P And T when using different numbers of E-coresproc The minimum number of cores k within the E-ICG can be determined. E Meanwhile, considering issues such as heat dissipation and power supply during high-load multi-core operation, T frame , with T proc A certain amount of time T should be reserved between them. y This makes T proc +T y Less than T frame This satisfies the real-time requirements. Let S be... ICG Given the number of time slots that can be processed per frame in an ICG, then For different ICGs, there are Among them, ICG j This represents the Jth thread pool. This represents the number of time slots that the j-th thread pool can handle. This represents the time it takes for the j-th thread pool to process one time slot, with the symbol... This indicates rounding down. When an ICG has strong processing power, its corresponding S... ICG It may be greater than 1, which means that under real-time requirements, the number of time slots that this ICG can process per frame can be greater than 1.

[0048] T13: ICG Division. Based on the results of steps T11 and T12, the isolated cores are divided into P-ICG and E-ICG, with the number of cores in each P-ICG being no less than k. P The number of cores in the E-ICG is no less than k E A maximum of N can be formed. P P-ICG and N E Each E-ICG, with P-ICG index from 0 to N. P -1, E-ICG index is N P To N p +N E -1;

[0049] T14: Create a thread pool. Create N pool =N P +N E There are N thread pools, and each thread pool is bound to an ICG (Interactive Group Classification). Different thread pools are bound to different ICGs. Let L be the number of time slots to be processed per frame, and let i ∈ [0, L-1] in chronological order. If N... pool If the value is less than L, then this processor cannot meet the real-time parallel processing requirement of L time slots per frame. In this case, batch processing is performed, so that when some thread pools process multiple times in batches, as long as the real-time requirement is met, it is acceptable.

[0050] T15: Time slots are bound to thread pools. The time slots in each frame of data are bound to a thread pool, expressed as: P(i) = i mod N, N = N pool P(i) represents the thread pool bound to the i-th time slot, and N represents the total number of thread pools. pool This binding method ensures that when each thread pool processes multiple time slots, the time interval between processing time slots is maximized. To guarantee parallel processing of all time slots (data blocks), it is preferable to have one thread per time slot. Figure 4 As shown, the number of time slots per frame is L = N. pool The time slot index is 0 to L-1, and the number of threads constructed is N. pool Thread indices are 0 to N pool -1, because L = N pool This ensures that one time slot can be bound to one thread pool. Preferably, time slots with the same sequence number in different frames are bound to the same thread pool. Figure 4 It can be seen that time slot 0 in frame 1, time slot 0 in frame 2, and time slot 0 in other frames are all bound to thread pool P(0), and so on. Furthermore, it should be noted that when the number of time slots in each frame is greater than the total number of thread pools constructed, i.e., L is greater than N... pool At times, such as Figure 5 As shown, according to the binding rule P(i) = i mod N, a thread pool will bind multiple time slots, from... Figure 5 It can be seen that time slot 0 and time slot N of frame 1 pool All are bound to thread pool 0, and time slot 1 and time slot N of frame 1. pool +1 is bound to thread pool 1, and so on.

[0051] T2. After completing the above pre-configuration, proceed with data processing:

[0052] T21: Time Slot Data Reading and Buffering. After the UE completes the initial synchronization process, it reads and buffers data every T... slot Read N from the SDR read interface sample One IQ data. For any time slot i to be processed, the UE every T frame Save the read IQ data to the corresponding memory location, still referring to... Figure 4 The IQ data corresponding to time slot i is represented as the variable rxdata[i].

[0053] T21. Time Slot Data Processing. After receiving the data rxdata[i] from time slot i, the previous task is first retrieved from the thread pool P(i), and the computation task for the current time slot i is added to P(i). There are two cases here: if the previous task has been completed, it can be successfully retrieved and processed in the current time slot i; if the previous task cannot be successfully retrieved (an exception occurs), the previous task is forcibly abandoned, and a new task is added to the thread pool. The time slot performs baseband data processing in the allocated thread pool, and after completion, the data is delivered to the MAC layer.

[0054] T23. Real-time processing capability check. If an exception occurs in step T22, indicating that the previous task could not be successfully retrieved from the thread pool, it means that the current configuration does not meet the real-time processing requirements, and the reserved time T needs to be adjusted in step T12. y Repeat step T1 until the real-time requirements are met.

[0055] To more intuitively understand the implementation process of the present invention, the following example of UE physical layer processing in 5G NR illustrates the implementation process using different processors.

[0056] 5G NR parameters: Designed according to the 5G NR physical layer protocol, a frame length is 10ms, a subcarrier spacing of 30kHz is selected, a frame contains 20 time slots, each time slot is 0.5ms, so a frame duration is 10ms; the communication bandwidth is configured as 40MHz, and the PDSCH channel contains 106 resource blocks. The system consists of blocks (RBs), each containing 12 subcarriers; each OFDM symbol uses 2048 points, and each time slot samples 15 OFDM symbol lengths (including the cyclic prefix), i.e., 30720 IQ data points are sampled; the time slots are set to the same modulation scheme 64QAM, with a target code rate of 948 / 1024; channel estimation uses the least squares algorithm, demodulation uses the LLR algorithm, and LDPC decoding uses a lookup table-based belief propagation algorithm; simultaneously, FFT, channel estimation, channel compensation, LLR demodulation, and LDPC decoding algorithms are implemented using the AVX2 instruction set; the system adopts TDD duplex mode, with a downlink time slot, flexible time slot, and uplink time slot ratio of 8:1:1, i.e., the UE has 16 downlink time slots to process per frame, with indices arranged 0-15 in chronological order.

[0057] If adopted Core TM The i9-10900X is a homogeneous multi-core processor containing 20 homogeneous logical cores (referred to as P-cores) with a base frequency of 3.7GHz. According to the above embodiment, 17 cores (3-19) are initially isolated for physical layer processing. The average single-core time slot processing time is measured to be 5.6ms (5.6ms is less than 10ms), therefore k P =1, each SICG A value of 1 is sufficient to meet real-time requirements. Sixteen thread pools are created, each bound to cores 4-19. After completing the above configuration, the UE downlink receiving procedure is started. The UE first performs initial synchronization to determine the frame header position, frame number, and other information. Then, it reads 30720 data points from the SDR every 0.5ms. According to the frame number and time slot number, the corresponding IQ data is saved to the memory address of rxdata[i], and the processing task for each time slot is added to the corresponding thread pool. The thread pool performs the calculation, delivers the data of this time slot to the MAC layer after completion, and checks whether the real-time requirements are met. With this processing structure, the UE can achieve continuous data reception and processing.

[0058] If adopted Core TM The i9-12900K heterogeneous multi-core processor contains 16 P-cores and 8 E-cores. The P-cores have a clock speed of 3.2GHz, and the E-cores have a clock speed of 2.4GHz. The average time slot processing time per P-core is measured to be 2.8ms, and the average time slot processing time per E-core is 4.3ms. Therefore, k P and k E All values ​​are 1. Therefore, a total of 12 thread pools based on P-ICG and 4 thread pools based on E-ICG were created. Each P-ICG contains 1 P-Core, and each E-ICG contains 1 E-Core. Then, steps T21 to T23 are performed similarly to achieve data processing.

[0059] As can be seen from the above embodiments, the solution of the present invention is applicable not only to broadband communication systems using homogeneous processors, but also to broadband communication systems using heterogeneous multi-core general-purpose processors (such as software radio platforms). The solution of the present invention, by constructing a method for determining the number of isolated cores, a method for dividing isolated core groups, and methods for binding them to thread pools, binding thread pools to time slots, and a parallel data block processing structure on a heterogeneous multi-core general-purpose processor to meet real-time processing requirements, can determine the minimum number of cores required to process a data block, ensuring that the data block processing time is less than the data block arrival time, thus meeting real-time processing requirements; it can also determine the core type used by each isolated computing core, the isolated core group used by the thread pool, and the thread pool used for each time slot to be processed in each frame; simultaneously, according to the allocated thread pool, data blocks can be processed in parallel, and after completion, the data is delivered to the upper-layer protocol. The solution of the present invention, while meeting real-time processing requirements, avoids the need to design complex load balancing mechanisms and processing startup mechanisms, and is easy to deploy and implement.

[0060] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0061] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0062] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0063] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for processing physical layer signals in a communication system, wherein the physical layer signals comprise multiple frames of data, each frame containing multiple consecutive time slots, characterized in that, The method includes sequentially numbering multiple consecutive time slots in each frame of data according to time order and performing the following steps: S1. The communication system is pre-configured based on the number of time slots and the transmission time of each frame of data in the physical layer signal to build multiple isolated thread pools, and the thread pools are numbered consecutively according to the time slot numbering rule; each thread pool is bound to an ICG, and the ICGs bound to different thread pools are different. The CPU group exclusively used by a thread pool is an ICG. S2. Bind each time slot in each frame of data to a thread pool, and bind time slots with the same number in different frames to the same thread pool. S3. Process the physical layer signals based on the binding relationship between the time slot and the thread pool in step S2.

2. The method according to claim 1, characterized in that, The communication system includes multiple processing cores, with multiple of each type. In step S1, each thread pool contains only one type of processing core.

3. The method according to claim 2, characterized in that, In step S1, each thread pool is bound to one or more processing cores, and the number of processing cores bound to each thread pool is determined in the following way: Obtain the core configuration parameters of the communication system to determine the data processing capabilities of each core; Get the transmission time of each frame of data; Configure the number of cores bound to each thread pool so that the time for each thread pool to process a time slot is less than or equal to the difference between the transmission time of each frame of data and the preset reserved time.

4. The method according to claim 3, characterized in that, S2 includes: S21. Calculate the number of time slots that each thread pool can handle in the following manner: in, Indicates the first A thread pool, Indicates the first The number of time slots that a thread pool can handle. Indicates the transmission time of each frame of data. Indicates the first Each thread pool processes the time of one time slot, symbol "". " indicates rounding down; S22. Bind time slots to thread pools as follows: in, Indicates the first A thread pool bound to a time slot, where N represents the total number of thread pools.

5. The method according to claim 3, characterized in that, In step S1, the number of threads in the thread pool is the same as the number of time slots in each frame of data.

6. The method according to claim 5, characterized in that, In step S2, each time slot in each frame of data is bound to a thread pool with the same number.

7. The method according to claim 1, characterized in that, Step S3 includes: S31. The communication system reads the time slot data of each frame sequentially and saves it to the corresponding thread pool memory according to the time slot interval; S32. When each thread pool receives a new time slot task, it determines whether the previous time slot task has been completed. If it has been completed, it retrieves the data of the previous time slot task and adds the new time slot task to the thread pool. If it has not been completed, it abandons the previous time slot task and processes the new time slot task.

8. The method according to claim 7, characterized in that, Step S3 further includes: S33. If the previous time slot task is not completed in step S32, the reserved time is adjusted so that the time for each thread pool to process the time slot task meets the real-time requirement. The real-time requirement means that the time for the thread pool to process the time slot task is less than or equal to the difference between the transmission time of each frame of data and the preset reserved time.

9. A communication system, characterized in that, The communication system is configured to process physical layer signals using the method described in any one of claims 1-8.

10. An SDR system, characterized in that, The SDR system is configured to perform analog processing of the physical layer signals of the communication system using the method described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1 to 8.

12. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the electronic device to perform the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data transmission method and device and electronic equipment

    CN115632752A

  • Thread scheduling method and device, chip, electronic equipment and storage medium

    CN115858132A