Special signal generating device based on software defined radio and GPU server
By using a dedicated signal generation device based on software-defined radio and GPU servers, efficient collaboration between the central processing unit and the graphics processing unit is achieved, solving the problem of low collaboration efficiency of signal generation devices in high-speed data stream processing, and improving signal processing speed and system adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-03
AI Technical Summary
Existing signal generation equipment cannot achieve efficient collaboration between the central processing unit and the graphics processing unit in high-speed data stream processing, resulting in improper task scheduling and resource waste, making it difficult to cope with diverse communication needs.
It employs a dedicated signal generation device based on software-defined radio and GPU servers, generates high-speed orthogonal data streams through the graphics processing unit, combines task scheduling and memory configuration of the central processing unit, uses the PCIe bus to transmit data, and achieves high-speed, low-latency data transmission through fiber optic transmission technology. It also works with the software-defined radio platform to enable dynamic switching of various communication systems and flexible configuration of modulation methods.
It improves signal processing speed and system adaptability, ensures the real-time performance and reliability of signal generation and reception, and provides an efficient, flexible and scalable solution.
Smart Images

Figure CN121785992A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of information technology, specifically to a dedicated signal generation device based on software-defined radio and GPU servers. Background Technology
[0002] In the field of modern wireless communication, the importance of signal generation and processing technology is self-evident, as it directly affects the efficiency, stability, and adaptability of communication systems. With increasingly complex communication demands, signal processing technology has become a core driving force for industry development, especially in scenarios involving multiple communication systems and modulation methods, where the research and application of related technologies are particularly crucial.
[0003] However, existing methods often face challenges in resource scheduling flexibility and computational efficiency when dealing with complex communication scenarios. These methods struggle to balance task allocation and computational speed when handling high-speed data streams, leading to sluggish system response to diverse demands and even performance bottlenecks. This limitation significantly restricts the adaptability of communication systems in dynamic environments.
[0004] At a deeper level, a core technical challenge in this field lies in achieving efficient collaboration between different computing units, especially when processing high-speed data streams, where the collaboration efficiency between the CPU and the graphics processing unit (GPU) becomes crucial. Due to differences in their computational characteristics and task division, without an effective coordination mechanism, the CPU may fail to complete task scheduling in a timely manner, while the GPU may struggle to fully leverage its parallel computing advantages. This lack of collaboration efficiency directly leads to delays or resource waste during signal generation and processing. For example, in a business scenario requiring rapid switching of communication systems, if the CPU cannot allocate tasks in a timely manner, and the GPU is idle due to data transmission bottlenecks, the real-time performance of signal generation will be severely affected, potentially even missing critical communication windows.
[0005] Therefore, how to achieve efficient collaboration between the central processing unit and the graphics processing unit in high-speed data stream processing, and ensure seamless integration of task scheduling and parallel computing, has become a key issue that urgently needs to be addressed. Summary of the Invention
[0006] This invention provides a dedicated signal generating device based on software-defined radio and GPU servers, aiming to solve the problem that existing signal generating devices cannot achieve efficient collaboration between the central processing unit and the graphics processing unit in high-speed data stream processing.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A dedicated signal generating device based on software-defined radio and a GPU server, comprising: The graphics processing unit is used to generate high-speed orthogonal data streams corresponding to various communication systems and modulation methods. The high-speed orthogonal data streams are transmitted to the intermediate frequency signal unit through optical fiber. The intermediate frequency signal unit is used to receive the high-speed orthogonal data stream and perform signal modulation to output an intermediate frequency signal; A software-defined radio platform is used for secondary development to realize signal generation and reception. The secondary development is based on a heterogeneous system integrating a central processing unit and a graphics processing unit built on general radio software. The central processing unit is used for task scheduling and memory configuration, and the graphics processing unit is used for processing the high-speed orthogonal data stream. The main thread of the central processing unit is used to transfer the data to be processed to the memory of the graphics processing unit via PCIe (Peripheral Component Interconnect Standard Extended Bus); The graphics processing unit is used to call multiple threads to process the data to be processed, and the processed data is then transmitted back to the central processing unit via PCIe (Peripheral Component Interconnect Standard Extended Bus).
[0008] In one aspect of the invention, the graphics processing unit generates high-speed orthogonal data streams corresponding to various communication systems and modulation schemes, including: General-purpose radio software creates custom modules and links CUDA (Computational Unified Device Architecture) signal processing programs, which use CUDA programming language to design dynamic link libraries. The custom module sets up a functional module linked to the graphics processing unit in the general radio software flowchart. When the functional module runs, it enters the global function of the CUDA architecture (Compute Unified Device Architecture) to execute the high-speed orthogonal data stream processing. The central processing unit processes signal data sequentially according to the system flowchart, and the system flowchart includes the functional modules linked to the graphics processing unit. The PCIe (Peripheral Component Interconnect Standard Extended Bus) transmits the data to be processed to the memory of the graphics processing unit (GPU), which stores the data for processing by multiple threads.
[0009] In one aspect of the invention, the software-defined radio platform is further developed to achieve signal generation and reception, including: The software-defined radio platform is connected to a radio frequency front-end, which includes a low-noise amplifier and a power amplifier. The radio frequency front end receives radio signals and converts them into intermediate frequency signals or baseband signals, and the intermediate frequency signals or baseband signals are transmitted to the analog-to-digital converter. The analog-to-digital converter converts the analog signal into a digital signal, and the digital signal is transmitted to the central processing unit-graphics processing unit heterogeneous system.
[0010] In one aspect of the invention, the intermediate frequency signal unit receives the high-speed quadrature data stream and performs signal modulation to output an intermediate frequency signal, including: The intermediate frequency signal unit acquires the high-speed orthogonal data stream transmitted through the optical fiber; The intermediate frequency signal unit modulates the high-speed orthogonal data stream, and the modulation process generates an intermediate frequency signal. The intermediate frequency signal unit outputs the intermediate frequency signal to the radio frequency front end, and the radio frequency front end includes a filter and an antenna.
[0011] In one aspect of the invention, the central processing unit performs task scheduling and memory allocation, and the graphics processing unit performs the high-speed orthogonal data stream processing, including: After receiving the high-speed wireless communication signal, the central processing unit processes the signal data according to the system flowchart. When the system flowchart reaches the functional module linked to the graphics processing unit, the graphics processing unit is invoked. The central processing unit's main thread enters the CUDA architecture (Compute Unified Device Architecture) global function, which identifies the graphics processing unit's execution function. After the graphics processing unit finishes processing, the data is transmitted back to the main thread of the central processing unit to continue running.
[0012] In one aspect of the invention, the software-defined radio platform is connected to a radio frequency front end, including: The radio frequency front end converts radio signals into intermediate frequency signals or baseband signals suitable for subsequent processing; The digital-to-analog converter converts digital signals into analog signals, which are then transmitted to the power amplifier. The general-purpose radio software operation control software coordinates the operation of the central processing unit and the graphics processing unit module, and the control software manages system resources.
[0013] In one aspect of the invention, the high-speed orthogonal data stream corresponds to various communication systems and modulation methods, including: The various communication systems mentioned include Long LTE, 5G radio, Digital Video Broadcasting Satellite 2, Digital Video Broadcasting Return Link 2, Discrete Fourier Transform Single Carrier Orthogonal Frequency Division Multiplexing, and Cyclic Prefix Orthogonal Frequency Division Multiplexing; The various modulation methods include binary phase shift keying, π / 2 binary phase shift keying, quadrature phase shift keying, octet phase shift keying, 16 amplitude phase keying, 32 amplitude phase keying, 16 quadrature amplitude modulation, and 64 quadrature amplitude modulation.
[0014] In one aspect of the invention, the central processing unit's main thread enters a CUDA architecture (Compute Unified Device Architecture) global function, including: The CUDA architecture (Computation Unified Device Architecture) programming language builder compilation tool handles the digital signal processing part; The build program compilation tools are compatible with general-purpose radio software and the CUDA architecture (Compute Unified Device Architecture) platform; The graphics processing unit performs filtering, modulation, and encoding processing on the high-speed orthogonal data stream, and the filtering, modulation, and encoding processing generates processed data.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention primarily involves a heterogeneous computing architecture that efficiently coordinates a central processing unit (CPU) and a graphics processing unit (GPU). The CPU handles task scheduling and system control, while the GPU focuses on parallel processing of high-speed orthogonal data streams. Simultaneously, it incorporates fiber optic transmission technology to achieve high-speed, low-latency data transmission. Furthermore, a software-defined radio platform enables dynamic switching between various communication systems and flexible configuration of modulation methods. This invention also optimizes the execution efficiency of signal processing algorithms through deep integration of general-purpose radio software and the CUDA (Computational Unified Device Architecture) architecture, ensuring the real-time performance and reliability of signal generation and reception. Ultimately, this invention significantly improves signal processing speed and system adaptability, providing an efficient, flexible, and scalable solution for modern wireless communication systems, demonstrating outstanding technical effectiveness and application value. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a simplified structural diagram of the signal generating device of the present invention.
[0018] Figure 2 This is a flowchart of the present invention. Detailed Implementation
[0019] The present invention will be further described below with reference to embodiments. These embodiments are merely some, not all, of the embodiments of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the protection scope of the present invention.
[0020] Example 1 Please see Figure 1 as well as Figure 2 As shown, this embodiment discloses a dedicated signal generating device based on software-defined radio and a GPU server, comprising: The graphics processing unit is used to generate high-speed orthogonal data streams corresponding to various communication systems and modulation methods. The high-speed orthogonal data streams are transmitted to the intermediate frequency signal unit through optical fiber. The intermediate frequency signal unit is used to receive the high-speed orthogonal data stream and perform signal modulation to output an intermediate frequency signal; A software-defined radio platform is used for secondary development to realize signal generation and reception. The secondary development is based on a heterogeneous system integrating a central processing unit and a graphics processing unit built on general radio software. The central processing unit is used for task scheduling and memory configuration, and the graphics processing unit is used for processing the high-speed orthogonal data stream. The main thread of the central processing unit is used to transfer the data to be processed to the memory of the graphics processing unit via PCIe (Peripheral Component Interconnect Standard Extended Bus); The graphics processing unit is used to call multiple threads to process the data to be processed, and the processed data is then transmitted back to the central processing unit via PCIe (Peripheral Component Interconnect Standard Extended Bus).
[0021] In some embodiments, the graphics processing unit generates high-speed orthogonal data streams corresponding to various communication systems and modulation schemes, including: General-purpose radio software creates custom modules and links CUDA (Computational Unified Device Architecture) signal processing programs, which use CUDA programming language to design dynamic link libraries. The custom module sets up a functional module linked to the graphics processing unit in the general radio software flowchart. When the functional module runs, it enters the global function of the CUDA architecture (Compute Unified Device Architecture) to execute the high-speed orthogonal data stream processing. The central processing unit processes signal data sequentially according to the system flowchart, and the system flowchart includes the functional modules linked to the graphics processing unit. The PCIe (Peripheral Component Interconnect Standard Extended Bus) transmits the data to be processed to the memory of the graphics processing unit (GPU), which stores the data for processing by multiple threads.
[0022] In some embodiments, the software-defined radio platform is further developed to achieve signal generation and reception, including: The software-defined radio platform is connected to a radio frequency front-end, which includes a low-noise amplifier and a power amplifier. The radio frequency front end receives radio signals and converts them into intermediate frequency signals or baseband signals, and the intermediate frequency signals or baseband signals are transmitted to the analog-to-digital converter. The analog-to-digital converter converts the analog signal into a digital signal, and the digital signal is transmitted to the central processing unit-graphics processing unit heterogeneous system.
[0023] In some embodiments, the intermediate frequency signal unit receives the high-speed quadrature data stream and performs signal modulation to output an intermediate frequency signal, including: The intermediate frequency signal unit acquires the high-speed orthogonal data stream transmitted through the optical fiber; The intermediate frequency signal unit modulates the high-speed orthogonal data stream, and the modulation process generates an intermediate frequency signal. The intermediate frequency signal unit outputs the intermediate frequency signal to the radio frequency front end, and the radio frequency front end includes a filter and an antenna.
[0024] In some embodiments, the central processing unit performs task scheduling and memory allocation, and the graphics processing unit performs the high-speed orthogonal data stream processing, including: After receiving the high-speed wireless communication signal, the central processing unit processes the signal data according to the system flowchart. When the system flowchart reaches the functional module linked to the graphics processing unit, the graphics processing unit is invoked. The central processing unit's main thread enters the CUDA architecture (Compute Unified Device Architecture) global function, which identifies the graphics processing unit's execution function. After the graphics processing unit finishes processing, the data is transmitted back to the main thread of the central processing unit to continue running.
[0025] In some embodiments, the software-defined radio platform is connected to a radio frequency front end, including: The radio frequency front end converts radio signals into intermediate frequency signals or baseband signals suitable for subsequent processing; The digital-to-analog converter converts digital signals into analog signals, which are then transmitted to the power amplifier. The general-purpose radio software operation control software coordinates the operation of the central processing unit and the graphics processing unit module, and the control software manages system resources.
[0026] In some embodiments, the high-speed orthogonal data stream corresponds to multiple communication systems and modulation methods, including: The various communication systems mentioned include Long LTE, 5G radio, Digital Video Broadcasting Satellite 2, Digital Video Broadcasting Return Link 2, Discrete Fourier Transform Single Carrier Orthogonal Frequency Division Multiplexing, and Cyclic Prefix Orthogonal Frequency Division Multiplexing; The various modulation methods include binary phase shift keying, π / 2 binary phase shift keying, quadrature phase shift keying, octet phase shift keying, 16 amplitude phase keying, 32 amplitude phase keying, 16 quadrature amplitude modulation, and 64 quadrature amplitude modulation.
[0027] In some embodiments, the central processing unit's main thread enters a CUDA architecture (Compute Unified Device Architecture) global function, including: The CUDA architecture (Computation Unified Device Architecture) programming language builder compilation tool handles the digital signal processing part; The build program compilation tools are compatible with general-purpose radio software and the CUDA architecture (Compute Unified Device Architecture) platform; The graphics processing unit performs filtering, modulation, and encoding processing on the high-speed orthogonal data stream, and the filtering, modulation, and encoding processing generates processed data.
[0028] This invention primarily involves a heterogeneous computing architecture that efficiently coordinates a central processing unit (CPU) and a graphics processing unit (GPU). The CPU handles task scheduling and system control, while the GPU focuses on parallel processing of high-speed orthogonal data streams. Simultaneously, it incorporates fiber optic transmission technology to achieve high-speed, low-latency data transmission. Furthermore, a software-defined radio platform enables dynamic switching between various communication systems and flexible configuration of modulation methods. This invention also optimizes the execution efficiency of signal processing algorithms through deep integration of general-purpose radio software and the CUDA (Computational Unified Device Architecture) architecture, ensuring the real-time performance and reliability of signal generation and reception. Ultimately, this invention significantly improves signal processing speed and system adaptability, providing an efficient, flexible, and scalable solution for modern wireless communication systems, demonstrating outstanding technical effectiveness and application value.
[0029] Example 2 This embodiment is a further optimization based on Embodiment 1. In this embodiment, the present invention provides a dedicated signal generation method based on software-defined radio and GPU server. This method achieves high-speed signal processing through heterogeneous computing architecture, and combines optical fiber transmission and software-defined radio technology to support flexible configuration of multiple communication systems and modulation methods.
[0030] Step 1: The graphics processing unit generates high-speed orthogonal data streams corresponding to various communication systems and modulation methods. The high-speed orthogonal data streams are transmitted to the intermediate frequency signal unit through optical fiber.
[0031] In one embodiment, the graphics processing unit (GPU) generates a digital baseband signal containing in-phase and quadrature components using a parallel computing architecture. Specifically, the GPU is internally configured with multiple computing units, each independently handling signal generation tasks under a specific communication architecture. For example, for LTE (Long Term Evolution) architecture, the computing units generate orthogonal frequency division multiplexing (OFDM) symbols containing cyclic prefixes according to protocol specifications; for 2G digital video broadcasting satellite architecture, they generate modulation symbols conforming to the frame structure. The generated data stream is converted into an optical signal via a high-speed serial interface and transmitted to the intermediate frequency (IF) signal unit via optical fiber, achieving a transmission rate of over 10 Gbps.
[0032] Step 11: The general-purpose radio software creates a custom module and links it to the CUDA architecture (Computation Unified Device Architecture) signal processing program.
[0033] In one possible implementation, developers create functional modules with specific interfaces within a general-purpose radio software environment. These modules establish a data channel with the computational program of the graphics processing unit (GPU) through an application programming interface (API). The computational program employs a parallel computing architecture and includes three functional units: signal generation, modulation mapping, and shaping filtering. Specifically, the signal generation unit generates the raw bitstream based on communication system parameters; the modulation mapping unit converts the bitstream into complex symbols; and the shaping filtering unit performs pulse shaping and sampling rate conversion.
[0034] Step 12: Customize the functional modules linked to the graphics processing unit in the general radio software flowchart.
[0035] For example, in the software design interface, developers drag and drop the graphics processing unit (GPU) module to a designated location in the signal processing chain. This module has data input and output ports, which connect to the upstream signal source module and the downstream intermediate frequency processing module, respectively. When the software runs, the system automatically detects the GPU resource status and dynamically allocates computing tasks to the GPU device.
[0036] Step 13: The central processing unit processes the signal data sequentially according to the system flowchart.
[0037] Specifically, the central processing unit (CPU) first parses the system configuration file and loads the parameter set corresponding to the communication system. These parameters include key indicators such as symbol rate, carrier frequency, and modulation order. Subsequently, the CPU initializes the graphics processing unit's computing environment, allocates device memory space, and establishes data transmission channels. During the signal processing phase, the CPU is responsible for scheduling the execution order of various functional modules to ensure the timing correctness of the data flow.
[0038] Step 14: PCIe (Peripheral Component Interconnect Standard Extended Bus) transfers the data to be processed to the graphics processing unit device memory storage.
[0039] In one embodiment, the central processing unit (CPU) transmits signal data from system memory to the graphics processing unit (GPU) in batches via direct memory access. A double-buffering mechanism is employed during the transmission, allowing the GPU to process data already stored in device memory while the current buffer is processing data. This design effectively hides data transmission latency and improves overall processing efficiency.
[0040] Step 2: The intermediate frequency signal unit receives the high-speed quadrature data stream and completes signal modulation to output the intermediate frequency signal.
[0041] In one implementation, the intermediate frequency (IF) signal unit includes a digital up-conversion module and a digital-to-analog converter (DAC) module. The digital up-conversion module shifts the baseband signal to the IF band, a process achieved through carrier modulation using a complex multiplier. The DAC module employs a high-speed converter to convert the digital signal into an analog signal, with a sampling rate exceeding 100 MS / s. The IF signal unit also integrates an automatic gain control circuit to ensure the stability of the output signal level.
[0042] Step 21: The intermediate frequency signal unit acquires the high-speed orthogonal data stream transmitted through the optical fiber.
[0043] Specifically, the photoelectric conversion module of the intermediate frequency signal unit converts the optical signal into an electrical signal and extracts the synchronization clock through a clock data recovery circuit. The data parsing module parses the data frame structure according to a predefined protocol and extracts the payload portion. The data verification module ensures the integrity of data transmission through cyclic redundancy check and requests retransmission when an error is detected.
[0044] Step 22: The intermediate frequency signal unit modulates the high-speed quadrature data stream.
[0045] In one embodiment, the modulation process includes three steps: digital up-conversion, pulse shaping, and interpolation filtering. Digital up-conversion shifts the baseband signal spectrum to the intermediate frequency band; pulse shaping uses a root-raised cosine filter to control out-of-band signal radiation; and interpolation filtering increases the signal sampling rate to meet the input requirements of the digital-to-analog converter. A polyphase filtering structure is employed during the processing to reduce computational complexity.
[0046] Step 23: The intermediate frequency signal unit outputs the intermediate frequency signal to the radio frequency front end.
[0047] For example, the intermediate frequency (IF) signal, after analog filtering and amplification, is transmitted to the radio frequency (RF) front-end via a shielded cable. The analog filter employs an LC resonant circuit with a center frequency of 70MHz and adjustable bandwidth. The amplifier circuit provides a 20dB gain adjustment range, and the output level is continuously adjustable between -20dBm and +10dBm. The signal output port integrates a standing wave ratio (VSWR) detection circuit to monitor the load matching status in real time.
[0048] Step 3: The software-defined radio platform is used for secondary development to realize signal generation and reception.
[0049] In one embodiment, the software-defined radio platform adopts a modular design, comprising three parts: a radio frequency (RF) front-end interface, a digital signal processing unit, and a system control unit. The RF front-end interface supports multiple standard protocols, enabling interoperability with equipment from different manufacturers. The digital signal processing unit, based on a heterogeneous computing architecture, works collaboratively with a central processing unit and a graphics processing unit to complete the entire signal processing workflow. The system control unit is responsible for resource allocation, task scheduling, and device status monitoring.
[0050] Step 31: Connect the software-defined radio platform to the radio frequency front end.
[0051] Specifically, the platform establishes a physical connection with the RF front-end through standard interfaces, including but not limited to SMA and BNC. The low-noise amplifier in the RF front-end employs a two-stage amplification structure: the first stage achieves noise matching, and the second stage provides gain compensation. The power amplifier operates in Class AB mode, striking a balance between efficiency and linearity. Configuration parameters, including gain settings and frequency selection, are transmitted between the platform and the RF front-end via a control bus.
[0052] Step 32: The radio frequency front end receives radio signals and converts them into intermediate frequency signals or baseband signals.
[0053] For example, when operating in the downlink direction, the RF front-end amplifies, mixes, and filters the received RF signal, outputting a 70MHz intermediate frequency (IF) signal. The IF signal bandwidth can be dynamically adjusted according to system requirements, supporting a maximum instantaneous bandwidth of 20MHz. In the uplink direction, the RF front-end upconverts the IF signal to the target frequency band, amplifies it, and then radiates it through the antenna.
[0054] Step 33: The analog-to-digital converter converts the analog signal into a digital signal.
[0055] In one possible implementation, the analog-to-digital converter (ADC) employs a pipelined architecture, including a sample-and-hold circuit, a comparator array, and an encoder. The sample-and-hold circuit ensures signal stability during the conversion process; the comparator array performs the analog-to-digital conversion; and the encoder converts the comparison results into a binary code stream. The converter supports 14-bit resolution, with an effective bit depth of at least 12 bits, meeting the processing requirements of high dynamic range signals.
[0056] Step 4: The central processing unit performs task scheduling and memory allocation, while the graphics processing unit performs high-speed orthogonal data stream processing.
[0057] In one embodiment, the central processing unit (CPU) employs a multi-level scheduling strategy to manage computing resources. Specifically, the CPU first analyzes task dependencies and establishes a directed acyclic graph (DAG) for task execution. Then, it dynamically allocates computing resources based on task computation volume and real-time system load. Regarding memory configuration, the CPU maintains three memory pools: a system memory pool for storing raw signal data, a graphics processing unit (GPU) memory pool for storing data to be processed, and a shared memory pool for facilitating data exchange between the CPU and the GPU. Each memory pool uses a paging management mechanism, supporting on-demand allocation and release.
[0058] Step 41: After receiving the high-speed wireless communication signal, the central processing unit processes the signal data according to the system flowchart.
[0059] For example, the central processing unit (CPU) acquires the digital signal from the radio frequency (RF) front-end via interrupts. The signal processing flow includes four main stages: frame synchronization, channel estimation, equalization, and demodulation. The frame synchronization stage uses a correlation detection algorithm to determine the signal start position; the channel estimation stage extracts the channel response using pilot symbols; the equalization stage uses a minimum mean square error (MMS) algorithm to compensate for channel distortion; and the demodulation stage recovers the original bitstream based on the modulation scheme. The entire processing adopts a pipelined architecture, with each stage executing in parallel.
[0060] In one implementation, the central processing unit (CPU) first performs frame synchronization detection on the received signal to determine the signal start position. Then, it performs channel estimation to calculate the impulse response of the multipath channel. Based on the estimation results, the CPU configures equalizer parameters to compensate for distortion introduced by the channel. A sliding window mechanism is used during processing to achieve continuous signal processing.
[0061] Step 42: When the system flowchart reaches the functional module linked to the graphics processing unit, the graphics processing unit is invoked.
[0062] In one possible implementation, when the processing flow reaches the Fast Fourier Transform (FFT) computation node, the CPU activates the graphics processing unit (GPU) computing resources. The invocation process involves three steps: first, the CPU transfers the data to be processed from system memory to GPU memory; second, it configures the GPU's computation grid and thread block parameters; and finally, it starts the GPU's computation kernel. The entire process is implemented through a driver interface, ensuring seamless integration of computational tasks.
[0063] Specifically, when the processing flow reaches a computationally intensive task node, the central processing unit (CPU) activates the graphics processing unit's computing resources through a driver. The task allocation module dynamically allocates the dimensions of the graphics processing unit's computing grid based on computational complexity. The data management module is responsible for data transfer between host memory and device memory, employing an asynchronous transfer mechanism to improve efficiency.
[0064] Step 43: The central processing unit's main thread enters the CUDA architecture (Compute Unified Device Architecture) global function.
[0065] Specifically, global functions contain the core algorithms for signal processing. For example, in an orthogonal frequency division multiplexing (OFDM) system, global functions implement Fast Fourier Transform (FFT), channel equalization, and symbol demapping. During function execution, the graphics processing unit (GPU) divides the computational task into multiple thread blocks, each containing 256 threads. Shared memory is used within each thread block to optimize data access and reduce global memory access latency. Computation results are kept consistent through atomic operations.
[0066] In one embodiment, the global function includes three core functions: signal detection, channel decoding, and bit error rate statistics. The signal detection function performs coherent demodulation of the modulated signal; the channel decoding function performs forward error correction decoding; and the bit error rate statistics function calculates the bit error rate of the received signal. During function execution, the graphics processing unit uses multi-level caching to optimize data access and reduce the impact of memory latency.
[0067] Step 44: After the graphics processing unit finishes processing, the data is transmitted back to the central processing unit's main thread to continue running.
[0068] In one embodiment, data transmission is implemented asynchronously. After the graphics processing unit completes its calculations, it triggers a direct memory access transfer via an event notification mechanism. While waiting for data transmission, the central processing unit can perform other computational tasks, improving the overall system throughput. After the data is returned to system memory, the main thread checks the data integrity flag; if it is confirmed to be correct, it continues with subsequent processing.
[0069] Step 5: The main thread of the central processing unit transfers the data to be processed to the memory of the graphics processing unit via PCIe (Peripheral Component Interconnect Standard Extension Bus).
[0070] It should be noted that the data transmission process adopts a distributed-aggregated mode, supporting batch transfers of non-contiguous memory regions. The bus controller is configured in 64-bit addressing mode, with a maximum transmission bandwidth of 16GB / s. To improve transmission efficiency, the system employs a double-buffering mechanism: while data is being transferred in the current buffer, the graphics processing unit can simultaneously process data already stored in another buffer. Automatic error detection and retransmission are performed during transmission to ensure data reliability.
[0071] Step 51: The graphics processing unit calls multiple threads to process the data to be processed.
[0072] For example, for an orthogonal frequency division multiplexing (OFDM) signal processing task, the graphics processing unit (GPU) initiates 1024 threads to perform Fast Fourier Transform (FFT) in parallel. Each thread is responsible for calculating the transform result for one frequency point. The threads are organized using a two-dimensional grid structure, with the first dimension corresponding to the symbol index and the second dimension corresponding to the subcarrier index. During the calculation, intermediate results are exchanged within shared memory within each thread block, reducing the number of global memory accesses. Special function units are responsible for performing complex multiplication and addition operations to improve computational efficiency.
[0073] Step 52: The data is transmitted back to the central processing unit via PCIe (Peripheral Component Interconnect Standard Extended Bus).
[0074] In one implementation, the transfer process is initiated by the memory controller of the graphics processing unit (GPU). The controller first checks the validity of the target memory address and then initiates a direct memory access transfer. A credit mechanism is used for flow control during the transfer to prevent buffer overflows at the receiving end. To improve transfer efficiency, the system supports a linked descriptor mode, which can automatically execute the transfer of multiple non-contiguous data blocks. After the transfer is complete, the GPU notifies the central processing unit (CPU) via a doorbell register.
[0075] Step 6: Connect the software-defined radio platform to the RF front end.
[0076] Specifically, the platform interconnects with the RF front-end via standardized hardware interfaces. These interfaces employ a modular design, supporting hot-swapping and automatic identification. The RF front-end's low-noise amplifier utilizes a two-stage structure: the first stage performs impedance matching and noise optimization, while the second stage provides programmable gain. The power amplifier operates in Class AB mode, with output power adjustable from 10mW to 10W. The platform configures RF front-end parameters, including operating frequency band, bandwidth, and gain settings, via a digital control bus.
[0077] Step 61: The radio frequency front end converts the radio signal into an intermediate frequency signal or a baseband signal suitable for subsequent processing.
[0078] In one embodiment, the down-conversion process employs a superheterodyne structure. The RF signal first undergoes a bandpass filter to suppress out-of-band interference; then it is mixed with the local oscillator signal to generate an intermediate frequency (IF) signal; finally, it passes through anti-aliasing filtering and automatic gain control to output a stable IF signal. The up-conversion process is the reverse, shifting the IF signal to the target RF band. The entire conversion process uses digital phase-locked loop (PLL) technology to ensure frequency stability and phase noise performance.
[0079] Step 62: The digital-to-analog converter converts the digital signal into an analog signal.
[0080] For example, the converter employs a segmented architecture, comprising both thermometer encoding and binary encoding. It achieves 14-bit conversion accuracy and supports sampling rates up to 200 MS / s. The converter integrates a reconstruction filter and output driver, allowing direct connection to a power amplifier. Key performance parameters include a spurious-free dynamic range greater than 80 dB and a signal-to-noise ratio exceeding 72 dB. The conversion process is compensated for using digital predistortion technology to improve linearity.
[0081] Step 63: The general radio software operation control software coordinates the operation of the central processing unit and graphics processing unit module.
[0082] In one possible implementation, the control software employs a three-tier architecture: the application layer implements the user interface and business logic; the middleware layer provides device abstraction and resource management; and the driver layer directly manipulates hardware resources. During software runtime, the resource manager dynamically monitors the load of the CPU and GPU, allocating computing resources based on task characteristics. The task scheduler uses a priority queue to manage pending tasks, ensuring that tasks with high real-time requirements are executed first.
[0083] Step 7: High-speed orthogonal data streams correspond to various communication systems and modulation methods.
[0084] Specifically, the system supports mainstream wireless communication standards, including mobile communication, broadcasting, and private network communication. Each communication system corresponds to a specific frame structure, coding scheme, and modulation parameters. For example, the LTE system uses orthogonal frequency division multiplexing (OFDM) with a subcarrier spacing of 15kHz; the second-generation digital video broadcasting satellite system uses orthogonal phase shift keying (QPSK) modulation with a symbol rate of up to 30MBaud. The system defines parameter sets for various systems through configuration files, supporting fast switching and dynamic reconfiguration.
[0085] Step 71 involves various communication systems, including LTE, 5G radio, and second-generation digital video broadcasting satellites.
[0086] In one embodiment, the LTE (Long Term Evolution) architecture supports variable bandwidth from 1.4MHz to 20MHz and employs Orthogonal Frequency Division Multiple Access (OFDMA). The fifth-generation radio architecture supports millimeter-wave bands with a maximum bandwidth of 400MHz and uses cyclic prefix OFDMA waveforms. The second-generation digital video broadcasting satellite architecture uses QPSK and 8PSK modulation and supports forward error correction coding and interleaving. The system switches between different architectures through software configuration while maintaining a unified hardware platform.
[0087] Step 72, various modulation methods including binary phase shift keying, quadrature phase shift keying, quadrature amplitude modulation, etc.
[0088] For example, binary phase shift keying uses differential coding and supports data rates from 1 Mbps to 10 Mbps. Quadrature phase shift keying has higher spectral efficiency, with each symbol carrying 2 bits of information. Quadrature amplitude modulation supports various configurations such as 16QAM and 64QAM, and reduces the peak-to-average power ratio through constellation diagram rotation technology. The system dynamically selects the optimal modulation scheme based on channel conditions and service requirements, balancing transmission efficiency and reliability.
[0089] Step 8: The central processing unit's main thread enters the CUDA architecture (Compute Unified Device Architecture) global function.
[0090] In one implementation, global functions implement the core algorithms of the signal processing chain. For example, in the Fast Fourier Transform calculation, the function uses a radix-2 algorithm, optimizing data access patterns through shared memory. Each thread is responsible for calculating one butterfly operation unit, and synchronization operations within the thread block ensure the calculation order. During function execution, the texture buffer of the graphics processing unit is automatically used to accelerate data reading, while special function units accelerate transcendental function calculations. The calculation results ensure consistency across multiple threads through atomic operations.
[0091] Step 81: The CUDA architecture (Compute Unified Device Architecture) programming language compiler builder processes the digital signal processing part.
[0092] Specifically, the compiler translates algorithms described in high-level languages into machine code executable by the graphics processing unit (GPU). The compilation process comprises three stages: front-end analysis, intermediate optimization, and back-end code generation. Front-end analysis extracts data parallelism from the algorithm; intermediate optimization implements register allocation and instruction scheduling; and back-end code generation optimizes the instruction sequence for the specific GPU architecture. The compiler supports various optimization options, enabling the generation of highly optimized code for computationally intensive tasks.
[0093] Step 82: Build program compilation tools compatible with general radio software and CUDA architecture (Compute Unified Device Architecture) platform.
[0094] In one embodiment, the compilation tool runs as a plug-in to general-purpose radio software. Developers write signal processing algorithms in the software environment, marking graphics processing unit (GPU) acceleration regions with specific annotations. The compilation tool automatically extracts these regions to generate GPU computational kernels while preserving the central processing unit (CPU) control logic. The resulting executable file contains both CPU code and GPU code, which are executed in a coordinated manner by the runtime system.
[0095] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A dedicated signal generating device based on software-defined radio and a GPU server, characterized in that, include: The graphics processing unit is used to generate high-speed orthogonal data streams corresponding to various communication systems and modulation methods. The high-speed orthogonal data streams are transmitted to the intermediate frequency signal unit through optical fiber. The intermediate frequency signal unit is used to receive the high-speed orthogonal data stream and perform signal modulation to output an intermediate frequency signal; A software-defined radio platform is used for secondary development to realize signal generation and reception. The secondary development is based on a heterogeneous system integrating a central processing unit and a graphics processing unit built on general radio software. The central processing unit is used for task scheduling and memory configuration, and the graphics processing unit is used for processing the high-speed orthogonal data stream. The main thread of the central processing unit is used to transfer the data to be processed to the memory of the graphics processing unit via PCIe. The graphics processing unit is used to call multiple threads to process the data to be processed, and the processed data is then transmitted back to the central processing unit via PCIe.
2. The dedicated signal generating device based on software-defined radio and GPU server according to claim 1, characterized in that, The graphics processing unit generates high-speed orthogonal data streams corresponding to various communication systems and modulation methods, including: General-purpose radio software creates custom modules and links CUDA architecture signal processing programs, which use CUDA architecture programming languages to design dynamic link libraries. The custom module sets up a functional module linked to the graphics processing unit in the general radio software flowchart. When the functional module runs, it enters the CUDA architecture global function to execute the high-speed orthogonal data stream processing. The central processing unit processes signal data sequentially according to the system flowchart, and the system flowchart includes the functional modules linked to the graphics processing unit. The PCIe transfers the data to be processed to the memory of the graphics processing unit (GPU), and the GPU memory stores the data to be processed for use by multiple threads.
3. The dedicated signal generating device based on software-defined radio and GPU server according to claim 1, characterized in that: The software-defined radio platform is further developed to achieve signal generation and reception, including: The software-defined radio platform is connected to a radio frequency front-end, which includes a low-noise amplifier and a power amplifier. The radio frequency front end receives radio signals and converts them into intermediate frequency signals or baseband signals, and the intermediate frequency signals or baseband signals are transmitted to the analog-to-digital converter. The analog-to-digital converter converts the analog signal into a digital signal, and the digital signal is transmitted to the central processing unit-graphics processing unit heterogeneous system.
4. A dedicated signal generating device based on software-defined radio and a GPU server according to claim 1, characterized in that: The intermediate frequency (IF) signal unit receives the high-speed quadrature data stream and performs signal modulation to output an IF signal, including: The intermediate frequency signal unit acquires the high-speed orthogonal data stream transmitted through the optical fiber; The intermediate frequency signal unit modulates the high-speed orthogonal data stream, and the modulation process generates an intermediate frequency signal. The intermediate frequency signal unit outputs the intermediate frequency signal to the radio frequency front end, and the radio frequency front end includes a filter and an antenna.
5. A dedicated signal generating device based on software-defined radio and a GPU server according to claim 1, characterized in that: The central processing unit performs task scheduling and memory allocation, and the graphics processing unit performs the high-speed orthogonal data stream processing, including: After receiving the high-speed wireless communication signal, the central processing unit processes the signal data according to the system flowchart. When the system flowchart reaches the functional module linked to the graphics processing unit, the graphics processing unit is invoked. The central processing unit's main thread enters the CUDA architecture global function, which identifies the graphics processing unit's execution function. After the graphics processing unit finishes processing, the data is transmitted back to the main thread of the central processing unit to continue running.
6. A dedicated signal generating device based on software-defined radio and a GPU server according to claim 3, characterized in that: The software-defined radio platform is connected to the radio frequency front end, including: The radio frequency front end converts radio signals into intermediate frequency signals or baseband signals suitable for subsequent processing; The digital-to-analog converter converts digital signals into analog signals, which are then transmitted to the power amplifier. The general-purpose radio software operation control software coordinates the operation of the central processing unit and the graphics processing unit module, and the control software manages system resources.
7. A dedicated signal generating device based on software-defined radio and a GPU server according to claim 1, characterized in that: The high-speed orthogonal data stream corresponds to various communication systems and modulation methods, including: The various communication systems mentioned include Long Term Evolution (LTE), fifth-generation radio, second-generation digital video broadcasting satellite, second-generation digital video broadcasting return link, discrete Fourier transform single-carrier orthogonal frequency division multiplexing (DFM), and cyclic prefix orthogonal frequency division multiplexing (CFM). The various modulation methods include binary phase shift keying, π / 2 binary phase shift keying, quadrature phase shift keying, octet phase shift keying, 16 amplitude phase keying, 32 amplitude phase keying, 16 quadrature amplitude modulation, and 64 quadrature amplitude modulation.
8. A dedicated signal generating device based on software-defined radio and a GPU server according to claim 5, characterized in that: The central processing unit's main thread enters CUDA architecture global functions, including: The CUDA architecture programming language builder compilation tool handles the digital signal processing part; The build program compilation tool is compatible with general-purpose radio software and CUDA architecture platforms; The graphics processing unit performs filtering, modulation, and encoding processing on the high-speed orthogonal data stream, and the filtering, modulation, and encoding processing generates processed data.