A heterogeneous computing signal processing method and system for high-speed communication scenarios

CN122554288BActive Publication Date: 2026-09-25SHANGHAI FUHUA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611025945.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25
Estimated Expiration
2046-07-10

AI Technical Summary

Technical Problem

这种静态划分方式无法根据信道条件的动态变化以及各计算单元实时负载的差异,灵活调整子任务与计算单元之间的对应关系,导致当信道特征与固定分配规则不匹配或某一计算单元出现拥塞时,整体处理效率和资源利用率下降

Benefits of technology

接收高速通信场景下的原始信号流,对所述原始信号流进行帧同步检测与信道粗估计,生成待处理的符号序列,将所述符号序列切分为多个子任务;获取各计算单元的实时负载状态值,将各个子任务分别派发到现场可编程门阵列、图形处理器和中央处理器上执行;在现场可编程门阵列上完成物理层前端运算,在图形处理器上进行信道补偿与软比特提取,在中央处理器上处理协议相关逻辑;将现场可编程门阵列、图形处理器和中央处理器各自得到的结果进行拼接,输出最终解调数据流。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554288B_ABST
    Figure CN122554288B_ABST
Patent Text Reader

Abstract

The application provides a heterogeneous computing signal processing method and system for a high-speed communication scene, relates to the technical field of signal processing, receives an original signal stream in a high-speed communication scene, performs frame synchronization detection and channel coarse estimation on the original signal stream, generates a symbol sequence, and cuts the symbol sequence into multiple subtasks; each subtask is respectively dispatched to a field programmable gate array, a graphics processor and a central processing unit for execution; the physical layer front-end operation is completed on the field programmable gate array, channel compensation and soft bit extraction are performed on the graphics processor, and protocol-related logic is processed on the central processing unit; the results obtained by the field programmable gate array, the graphics processor and the central processing unit are spliced, and the final demodulation data stream is output. The application can realize efficient collaborative use of heterogeneous resources, and improve system throughput and demodulation reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and more specifically, to a heterogeneous computing signal processing method and system for high-speed communication scenarios. Background Technology

[0002] In high-speed communication scenarios, receivers typically employ heterogeneous computing architectures to process signals, distributing different types of computational tasks across different processing units. Common heterogeneous computing units include Field-Programmable Gate Arrays (FPGAs), Graphics Processing Units (GPUs), and Central Processing Units (CPUs). FPGAs are commonly used to perform physical layer front-end operations such as digital down-conversion, matched filtering, and Fast Fourier Transform. Graphics Processing Units are often used to perform massively parallel computing tasks such as channel equalization, Multiple-Input Multiple-Output (MIMO) detection, and soft bit extraction. CPUs are used to perform control-intensive and highly serial processing logic such as descrambling, rate matching descrambling, hybrid automatic repeat request merging, turbo decoding, and protocol stack parsing.

[0003] In high-speed communication scenarios, existing heterogeneous computing signal processing methods typically employ static or fixed-rule task partitioning. This means that specific types of processing tasks are pre-assigned to Field-Programmable Gate Arrays (FPGAs), Graphics Processing Units (GPUs), or Central Processing Units (CPUs). For example, Fast Fourier Transform (FFT) is pre-assigned to FPGAs, channel equalization to GPUs, and protocol parsing to CPUs. This static partitioning approach cannot flexibly adjust the correspondence between subtasks and computing units based on dynamic changes in channel conditions and differences in real-time load across computing units. Consequently, when channel characteristics do not match the fixed allocation rules or when a computing unit experiences congestion, overall processing efficiency and resource utilization decrease. Therefore, achieving efficient collaborative utilization of heterogeneous resources to improve system throughput and demodulation reliability is a challenge facing the industry. Summary of the Invention

[0004] This application provides a heterogeneous computing signal processing method and system for high-speed communication scenarios, which can realize the efficient collaborative utilization of heterogeneous resources and improve system throughput and demodulation reliability.

[0005] In a first aspect, this application provides a heterogeneous computing signal processing method for high-speed communication scenarios, the signal processing method comprising the following steps: Receive the raw signal stream in a high-speed communication scenario, perform frame synchronization detection and coarse channel estimation on the raw signal stream, generate a symbol sequence to be processed, and divide the symbol sequence into multiple sub-tasks; The real-time load status values ​​of each computing unit are obtained, and each subtask is dispatched to the field programmable gate array, graphics processor and central processing unit for execution. The physical layer front-end operations are performed on the field-programmable gate array (FPGA), channel compensation and soft bit extraction are performed on the graphics processor (GPU), and protocol-related logic is processed on the central processing unit (CPU). The results obtained from the field-programmable gate array, graphics processor, and central processing unit are combined to output the final demodulated data stream.

[0006] In this embodiment, performing frame synchronization detection and coarse channel estimation on the original signal stream to generate the symbol sequence to be processed specifically includes: The original signal stream is correlated with the local frame header sequence, and the correlation peak position is used as the frame start boundary. A complete frame length data is extracted from the original signal stream, and the pilot symbol segment in the complete frame length data is estimated by least squares to obtain the coarse channel response value on each subcarrier. The coarse channel response value is used to perform single-tap equalization on data symbols within the same complete frame length to obtain pre-equalized symbols. The symbols after preliminary equalization are hard-determined to obtain decision symbols. The residual phase error is extracted based on the decision symbols and the data symbols before single-tap equalization. Then, the common phase rotation correction is performed on the symbols after preliminary equalization based on the residual phase error. The phase-corrected symbol sequence is split into orthogonal frequency division multiplexing symbol blocks to generate the symbol sequence to be processed.

[0007] In this embodiment, extracting the residual phase error based on the decision symbol and the data symbol before single-tap equalization, and then performing common phase rotation correction on the symbol after preliminary equalization based on the residual phase error specifically includes: The data symbols before single-tap equalization are multiplied by the decision symbols using conjugate multiplication to obtain the phase difference measurement value corresponding to each data symbol; The phase difference measurements of all data symbols within the same orthogonal frequency division multiplexing symbol block are weighted and averaged. The weighted average is used as the residual phase error estimate of the corresponding orthogonal frequency division multiplexing symbol block; Based on the estimated residual phase error, a common rotation factor is constructed. All data symbols belonging to the same orthogonal frequency division multiplexing symbol block in the pre-equalized symbols are multiplied by the common rotation factor to complete the common phase rotation correction.

[0008] In this embodiment, dividing the symbol sequence into multiple sub-tasks specifically includes: Obtain the signal-to-noise ratio and phase change rate within a time window near each symbol in the symbol sequence; The signal-to-noise ratio and the phase change rate are weighted together to form a tendency coefficient; Based on the change curve of the tendency coefficient along the time axis, a splitting point is set at the position where the tendency coefficient crosses the threshold, and the symbol sequence is cut into blocks of unequal length. The tendency coefficient fluctuation range within each block is kept within a preset tolerance value, and each block is treated as an independent subtask.

[0009] In this embodiment, obtaining the real-time load status values ​​of each computing unit and dispatching each subtask to the field-programmable gate array, graphics processor, and central processing unit for execution specifically includes: At each dispatch time, the real-time load status values ​​of the field-programmable gate array, graphics processor, and central processing unit are obtained respectively; When the tendency coefficient is higher than a preset threshold and the real-time load state value of the field-programmable gate array is lower than the real-time load state value of the graphics processor, the data is dispatched to the field-programmable gate array. When the tendency coefficient is lower than a preset threshold and the real-time load state value of the graphics processor is lower than the real-time load state value of the field-programmable gate array, the data is dispatched to the graphics processor. If neither of the above two conditions is met, the task is assigned to the computing unit with the lowest real-time load status value. When the execution of the current subtask depends on control plane information in the central processing unit, it is dispatched directly to the central processing unit.

[0010] In this embodiment, the real-time load status value refers to the ratio of the number of subtasks currently being processed by each computing unit to the maximum concurrent processing capacity of the corresponding computing unit.

[0011] In this embodiment, performing physical layer front-end operations on a field-programmable gate array specifically includes: The symbol data corresponding to the subtasks dispatched to the field programmable gate array are sequentially subjected to digital downconversion, matched filtering, and fast Fourier transform to obtain a frequency domain symbol sequence. Perform a detection operation on the frequency domain symbol sequence and set the completion flag.

[0012] In this embodiment, channel compensation and soft bit extraction on the graphics processor specifically include: The symbolic data corresponding to the subtasks dispatched to the graphics processor is copied from main memory to the graphics processor's global memory. The graphics processor starts multiple thread blocks, each thread block is responsible for channel compensation calculation on a subcarrier, and equalizes all symbols on the same subcarrier according to the coarse channel response value; The log-likelihood ratio of each bit of the equalized symbol is determined according to the modulation order. The log-likelihood ratios of all bits are then arranged by subcarrier and symbol number and written back to main memory.

[0013] In this embodiment, the final demodulated data stream refers to a binary bit sequence that can be directly delivered to the media access control layer for parsing after being demodulated and decoded by the physical layer.

[0014] Secondly, this application provides a heterogeneous computing signal processing system for high-speed communication scenarios, used to execute a heterogeneous computing signal processing method for high-speed communication scenarios, the signal processing system comprising: The signal receiving module is used to receive the raw signal stream in a high-speed communication scenario, perform frame synchronization detection and coarse channel estimation on the raw signal stream, generate a symbol sequence to be processed, and divide the symbol sequence into multiple sub-tasks. The task dispatch module is used to obtain the real-time load status value of each computing unit and dispatch each subtask to the field programmable gate array, graphics processor and central processing unit for execution. The task processing module is used to perform physical layer front-end operations on the field-programmable gate array, perform channel compensation and soft bit extraction on the graphics processor, and process protocol-related logic on the central processing unit. The fusion output module is used to combine the results obtained from the field-programmable gate array, graphics processor and central processing unit to output the final demodulated data stream.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The system receives the raw signal stream from a high-speed communication scenario, performs frame synchronization detection and coarse channel estimation on the raw signal stream, generates a symbol sequence to be processed, and divides the symbol sequence into multiple sub-tasks; obtains the real-time load status values ​​of each computing unit, and dispatches each sub-task to a field-programmable gate array (FPGA), a graphics processing unit (GPU), and a central processing unit (CPU) for execution; performs physical layer front-end operations on the FPGA, performs channel compensation and soft bit extraction on the GPU, and processes protocol-related logic on the CPU; and concatenates the results obtained by the FPGA, GPU, and CPU to output the final demodulated data stream.

[0016] Therefore, this application first generates a symbol sequence to be processed by performing frame synchronization detection and coarse channel estimation on the original signal stream, and then divides the symbol sequence into multiple sub-tasks to provide task units for subsequent heterogeneous computing, laying the foundation for heterogeneous parallel processing. Second, by obtaining the real-time load status values ​​of each computing unit, each sub-task is dispatched to the field-programmable gate array (FPGA), graphics processing unit (GPU), and central processing unit (CPU) for execution, realizing dynamic matching between sub-tasks and heterogeneous computing units, improving the overall resource utilization efficiency and task processing throughput of the heterogeneous computing system. Then, by executing the FPGA, GPU, and CPU in parallel, the performance bottleneck of a single processor handling all types of tasks is avoided, reducing the overall processing latency and improving the system throughput. Finally, by splicing and fusing the hard-decision bits, soft-decision bits, and protocol data units output by the FPGA, GPU, and CPU, the final demodulated data stream is generated by using cross-validation and confidence level judgment of multi-source results, reducing the demodulation error rate and improving the reliability of the final data stream.

[0017] In summary, the technical solution adopted in this application can achieve efficient collaborative utilization of heterogeneous resources and improve system throughput and demodulation reliability. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is an exemplary flowchart of a heterogeneous computing signal processing method for high-speed communication scenarios provided in this application; Figure 2 This is a schematic diagram of the signal flow of a heterogeneous computing signal processing system for high-speed communication scenarios provided in this application; Figure 3 This is a module structure diagram of a heterogeneous computing signal processing system for high-speed communication scenarios provided in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0021] This application provides a heterogeneous computing signal processing method and system for high-speed communication scenarios. Its core is to receive the raw signal stream from the high-speed communication scenario, perform frame synchronization detection and coarse channel estimation on the raw signal stream, generate a symbol sequence to be processed, and divide the symbol sequence into multiple sub-tasks; obtain the real-time load status values ​​of each computing unit, and dispatch each sub-task to a field-programmable gate array (FPGA), a graphics processing unit (GPU), and a central processing unit (CPU) for execution; perform physical layer front-end operations on the FPGA, perform channel compensation and soft bit extraction on the GPU, and process protocol-related logic on the CPU; and concatenate the results obtained from the FPGA, GPU, and CPU to output the final demodulated data stream.

[0022] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown in the figure, this is an exemplary flowchart of a heterogeneous computing signal processing method for high-speed communication scenarios according to this embodiment of the present application. The signal processing method includes the following steps: In step S1, the raw signal stream in a high-speed communication scenario is received, frame synchronization detection and coarse channel estimation are performed on the raw signal stream, a symbol sequence to be processed is generated, and the symbol sequence is divided into multiple sub-tasks.

[0023] In practical implementation, the receiver receives the original signal stream in a high-speed communication scenario. Specifically, in a high-speed communication scenario, after the receiver receives the radio frequency (RF) signal through the antenna, it first amplifies the signal using a low-noise amplifier, and then converts the RF signal into an intermediate frequency (IF) signal or a baseband signal using a down-converter. The down-converted analog signal is then sent to an analog-to-digital converter (ADC) for sampling at a rate higher than the Nyquist sampling rate, converting the analog signal into a digital signal to obtain the original signal stream. The original signal stream is a complex sequence arranged in time order, with each complex number consisting of in-phase and quadrature components.

[0024] In this embodiment, frame synchronization detection and coarse channel estimation are performed on the original signal stream to generate a symbol sequence to be processed. Specifically, this can be achieved through the following steps: The original signal stream is correlated with the local frame header sequence, and the correlation peak position is used as the frame start boundary. A complete frame length data is extracted from the original signal stream, and the pilot symbol segment in the complete frame length data is estimated by least squares to obtain the coarse channel response value on each subcarrier. The coarse channel response value is used to perform single-tap equalization on data symbols within the same complete frame length to obtain pre-equalized symbols. The symbols after initial equalization are hard-determined to obtain decision symbols. The residual phase error is extracted based on the decision symbols and the data symbols before single-tap equalization. Then, the common phase rotation correction is performed on the symbols after initial equalization based on the residual phase error. The phase-corrected symbol sequence is split into orthogonal frequency division multiplexing symbol blocks to generate the symbol sequence to be processed.

[0025] In practical implementation, firstly, a sliding correlation is performed between the original signal stream and the local frame header sequence, using the correlation peak position as the frame start boundary. That is, the local frame header sequence is pre-stored in the receiver's read-only memory and is completely consistent with the frame header agreed upon by the transmitter, employing a pseudo-random sequence with good autocorrelation characteristics. At each sampling time, the sliding correlator performs a complex correlation operation between a continuous segment of data from the original signal stream starting at the current sampling point and the local frame header sequence, outputting a correlation value. When the magnitude of the correlation value exceeds a preset threshold, the position of the correlation peak is taken as the frame start boundary. The preset threshold can be set through expert experience. Next, a complete frame length is extracted from the original signal stream. Least squares estimation is performed on the pilot symbol segments in the complete frame length data to obtain a coarse channel response value on each subcarrier. That is, a complete frame length is extracted from the original signal stream, starting from the frame start boundary and extending forward by a predefined frame length. The complete frame length data includes pilot symbol segments and data symbol segments. The position of the pilot symbol segments in the frame structure and the specific value of each pilot symbol are predetermined by the communication protocol. For each pilot subcarrier, the received pilot symbol is divided by the locally known transmitted pilot symbol, and the quotient is used as the coarse channel response value on that subcarrier. After collecting the coarse channel response values ​​on all pilot subcarriers, for data subcarriers in non-pilot positions, a linear interpolation method is used to calculate the coarse channel response value on each data subcarrier from the coarse channel response values ​​of adjacent pilot subcarriers. Next, single-tap equalization is performed on the data symbols within the same complete frame length using the coarse channel response values ​​to obtain the pre-equalized symbols. That is, the received symbol on each data subcarrier is divided by its corresponding coarse channel response value, and the quotient is the pre-equalized symbol. Then, a hard decision is made on the pre-equalized symbols to obtain the decision symbol. The residual phase error is extracted based on the decision symbol and the data symbol before single-tap equalization. Then, a common phase rotation correction is performed on the pre-equalized symbols based on the residual phase error. That is, the Euclidean distance from the pre-equalized symbol to each standard constellation point on the constellation diagram is calculated based on the modulation order of the current data symbol, and the constellation point with the closest distance is selected as the decision symbol. This decision symbol can be regarded as a preliminary estimate of the original transmitted symbol. The modulation order can be quadrature phase shift keying, sixteen quadrature amplitude modulation, or sixty-four quadrature amplitude modulation, etc. The decision process uses the corresponding constellation diagram according to different modulation orders.The residual phase error is extracted from the decision symbol and the data symbol before single-tap equalization. Then, the common phase rotation correction is performed on the symbols after preliminary equalization based on the residual phase error, which will be explained in detail in subsequent steps. Finally, the phase-corrected symbol sequence is split into orthogonal frequency division multiplexing (OFDM) symbol blocks to generate the symbol sequence to be processed. That is, the phase-corrected symbol sequence is split into OFDM symbol blocks, the cyclic prefix at the beginning of each OFDM symbol block is removed, and then the remaining valid symbols are arranged into a two-dimensional matrix according to the subcarrier index from smallest to largest and according to the reception time of the OFDM symbol blocks. The rows of the two-dimensional matrix correspond to the subcarriers, and the columns correspond to the sequence numbers of the OFDM symbol blocks. This two-dimensional matrix is ​​the symbol sequence to be processed.

[0026] In this embodiment, the residual phase error is extracted based on the decision symbol and the data symbol before single-tap equalization. Then, a common phase rotation correction is performed on the symbol after preliminary equalization based on the residual phase error. Specifically, this can be achieved through the following steps: The data symbols before single-tap equalization are multiplied by the decision symbols using conjugate multiplication to obtain the phase difference measurement value corresponding to each data symbol; The phase difference measurements of all data symbols within the same orthogonal frequency division multiplexing symbol block are weighted and averaged. The weighted average is used as the residual phase error estimate of the corresponding orthogonal frequency division multiplexing symbol block; Based on the estimated residual phase error, a common rotation factor is constructed. All data symbols belonging to the same orthogonal frequency division multiplexing symbol block in the pre-equalized symbols are multiplied by the common rotation factor to complete the common phase rotation correction.

[0027] In the specific implementation, firstly, the data symbols before single-tap equalization and the decision symbols are multiplied by their conjugates to obtain the phase difference measurement value corresponding to each data symbol. That is, for each data symbol, the receiver stores the original data symbol before single-tap equalization. The original data symbol is a complex number extracted from the complete frame length data of the original signal stream, without any equalization processing. At the same time, the decision symbol obtained from the previous hard decision is also a complex number, representing an estimate of the original transmitted symbol. The original data symbol and the decision symbol are multiplied by their conjugates, and the result of the conjugate multiplication is a new complex number. The argument of this complex number is the phase difference measurement value between the original data symbol and the decision symbol. Secondly, the phase difference measurements corresponding to all data symbols within the same orthogonal frequency division multiplexing (OFDM) symbol block are weighted and averaged. That is, for each data symbol within the same OFDM symbol block, the magnitude of the coarse channel response value corresponding to the symbol is taken as a weighting coefficient. The phase difference measurement value of the symbol is multiplied by the weighting coefficient to obtain the weighted phase difference of the symbol. The weighted phase differences of all data symbols are summed, and the sum is divided by the sum of all weighting coefficients to obtain the weighted average phase difference. Then, the weighted average is used as the residual phase error estimate for the corresponding orthogonal frequency division multiplexing (OFDM) symbol block. This residual phase error estimate is a scalar angle value representing the average phase rotation experienced by all data symbols within the current OFDM symbol block. Finally, a common rotation factor is constructed based on the residual phase error estimate. All data symbols belonging to the same OFDM symbol block in the pre-equalized symbols are multiplied by this common rotation factor to complete the common phase rotation correction. This means the residual phase error estimate is an angle value, and the common rotation factor is constructed as a complex number with a magnitude of 1 and an argument equal to the negative residual phase error estimate. In other words, the rotation direction of the common rotation factor in the complex plane is opposite to the direction of the residual phase error and equal in magnitude. All pre-equalized symbols belonging to the same orthogonal frequency division multiplexing (OFDM) symbol block are extracted one by one. Each symbol is multiplied by a common rotatant factor. The real part of the common rotatant factor is multiplied by the real part of the symbol, and then the imaginary part of the common rotatant factor is multiplied by the imaginary part of the symbol to obtain the real part of the new symbol. The real part of the common rotatant factor is multiplied by the imaginary part of the symbol, and then the imaginary part of the common rotatant factor is multiplied by the real part of the symbol to obtain the imaginary part of the new symbol. After the above multiplication operation, the common phase rotation existing in the original symbol is subtracted, and the constellation diagram of the symbol is rotated back to the correct position, thus completing the common phase rotation correction.

[0028] In this embodiment, the symbol sequence is divided into multiple sub-tasks, which can be achieved through the following steps: Obtain the signal-to-noise ratio and phase change rate within a time window near each symbol in the symbol sequence; The signal-to-noise ratio and the phase change rate are weighted together to form a tendency coefficient; Based on the change curve of the tendency coefficient along the time axis, a splitting point is set at the position where the tendency coefficient crosses the threshold, and the symbol sequence is cut into blocks of unequal length. The tendency coefficient fluctuation range within each block is kept within a preset tolerance value, and each block is treated as an independent subtask.

[0029] In practical implementation, firstly, the signal-to-noise ratio (SNR) and phase change rate within a time window near each symbol in the symbol sequence are obtained. That is, the segmentation operation is performed along the time axis, processing each symbol sequentially in the direction of increasing Orthogonal Frequency Division Multiplexing (OFDM) symbol block number. For the currently processed symbol, a fixed length is extended forward and backward from it, forming a time window. The window length can be dynamically adjusted according to the channel fading rate: a smaller window length is used when the channel fading rate is fast to capture rapid fluctuations in SNR and phase change rate; a larger window length is used when the channel fading rate is slow to improve the stability of the statistical estimation. The specific setting can be based on expert recommendations. Within this time window, the SNR is calculated by averaging the signal power and noise power of all symbols within the window. Dividing the average signal power by the average noise power yields the estimated SNR. Within this time window, the phase change rate is calculated by taking the phase difference between two adjacent symbols within the window, averaging the absolute values ​​of all adjacent phase differences, and then dividing the average phase difference by the symbol period to obtain the phase change rate. The phase change rate refers to the average magnitude of phase change per unit time, reflecting the degree of time-selective fading in the channel. A larger phase change rate indicates faster channel changes. Next, the signal-to-noise ratio (SNR) and phase change rate are weighted to form a bias coefficient. That is, the bias coefficient equals the SNR multiplied by a first weight minus the phase change rate multiplied by a second weight, plus an offset constant. The first and second weights can be dynamically adjusted according to the receiver's operating mode. When the receiver is in a low SNR mode, the first weight is increased to make the SNR dominate the bias coefficient; when the receiver is in a high mobility mode, the second weight is increased to make the phase change rate dominate the bias coefficient. The offset constant is used to adjust the bias coefficient to an appropriate numerical range. The first weight, second weight, and offset constant can be set with expert advice. The bias coefficient is a dimensionless scalar value used to quantitatively reflect whether the current symbol or its associated data block is more suitable for processing by a field-programmable gate array (FPGA) or a graphics processor (GPU).

[0030] Furthermore, in practical implementation, based on the trend curve of the tendency coefficient along the time axis, split points are set at the locations where the tendency coefficient crosses a threshold, dividing the symbol sequence into blocks of unequal length. That is, each symbol is arranged according to its chronological order in the symbol sequence, and a trend curve of the tendency coefficient over time is plotted with the symbol number as the horizontal axis and the tendency coefficient of each symbol as the vertical axis. This curve reflects the changing trend of channel characteristics on the time axis. A preset threshold value is used; the specific size of the threshold value depends on the performance comparison between the field-programmable gate array (FPGA) and the graphics processor (GPU) and can be set by expert experience. The tendency coefficient curve is scanned from left to right along the time axis. When the tendency coefficient changes from above the threshold to below the threshold, a split point is set at the point of change; when the tendency coefficient changes from below the threshold to above the threshold, another split point is set at the point of change. The location of the split point is the boundary between the two subtasks. Based on the predefined splitting points, the original symbol sequence is divided into several continuous blocks. Finally, the fluctuation range of the tendency coefficient within each block is ensured to not exceed a preset tolerance value. Each block is treated as an independent subtask; that is, after splitting, each block is checked, and the difference between the maximum and minimum tendency coefficients of all symbols within the block is calculated as the fluctuation range of the tendency coefficient for that block. If the fluctuation range is less than or equal to the preset tolerance value, the block passes the check and is directly output as a subtask. If the fluctuation range of a block's tendency coefficient exceeds the preset tolerance value, the block needs to be further subdivided. The maximum and minimum tendency coefficient points within the block are found, and one or more additional splitting points are set between these two extreme points, splitting the original block into smaller sub-blocks until the fluctuation range of each sub-block meets the tolerance value requirement. The tolerance value can be configured according to the system's requirements for task granularity. The smaller the tolerance value is set, the finer the granularity of the subtasks will be, and the higher the scheduling flexibility will be, but the overhead of splitting and scheduling will also increase accordingly. The larger the tolerance value is set, the coarser the granularity of the subtasks will be, and the lower the overhead will be, but the load balancing effect may be worse. The specific setting can be determined based on experimental data analysis.

[0031] In step S2, the real-time load status value of each computing unit is obtained, and each subtask is dispatched to the field programmable gate array, graphics processor and central processing unit for execution.

[0032] In this embodiment, the real-time load status values ​​of each computing unit are obtained, and each subtask is dispatched to the field-programmable gate array, graphics processor, and central processing unit for execution. This can be achieved through the following steps: At each dispatch time, the real-time load status values ​​of the field-programmable gate array, graphics processor, and central processing unit are obtained respectively; When the tendency coefficient is higher than a preset threshold and the real-time load state value of the field-programmable gate array is lower than the real-time load state value of the graphics processor, the data is dispatched to the field-programmable gate array. When the tendency coefficient is lower than a preset threshold and the real-time load state value of the graphics processor is lower than the real-time load state value of the field-programmable gate array, the data is dispatched to the graphics processor. If neither of the above two conditions is met, the task is assigned to the computing unit with the lowest real-time load status value. When the execution of the current subtask depends on control plane information in the central processing unit, it is dispatched directly to the central processing unit.

[0033] In practical implementation, firstly, at each dispatch time, the real-time load status values ​​of the Field-Programmable Gate Array (FPGA), Graphics Processing Unit (GPU), and Central Processing Unit (CPU) are obtained respectively. That is, the real-time load status value refers to the ratio of the number of subtasks currently being processed by each computing unit to its maximum concurrent processing capacity. For the FPGA, the scheduler reads the internal status register of the FPGA, which records the number of currently executing subtasks, the digit signal processing unit (DSP) utilization rate, the lookup table (Lookup Table) utilization rate, and the block random access memory (BRAM) utilization rate. The scheduler takes the ratio of the number of currently executing subtasks to the maximum concurrent processing capacity as the real-time load status value of the FPGA. For the GPU, the scheduler queries the current device occupancy through the application programming interface provided by the GPU driver. The queried information includes the number of currently executing thread blocks, stream processor utilization rate, video memory utilization rate, and computing unit utilization rate. The scheduler takes the ratio of the number of currently executing thread blocks to the maximum number of thread blocks as the real-time load status value of the GPU. For the CPU, the scheduler reads the process scheduling information provided by the operating system. The information read includes the number of processes in the current run queue, the utilization rate of each logical core, the context switching frequency, and the interrupt handling load. The scheduler takes the ratio of the number of processes in the current run queue to the total number of logical cores as the real-time load status value of the CPU. Secondly, when the tendency coefficient is higher than a preset threshold and the real-time load status value of the FPGA is lower than that of the GPU, the task is dispatched to the FPGA. That is, when the tendency coefficient of the current subtask is higher than the threshold and the real-time load status value of the FPGA is lower than that of the GPU, dispatching the subtask to the FPGA can achieve the shortest waiting time and the highest execution efficiency.

[0034] In addition, in specific implementation, when the tendency coefficient is lower than a preset threshold and the real-time load value of the graphics processor is lower than the real-time load value of the field-programmable gate array (FPGA), the task is assigned to the graphics processor. That is, when the tendency coefficient of the current subtask is lower than the preset threshold and the real-time load value of the graphics processor is lower than the real-time load value of the FPGA, assigning the subtask to the graphics processor can achieve the shortest waiting time and the highest execution efficiency. The preset threshold can be set through expert experience. Then, when neither of the above two conditions is met, the task is assigned to the computing unit with the lowest real-time load value. That is, when neither of the above two conditions is met, the real-time load values ​​of the FPGA, the graphics processor, and the central processing unit are compared, and the computing unit with the lowest load value is selected, and the current subtask is assigned to that unit for execution. If multiple computing units have the same and minimum real-time load status values, they are selected according to a preset priority order. This priority order can be pre-set based on system energy consumption requirements and expert experience. Finally, when the execution of the current subtask depends on control plane information in the CPU, it is directly dispatched to the CPU. That is, in the communication protocol stack, some subtasks involve control plane operations such as protocol parsing, Hybrid Automatic Repeat Request (HARPR) status maintenance, and Radio Resource Control (RRC) signaling processing. These operations require access to protocol stack context information in the CPU's memory, including the status of the HARPR process, radio bearer configuration parameters, and scheduling authorization information. For these subtasks, the scheduler does not consider preference coefficients or load status during dispatch and directly dispatches them to the CPU for execution. After a subtask is dispatched to a selected computing unit, the scheduler writes the subtask's description information into the corresponding computing unit's task queue. The task queue adopts a first-in-first-out (FIFO) order to ensure the consistency of order among subtasks of the same type.

[0035] In step S3, physical layer front-end operations are performed on a field-programmable gate array (FPGA), channel compensation and soft bit extraction are performed on a graphics processor (GPU), and protocol-related logic is processed on a central processing unit (CPU).

[0036] In this embodiment, the physical layer front-end computation is performed on a field-programmable gate array (FPGA), which can be achieved through the following steps: The symbol data corresponding to the subtasks dispatched to the field programmable gate array are sequentially subjected to digital downconversion, matched filtering, and fast Fourier transform to obtain a frequency domain symbol sequence. Perform a detection operation on the frequency domain symbol sequence and set the completion flag.

[0037] In practice, the symbol data corresponding to the subtasks assigned to the field-programmable gate array (FPGA) is first subjected to digital down-conversion, matched filtering, and fast Fourier transform sequentially to obtain a frequency domain symbol sequence. Specifically, the FPGA integrates a numerically controlled oscillator (CNC), which generates sine and cosine digital sequences with the same carrier frequency as the transmitter. The input symbol data is multiplied by the sine and cosine sequences generated by the CNC, respectively, to obtain two mixing results. After the two mixing results are filtered by a low-pass filter to remove high-frequency components, a zero-IF baseband complex signal is output. The digitally down-converted baseband complex signal enters a matched filter. The impulse response of the matched filter is matched to the shaping filter at the transmitter, typically a root-raised cosine roll-off filter. In the FPGA, the matched filter is implemented using a finite impulse response (FIR) filter structure. The filter coefficients are pre-stored in read-only memory and selected and loaded according to the roll-off factor parameters of the current system. The symbol data sequentially passes through the tapped delay lines of the FIR filter. The output of each tap is multiplied by the corresponding filter coefficient, and all multiplications are accumulated to obtain the filtered output. The matched-filtered baseband signal enters the Fast Fourier Transform (FFT) module, which converts the time-domain symbols to the frequency domain. The FFT processor first removes the cyclic prefix from the input time-domain symbols, then performs radix-2 or radix-4 butterfly operations. Each butterfly operation includes complex multiplication, complex addition, and complex subtraction. After logarithmic-level calculations, a frequency-domain symbol sequence is output, with each sequence corresponding to a received symbol on a subcarrier. Then, a detection operation is performed on the frequency-domain symbol sequence, and a completion flag is set. The specific content of the detection operation depends on the physical channel type corresponding to the current subtask. For data channels, the detection operation refers to Multiple-Input Multiple-Output (MIMO) detection, which estimates the symbols transmitted by the transmitter on each antenna and subcarrier based on the frequency-domain symbols on multiple receiving antennas. For control channels, the detection operation includes sequence correlation detection. In MIMO detection, the Field-Programmable Gate Array (FPGA) performs minimum mean square error (MME) detection based on the channel matrix provided by the channel estimation module. After completing the detection operation, the FPGA writes the detection result back to a specified location in the Block Random Access Memory (BRAM) and then sets the completion flag in the status register.

[0038] In this embodiment, channel compensation and soft bit extraction are performed on the graphics processor, which can be implemented using the following steps: The symbolic data corresponding to the subtasks dispatched to the graphics processor is copied from main memory to the graphics processor's global memory. The graphics processor starts multiple thread blocks, each thread block is responsible for channel compensation calculation on a subcarrier, and equalizes all symbols on the same subcarrier according to the coarse channel response value; The log-likelihood ratio of each bit of the equalized symbol is determined according to the modulation order. The log-likelihood ratios of all bits are then arranged by subcarrier and symbol number and written back to main memory.

[0039] In the specific implementation, firstly, the symbol data corresponding to the subtask dispatched to the graphics processor is copied from main memory to the graphics processor's global memory. That is, after the scheduler prepares the symbol data corresponding to the subtask, it calls the data copy application programming interface through the graphics processor driver. This application programming interface triggers direct memory access transfer, transferring the symbol data from the designated buffer in the CPU's main memory to the graphics processor's global memory via the Fast Peripheral Component Interconnect Standard Bus. Then, the graphics processor starts multiple thread blocks, each thread block is responsible for channel compensation calculation on one subcarrier, and equalizes all symbols on the same subcarrier based on the coarse channel response value. That is, the number of thread blocks is equal to the total number of subcarriers contained in the corresponding subtask, and multiple threads are started within each thread block. The number of threads is equal to the number of symbols carried on each subcarrier. Each thread performs single-tap equalization on the corresponding received symbol based on the coarse channel response value, that is, divides the received symbol by the coarse channel response value on that subcarrier to obtain the equalized symbol. Finally, the log-likelihood ratio of each bit of the equalized symbol is determined according to the modulation order. The log-likelihood ratios of all bits are then arranged by subcarrier and symbol number and written back to main memory. That is, for quadrature phase shift keying modulation, the real part of the equalized symbol is taken as the log-likelihood ratio of the first bit, and the imaginary part is taken as the log-likelihood ratio of the second bit. For hexagonal amplitude modulation and higher orders, the maximum logarithmic approximation method is used to calculate the log-likelihood ratio of each bit. After all threads have completed the calculation, the log-likelihood ratio of each bit corresponding to each symbol on each subcarrier is arranged in order of subcarrier index and symbol number, written into the global memory of the graphics processor, and then transferred back to main memory through direct memory access.

[0040] In practical implementation, protocol-related logic is processed on the central processing unit (CPU). Specifically, the symbol data or soft bit data corresponding to the subtasks dispatched to the CPU is placed into a dedicated buffer within the CPU. This dedicated buffer uses a circular buffer structure and is managed by write and read pointers. The write pointer is updated by data write operations, and the read pointer is updated by data processing operations. The CPU retrieves data from the dedicated buffer and sequentially performs descrambling, rate matching descrambling, and hybrid automatic repeat request merging. The descrambling operation uses the same scrambling code sequence as the sender to perform an XOR operation on the data, recovering the bit sequence before scrambling. The rate matching descrambling operation restores the descrambled bit sequence to the original length output by the encoder based on the sender's rate matching parameters, inserting placeholders for punctured bit positions. The hybrid automatic repeat request merging operation merges the soft bit data of the current transport block with the historical soft bit data of the same transport block stored in the buffer. Merging methods include catch-up merging and incremental redundancy merging. The combined data from the Hybrid Automatic Repeat Request (HARRequest) is fed into a turbine decoder for iterative decoding. The turbine decoder uses a maximum a posteriori (MAP) algorithm for multiple iterations. Each iteration includes decoding by two component decoders and interleaving / deinterleaving between them. After reaching a preset limit, the decoder outputs a decided bit sequence. This decided bit sequence is assembled into a Media Access Control (MAC) protocol data unit (MAG) according to the protocol stack format. The MAG header is extracted from the decoded bitstream, and the logical channel identifier in the header determines which logical channel the MAG belongs to. The payload is then concatenated to form a complete MAG service data unit. The assembled MAG service data unit is delivered to upper-layer protocol processing, which includes: reordering and reassembly at the Radio Link Control (RANC) layer, header compression and decompression at the Packet Data Convergence (PaDC) layer, and signaling parsing at the Radio Resource Control (RRC) layer.

[0041] In step S4, the results obtained by the field-programmable gate array, the graphics processor, and the central processing unit are spliced ​​together to output the final demodulated data stream.

[0042] In practical implementation, the results obtained from the Field-Programmable Gate Array (FPGA), Graphics Processing Unit (GPU), and Central Processing Unit (CPU) are concatenated to output the final demodulated data stream. Specifically, the computational results are collected from the FPGA, GPU, and CPU respectively. The FPGA outputs a hard-determined bit sequence, with each subtask's output data carrying a timestamp recording the start and end positions of the corresponding symbol in the original signal stream. The GPU outputs a sequence of log-likelihood ratios for each bit, with each subtask's log-likelihood ratio sequence also carrying a timestamp in the same format to ensure time correspondence with the FPGA output. The CPU outputs Media Access Control (MAC) protocol data units, which are bit sequences that have passed cyclic redundancy check (CARC). Each protocol data unit carries a frame number or transport block number as a basis for time alignment. For the same data segment covered by all three outputs, the FPGA's hard-determined result is used as a benchmark, and the hard-determined direction corresponding to the GPU's soft information is compared bit by bit. If the two values ​​match, the result is used directly; otherwise, the absolute value of the log-likelihood ratio output by the graphics processor is checked. If the absolute value is higher than the preset threshold, the graphics processor's decision is used; if it is lower, the data block is returned to the central processing unit for turbo decoding. If the turbo decoding still fails the cyclic redundancy check, the data block, along with its adjacent blocks, is returned to the scheduler, forcing both the field-programmable gate array (FPGA) and the graphics processor to recalculate the data block. The two new results are then compared again before being output. All bit sequences that pass the check are concatenated in chronological order to form the final demodulated data stream. The preset threshold is obtained through offline simulation. Under typical channel conditions, soft decision-making is performed on the received data using different absolute value thresholds. The curve of bit error rate changing with the threshold is statistically analyzed, and the threshold corresponding to the lowest bit error rate is selected as the preset threshold.

[0043] It should be noted that adjacent blocks refer to the data block before and after the current data block, that is, the preceding and following subtasks that are temporally consecutive to the current data block. When dividing into subtasks, each subtask carries a timestamp and start and end position markers. The adjacent blocks can be uniquely identified by finding the predecessor and successor of the current data block's timestamp. When at the start or end of a frame, only one existing adjacent block is considered.

[0044] Additionally, it's important to note that the scheduler maintains a global task queue, while the task queues of each computing unit are independent. When recalculation is required, the scheduler removes the data block and its adjacent blocks from the original computing unit's task queue and inserts them as high-priority tasks at the head of each computing unit's task queue. This does not affect the normal enqueuing and scheduling of subsequent data blocks. After recalculation, the fusion output module inserts the results into the corresponding positions according to the timestamps before continuing the concatenation process.

[0045] Additionally, it's important to note that: First, the rollback operation is triggered only when the hard comparison and soft information are inconsistent and the confidence level is insufficient. This is an occasional event, not a systemic behavior, and will not occur frequently enough to cause continuous pipeline pauses. Second, recalculation will not block subsequent data processing. After the data block is marked as a high-priority task, it is recalculated in parallel by the FPGA and GPU. During this period, other data blocks in the normal pipeline continue to be processed on their respective computing units, with both processes running in parallel. While waiting for the recalculation result of this data block, the fusion output module skips that timestamp position and continues to assemble subsequent confirmed data blocks. Finally, regarding the avoidance of data conflicts: each data block has a unique timestamp identifier and start and end position information. The recalculation by the FPGA and GPU each has independent storage space. When the calculation result is written back, it is written to the corresponding position according to the timestamp, without causing write conflicts with normal pipeline tasks.

[0046] like Figure 2 As shown in the diagram, the heterogeneous computing signal processing system for high-speed communication scenarios provided by this solution includes a receiving end, a computing unit, and an output end. The original signal stream enters the receiving end, which performs frame synchronization detection, coarse channel estimation, and symbol sequence segmentation on the original signal stream to generate multiple sub-tasks, and outputs the multiple sub-tasks to the computing unit. The computing unit includes a field-programmable gate array, a graphics processor, and a central processing unit. Each computing unit receives its corresponding sub-task, performs processing, and outputs its processing results to the output end. The output end splices and fuses the processing results to output the final demodulated data stream.

[0047] Therefore, this application first generates a symbol sequence to be processed by performing frame synchronization detection and coarse channel estimation on the original signal stream, and then divides the symbol sequence into multiple sub-tasks to provide task units for subsequent heterogeneous computing, laying the foundation for heterogeneous parallel processing. Second, by obtaining the real-time load status values ​​of each computing unit, each sub-task is dispatched to the field-programmable gate array (FPGA), graphics processing unit (GPU), and central processing unit (CPU) for execution, realizing dynamic matching between sub-tasks and heterogeneous computing units, improving the overall resource utilization efficiency and task processing throughput of the heterogeneous computing system. Then, by executing the FPGA, GPU, and CPU in parallel, the performance bottleneck of a single processor handling all types of tasks is avoided, reducing the overall processing latency and improving the system throughput. Finally, by splicing and fusing the hard-decision bits, soft-decision bits, and protocol data units output by the FPGA, GPU, and CPU, the final demodulated data stream is generated by using cross-validation and confidence level judgment of multi-source results, reducing the demodulation error rate and improving the reliability of the final data stream.

[0048] In summary, the technical solution adopted in this application can achieve efficient collaborative utilization of heterogeneous resources and improve system throughput and demodulation reliability.

[0049] Example 2: This application provides a heterogeneous computing signal processing system for high-speed communication scenarios, referencing... Figure 3 As shown in the figure, this is a module structure diagram of heterogeneous computing signal processing for high-speed communication scenarios according to this embodiment of the present application. The signal processing system includes: The signal receiving module 100 is used to receive the raw signal stream in a high-speed communication scenario, perform frame synchronization detection and coarse channel estimation on the raw signal stream, generate a symbol sequence to be processed, and divide the symbol sequence into multiple sub-tasks. The task dispatch module 200 is used to obtain the real-time load status value of each computing unit and dispatch each subtask to the field programmable gate array, graphics processor and central processing unit for execution. The task processing module 300 is used to perform physical layer front-end operations on the field-programmable gate array, perform channel compensation and soft bit extraction on the graphics processor, and process protocol-related logic on the central processing unit. The fusion output module 400 is used to combine the results obtained by the field-programmable gate array, the graphics processor and the central processing unit to output the final demodulated data stream.

[0050] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0051] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0052] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A heterogeneous computing signal processing method for high-speed communication scenarios, characterized in that, The signal processing method includes the following steps: The system receives the raw signal stream in a high-speed communication scenario, performs frame synchronization detection and coarse channel estimation on the raw signal stream, generates a symbol sequence to be processed, and obtains the signal-to-noise ratio and phase change rate within a time window near each symbol in the symbol sequence. The signal-to-noise ratio and the phase change rate are weighted and synthesized into a trend coefficient. Based on the trend coefficient's change curve along the time axis, a segmentation point is set at the position where the trend coefficient crosses a threshold, and the symbol sequence is divided into blocks of unequal length. The fluctuation range of the trend coefficient within each block does not exceed a preset tolerance value, and each block is treated as an independent subtask. At each dispatch time, the real-time load status values ​​of the Field Programmable Gate Array (FPGA), Graphics Processing Unit (GPU), and Central Processing Unit (CPU) are acquired respectively. When the tendency coefficient is higher than a preset threshold and the real-time load status value of the FPGA is lower than that of the GPU, the task is dispatched to the FPGA. When the tendency coefficient is lower than a preset threshold and the real-time load status value of the GPU is lower than that of the FPGA, the task is dispatched to the GPU. When neither of the above two conditions is met, the task is dispatched to the computing unit with the lowest real-time load status value. When the execution of the current subtask depends on the control plane information in the CPU, the task is directly dispatched to the CPU. The physical layer front-end operations are performed on the field-programmable gate array (FPGA), channel compensation and soft bit extraction are performed on the graphics processor (GPU), and protocol-related logic is processed on the central processing unit (CPU). The results obtained from the field-programmable gate array, graphics processor, and central processing unit are combined to output the final demodulated data stream.

2. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 1, characterized in that, Performing frame synchronization detection and coarse channel estimation on the original signal stream to generate the symbol sequence to be processed specifically includes: The original signal stream is correlated with the local frame header sequence, and the correlation peak position is used as the frame start boundary. A complete frame length data is extracted from the original signal stream, and the pilot symbol segment in the complete frame length data is estimated by least squares to obtain the coarse channel response value on each subcarrier. The coarse channel response value is used to perform single-tap equalization on data symbols within the same complete frame length to obtain pre-equalized symbols. The symbols after initial equalization are hard-determined to obtain decision symbols. The residual phase error is extracted based on the decision symbols and the data symbols before single-tap equalization. Then, the common phase rotation correction is performed on the symbols after initial equalization based on the residual phase error. The phase-corrected symbol sequence is split into orthogonal frequency division multiplexing symbol blocks to generate the symbol sequence to be processed.

3. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 2, characterized in that, Extracting the residual phase error based on the decision symbol and the data symbol before single-tap equalization, and then performing common phase rotation correction on the initially equalized symbol based on the residual phase error, specifically includes: The data symbols before single-tap equalization are multiplied by the decision symbols using conjugate multiplication to obtain the phase difference measurement value corresponding to each data symbol; The phase difference measurements of all data symbols within the same orthogonal frequency division multiplexing symbol block are weighted and averaged. The weighted average is used as the residual phase error estimate of the corresponding orthogonal frequency division multiplexing symbol block; Based on the estimated residual phase error, a common rotation factor is constructed. All data symbols belonging to the same orthogonal frequency division multiplexing symbol block in the pre-equalized symbols are multiplied by the common rotation factor to complete the common phase rotation correction.

4. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 1, characterized in that, The real-time load status value refers to the ratio of the number of subtasks currently being processed by each computing unit to the maximum concurrent processing capacity of the corresponding computing unit.

5. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 1, characterized in that, Performing physical layer front-end operations on a field-programmable gate array (FPGA) specifically includes: The symbol data corresponding to the subtasks dispatched to the field programmable gate array are sequentially subjected to digital downconversion, matched filtering, and fast Fourier transform to obtain a frequency domain symbol sequence. Perform a detection operation on the frequency domain symbol sequence and set the completion flag.

6. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 2, characterized in that, The channel compensation and soft bit extraction performed on the graphics processor specifically include: The symbolic data corresponding to the subtasks dispatched to the graphics processor is copied from main memory to the graphics processor's global memory. The graphics processor starts multiple thread blocks, each thread block is responsible for channel compensation calculation on a subcarrier, and equalizes all symbols on the same subcarrier according to the coarse channel response value; The log-likelihood ratio of each bit of the equalized symbol is determined according to the modulation order. The log-likelihood ratios of all bits are then arranged by subcarrier and symbol number and written back to main memory.

7. The heterogeneous computing signal processing method for high-speed communication scenarios as described in claim 1, characterized in that, The final demodulated data stream refers to the binary bit sequence that, after being demodulated and decoded by the physical layer, can be directly delivered to the media access control layer for parsing.

8. A heterogeneous computing signal processing system for high-speed communication scenarios, used to execute the heterogeneous computing signal processing method for high-speed communication scenarios as described in any one of claims 1 to 7, characterized in that, The signal processing system includes: The signal receiving module is used to receive the raw signal stream in a high-speed communication scenario, perform frame synchronization detection and coarse channel estimation on the raw signal stream, generate a symbol sequence to be processed, and obtain the signal-to-noise ratio and phase change rate within a time window near each symbol in the symbol sequence; weight the signal-to-noise ratio and the phase change rate to synthesize a tendency coefficient; based on the change curve of the tendency coefficient along the time axis, set a segmentation point at the position where the tendency coefficient crosses a threshold, and cut the symbol sequence into blocks of unequal length; ensure that the fluctuation range of the tendency coefficient within each block does not exceed a preset tolerance value, and each block is treated as an independent subtask; The task dispatch module is used to acquire the real-time load status values ​​of the Field Programmable Gate Array (FPGA), Graphics Processing Unit (GPU), and Central Processing Unit (CPU) at each dispatch time. When the tendency coefficient is higher than a preset threshold and the real-time load status value of the FPGA is lower than that of the GPU, the task is dispatched to the FPGA. When the tendency coefficient is lower than a preset threshold and the real-time load status value of the GPU is lower than that of the FPGA, the task is dispatched to the GPU. When neither of the above two conditions is met, the task is dispatched to the computing unit with the lowest real-time load status value. When the execution of the current subtask depends on the control plane information in the CPU, the task is directly dispatched to the CPU. The task processing module is used to perform physical layer front-end operations on the field-programmable gate array, perform channel compensation and soft bit extraction on the graphics processor, and process protocol-related logic on the central processing unit. The fusion output module is used to combine the results obtained from the field-programmable gate array, graphics processor and central processing unit to output the final demodulated data stream.

Citation Information

Patent Citations

  • Hardware acceleration system and method based on multi-core architecture and related equipment

    CN120336242A

  • Target positioning method and system based on single reference station

    CN120446934A