A method and apparatus for capturing a direct sequence spread spectrum signal
By working together between the CPU and the GPU device and performing multi-threaded parallel capture operations using the GPU device, the problem of real-time demodulation of direct sequence spread spectrum signals in the prior art is solved, and fast and real-time signal capture and demodulation are achieved.
Patent Information
- Application Number
- CN202510458817.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing software demodulation technology cannot meet the real-time demodulation requirements of direct sequence spread spectrum signals, especially when processing long signals, the calculation complexity is too high, and the demodulation rate cannot be effectively improved.
By working together between the CPU device and the GPU device, multi-threaded parallel capture operations are used to improve computing efficiency. The specific steps include receiving the spread spectrum signal, performing pre-processing, selecting candidate combinations, sending data to the GPU device, performing parallel capture operations, obtaining local energy maximum, searching for global energy maximum, and determining the candidate combination as the capture result.
It realizes fast and real-time capture of direct sequence spread spectrum signals, improves the understanding of the regulation rate, and meets the requirements of real-time demodulation.
Smart Images

Figure CN120017093B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of signal processing, and in particular, to a method and device for capturing a direct sequence spread spectrum signal. Background Art
[0002] A direct sequence spread spectrum system is a communication technology with advantages such as anti-interference and low probability of intercept. It spreads the signal bandwidth by using a specific spreading code, enabling the signal to have excellent anti-point interference ability during propagation, and is widely used in military communications, satellite communications, underwater acoustic communications, and other fields. In the demodulation process of spread spectrum signals, how to quickly, real-time, and accurately demodulate spread spectrum signals is a key problem to be solved in the design of spread spectrum demodulation devices at the receiving end. The capture of spread spectrum signals during the demodulation process is a time-consuming process.
[0003] In the prior art, although real-time demodulation rates have been achieved based on peripheral boards such as DSP and FPGA, it cannot adapt to the current development trend of software-based signal demodulation, and has problems such as high development difficulty, difficult upgrade, and limited usage scenarios. In the current software-based signal demodulation process, the PMF-FFT algorithm is used to capture spread spectrum signals. Since the length of the received spread spectrum signal is relatively long, the received signal and the local spreading code are divided into multiple segments, and each segment is subjected to a correlation operation to reduce the length of a single operation and reduce the computational complexity; after obtaining multiple partial correlation results, these results are combined and subjected to an FFT transform to convert the correlation operation in the time domain into a multiplication operation in the frequency domain, and then an IFFT operation is performed to obtain the captured result. Since current software-based demodulation devices are often developed based on multi-core CPUs, they cannot meet the requirements of real-time demodulation when dealing with a large amount of calculations in the above-mentioned spread spectrum signal capture process. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and device for capturing a direct sequence spread spectrum signal to solve the problem that existing software demodulation technologies cannot meet the requirements of real-time demodulation.
[0005] To achieve the above object, embodiments of the present invention provide the following technical solutions:
[0006] A first aspect of an embodiment of the present invention discloses a method for capturing a direct sequence spread spectrum signal, which is applied to a CPU device, and the CPU device is connected to a GPU device. The method includes:
[0007] Receiving a spread spectrum signal sent by a sending end, and preprocessing the spread spectrum signal;
[0008] Select a candidate combination from the candidate combination set and remove the candidate combination from the candidate combination set; the candidate combination set includes a plurality of the candidate combinations, and each of the candidate combinations is obtained by sequentially combining a preset plurality of carrier frequency ranges with each preset pseudo-code phase range;
[0009] Send the candidate combination and the preprocessed spread spectrum signal as operation data to the GPU device;
[0010] Receive the local energy maximum value obtained by the GPU device through parallel acquisition operations based on the operation data;
[0011] Return to execute the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set until there is no such candidate combination in the candidate combination set, and obtain a plurality of the local energy maximum values;
[0012] Search for the global energy maximum value from each of the local energy maximum values;
[0013] If the global energy maximum value is greater than the acquisition threshold, determine the candidate combination corresponding to the global energy maximum value as the acquisition result.
[0014] Preferably, the preprocessing of the spread spectrum signal includes:
[0015] Convert the data type of the spread spectrum signal to the float type.
[0016] Preferably, the searching for the global energy maximum value from each of the local energy maximum values includes:
[0017] Use a preset reduction algorithm to search for the global energy maximum value from each of the local energy maximum values.
[0018] Preferably, after determining the candidate combination corresponding to the global energy maximum value as the acquisition result, the method further includes:
[0019] If the accuracy of the acquisition result does not meet the requirements, divide the carrier frequency range in the candidate combination corresponding to the global energy maximum value to obtain a plurality of new carrier frequency ranges, and divide the pseudo-code phase range in the candidate combination corresponding to the global energy maximum value to obtain a plurality of new pseudo-code phase ranges;
[0020] Combine each of the new carrier frequency ranges with each of the new pseudo-code phase ranges in sequence to obtain a candidate combination set composed of a plurality of new candidate combinations, and return to execute the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set.
[0021] In a second aspect of the embodiments of the present invention, a method for capturing a direct sequence spread spectrum signal is disclosed, which is applied to a GPU device. The GPU device is connected to a CPU device, and the method includes:
[0022] Receiving the operation data sent by the CPU; the operation data includes a candidate combination and a preprocessed spread spectrum signal; the candidate combination includes a carrier frequency range and a pseudo-code phase range;
[0023] Dividing the carrier frequency range to obtain a plurality of carrier frequency blocks, and dividing the pseudo-code phase range to obtain a plurality of pseudo-code phase blocks;
[0024] Combining each of the carrier frequency blocks with each of the pseudo-code phase blocks to obtain a plurality of operation combinations;
[0025] For each of the operation combinations, using a plurality of threads in a thread block to perform parallel capture operations on the operation combination and the spread spectrum signal to obtain energy values corresponding to each of the operation combinations;
[0026] Taking the maximum value among the energy values corresponding to each of the operation combinations as the local energy maximum value, and sending the local energy maximum value to the CPU device.
[0027] Preferably, for each of the operation combinations, using a plurality of threads in a thread block to perform parallel capture operations on the operation combination and the spread spectrum signal to obtain energy values corresponding to each of the operation combinations, includes:
[0028] Determining a current operation combination from each of the operation combinations;
[0029] Generating an in-phase component and a quadrature component based on the carrier frequency block in the current operation combination;
[0030] Generating a plurality of down-conversion calculation tasks based on each frequency point in the spread spectrum signal, the in-phase component, and the quadrature component, and allocating each of the down-conversion calculation tasks to each thread in the thread block for parallel calculation to obtain a down-converted spread spectrum signal;
[0031] Generating a plurality of local pseudo-code sub-blocks with a preset phase difference based on the pseudo-code phase block in the current operation combination;
[0032] Generating a plurality of despreading calculation tasks based on the down-converted spread spectrum signal and each of the local pseudo-code sub-blocks, and allocating each of the despreading calculation tasks to each of the threads in the thread block for parallel calculation to obtain a despread signal;
[0033] Performing integral cleaning on the despread signal to obtain an integrally cleaned signal;
[0034] Perform an FFT transform operation on the signal after integral cleaning to obtain a transform result;
[0035] Based on each frequency point in the spread spectrum signal and the transform result, generate a plurality of energy calculation tasks, and allocate each of the energy calculation tasks to each of the threads in the thread block for parallel calculation to obtain the energy values of each of the frequency points;
[0036] Incoherently accumulate the energy values of each of the frequency points to obtain the energy value corresponding to the current operation combination;
[0037] Determine the next current operation combination from the remaining operation combinations, and return to execute the step of generating the in-phase component and the quadrature component based on the carrier frequency block in the current operation combination until the energy values corresponding to each of the operation combinations are obtained.
[0038] Preferably, the incoherently accumulating the energy values of each of the frequency points to obtain the energy value corresponding to the current operation combination includes:
[0039] Use a preset reduction algorithm to incoherently accumulate the energy values of each of the frequency points to obtain the energy value corresponding to the current operation combination.
[0040] Preferably, the performing an FFT transform operation on the signal after integral cleaning to obtain a transform result includes:
[0041] Use the cuFFT library in the CUDA platform to perform an FFT transform operation on the signal after integral cleaning to obtain a transform result.
[0042] A third aspect of the embodiments of the present invention discloses a direct sequence spread spectrum signal acquisition device, which is applied to a CPU device, and the CPU device is connected to a GPU device. The device includes:
[0043] A preprocessing unit, configured to receive a spread spectrum signal sent by a sending end and perform preprocessing on the spread spectrum signal;
[0044] A selection unit, configured to select a candidate combination from a candidate combination set and remove the candidate combination from the candidate combination set; the candidate combination set includes a plurality of candidate combinations, and each of the candidate combinations is obtained by sequentially combining a preset plurality of carrier frequency ranges with a preset each pseudo-code phase range;
[0045] A sending unit, configured to send the candidate combination and the preprocessed spread spectrum signal as operation data to the GPU device;
[0046] A first receiving unit, configured to receive the local energy maximum value obtained by the GPU device through parallel capture operations based on the operation data;
[0047] A return unit, configured to return the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set until there is no such candidate combination in the candidate combination set, so as to obtain multiple local energy maximum values;
[0048] A search unit, configured to search for the global energy maximum value from each of the local energy maximum values;
[0049] A determination unit, configured to determine the candidate combination corresponding to the global energy maximum value as the capture result if the global energy maximum value is greater than the capture threshold.
[0050] A fourth aspect of the embodiments of the present invention discloses a direct sequence spread spectrum signal capture device, which is applied to a GPU device, and the GPU device is connected to a CPU device. The device includes:
[0051] A second receiving unit, configured to receive the operation data sent by the CPU; the operation data includes a candidate combination and a preprocessed spread spectrum signal; the candidate combination includes a carrier frequency range and a pseudo-code phase range;
[0052] A division unit, configured to divide the carrier frequency range to obtain a plurality of carrier frequency blocks, and divide the pseudo-code phase range to obtain a plurality of pseudo-code phase blocks;
[0053] A combination unit, configured to combine each of the carrier frequency blocks with each of the pseudo-code phase blocks respectively to obtain a plurality of operation combinations;
[0054] An operation unit, configured to perform parallel capture operations on each of the operation combinations and the spread spectrum signal by using a plurality of threads in a thread block, so as to obtain the energy value corresponding to each of the operation combinations;
[0055] A feedback unit, configured to use the maximum value among the energy values corresponding to each of the operation combinations as the local energy maximum value, and send the local energy maximum value to the CPU device.
[0056] Based on the method and device for capturing a direct-sequence spread-spectrum signal provided by the embodiments of the present invention above, a spread-spectrum signal sent by a sending end is received, and the spread-spectrum signal is preprocessed; a candidate combination is selected from a candidate combination set, and the candidate combination is removed from the candidate combination set; the candidate combination set includes a plurality of the candidate combinations, and each of the candidate combinations is obtained by sequentially combining a plurality of preset carrier frequency ranges with each of a plurality of preset pseudo-code phase ranges; the candidate combination and the preprocessed spread-spectrum signal are sent to the GPU device as operation data; the local energy maximum value obtained by the GPU device performing parallel capture operations based on the operation data is received; the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set is returned to be executed until there is no candidate combination in the candidate combination set, and a plurality of the local energy maximum values are obtained; a global energy maximum value is searched from each of the local energy maximum values; if the global energy maximum value is greater than a capture threshold, the candidate combination corresponding to the global energy maximum value is determined as a capture result. In this solution, the CPU device and the GPU device are used in cooperation to complete the capture, and the parallel capture operation method using multiple threads in the GPU device improves the calculation efficiency, and solves the problem that the requirements of real-time demodulation cannot be met in the existing software demodulation technology. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0058] Figure 1 It is a schematic diagram of anti-interference principle of a direct-sequence spread-spectrum system disclosed in an embodiment of the present invention;
[0059] Figure 2 It is a structural diagram of a direct-sequence spread-spectrum system disclosed in an embodiment of the present invention;
[0060] Figure 3 It is a schematic diagram of the structure of a PMF-FFT capture algorithm disclosed in an embodiment of the present invention;
[0061] Figure 4 It is a schematic diagram of two-dimensional capture search of a spread-spectrum signal disclosed in an embodiment of the present invention;
[0062] Figure 5 It is a flowchart of a serial capture algorithm executed by a CPU device disclosed in an embodiment of the present invention;
[0063] Figure 6Interaction diagram of the direct sequence spread spectrum signal acquisition system disclosed in the embodiments of the present invention;
[0064] Figure 7 Flowchart of a method for acquiring a direct sequence spread spectrum signal disclosed in the embodiments of the present invention;
[0065] Figure 8 Schematic diagram of finding the maximum value by a reduction algorithm disclosed in the embodiments of the present invention;
[0066] Figure 9 Flowchart of another method for acquiring a direct sequence spread spectrum signal disclosed in the embodiments of the present invention;
[0067] Figure 10 Schematic diagram of a parallel down-conversion algorithm disclosed in the embodiments of the present invention;
[0068] Figure 11 Schematic diagram of the structure of a capture unit disclosed in the embodiments of the present invention;
[0069] Figure 12 Structure diagram of a parallel despreading scheme disclosed in the embodiments of the present invention;
[0070] Figure 13 Structure diagram of a parallel energy calculation scheme disclosed in the embodiments of the present invention;
[0071] Figure 14 Schematic diagram of summing by a reduction algorithm disclosed in the embodiments of the present invention;
[0072] Figure 15 Structure diagram of a direct sequence spread spectrum signal acquisition device disclosed in the embodiments of the present invention;
[0073] Figure 16 Structure diagram of another direct sequence spread spectrum signal acquisition device disclosed in the embodiments of the present invention. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0075] In this application, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0076] As can be seen from the background art, the direct sequence spread spectrum system is a communication technology with advantages such as anti-interference and low probability of intercept.
[0077] As Figure 1 shown, it is a schematic diagram of the anti-interference principle of a direct sequence spread spectrum system disclosed in an embodiment of the present invention. The specific anti-interference principle is as follows:
[0078] During the propagation process, the dot interference only occupies a small part of the entire bandwidth for the spread spectrum signal. After the original signal is restored after despreading at the receiving end, the original dot interference is equivalent to being spread spectrum. After low-pass filtering, most of the interference is filtered out. This is the anti-interference principle of spread spectrum communication.
[0079] As Figure 2 shown, it is a structural diagram of a direct sequence spread spectrum system disclosed in an embodiment of the present invention.
[0080] Among them, the left half is the transmitting end and the right half is the receiving end. In the overall design of the spread spectrum system, how to quickly, real-time and accurately demodulate the spread spectrum signal at the receiving end of the signal is the key problem to be solved in the design of the receiving end spread spectrum demodulation device. In the process of demodulating the spread spectrum signal at the receiving end, the more time-consuming part is the signal acquisition and tracking part. This application mainly focuses on the acquisition part.
[0081] As Figure 3 shown, it is a schematic diagram of the PMF-FFT acquisition algorithm structure disclosed in an embodiment of the present invention.
[0082] The currently commonly used acquisition algorithm is the PMF-FFT algorithm. Since the length of the received spread spectrum signal is relatively long, the received signal and the local spread spectrum code are first divided into multiple segments, and each segment is subjected to a correlation operation to reduce the length of a single operation and reduce the computational complexity; after obtaining multiple partial correlation results, these results are combined and subjected to an FFT transform to convert the correlation operation in the time domain into a multiplication operation in the frequency domain, and then an IFFT operation is performed to obtain the acquisition result. This algorithm combines a matched filter and a fast Fourier transform and reduces the computational complexity. It is an efficient algorithm. However, when using a CPU device, it still cannot meet the requirements for real-time demodulation of direct sequence spread spectrum signals.
[0083] As shown in Figure 4 Fig. 4, it is a schematic diagram of two-dimensional capture search of a spread spectrum signal disclosed in an embodiment of the present invention.
[0084] It should be noted that the capture of a spread spectrum signal is a two-dimensional search process in the pseudo-code domain and the frequency domain. When serially executing the direct spread signal capture process, for each candidate frequency point, operations such as down-conversion, despreading, integration and cleaning, non-coherent accumulation, energy statistics, and peak detection need to be completed in sequence.
[0085] In the prior art, software demodulation devices are often developed based on multi-core CPUs, and the part involving the above-mentioned large number of capture operations is serially executed by CPU devices.
[0086] As shown in Figure 5 Fig. 5, it is a flow chart of a serial capture algorithm executed by a CPU device disclosed in an embodiment of the present invention.
[0087] When a serial program is executed, the external main loop traverses all frequency offset indexes, and for each frequency offset point, a search in the pseudo-code dimension is performed. The pseudo-code dimension search includes:
[0088] 1. Down-conversion and quadrature decomposition, and the quadrature mixing separates into I / Q two paths of signals.
[0089] 2. Despreading, the received signal is circularly shifted and aligned with the local pseudo-code. When the maximum correlation value is obtained, the signal despreading is achieved.
[0090] 3. Integration and cleaning and non-coherent accumulation to improve the signal-to-noise ratio of the signal.
[0091] 4. Peak detection and frequency estimation, traverse all candidate frequencies, select the frequency corresponding to the maximum energy value as the carrier Doppler frequency offset estimation obtained by capture, and use it as the basis for the next tracking process.
[0092] It can be seen that when the existing software demodulation technology performs spread spectrum signal capture, the CPU device is used for operations, and the overall calculation process is serial. The spread spectrum signal capture involves a large number of basic multiplication and addition operations and FFT operations. Although these operations have simple logic, they have a large amount of calculation and cannot meet the requirements of real-time demodulation.
[0093] Therefore, an embodiment of the present invention discloses a method and device for capturing a direct sequence spread spectrum signal. In this solution, the CPU device and the GPU device are used in cooperation to complete the capture. In the GPU device, a multi-threaded parallel capture operation method is adopted to improve the calculation efficiency, and the problem that the requirements of real-time demodulation cannot be met in the existing software demodulation technology is solved.
[0094] As shown in Figure 6As shown in the figure, it is an interaction diagram of the direct sequence spread spectrum signal acquisition system disclosed in the embodiment of the present invention. Among them, the acquisition system includes: a CPU device and a GPU device, and the CPU device and the GPU device are connected through a PCIe bus.
[0095] It should be noted that the GPU device is naturally good at processing large-scale and simple repetitive calculations. Different from hardware peripherals such as DSP and FPGA, the programming mode of the GPU device is similar to that of the CPU device, and the acquisition program is easy to adjust, which conforms to the development trend of signal processing softwareization.
[0096] Figure 6 The interaction process of the acquisition system shown is a typical heterogeneous architecture computing method. The CPU device is responsible for the logical processing of the entire acquisition operation. The CPU device transmits the operation data to the GPU device through the PCIe bus. The GPU device is responsible for giving full play to the advantages of large-scale parallel computing and performing parallel acquisition operations such as down-conversion, despreading, FFT operation, signal energy calculation, non-coherent accumulation, taking the maximum value, etc. Then, the results obtained from the operation are transmitted back to the CPU device for final judgment. The specific interaction process is as follows:
[0097] 1. Receive data, that is, the CPU device receives the spread spectrum signal sent by the sending end.
[0098] 2. Preprocess the data and allocate memory for the following GPU algorithms.
[0099] Among them, data preprocessing refers to converting the received char type or short type data into float type data suitable for processing by the GPU device.
[0100] Allocating memory for the GPU algorithm means: allocating the video memory required for each parallel acquisition operation step in the GPU device to provide the corresponding storage space.
[0101] 3. Use cudaMemcpy to transfer the preprocessed data into the acquisition operation memory.
[0102] Among them, cudaMemcpy is a function used to transfer data between the CPU memory and the GPU memory. The acquisition operation memory refers to the memory pre-created in the CPU device.
[0103] 4. Determine whether the accumulated data is sufficient for the spread spectrum signal acquisition operation.
[0104] Among them, the accumulated data refers to the received and accumulated spread spectrum signals and the candidate combination set. The candidate combination set includes multiple candidate combinations, and each candidate combination is obtained by combining a preset multiple carrier frequency ranges with each preset pseudo-code phase range in sequence.
[0105] Preferably, each candidate combination is also pre - processed data, which is converted into float - type data suitable for GPU device processing.
[0106] If the accumulated data is not sufficient, return and continue to accumulate until it is sufficient; if the accumulated data is sufficient, start the next step.
[0107] In this application, the carrier frequency is the Doppler frequency, that is, the carrier Doppler frequency.
[0108] 5. Nest two - layer loops to traverse each possible combination of carrier frequency ranges and pseudo - code phase ranges.
[0109] Specifically, select a candidate combination from the candidate combination set, and remove the candidate combination from the candidate combination set; send the candidate combination and the pre - processed spread - spectrum signal as operation data to the GPU device; receive the local energy maximum value obtained by the GPU device based on the operation data for parallel acquisition operations; return to select a candidate combination from the candidate combination set again and remove the candidate combination from the candidate combination set until there is no candidate combination in the candidate combination set, and obtain multiple local energy maximum values.
[0110] Among them, the parallel acquisition operations are a series of operations such as down - conversion, despreading, FFT operation, signal energy calculation, non - coherent accumulation, taking the maximum value, etc., which will be specifically explained in the following embodiments of the present invention.
[0111] 6. Compare the global maximum value and obtain the data index.
[0112] Specifically, search for the global energy maximum value from each local energy maximum value.
[0113] Among them, obtaining the data index means obtaining the index of the candidate combination corresponding to the global energy maximum value.
[0114] 7. Compare whether the global energy maximum value is greater than a pre - set acquisition threshold.
[0115] If it is greater, it can be determined that the candidate combination corresponding to the global energy maximum value is the acquisition result, that is, the frequency offset of the spread - spectrum signal and the pseudo - code phase are within the carrier frequency range and pseudo - code phase range in the candidate combination.
[0116] If it is less than or equal to, it is necessary to receive the spread - spectrum signal again and perform acquisition again.
[0117] After obtaining the acquisition result in the previous step, the approximate ranges of the spread - spectrum signal frequency offset and pseudo - code phase are determined, that is, relatively rough values of the carrier Doppler frequency shift and pseudo - code phase are obtained. In the subsequent steps, the carrier Doppler frequency shift and pseudo - code phase are further accurately searched within this approximate range.
[0118] Specifically, divide the carrier frequency range in the capture result of the previous step to obtain multiple new carrier frequency ranges, and divide the pseudo-code phase range to obtain multiple new pseudo-code phase ranges. Combine each new carrier frequency range with each new pseudo-code phase range in sequence to obtain a candidate combination set composed of multiple new candidate combinations. Repeatedly execute the above steps based on this candidate combination set, and a more accurate capture result can be obtained.
[0119] Preferably, the carrier frequency range needs to be accurate to within dozens to hundreds of hertz, and the pseudo-code phase range needs to be accurate to within half a chip.
[0120] Based on the interaction process of a capture system disclosed in the embodiments of the present invention as described above, as Figure 7 shown, it is a flowchart of a method for capturing a direct-sequence spread-spectrum signal disclosed in the embodiments of the present invention. This capture method is applied to a CPU device, and the CPU device is connected to a GPU device, including the following steps:
[0121] Step S101: Receive the spread-spectrum signal sent by the sending end and preprocess the spread-spectrum signal.
[0122] Specifically, convert the data type of the spread-spectrum signal to the float type suitable for processing by the GPU device.
[0123] Step S102: Select a candidate combination from the candidate combination set and remove the candidate combination from the candidate combination set.
[0124] Among them, the candidate combination set includes multiple candidate combinations, and each candidate combination is obtained by combining a preset multiple carrier frequency ranges with each preset pseudo-code phase range in sequence.
[0125] It should be noted that the spread-spectrum signal is the original BPSK (Binary Phase Shift Keying) signal plus the modulation of the spread-spectrum code (which can be understood as multiplying by the spread-spectrum code). This signal will be affected by the Doppler effect during the propagation process (the carrier frequency will deviate by about ±1.5 kHz), and the pseudo-code will generate time delay (due to the influence of the atmosphere, ionosphere, etc. during the propagation process). This will cause offsets in two dimensions, so searches need to be performed in two dimensions (carrier frequency and pseudo-code phase). Only when they are aligned can the carrier and pseudo-code be removed.
[0126] It can be understood that among the multiple candidate combinations, all possibilities of one-to-one combination of multiple carrier frequency ranges and multiple pseudo-code phase ranges are included.
[0127] Refer to Figure 4 , each small square corresponds to a carrier frequency range and a pseudo-code phase range, so each small square can represent a candidate combination.
[0128] Exemplarily, assume that the carrier frequency range is from 1 to 2 and the pseudo-code phase range is from 1 to 2. Then the obtained candidate combinations are: carrier frequency range 1 and pseudo-code phase range 1; carrier frequency range 1 and pseudo-code phase range 2; carrier frequency range 2 and pseudo-code phase range 1; carrier frequency range 2 and pseudo-code phase range 2.
[0129] Step S103: Send the candidate combinations and the preprocessed spread-spectrum signal as operation data to the GPU device.
[0130] It should be noted that the operation data sent to the GPU device is hardly processed and is directly sent using the PCIe bus.
[0131] Step S104: Receive the local energy maximum value obtained by the GPU device through parallel capture operations based on the operation data.
[0132] It should be noted that during the spread-spectrum signal capture process, the local energy maximum value refers to a significant peak reached by the signal energy under the combination of a specific carrier frequency range and pseudo-code phase range. This peak is used to determine the initial synchronization parameters of the signal, namely the pseudo-code phase and frequency offset. Specifically, during the capture process, the receiving end searches for combinations of different carrier frequency ranges and pseudo-code phase ranges, and calculates the signal energy under each combination, that is, the local energy maximum value.
[0133] For the specific parallel capture operation process, please refer to Figure 9 the corresponding embodiment of the present invention.
[0134] Step S105: Return to execute step S102 until there are no candidate combinations in the candidate combination set, and obtain multiple local energy maximum values.
[0135] It can be understood that by using the GPU device, the local energy maximum values are calculated for multiple candidate combinations in sequence. Therefore, after obtaining the local energy maximum value corresponding to each candidate combination, return to step S102 to calculate the local energy maximum value corresponding to the next candidate combination.
[0136] Step S106: Search for the global energy maximum value from each local energy maximum value.
[0137] In the specific implementation process of step S106, use a preset reduction algorithm to search for the global energy maximum value from each local energy maximum value.
[0138] As Figure 8 shown, it is a schematic diagram of finding the maximum value using a reduction algorithm disclosed in the embodiment of the present invention.
[0139] This method is similar to gradually narrowing down the range through multiple rounds of comparison in a set of data, and finally finding the global energy maximum. The specific principle is as follows:
[0140] 1. Initial data layer (Index1)
[0141] In this layer, the data is organized into a linear array containing 1024 elements (equivalent to multiple local energy maxima), from a1 to a 1024 .
[0142] These elements will be paired up and compared. For example, a1 and a2 are compared, a3 and a4 are compared, and so on.
[0143] 2. First comparison layer (Index2)
[0144] The amount of data in this layer is half of the previous layer, that is, 512 elements, from b1 to b 512 .
[0145] Each element b i is the larger one after comparing two paired elements in the previous layer. For example, b1 is the larger one of a1 and a2, b2 is the larger one of a3 and a4, and so on.
[0146] 3. Subsequent comparison layers (Index3 to Index9)
[0147] The amount of data in each layer is half of the previous layer. For example, the Index3 layer has 256 elements, the Index4 layer has 128 elements, and so on, until the Index9 layer has 4 elements.
[0148] In each layer, the elements are the larger ones after comparing two adjacent elements in the previous layer. This process continues until the amount of data is reduced to a certain extent.
[0149] 4. Final comparison layer (Index10)
[0150] The Index10 layer contains 2 elements, namely d1 and d2, which are the larger ones after comparing two adjacent elements in the Index9 layer.
[0151] Compare d1 and d2 in the Index10 layer to get the larger value.
[0152] 5. Result layer (Index11)
[0153] In the Index11 layer, the larger value obtained in the Index10 layer is determined as the maximum value in the entire dataset (corresponding to the global energy maximum).
[0154] Compared with the method of the CPU device to sequentially compare and obtain the maximum value in a loop, this method greatly reduces the number of comparisons through parallel comparison and hierarchical elimination, thereby improving the operation efficiency. In large-scale data processing, the advantages of this method are particularly obvious.
[0155] Step S107: If the global energy maximum value is greater than the capture threshold, determine the candidate combination corresponding to the global energy maximum value as the capture result.
[0156] It can be understood that if the global energy maximum value is less than or equal to the capture threshold, it is determined that the frequency offset and pseudo-code phase of the signal are not within the carrier frequency range and pseudo-code phase range corresponding to the global energy maximum value.
[0157] In one embodiment, whenever the CPU device receives the local energy maximum value corresponding to a candidate combination, it determines whether the local energy maximum value is greater than the capture threshold. If not, the local energy maximum value is eliminated to reduce the time for subsequent searching for the global energy maximum value.
[0158] The accuracy requirements for the carrier frequency range and pseudo-code phase range are as follows: the carrier frequency range needs to be accurate within dozens to hundreds of hertz, and the pseudo-code phase range needs to be accurate within half a chip.
[0159] If the carrier frequency range and pseudo-code phase range of the capture result do not meet the above accuracy requirements, divide the carrier frequency range in the candidate combination corresponding to the global energy maximum value to obtain multiple new carrier frequency ranges, and divide the pseudo-code phase range in the candidate combination corresponding to the global energy maximum value to obtain multiple new pseudo-code phase ranges.
[0160] Then, sequentially combine each new carrier frequency range with each new phase range to obtain a candidate combination set composed of multiple new candidate combinations, and return to execute step S102 to further accurately search for the carrier frequency shift and pseudo-code phase within a smaller range, and determine the smaller range where the carrier frequency shift and pseudo-code phase are located.
[0161] Based on the direct sequence spread spectrum signal capture method disclosed in the above embodiments of the present invention, in this solution, the CPU device and the GPU device are used in cooperation to complete the capture. Through the relatively powerful logical processing ability of the CPU device and the parallel capture operation method using multi-threads in the GPU device, the calculation efficiency is improved, and the problem that the existing software demodulation technology cannot meet the requirements of real-time demodulation is solved.
[0162] Based on the interaction process of a capture system disclosed in the above embodiments of the present invention, as Figure 9As shown in the figure, it is a flowchart of another method for capturing a direct sequence spread spectrum signal disclosed in an embodiment of the present invention. This capturing method is applied to a GPU device, and the GPU device is connected to a CPU device. The method includes the following steps:
[0163] Step S201: Receive the operation data sent by the CPU.
[0164] Among them, the operation data includes a candidate combination and a preprocessed spread spectrum signal; the candidate combination includes a carrier frequency range and a pseudo-code phase range.
[0165] Step S202: Divide the carrier frequency range to obtain multiple carrier frequency blocks, and divide the pseudo-code phase range to obtain multiple pseudo-code phase blocks.
[0166] Step S203: Combine each carrier frequency block with each pseudo-code phase block respectively to obtain multiple operation combinations.
[0167] Specifically, the carrier frequency range is divided into 7 carrier frequency blocks, and the pseudo-code phase range is divided into 6 pseudo-code phase blocks, for a total of 6 * 7 = 42 operation combinations. That is to say, for the operation data sent by the CPU each time, a total of 42 parallel capture operation loops corresponding to the operation combinations are processed, and 433088 float point data are processed in each loop.
[0168] Step S204: For each operation combination, use multiple threads in the thread block to perform parallel capture operations on the operation combination and the spread spectrum signal to obtain the energy value corresponding to each operation combination.
[0169] The specific implementation process of Step S204 is divided into the following steps:
[0170] Step S301: Determine the current operation combination from each operation combination.
[0171] It should be noted that the same parallel capture operation process is looped for each operation combination, and after the parallel capture operation process of each operation combination is completed, the parallel capture operation process of the next operation combination is looped and executed. Therefore, it is first necessary to determine the current operation combination from each operation combination.
[0172] Step S302: Generate in-phase and quadrature components based on the carrier frequency block in the current operation combination.
[0173] In the specific implementation process of Step S302, use a carrier NCO (Numerically Controlled Oscillator) to generate the in-phase component I NCO and the quadrature component Q NCO .
[0174] Step S303: Based on each frequency point, in-phase component, and quadrature component in the spread-spectrum signal, generate multiple down-conversion calculation tasks, and allocate each down-conversion calculation task to each thread in the thread block for parallel calculation to obtain the down-converted spread-spectrum signal.
[0175] As Figure 10 shown, it is a schematic diagram of a parallel down-conversion algorithm disclosed in an embodiment of the present invention.
[0176] Among them, Block0, Block1, Block2,..., Block i ,..., represent multiple thread blocks in the GPU device. Each thread block contains multiple threads (Thread), marked as Th0, Th1, Th2,..., Th k .
[0177] The carrier NCO generates the in-phase component and quadrature component of the local carrier for mixing with the intermediate-frequency signal to complete the down-conversion operation.
[0178] Specifically, the down-conversion includes analog down-conversion and digital down-conversion. The spread-spectrum signal is subjected to analog down-conversion to obtain an intermediate-frequency signal, and the intermediate-frequency signal is subjected to digital down-conversion, that is, based on each frequency point, in-phase component, and quadrature component in the intermediate-frequency signal, multiple down-conversion calculation tasks are generated. Each down-conversion calculation task is used to calculate the product of the frequency point and the in-phase component, and the product of the frequency point and the quadrature component to obtain the down-converted spread-spectrum signal.
[0179] Among them, the down-converted spread-spectrum signal includes two paths of signals, namely the I-channel signal and the Q-channel signal, and these two paths of signals will be despread separately later.
[0180] Allocate each down-conversion calculation task to each thread in the thread block for parallel calculation, so that each thread corresponds to processing a down-conversion calculation task, and the entire operation is executed concurrently, improving the digital down-conversion efficiency.
[0181] Step S304: Based on the pseudo-code phase block in the current operation combination, generate multiple local pseudo-code sub-blocks with a preset phase difference.
[0182] It should be noted that after the down-conversion of the spread-spectrum signal, the most important and computationally intensive step is to despread the down-converted spread-spectrum signal. During the despreading process, it must be considered that the influence of the Doppler effect is not only in the frequency domain, but also has a Doppler effect on the signal pseudo-code. Different from the carrier Doppler that only affects the coherent accumulation length, the pseudo-code Doppler affects the entire accumulated length of the signal. As time goes by, the code phase difference between the received pseudo-code and the local pseudo-code gradually becomes larger, and after a period of time, a second-order phase difference appears, that is, the phase difference is greater than one chip again, and the signal cannot be despread, resulting in the inability to accumulate energy in the accumulation operation.
[0183] Considering this influence, in order to make the running of the pseudo-code phase difference between the received signal and the local pseudo-code phase less than 0.3 chips to ensure successful despreading. When designing the acquisition unit, it is necessary to remove the accumulation caused by the running of the pseudo-code phase and block the locally generated pseudo-code.
[0184] As Figure 11 shown, it is a schematic structural diagram of an acquisition unit disclosed in an embodiment of the present invention.
[0185] Specifically, a preset pseudo-code generation module is used to generate local pseudo-code blocks. When the pseudo-code generation module generates local pseudo-code blocks, within the pseudo-code phase block range, each local pseudo-code block is delayed by half a chip to remove the accumulated phase difference caused by pseudo-code Doppler to ensure successful despreading.
[0186] Step S305: Based on the down-converted spread-spectrum signal and each local pseudo-code block, generate multiple despreading calculation tasks, and allocate each despreading calculation task to each thread in the thread block for parallel calculation to obtain the despread signal.
[0187] It should be noted that the pseudo-code is known, the pseudo-code at the receiving end is the same as that at the transmitting end, the pseudo-code at the receiving end will be set in advance, and the down-converted spread-spectrum signal will be multiplied by each local pseudo-code block. When the maximum correlation value is obtained, it means that the pseudo-code of the down-converted spread-spectrum signal is aligned with the local pseudo-code block, and despreading is completed.
[0188] In specific implementation, each local pseudo-code block is a complete pseudo-code phase search process. In one search, the multi-thread parallelization advantage of the GPU device is also utilized to decompose the multiplication in the despreading process into different threads for operation, greatly improving the calculation efficiency of the despreading process.
[0189] Each thread multiplies the signal pseudo-code of the down-converted spread-spectrum signal by the local pseudo-code block to obtain the despread signal, that is, each thread corresponds to executing a despreading calculation task, greatly improving the despreading rate.
[0190] As Figure 12 shown, it is a structural diagram of a parallel despreading scheme disclosed in an embodiment of the present invention.
[0191] Among them, the intermediate-frequency sampling signal is the sampled intermediate-frequency signal (i.e., the down-converted spread-spectrum signal), expressed as a discrete digital signal. In the despreading process, this signal is the object to be processed, and the purpose is to recover the original narrowband information from it.
[0192] The sampling clock provides the clock signal required for sampling to ensure that the signal is sampled at the correct time point. The synchronization of the clock signal is crucial for accurate despreading.
[0193] Z -1 It means that the modern matched filter is implemented on digital signals, so the intermediate frequency sampled signal is converted into the discrete domain.
[0194] h(k)e jwk represents the complex coefficients of each matched filter, where h(k) is the impulse response of the filter, and e jwk is the complex exponential factor. h(k)e jwk is preset according to the corresponding local pseudo-code block. During the despreading process, these complex coefficients are used to perform correlation operations with the intermediate frequency sampled signal to recover the original signal.
[0195] The carrier generator is used to generate the required carrier signal and supply it to each matched filter.
[0196] During the despreading process, the multiplier is used to multiply the intermediate frequency sampled signal by the complex coefficients set according to the local pseudo-code block, and the despread signal is obtained when the maximum value is obtained.
[0197] Step S306: Integrate and clean the despread signal to obtain the integrated and cleaned signal.
[0198] It should be noted that integration and cleaning is a method of accumulating the despread signal to improve the signal-to-noise ratio. The scope involved in this application is digital signals. Integration and cleaning is to accumulate the number of points after despreading to improve the signal-to-noise ratio. Generally speaking, when the signal-to-noise ratio condition is relatively poor, the accumulation points of integration and cleaning will be increased accordingly.
[0199] In one embodiment, after integration and cleaning, the integrated and cleaned signal is first stored and then FFT transformation operation is performed.
[0200] It should be noted that the entire acquisition process is carried out in several kernel functions in sequence, including downconversion and despreading, integration and cleaning, FFT transformation operation, non-coherent accumulation and taking the maximum value. Each kernel function has its own shared memory. After a kernel function completes the calculation, it will first return the calculated data to its own shared memory. The next kernel function will take the operation result of the previous kernel function from the shared memory as input, and store it in the shared memory again after the calculation for the next kernel function to use.
[0201] Step S307: Perform FFT transformation operation on the integrated and cleaned signal to obtain the transformation result.
[0202] In the specific implementation process of step S307, the cuFFT library in the CUDA (Compute Unified Device Architecture) platform is used to perform FFT transformation operations on the signal after integral cleaning to obtain the transformation result.
[0203] After the FFT transformation operation, the frequency spectrum diagram of the time-domain waveform of the original signal will be obtained. The purpose of spread-spectrum signal acquisition is to obtain the carrier Doppler frequency shift and the pseudo-code phase. After the FFT transformation operation, the signal is transformed into the frequency domain, and the energy of the signal will be concentrated at the frequency corresponding to the Doppler frequency shift. Note that the spectrum after the FFT transformation at this time is the spectrum of the I / Q two branches. After subsequent amplitude calculation and non-coherent accumulation, peak searching will be performed, and the frequency corresponding to the Doppler frequency shift is at the peak.
[0204] It should be noted that the GPU device has great advantages in performing FFT transformation operations. The unique cuFFT library in the CUDA platform can produce an acceleration ratio dozens of times that of the CPU device when performing FFT transformation operations, greatly improving the overall operation efficiency.
[0205] Step S308: Based on each frequency point in the spread-spectrum signal and the transformation result, generate multiple energy calculation tasks, and allocate each energy calculation task to each thread in the thread block for parallel calculation to obtain the energy values of each frequency point.
[0206] In step S308, energy calculation also involves a large number of multiplication and addition operations. Using the concurrent advantage of the GPU device for parallel energy calculation, each thread performs the energy calculation of a single frequency point, and the energy calculation of all frequency points can be completed with a great acceleration ratio.
[0207] As Figure 13 shown, it is a structural diagram of a parallel energy calculation scheme disclosed in an embodiment of the present invention.
[0208] As can be seen from the above embodiments of the present invention, the down-converted spread-spectrum signal includes two signals, namely the I-channel signal and the Q-channel signal. After despreading, integral cleaning, and FFT transformation operations on the down-converted spread-spectrum signal, two FFT transformation results will be obtained, corresponding to the I-channel signal and the Q-channel signal respectively. Therefore Figure 13 I represents the FFT transformation result corresponding to the I-channel signal, and Q represents the FFT transformation result corresponding to the Q-channel signal.
[0209] Among them, each thread processes the energy calculation task of one frequency point, and each energy calculation task is used to calculate the result of taking the square root of the sum of the squares of I and Q to obtain the energy value of this frequency point.
[0210] Step S309: Incoherently accumulate the energy values of each frequency point to obtain the energy value corresponding to the current operation combination.
[0211] In the incoherently accumulating stage, a reduction algorithm commonly used in GPU devices is adopted to accelerate the rate of accumulation calculation. Reduction is a common operation in GPU programming, which is used to merge a data set into a single value through specific operations (such as summation, finding the maximum value, etc.). The reduction algorithm has natural data parallelism.
[0212] As Figure 14 shown, it is a schematic diagram of summation of a reduction algorithm disclosed in an embodiment of the present invention. In the preset reduction algorithm, each thread block processes a part of the data, and adjacent elements are merged through iteration, gradually reducing the data scale. This greatly improves the processing rate during incoherent accumulation. The specific principle is as follows:
[0213] 1. Initial data layer (Index1)
[0214] Contains 1024 data items, from a1 to a 1024 which is equivalent to the energy values of each frequency point.
[0215] 2. First accumulation layer (Index2)
[0216] Contains 512 data items, from b1 to b 512 .
[0217] Each data item is obtained by adding two adjacent data items in the initial data layer. For example, b1 = a1 + a2, b2 = a3 + a4, and so on.
[0218] 3. Subsequent accumulation layers (Index3 to Index10)
[0219] Continue to aggregate the data in the above - mentioned manner, halving the number of data items each time.
[0220] 4. Accumulation result layer (Index11)
[0221] Contains 1 data item, d1. This is the final result obtained by adding two data items in Index10 (i.e., the energy value corresponding to the current operation combination).
[0222] Step S310: Determine the next current operation combination from the remaining operation combinations, and return to execute Step S302 until the energy values corresponding to all operation combinations are obtained.
[0223] Among them, the remaining operation combinations refer to the operation combinations that have not participated in the parallel acquisition operation loop.
[0224] Step S205: Take the maximum value among the energy values corresponding to each operation combination as the local energy maximum value, and send the local energy maximum value to the CPU device.
[0225] It should be noted that after finding the local energy maximum value, it will be sent back to the CPU device, and then the global energy maximum value will be found in the CPU device. If the obtained global energy maximum value exceeds the set capture threshold, it indicates that the capture of the spread spectrum signal is completed, and the next stage of demodulation can be entered to perform signal tracking.
[0226] Preferably, for each operation combination in turn, use multiple threads in the thread block and the spread spectrum signal to perform parallel capture operations to obtain the energy value corresponding to the operation combination, and when the energy value corresponding to the operation combination is greater than the historical maximum value, use the energy value corresponding to the operation combination to update the historical maximum value; when the parallel capture operations for each operation combination are completed, take the historical maximum value as the local energy maximum value, and send the local energy maximum value to the CPU device.
[0227] Based on the above-mentioned method for capturing a direct sequence spread spectrum signal disclosed in the embodiments of the present invention, in this solution, the CPU device and the GPU device are used in cooperation to complete the capture. Through the relatively powerful logical processing ability of the CPU device and the parallel capture operation method using multiple threads in the GPU device, the calculation efficiency is improved, and the problem that the existing software demodulation technology cannot meet the requirements of real-time demodulation is solved.
[0228] As Figure 15 shown, it is a structural diagram of a device for capturing a direct sequence spread spectrum signal disclosed in the embodiments of the present invention. This device is applied to the CPU device, and the CPU device is connected to the GPU device, including: a preprocessing unit 1501, a selection unit 1502, a sending unit 1503, a first receiving unit 1504, a return unit 1505, a search unit 1506, and a determination unit 1507.
[0229] Among them, the preprocessing unit 1501 is used to receive the spread spectrum signal sent by the sending end and perform preprocessing on the spread spectrum signal.
[0230] In an embodiment, the preprocessing unit 1501 is specifically used for:
[0231] Convert the data type of the spread spectrum signal to the float type.
[0232] The selection unit 1502 is used to select a candidate combination from the candidate combination set and remove the candidate combination from the candidate combination set; the candidate combination set includes multiple candidate combinations, and each candidate combination is obtained by combining a preset plurality of carrier frequency ranges with each preset pseudo-code phase range in turn.
[0233] A transmitting unit 1503, configured to transmit the candidate combination and the preprocessed spread spectrum signal as operation data to a GPU device.
[0234] A first receiving unit 1504, configured to receive the local energy maximum value obtained by the GPU device through parallel acquisition operations based on the operation data.
[0235] A returning unit 1505, configured to return and execute the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set until there is no candidate combination in the candidate combination set, so as to obtain multiple local energy maximum values.
[0236] A searching unit 1506, configured to search for the global energy maximum value from each local energy maximum value.
[0237] In an embodiment, the searching unit 1506 is specifically configured to:
[0238] Search for the global energy maximum value from each local energy maximum value by using a preset reduction algorithm.
[0239] A determining unit 1507, configured to determine that the candidate combination corresponding to the global energy maximum value is the acquisition result if the global energy maximum value is greater than the acquisition threshold.
[0240] In an embodiment, the direct-sequence spread spectrum signal acquisition device further includes:
[0241] A fine acquisition unit, configured to, if the accuracy of the acquisition result does not meet the requirements, divide the carrier frequency range in the candidate combination corresponding to the global energy maximum value to obtain multiple new carrier frequency ranges, and divide the pseudo-code phase range in the candidate combination corresponding to the global energy maximum value to obtain multiple new pseudo-code phase ranges; combine each new carrier frequency range with each new pseudo-code phase range in sequence to obtain a candidate combination set composed of multiple new candidate combinations, and return and execute the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set.
[0242] Based on the direct-sequence spread spectrum signal acquisition device disclosed in the above embodiments of the present invention, in this solution, the CPU device and the GPU device are used in cooperation to complete the acquisition. Through the relatively powerful logical processing ability of the CPU device and the parallel acquisition operation method using multi-threads in the GPU device, the computing efficiency is improved, and the problem that the requirements of real-time demodulation cannot be met in the existing software demodulation technology is solved.
[0243] Such as Figure 16As shown in the figure, it is a structural diagram of another direct sequence spread spectrum signal acquisition device disclosed in an embodiment of the present invention. This device is applied to a GPU device, and the GPU device is connected to a CPU device. The device includes: a second receiving unit 1601, a dividing unit 1602, a combining unit 1603, an arithmetic unit 1604, and a feedback unit 1605.
[0244] Among them, the second receiving unit 1601 is used to receive the arithmetic data sent by the CPU; the arithmetic data includes candidate combinations and preprocessed spread spectrum signals; the candidate combinations include carrier frequency ranges and pseudo-code phase ranges.
[0245] The dividing unit 1602 is used to divide the carrier frequency range to obtain multiple carrier frequency blocks, and divide the pseudo-code phase range to obtain multiple pseudo-code phase blocks.
[0246] The combining unit 1603 is used to combine each carrier frequency block with each pseudo-code phase block respectively to obtain multiple arithmetic combinations.
[0247] The arithmetic unit 1604 is used to perform parallel acquisition arithmetic on the arithmetic combination and the spread spectrum signal for each arithmetic combination by using multiple threads in a thread block to obtain the energy value corresponding to each arithmetic combination.
[0248] In one embodiment, the arithmetic unit 1604 is specifically used for:
[0249] Determine the current arithmetic combination from each arithmetic combination;
[0250] Generate an in-phase component and a quadrature component based on the carrier frequency block in the current arithmetic combination;
[0251] Generate multiple down-conversion calculation tasks based on each frequency point, the in-phase component and the quadrature component in the spread spectrum signal, and allocate each down-conversion calculation task to each thread in the thread block for parallel calculation to obtain the down-converted spread spectrum signal;
[0252] Generate multiple local pseudo-code sub-blocks with a preset phase difference based on the pseudo-code phase block in the current arithmetic combination;
[0253] Generate multiple despreading calculation tasks based on the down-converted spread spectrum signal and each local pseudo-code sub-block, and allocate each despreading calculation task to each thread in the thread block for parallel calculation to obtain the despread signal;
[0254] Perform integral cleaning on the despread signal to obtain the integrally cleaned signal;
[0255] Perform FFT transform arithmetic on the integrally cleaned signal to obtain the transform result;
[0256] Based on each frequency point in the spread spectrum signal and the transformation result, generate multiple energy calculation tasks, and allocate each energy calculation task to each thread in the thread block for parallel calculation to obtain the energy values of each frequency point;
[0257] Incoherently accumulate the energy values of each frequency point to obtain the energy value corresponding to the current operation combination;
[0258] Determine the next current operation combination from the remaining operation combinations, and return to execute the step of generating the in-phase component and the quadrature component based on the carrier frequency block in the current operation combination until the energy values corresponding to each operation combination are obtained.
[0259] In one embodiment, the operation unit 1604 for incoherently accumulating the energy values of each frequency point to obtain the energy value corresponding to the current operation combination is specifically configured to:
[0260] Use a preset reduction algorithm to incoherently accumulate the energy values of each frequency point to obtain the energy value corresponding to the current operation combination.
[0261] In one embodiment, the operation unit 1604 for performing FFT transformation operation on the signal after integral cleaning to obtain the transformation result is specifically configured to:
[0262] Use the cuFFT library in the CUDA platform to perform FFT transformation operation on the signal after integral cleaning to obtain the transformation result.
[0263] The feedback unit 1605 is configured to use the maximum value among the energy values corresponding to each operation combination as the local energy maximum value, and send the local energy maximum value to the CPU device.
[0264] Based on the disclosed direct sequence spread spectrum signal acquisition device in the above embodiments of the present invention, in this solution, the CPU device and the GPU device are used in cooperation to complete the acquisition. Through the relatively powerful logical processing ability of the CPU device and the parallel acquisition operation method using multi-threading in the GPU device, the calculation efficiency is improved, and the problem that the existing software demodulation technology cannot meet the requirements of real-time demodulation is solved.
[0265] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding description in the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0266] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0267] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for capturing a direct sequence spread spectrum signal, characterized in that: Applied to a CPU device, the CPU device is connected to a GPU device, and the method includes: receiving a spread spectrum signal sent by a transmitting end, and preprocessing the spread spectrum signal; Selecting a candidate combination from a candidate combination set, and removing the candidate combination from the candidate combination set; the candidate combination set includes a plurality of candidate combinations, each of which is obtained by sequentially combining a plurality of preset carrier frequency ranges with each preset pseudo code phase range; Sending the candidate combination and the preprocessed spread spectrum signal as operation data to the GPU device; Receiving a local energy maximum value obtained by the GPU device performing a parallel capture operation based on the operation data; Returning to the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set, until the candidate combination does not exist in the candidate combination set, and obtaining a plurality of local energy maxima; Searching for a global energy maximum from each of the local energy maxima; If the global energy maximum is greater than the capture threshold, the candidate combination corresponding to the global energy maximum is determined as the capture result.
2. The method according to claim 1, characterized in that The preprocessing of the spread spectrum signal comprises: Convert the data type of the spread spectrum signal to float type.
3. The method according to claim 1, characterized in that The step of searching for the global energy maximum value from the local energy maxima comprises: The global energy maximum value is obtained by searching from each of the local energy maxima using a preset reduction algorithm.
4. The method according to any one of claims 1 to 3, characterized in that: After determining that the candidate combination corresponding to the global energy maximum value is a capture result, the method further includes: If the accuracy of the capture result does not meet the requirements, the carrier frequency range in the candidate combination corresponding to the global energy maximum value is divided to obtain multiple new carrier frequency ranges, and the pseudo code phase range in the candidate combination corresponding to the global energy maximum value is divided to obtain multiple new pseudo code phase ranges; Combine each of the new carrier frequency ranges with each of the new pseudo code phase ranges in turn to obtain a candidate combination set consisting of a plurality of new candidate combinations, and return to execute the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set.
5. A method for capturing a direct sequence spread spectrum signal, characterized in that: Applied to a GPU device, the GPU device is connected to a CPU device, and the method includes: receiving operation data sent by the CPU; the operation data includes a candidate combination and a preprocessed spread spectrum signal; the candidate combination includes a carrier frequency range and a pseudo code phase range; Dividing the carrier frequency range to obtain a plurality of carrier frequency blocks, and dividing the pseudo code phase range to obtain a plurality of pseudo code phase blocks; Combining each of the carrier frequency blocks with each of the pseudo code phase blocks to obtain a plurality of operation combinations; For each of the operation combinations, using multiple threads in a thread block, perform parallel capture operations on the operation combination and the spread spectrum signal to obtain energy values corresponding to each of the operation combinations; The maximum value among the energy values corresponding to the operation combinations is taken as the local energy maximum value, and the local energy maximum value is sent to the CPU device.
6. The method according to claim 5, characterized in that For each of the operation combinations, using multiple threads in a thread block to perform parallel capture operations on the operation combination and the spread spectrum signal to obtain energy values corresponding to each of the operation combinations, including: Determine a current operation combination from each of the operation combinations; generating an in-phase component and a quadrature component based on the carrier frequency block in the current operation combination; Based on each frequency point, the in-phase component and the orthogonal component in the spread spectrum signal, a plurality of down-conversion calculation tasks are generated, and each of the down-conversion calculation tasks is assigned to each thread in a thread block for parallel calculation to obtain a spread spectrum signal after down-conversion; Based on the pseudo code phase block in the current operation combination, generating a plurality of local pseudo code blocks with a preset phase difference; Based on the down-converted spread spectrum signal and each of the local pseudo code blocks, a plurality of despreading calculation tasks are generated, and each of the despreading calculation tasks is assigned to each of the threads in the thread block for parallel calculation to obtain a despread signal; Performing integral cleaning on the despread signal to obtain an integral cleaned signal; Performing FFT transformation operation on the signal after the integral cleaning to obtain a transformation result; Based on each frequency point in the spread spectrum signal and the transformation result, a plurality of energy calculation tasks are generated, and each of the energy calculation tasks is assigned to each of the threads in the thread block for parallel calculation to obtain an energy value of each of the frequency points; Incoherently summing the energy values of the respective frequency points to obtain an energy value corresponding to the current operation combination; The next current operation combination is determined from the remaining operation combinations, and the step of generating an in-phase component and an orthogonal component based on the carrier frequency block in the current operation combination is returned to be executed until the energy values corresponding to the respective operation combinations are obtained.
7. The method according to claim 6, characterized in that The non-coherently accumulating the energy values of the respective frequency points to obtain the energy value corresponding to the current operation combination includes: By using a preset reduction algorithm, the energy values of the various frequency points are incoherently accumulated to obtain the energy value corresponding to the current operation combination.
8. The method according to claim 6, characterized in that The performing of FFT transformation operation on the signal after the integral cleaning to obtain a transformation result includes: The cuFFT library in the CUDA platform is used to perform FFT transformation operation on the signal after the integral cleaning to obtain a transformation result.
9. A device for capturing a direct sequence spread spectrum signal, characterized in that: Applied to a CPU device, the CPU device is connected to a GPU device, and the device comprises: A preprocessing unit, used for receiving a spread spectrum signal sent by a transmitting end, and preprocessing the spread spectrum signal; A selection unit is used to select a candidate combination from a candidate combination set and remove the candidate combination from the candidate combination set; the candidate combination set includes a plurality of candidate combinations, each of which is obtained by sequentially combining a plurality of preset carrier frequency ranges with each preset pseudo code phase range; A sending unit, configured to send the candidate combination and the preprocessed spread spectrum signal as operation data to the GPU device; A first receiving unit is used to receive a local energy maximum value obtained by the GPU device through parallel capture operation based on the operation data; a returning unit, configured to return to the step of selecting a candidate combination from the candidate combination set and removing the candidate combination from the candidate combination set until the candidate combination does not exist in the candidate combination set, thereby obtaining a plurality of local energy maxima; A searching unit, used for searching for a global energy maximum value from each of the local energy maxima; A determination unit is used to determine that the candidate combination corresponding to the global energy maximum value is a capture result if the global energy maximum value is greater than a capture threshold.
10. A device for capturing a direct sequence spread spectrum signal, characterized in that: Applied to a GPU device, the GPU device is connected to a CPU device, and the device comprises: A second receiving unit is used to receive operation data sent by the CPU; the operation data includes a candidate combination and a pre-processed spread spectrum signal; the candidate combination includes a carrier frequency range and a pseudo code phase range; A dividing unit, used for dividing the carrier frequency range to obtain a plurality of carrier frequency blocks, and dividing the pseudo code phase range to obtain a plurality of pseudo code phase blocks; A combining unit, used for combining each of the carrier frequency blocks with each of the pseudo code phase blocks to obtain a plurality of operation combinations; An operation unit, configured to perform parallel capture operation on each of the operation combinations by using a plurality of threads in a thread block to obtain energy values corresponding to each of the operation combinations; The feedback unit is used to take the maximum value of the energy values corresponding to each of the operation combinations as the local energy maximum value, and send the local energy maximum value to the CPU device.
Citation Information
Patent Citations
Pseudo code capturing method and capturing device using multiple antennae of direct sequence spread spectrum system
CN101702628A
CPU-assisted GPU spread spectrum signal fast acquisition realization method
CN105577229A