Systems and methods for combining a plurality of downlink signals representing communication signals

By employing general-purpose processors with multiple cores and SIMD capabilities for distributed computing, the communication system achieves efficient high-speed processing of satellite communication signals, addressing the inefficiencies and costs associated with traditional systems.

JP7700123B2Active Publication Date: 2025-06-30KRATOS INTEGRAL HOLDINGS LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022536588
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-16
Filing Date
2020-12-16
Publication Date
2025-06-30
Estimated Expiration
2040-12-16

AI Technical Summary

Technical Problem

Existing communication systems require large-scale ground stations and dedicated hardware for processing satellite communication signals, which is costly and inefficient.

Method used

The use of general-purpose processors with multiple cores and SIMD capabilities for distributed computing, employing techniques such as feed-forward processing, pre-computation of metadata, and aggregation of functions to achieve high-speed signal processing without dedicated signal processing hardware.

Benefits of technology

This approach enables efficient high-speed processing of satellite communication signals, reducing the need for extensive ground stations and dedicated hardware, thereby lowering costs and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700123000001
    Figure 0007700123000001
  • Figure 0007700123000002
    Figure 0007700123000002
  • Figure 0007700123000003
    Figure 0007700123000003
Patent Text Reader

Abstract

Provided herein are embodiments of systems and methods for combining downlink signals representing communication signals. An exemplary method includes receiving a downlink signal from an antenna feed. The method includes, in a first processing block within a processor, performing a first blind detection operation on a first packet of a first signal and performing a first Doppler compensation operation on the first packet. The method also includes, in a second processing block within the processor parallel to the first processing block, performing a second blind detection operation on a second packet of a second signal and performing a second Doppler compensation operation on the second packet. The method also includes combining the first and second signals based on (i) aligning the first data packet with the second data packet and (ii) performing a weighted combiner operation that applies scaling to the first and second data packets based on corresponding signal qualities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 948,599, filed on December 16, 2019, which is hereby incorporated by reference as if fully set forth herein.

[0002] This disclosure relates to signal processing. More specifically, this disclosure relates to performing distributed computing using a general - purpose processor to achieve high - speed processing.

Background Art

[0003] In some examples, satellite communication signals may require large - scale ground stations and other equipment to transmit and / or receive and process data locally. This can include extensive antenna arrays, associated radio - frequency terminals (RFTs), and significant electronics (modems, signal processors, etc.) for receiving, processing, and using data received from associated satellites.

Summary of the Invention

[0004] This disclosure provides an improved communication system. The following summary is not intended to define all aspects of the invention, and other features and advantages of this disclosure will become apparent from the following detailed description, including the drawings. This disclosure is intended to be associated as a unified document, and it should be understood that all combinations of the features described herein are contemplated, even if the combinations of features are not found together in the same sentence, paragraph, or section of this disclosure. Further, this disclosure includes, as additional aspects, all embodiments of the invention that are narrower in scope in any form than the specifically recited variations described herein.

[0005] As disclosed herein, digital signal processing (DSP) can be performed in many different ways using a general-purpose processor or a central processing unit (CPU). Exemplary techniques for performing the disclosed functions on a general-purpose processor to achieve high-speed processing include, but are not limited to, using multiple CPUs and parallel processing on many cores of each CPU, using single instruction multiple data (SIMD) techniques, feed-forward processing to break up feedback loops, pre-computation of metadata (or state information) to split heavy processing among several CPUs, and aggregation of multiple functions into a single function to improve CPU performance or reduce memory bandwidth usage.

[0006] One way to improve the throughput of a general-purpose CPU is to utilize as many of the cores present on the CPU as possible. Care must be taken to ensure that data is properly shared among some of the cores within the CPU, but this enables an increase in processing throughput with the addition of more CPU cores. It is also possible to use multiple CPUs on the same system, with each CPU containing multiple cores. All embodiments within this disclosure utilize the use of multiple cores within a CPU, and some embodiments utilize having multiple CPUs for each system and / or group of systems in a server environment.

[0007] Another way to achieve high processing speeds is to utilize the single instruction multiple data (SIMD) capabilities of general-purpose CPUs. This enables a single CPU core to execute up to 16 floating-point operations for a single instruction, as in the case of AVX512 SIMD operations. One example of using SIMD is to use a finite impulse response (FIR) filter function where 16 floating-point results are calculated at once. Another example is when multiplying complex numbers. Instead of calculating a pair of orthogonal signals (IQ data), it is possible to calculate eight IQ pairs at once using AVX512. Complex multiplication is used in almost all processing algorithms described in this disclosure. Other examples of using SIMD include correlators in diversity combiners, decimation in signal analyzers, and again adjustment in channelizers / combiners.

[0008] Some processing systems implement various forms of feedback, often including a phase-locked loop (PLL) or a delay-locked loop (DLL). However, generally, as in the case of PLLs and DLLs, feedback can be problematic because the nature of the feedback itself causes a bottleneck. The feedback loop processes all incoming data in a single (e.g., linear) process that does not easily spill or split. In addition to feedback, there are other obstacles to overcome using PLLs and DLLs, including how often the error term is calculated. The feedback loop can be replaced with a feedforward loop, in which the error state can be processed over a block of data and then the calculated error term is feedforwarded to another block that applies the error term. When appropriate redundancy is used, the error calculation and application of that term can be split across several CPU cores to further increase throughput. An example of this is a diversity combiner, where timing and phase correction are calculated in one block, timing adjustment is applied in another block, and phase correction is applied in yet another block. This method as a set can then be parallelized across several CPU cores to further increase throughput.

[0009] In addition to the feed-forward approach for processing data, it may be beneficial to perform pre-computation of metadata within a single block and then divide the processing of the data across several CPU cores. This method is similar to the feed-forward method already described, but in this case, rather than splitting a loop (such as a feedback loop), it simply utilizes multiple CPU cores to increase the amount of data that can be processed. In this way, the block that performs the pre-computation does not perform CPU-intensive processing, but calculates the necessary steps such as the iterations within a for loop and the starting index, as well as the slope points between interpolated phase values. One such example is the Doppler compensation performed in a diversity combiner. The required phase adjustment is created in a first block, but the CPU-intensive calculations for performing the phase adjustment are handed off downstream to subsequent blocks. If the second part of the processing is the CPU-intensive part, this allows the utilization of any number of CPU cores, and thus can increase the processing speed that could not otherwise be achieved within a single block.

[0010] Another technique that can be used on general-purpose CPUs to achieve high throughput is the way in which a set of functions are used and the type of memory used. In some cases, the memory bandwidth becomes a limiting factor in performance. If so, the goal is to limit the amount of data that needs to be transferred between random access memory (RAM) (not high-speed memory such as the CPU cache). To do this, the functions need to be folded so that they are all executed together, rather than being executed individually with the aim of accessing slower RAM as little as possible compared to accessing the faster CPU cache. Another way to reduce memory bandwidth is to utilize an appropriate spatial memory type, for example, using int8 where possible instead of floating or double.

[0011] In one embodiment, a method for combining a plurality of downlink signals representing communication signals is provided herein. The method includes receiving a plurality of downlink signals from a plurality of antenna feeds. The method also includes performing a first blind detection operation on a first data packet of a first signal among the plurality of signals in a first one or more processing blocks within one or more processors, and performing a first Doppler compensation operation on the first data packet of the first signal among the plurality of signals. Further, the method includes performing a second blind detection operation on a second data packet of a second signal among the plurality of signals in a second one or more processing blocks within one or more processors in parallel with the first one or more processing blocks, and performing a second Doppler compensation operation on the second data packet of the second signal among the plurality of signals. The method also includes combining the first signal and the second signal based on (i) aligning the timing and phase of the first data packet with the second data packet, and (ii) performing a weighted combiner operation of applying scaling to each of the first and second data packets based on corresponding signal quality.

[0012] In another embodiment, a system for combining a plurality of downlink signals representing communication signals is provided. The system includes a plurality of antennas configured to receive a plurality of downlink signals, and one or more processors communicatively coupled to the plurality of antennas, the one or more processors having a plurality of processing blocks. The one or more processors are operable to perform the method described above.

Brief Description of the Drawings

[0013] Details of the present invention can be collected in part by considering the accompanying drawings with respect to both its structure and operation, in which like reference numerals refer to like parts.

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figures 6 - 8

Figures 9 - 10

Figure 11

Figure 12

Figure 13

Figures 14A - B

Figure 15

Figure 16

DETAILED DESCRIPTION OF THE INVENTION

[0015] Embodiments of an improved communication system for achieving high-speed processing using a general-purpose processor are disclosed. The embodiments disclosed herein provide an improved communication system that can efficiently achieve high-speed signal processing using a general-purpose processor. Reading this specification will make it apparent to those skilled in the art how to implement the present invention in various alternative embodiments and alternative uses. However, while various embodiments of the present invention are described herein, these embodiments are presented by way of example and illustration only, and not by way of limitation. Accordingly, this detailed description of the various embodiments should not be construed as limiting the scope or breadth of the present invention as set forth in the appended claims.

[0016] Throughout this specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0017] The communication system has been used as a main example throughout the description, but the application of the disclosed method is not so limited. For example, any wireless or wireline communication system that requires the use of digital signal processing, modems, etc. can implement the systems, methods, and computer-readable media described herein.

[0018] The present disclosure provides systems and methods for performing digital signal processing using a general-purpose central processing unit (CPU) in either a standard server environment or a virtualized cloud environment. In some examples, the system can use single instruction multiple data (SIMD) techniques to achieve high throughput, including SSE, SSE2, SSE3, SSE4.1, SSE4.2, AVX, AVX2, and AVX512 instruction sets. The present disclosure describes how data processing is managed across multiple processing cores of a processor (e.g., a CPU) to achieve the required throughput without using dedicated signal processing hardware such as a field programmable gate array (FPGA), or high-performance computing (HPC) hardware such as a graphics processing unit (GPU). The ability to perform this processing in general-purpose server CPUs, including but not limited to x86 architectures manufactured by Intel and AMD microprocessors, and ARM processors such as Cortex-A76, NEON, and AWS Graviton and Graviton2, enables the deployment of functionality within a virtualized processing architecture in a general-purpose cloud processing environment without the need for dedicated hardware. Processing in the general-purpose CPU is enabled by a digital IF device that samples an analog signal and supplies the digitized samples to the CPU via an Ethernet connection. The digital IF device can also receive the digitized samples and convert them to an analog signal, which is similar to that described in U.S. Patent No. 9,577,936, entitled "Packetized Radio Frequency Transport System," issued on February 21, 2017, the contents of which are incorporated herein by reference in their entirety.

[0019] FIG. 1 is a schematic representation of an embodiment of a communication system. A communication system (system) 100 can have a platform 110 and a satellite 111 that communicate with a plurality of ground stations. The platform 110 can be an aircraft (e.g., an airplane, a helicopter, or an unmanned aerial vehicle (UAV), etc.). The plurality of ground stations 120, 130, 140 can be associated with a ground radio frequency (RF) antenna 122 or one or more satellite antennas 132, 142. The ground station 120 can have an antenna 122 coupled to a digitizer 124. The digitizer 124 can have one or more analog / digital converters (A2D) for converting an analog signal received by the antenna 122 into a digital bit stream for transmission over a network. The digitizer 124 can also include a corresponding digital-to-analog converter (D2A) for operation on an uplink to the platform 110 and the satellite 111.

[0020] Similarly, the ground station 130 can have an antenna 132 and a digitizer 134, and the ground station 140 can have an antenna 142 and a digitizer 144.

[0021] The ground stations 120, 130, 140 can each receive, in a receive chain, downlink signals 160 (labeled 160a, 160b, 160c) from the platform 110 and downlink signals 170 (labeled 170a, 170b, 170c) from the satellite 111. The ground stations 120, 130, 140 can also transmit uplink signals via their respective antennas 122, 132, 142 within a transmit chain. The digitizers 124, 134, 144 can digitize the received downlink signals 160, 170 for transmission as a digital bit stream 152. The digital bit stream 134 can then be transmitted to a cloud processing system via a network 154.

[0022] In some examples, the terrestrial stations 120, 130, 140 can locally process all data (e.g., included in the downlink signal), but this can be very expensive in terms of time, resources, and efficiency. Thus, in some embodiments, the downlink signal can be digitized and transmitted as a digital bitstream 152 to a remote signal processing server (SPS) 150. In some implementations, the SPS 150 can be located at a physical location such as a data center located at an offsite facility accessible via a wide area network (WAN). Such a WAN can be, for example, the Internet. The SPS 150 can demodulate the downlink signal from the digital bitstream 152 and output data or information bits from the downlink signal. In some other implementations, the SPS 150 can use cloud computing or cloud processing to perform the signal processing and other methods described herein. The SPS 150 can also be referred to as a cloud server.

[0023] The SPS 150 can then provide the processed data to the user or transmit it to another site. The data and information can be mission-dependent. Additionally, the information contained in the data can be the main purpose of the satellite, including weather data, image data, and satellite communication (SATCOM) payload data. As described above, SATCOM is used as a main example in this specification, but any communication or signal processing system using DSP can implement the methods described herein.

[0024] To achieve high processing speeds using software, phase-locked loop (PLL) or delay-locked loop (DLL) approaches can be problematic due to feedback within the loop. The feedback loop processes all of the incoming data (e.g., downlink signal 132) in a single (e.g., linear) process that does not easily spill or divide. In addition to feedback, there are other obstacles to overcome using PLL / DLL, including, for example, how often to calculate the error term.

[0025] Figure 2 is a functional block diagram of a wired or wireless communication device for use as one or more components of the system of FIG. 1. Processing device (device) 200 may be implemented, for example, as SPS 150 of FIG. 1. Device 200 may be implemented as needed to perform one or more of the signal processing methods or steps disclosed herein.

[0026] Device 200 may include a processor 202 that controls the operation of device 200. Processor 202 may also be referred to as a CPU. Processor 202 can, for example, direct and / or perform functions resulting from SPS 150. Some aspects of device 200 that include processor 202 can be implemented as various cloud-based elements, such as cloud-based processing. Thus, processor 202 can represent cloud processing distributed across several different processors via a network (e.g., the Internet). Alternatively, some components can be implemented in hardware. Processor 202 can be implemented in any combination of one or more of a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gate logic, discrete hardware components, a dedicated hardware finite state machine, or any other suitable entity capable of performing calculations or other operations on information.

[0027] Processor 202 can have one or more cores 204 (shown as cores 204a through 204n) that can execute computations. In implementations that use cloud processing, cores 204 can represent multiple iterations of distributed cloud processing. In some embodiments, using hardware, processor 202 can be a complex integrated circuit in which all computations for the receiver are performed. As used herein, each core 204 can be one processing element of processor 202. Processor 202 can implement multiple cores 204 to perform the parallel processing required by the methods disclosed herein. In some embodiments, processor 202 can be distributed across multiple CPUs, such as in cloud computing.

[0028] Device 200 can further include a memory 206 operatively coupled to processor 202. Memory 206 can be cloud-based storage or local hardware storage. Memory 206 can include both read-only memory (ROM) and random access memory (RAM) that provide instructions and data to processor 202. A portion of memory 206 can also include non-volatile random access memory (NVRAM). Processor 202 typically executes logical and arithmetic operations based on program instructions stored in memory 206. The instructions in memory 206 may be executable to implement the methods described herein. Memory 206 can further include removable media or multiple distributed databases.

[0029] Memory 206 may also include a machine-readable medium for storing software. The software shall be broadly construed to mean any type of instruction, whether it is called software, firmware, middleware, microcode, a hardware description language, or the like. The instructions may include code (e.g., source code format, binary code format, executable code format, or any other suitable code format). When the instructions are executed by processor 202 or one or more cores 204, they cause device 200 (e.g., SPS150) to perform the various functions described herein.

[0030] Device 200 may also include a transmitter 210 and a receiver 212 to enable the transmission and reception of data between the communication device 200 and a remote location. Such communication may occur, for example, between ground station 120 and SPS150 via network 124. Such communication may be wireless or may occur via wired communication. Transmitter 210 and receiver 212 may be combined into a transceiver 214. Transceiver 214 can be communicatively coupled to network 124. In some examples, transceiver 214 can include or be part of a network interface card (NIC).

[0031] Device 200 may further include a user interface 222. User interface 222 may include a keypad, a microphone, a speaker, and / or a display. User interface 222 may include any element or component that conveys information to and / or receives input from a user of device 200.

[0032] The various components of device 200 described in the specification can be coupled to each other by bus system 226. Bus system 226 can include, for example, a data bus, and in addition to the data bus, a power bus, a control signal bus, and a status signal bus. In some embodiments, bus system 226 can be communicatively coupled to network 124. Network 124 can provide, for example, a communication link between device 200 (e.g., processor 202) and ground station 120. Those skilled in the art will understand that the components of device 200 can be coupled together, or can receive or provide inputs to each other using some other mechanism, such as a local area network or a wide area network for distributed processing.

[0033] FIG. 3 is a schematic block diagram illustration of feedforward or precomputed signal processing 300. Method 300 can be performed as a generalized process incorporating a plurality of functions by, for example, processor 202. Processor 202 can perform the plurality of functions in serial or parallel arrangements, as shown, to execute one or more desired processes. Each function can be executable by processor 202 and can refer to a block or set of instructions or software stored in memory 206.

[0034] The first function 302 can be executed by processor 202. In some embodiments, following the first function 302, a second function 304 can be executed in series. Accordingly, processor 202 can divide blocks of data for different functions to be processed on the plurality of cores 204 to execute the first function 302 and the second function 304.

[0035] Following the second function 304, the processor 202 can execute the distributed processing of the third function 306 (shown as 306a, 306b... 306n) in parallel. The parallel processing of the third function 306 can include, for example, dividing a block of data associated with the same function across several cores 204 (e.g., processing blocks) of the processor 202. For example, a "block of data" can mean a group of samples that need to be processed.

[0036] Next, the processor 202 can execute the fourth function 308 and the fifth function 309 in series. Similar to the first function 302 and the second function 304, the serial performance of the fourth function 308 and the fifth function 309 can include dividing blocks of data associated with different functions for processing across multiple cores 204. Generally, each of the first function 302, the second function 304, the third function 306, the fourth function 308, and the fifth function 309 can be executed in different processing blocks. As used herein, a processing block can refer to a specific task executed on a block of data. A processing block can be associated with, for example, one or more of the cores 204.

[0037] Accordingly, the method 300 can, for example, divide blocks of data regarding the same function for processing across multiple cores 204. Similarly, the method 300 can divide blocks of data regarding different functions for processing across multiple cores 204.

[0038] In some other implementations of the method 300, the same processing block (e.g., core 204) can execute the processing of data using single instruction multiple data (SIMD), regardless of whether it is the same function or different functions.

[0039] In other implementations, embodiments of method 300 can support data processing blocks with minimal state information by using overlapping data. As used herein, state information can include variables required during feedback (e.g., feedback processing), data frame boundaries, and the like. For example, in the case of a feedback loop, state information can include variables calculated within the loop that are required during feedback when processing a continuous data stream. State information can also include the position of frame boundaries within the data stream. Other examples can include things such as FIR filters, where state information includes values stored in buffers (e.g., perhaps many delay elements) required to maintain a continuous data flow.

[0040] By ignoring the overlapping portions of adjacent blocks of state information and data, the process can utilize parallel processing using a variable level of overlap between blocks of data.

[0041] FIG. 4 is a graphical representation of one embodiment of a method for feedforward or precomputed signal processing of FIG. 3. Method 400 can use the principles of method 300 for serial - parallel and / or parallel - serial processing for multiple functions. In one example, the first function 302 (FIG. 3) can be a data capture function 305 where the processor 202 receives data for processing. The second function 304 (FIG. 3) can be a data splitting function 310, and the processor 202 can analyze the data within the overlapping blocks of data. Then, the overlapping blocks of data can be processed in parallel in various parallel iterations of third functions 306a - 306n as processing blocks 315a - 315n. The overlap in the blocks of data can provide a level of redundancy that is less (or not at all) dependent on state information. The less state information required, the easier it is to process the blocks of data in parallel as opposed to a continuous stream.

[0042] Method 400 can further include a data combining function 320 similar to the fourth function 308 (FIG. 3) that combines processed data, and a data output function 325 similar to the fifth function 309 (FIG. 3).

[0043] In a further example, the adjustable serial - parallel or parallel - serial arrangements of the various functions of method 300 provide several ways to perform feed - forward processing to replace a feedback loop. This is advantageous because it can increase throughput and avoid bottlenecks caused by the latency of feedback processing.

[0044] A further advantage of the serial - parallel or parallel - serial processing provided by methods 300 and 400 is that by placing one or more of the desired algorithms within a processing block (e.g., one of the five processing blocks of method 300), the processor 202 can distribute the processing load (e.g., across multiple cores 204) without worrying about the speed of a given algorithm within the processing block (e.g., core 204). Thus, each core 204 shares exactly the same processing load and eliminates bottleneck problems caused by individual algorithms.

[0045] A further advantage of embodiments of method 300 can include customizing the specific order of algorithms (e.g., processing blocks) to reduce the computational load within the processor 202. As will be explained below, the overall multi - stage processing of a given process may not be constrained by the order of multiple sub - processes. Thus, in some examples, ordering the fourth function 308 may have an advantage if it is executed before the third function 306.

[0046] Method 300 can further implement different variable types for memory bandwidth optimization, such as, for example, int8, int16, and float. This can accelerate some algorithms (e.g., based on the type). Additionally, this can provide an increase in flexibility for maximizing memory bandwidth.

[0047] FIG. 5 is a functional block diagram of one embodiment of a digital signal diversity combiner. Method 500 for diversity combining can include feedforward block processing as described above. Method 500 includes a plurality of blocks. In some examples, each block can represent a processing block and can perform functions in a manner similar to processing blocks 315a, 315b... 315n (FIG. 4). In another example, the plurality of blocks can be grouped together as a single "processing block" that performs functions in a manner similar to processing blocks 315a, 315b... 315n (FIG. 4).

[0048] Figures 12 to 16 are functional block diagrams of various embodiments of a channelizer and a combiner. The methods shown in FIGS. 12 to 14B illustrate an exemplary process including pre-computed signal processing. Similar to method 500, one or more of methods 1200, 1300, 1400a, 1400b, 1500, and / or 1600 can include multiple processing blocks. In some examples, each block can represent a processing block and can perform functions in a similar manner to processing blocks 315a, 315b... 315n (FIG. 4). In another example, multiple blocks can be grouped together as a single "processing block" that performs functions in a similar manner to processing blocks 315a, 315b... 315n (e.g., FIG. 4). For example, FIG. 15 schematically shows an exemplary processing block implemented as a channelizer processing block 1500, and FIG. 16 schematically shows an exemplary pre-computed processing block implemented as a combiner processing block 1600. The sub-elements or blocks of block 1500 may be executed individually as shown in FIG. 15 or may be combined into a single block. Similarly, the sub-elements or blocks of block 1600 may be executed individually as shown in FIG. 16 or may be combined into a single block.

[0049] In block 305, SPS 150 can capture or otherwise receive digital bitstream 134 (e.g., via network 124). Data capture in block 305 can receive digital bitstream 134 data from a network connection (e.g., Ethernet).

[0050] In block 310, the data can be separated into parallel data streams by a data splitter. In some embodiments, the processor 202 can perform the data separation function required at block 310. In some other embodiments, a separate data separation component (e.g., a data splitter) can be included in the device 200 (FIG. 2). Separating the data into multiple parallel streams enables parallel processing of the downlink signal 132. Thus, method 300 can utilize feedforward or precomputation processing to break incoming digitized signal data into smaller fragments and then process them on multiple cores 204. The digital bitstream 134 can be separated to form overlapping packets in in-phase / quadrature (I / Q) pairs. In some embodiments, the "overlapping packets" can include data packets where consecutive packets overlap with adjacent data packets. In some embodiments, the data packets may all be of the same length but may overlap. The overlap of the data packets can be at the beginning or end of the data packet. Additionally, a data packet can overlap with both the preceding and the following data packets. Also, the data packets can have different lengths (e.g., different amounts of data). Thus, the first packet transmitted to processing block 315a may overlap with some of the data of the second packet transmitted to processing block 315b or may repeat otherwise.

[0051] The amount of duplication between packets, or the duplication size, is programmable and can be set as needed. In some examples, the duplication can be set to 1 percent (1%) of the packet size. This duplication size can be increased or decreased as needed. For example, one particular parameter that can affect the duplication size is the uncertainty in the symbol rate in data stream 134. For most signals, the worst-case uncertainty is less than 1%, and thus, 1% covers most cases. In some other embodiments, the duplication can be, as needed, on the order of 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%, or anywhere in between. Similarly, it is also possible to have a duplication of less than 1%. If the uncertainty in the data rate is less than 0.1%, the duplication can be 0.1% or less.

[0052] Processor 202 can perform single instruction multiple data (SIMD) processing on digital bitstream 134. In some examples, SIMD can include advanced vector extensions using 512 bits (AVX-512) that enable 16 floating-point operations on a single CPU core on a single CPU instruction. For example, AVX-512 can process a huge amount of data on a CPU (e.g., CPU 202). For example, processor 202 (and device 200) can receive a 500 MHz bandwidth data stream. The 500 MHz bandwidth is important in several respects because it is a generally accepted practical limit for a 10 gigabit Ethernet link. Sampling data at 500 MHz with 8-bit samples for the I / Q pair, including parity bits, can saturate a 10 gigabit Ethernet link. The 500 MHz example is not limiting of the present disclosure. Data pipes larger than a 10 gigabit Ethernet link are possible. Additionally, the processing can be separated into n parallel blocks (e.g., block 315) to handle any amount of data.

[0053] Block 315 is shown by the dashed line and represents the processing steps of method 300. Block 315 is shown as a plurality of parallel steps, namely, blocks 315a, 315b~315n. The term "parallel" is used herein to describe that the processing in processing blocks 315a~315n is performed simultaneously or at the same time. The packets to be processed may be of different lengths from one processing block 315 to another, and thus, the processing of the packets may have the same rate or speed from one processing block 315 to the next. As described below, some of the processing blocks 315 may proceed faster or slower than others. Therefore, the term "parallel" shall not be limited to simultaneous or parallel processing within processing block 315.

[0054] The processing block 315 used herein may refer to, for example, a set of processing functions executed by processor 202. The digital bitstream 134 can be sent to a plurality of parallel processing blocks 315a, 315b…315n to distribute the processing load among several cores 204. Each individual processing block 315a, 315b…315n can represent an individual iteration of cloud processing. Therefore, the processing of each of the processing blocks 315a~315n can be associated with (cloud-based) cores 204a~204n. The number of processing blocks 315 required varies based on the amount of data to be processed. In some embodiments, the number of processing blocks 315 can be limited by the number of logical cores available within processor 202 for network 154 or for local hardware processing. In some other embodiments, the memory bandwidth constraint may cause a bottleneck in signal processing. Memory bandwidth can refer to the rate at which data can be read from a semiconductor memory (e.g., memory 206) or stored in a semiconductor memory (e.g., memory 206) by a processor (e.g., processor 202).

[0055] In some embodiments, the number of processing blocks 315 can vary. Generally, the fewer processing blocks 315 there are, the better it is to limit the number of cores required for the overall process. This, in turn, allows the system to be adapted to smaller virtual private cloud (VPC) machines that operate more inexpensively. The VPC can include, for example, an SPS150 having several CPUs. In some embodiments, eight processing blocks 315 can be used for a 10 gigabit Ethernet link. Such embodiments may not include a forward error correction processing block. In some other embodiments, the only practical limitation on the number of processing blocks 315 required is the bit rate and bandwidth of the communication link (e.g., the size of the pipe). Thus, any number (n) of processing blocks 315 is possible. However, in some embodiments, there may be a practical limitation on the number (n) of processing blocks 315 based on the number of threads that can be executed on the CPU or the number of cores 204 within the processor 202. However, when the limit is reached within a single CPU, there are an unlimited number of cloud-based CPUs or cores 204 within the SPS150 (e.g., VPC) for multiple CPUs (e.g., processors 202) to execute processing together. Additionally, the processor 202 can create new processing blocks 315 as needed. The processing cores 204 can be distributed among multiple distributed processors (e.g., processors 202) as needed for throughput and efficiency.

[0056] Processing block 315 is configured such that it does not matter which of the processing blocks 315a, 315b... 315n is executed the slowest (or the fastest). Method 300 can share the processing load across the processing blocks 315 and thus can reduce any processing delays caused by bottleneck problems in the individual processing blocks 315. For example, the individual sub - processes of the processing block 315 (see the description of FIG. 4 below) can be either not executed or performed at an equal rate (e.g., some may be faster than others). Thus, for example, the larger process of method 400 (FIG. 4) can account for variations in performance or processing time. The processing blocks 315 can then be created as many times as necessary to process the incoming data.

[0057] In some embodiments, each processing block 315 can represent a set of signal - processing algorithms executed by the processor 202. As used herein, an algorithm can refer to the smallest set of functions or method steps that perform a desired function. Multiple exemplary algorithms are described herein.

[0058] An exemplary advantage of method 300 is the ability to create more processing blocks 315 as needed. Generally, the processing blocks 315 can be implemented in software and thus can be created or deleted as needed to suit a given data rate or processing load. Each processing block 315 can be reconfigured to accommodate the need for different received waveforms (e.g., downlink signal 132) and associated digital bitstreams 134.

[0059] In block 320, the processed signal data from the plurality of processing blocks 315 can be recombined again to form the original data encoded and modulated on the downlink signal 134. In some embodiments, the processor 202 can perform the function of a data recombiner. In other embodiments, the device 200 can have additional components for performing such functions. Each data packet or processed data block can have a timestamp. The data recombiner (e.g., processor 202) can order the data blocks based on the timestamps and compare the phases between the ordered blocks. The recombiner can further adjust the phases of adjacent blocks that reorder the data stream. In some embodiments, the phase of a subsequent data block can be adjusted to match the phase of the previous data block.

[0060] For all the processing blocks shown in 315, there are at least four options for execution.

[0061] 1) Multiple blocks are executed, and each sub - element within the processing block 315 (e.g., each block 315a - 315n) obtains its own core (e.g., cores 204a - 204n).

[0062] 2) Multiple blocks are executed, and the processing block 315 obtains only one dedicated core for the entire block.

[0063] 3) A single block is executed, and each sub - element within the processing block obtains its own core.

[0064] 4) A single block is executed, and the processing block obtains only one dedicated core for the entire block.

[0065] The more executable cores there are, the higher the achievable rate.

[0066] In block 325, device 200 can output data to a suitable receiver. In some examples, such a receiver can be one or more mission operation centers. This data can be mission-dependent (e.g., satellite purpose) and can include, among other things, weather data, image data, and SATCOM payload data.

[0067] In general-purpose CPUs, there are at least three major factors that can limit high-speed performance. 1) Data fetching, 2) CPU capacity, and 3) memory bandwidth utilization. Data fetching refers to how quickly data can be supplied to the CPU. CPU capacity is driven by the CPU clock speed and the number of cores within the CPU. Memory bandwidth refers to how quickly data can be transferred between the CPU and external DDR RAM (not the CPU cache). Memory bandwidth can be determined by the number of memory lanes and the DDR RAM clock speed. In some cases, the limiting factor for achieving high-speed processing is CPU capacity, while in other cases it is memory bandwidth. Care must be taken to determine which of the above cases is affecting performance, and if memory bandwidth is limited, the embodiments described below are non-limiting examples of ways to reduce memory bandwidth utilization within the proposed patent approach.

[0068] Function calls within a given processing block can be configured to optimize CPU computation or memory bandwidth utilization. For example, referring to the function calls shown in FIG. 10 (exemplarily shown as blocks), for a given example, various function calls (e.g., Raise to N Power block, mixing block, and decimation block) can be grouped to minimize memory bandwidth. These function calls can be called independently such that each function completes on a set of data before another function starts, thereby simplifying each function. In another example, multiple or all function calls can be combined into one block, such that data is not transferred to RAM after each executed function, and the memory bandwidth of the combined functions is much smaller than when called independently. In the case of functions called independently, the first function call (e.g., Raise to N Power) can be executed across the entire dataset before the second function call (e.g., mixing block) occurs. In the case of combination, only a portion of the data will be processed by the first function call before the second function is executed. In this way, the memory bandwidth is reduced. This method can be applied not only to those shown in FIG. 10, but also to any group of functions. For example, this method can be applied to the timing and phase error calculations shown in FIG. 11, or any other grouping for function calls executed in processing blocks as disclosed herein (e.g., various function call blocks shown in FIGS. 7 - 16).

[0069] Another way to improve memory bandwidth utilization may be to fold several function call blocks into one block, similar to the approach described above. For example, in the case of a channelizer, to separate one channel into N channels, three main functions, 1) a finite impulse response (FIR) filter, 2) a circular buffer, and 3) an inverse fast Fourier transform (IFFT) may be required. In the case of a combiner, to combine M channels into one channel, three main functions, 1) an IFFT, 2) a circular buffer, and 3) a finite impulse response (FIR) filter may be required. Usually, each function requires its own block, as shown in FIGS. 15 and 16, to facilitate operation and CPU optimization, but all functions can be combined into one block to reduce memory bandwidth utilization. This trade-off reduces memory bandwidth utilization at the cost of a hit to CPU performance. Exemplary embodiments of a diversity combiner with blind detection and Doppler compensation executed on a general-purpose CPU that uses parallel processing on multiple cores to achieve high-throughput operation in a cloud environment

[0070] As described above, FIG. 5 is a functional block diagram of one embodiment of method 500. In one example, method 500 may be referred to as a diversity combiner method 500 with blind detection and Doppler compensation. Diversity combining may be used to combine multiple antenna feeds together such that the signals are aligned in time and phase and each is weighted based on signal quality to optimize information transfer across multiple channels. Signal quality may be determined using one or more of, for example, but not limited to, signal-to-noise ratio, energy per symbol-to-noise power spectral density (Es / No), power estimate, received signal strength indicator (RSSI), etc. The multiple antenna feeds may be from one or more remote locations such as platform 110 or satellite 111. Although a satellite is used as an example herein, other wireless transmission systems such as a wireless antenna (e.g., antenna 122) or other type of transmitter may be implemented. Thus, the use of a satellite does not limit the present disclosure.

[0071] In the case of a satellite as shown in FIG. 1, although platform 110 and satellite 111 are visible from the same ground station (e.g., ground station 122), diversity combining may also be used during an antenna handover event, for example, when satellite 111 is descending below the horizon (e.g., to the east) and platform 110 is ascending above the horizon (e.g., to the west). In order to properly combine the downlink signals, several calculations must be performed. The disclosed system can digitize the signals and convert them to digital samples, which are then conveyed to signal processing elements. The system can further calculate and compensate for the Doppler effect. The system can also determine the residual phase and frequency delta (e.g., difference) between the downlink signals, as well as the time difference and estimated signal-to-noise ratio for each channel. Following these operations, the signals are then combined together.

[0072] As described above, the figures show a plurality of blocks that can each be implemented as one or more of elements 306a, 306b, ... 306c (FIG. 4) and / or as one or more of processing blocks 315a, 315b... 315n (FIG. 4). In another example, the plurality of blocks shown in FIG. 5 can be grouped as a single "processing block" that functions in a similar manner to processing blocks 315a, 315b… 315n (FIG. 4) and / or elements 306a, 306b... 306c (FIG. 3).

[0073] Method 500 is illustratively shown at a high level in FIG. 5. Method 500 includes a plurality of processing blocks, such as one or more signal analyzer processing blocks 510a - 510n (collectively referred to as signal analyzer processing block 510 or processing block 510), one or more Doppler compensator processing blocks 520a - 520n (collectively referred to as Doppler compensator processing block 520 or processing block 520), and a diversity combiner processing block 530. In the illustrated example, a plurality of signal analyzer processing blocks 510a - 510n and Doppler compensator processing blocks 520a - 520n for performing functions on a plurality of signals are shown, and each block is performed on a corresponding signal. Any number of signals is possible, but the examples herein are described with reference to two signals.

[0074] An example of a signal analyzer processing block 510 is schematically shown in FIG. 6, an example of a Doppler compensator processing block 520 is schematically shown in FIG. 7, and an example of a diversity combiner processing block 530 is schematically shown in FIG. 8. Method 300 and / or method 400 can be used to process downlink signals 160, 170 (e.g., FIG. 1) for each of processing blocks 510 - 530. One or more of processing blocks 510, 520, and / or 530 can be implemented as one or more of elements 306a - c of element 306 of method 300 described in connection with FIG. 3, or as one or more of elements 315a - n of element 315 of method 315 described in connection with FIG. 4.

[0075] It is also possible to execute each of the processing blocks 510, 520, and 530 separately. One example is, in contrast to the case of the aforementioned antenna handover, when there is one transmitting satellite having two independent downlink signals such as right-hand and left-hand polarization outputs, only using the diversity combiner processing block 530. In this case, the timing and Doppler effects can be ignored, and thus, the signal analyzer processing block 510 and the Doppler compensator processing block 520 may become unnecessary.

[0076] An exemplary signal analyzer processing block 510 is shown in FIG. 6, which can be used for blind detection where the symbol rate, modulation type (referred to herein as "mod type"), and / or center frequency can be estimated without any input from the user. The processing block 510 may include, for example, a coarse symbol rate estimator functional block 605, a timing recovery error calculator functional block 610, a Mod type and carrier estimator functional block 620, and may include a plurality of sub-elements such as an Es / No estimator functional block 625. The processing block 510 may also include a timing recovery functional block 615.

[0077] An exemplary coarse symbol rate estimator functional block 610 is shown schematically in FIG. 9. The illustrated example of functional block 610 includes a plurality of sub-elements or sub-functional blocks 905-920. As an exemplary method, the first sub-functional block 905 estimates the symbol rate by using Gardner calculations and / or performing differential conjugate calculations. Examples of Gardner calculations are described in more detail in U.S. Patent No. 10,790,920, the disclosure of which is incorporated herein by reference as if fully set forth. An example of differential conjugate is a vector calculation, y[n]=a[n]*conj(a[n+1]), where n ranges from 0 to the length of the input - 1. In any of the calculations, sub-functional block 905 outputs the estimated symbol rate to sub-functional block 910, the FFT of the output of sub-functional block 905 is obtained, and the maximum peak frequency is detected in sub-functional block 920. The detected maximum peak frequency corresponds to the symbol rate, and the symbol rate can be estimated based on the detected maximum frequency. In various embodiments, it is also possible to measure a rough carrier estimate value from the phase of the differential conjugate calculation y[n] with a phase calculation functional block 915.

[0078] Referring again to FIG. 6, examples of timing recovery error calculation 610 and timing recovery 615 are described in U.S. Patent No. 10,790,920, the disclosure of which is incorporated herein by reference as if fully set forth. For example, a Gardner timing error detector estimated at functional block 605 can be applied to the input data to create timing information, as is known in the art. In another embodiment, the incoming sample stream can be delayed by only 1 sample. Then, the non-delayed data can be multiplied by the conjugate of the delayed data (conjugate multiplication). Both have advantages and disadvantages, so it is an engineering trade-off that can be implemented.

[0079] An exemplary Mod type detection and carrier wave estimation functional block 620 is schematically shown in FIG. 10. Exemplary examples of the functional block 620 include, but are not limited to, a plurality of sub-functional blocks including a Raise to N Power functional block 1005, a Mix by Coarse carrier wave estimation functional block 1010, a Decimate functional block 1015, an FFT trial functional block 1020, and a Peak detection functional block 1025.

[0080] After the timing recovery 615 of FIG. 6, the signal output from block 615 is then symbol synchronized, and the mod type detection becomes more accurate in the symbol space rather than the sample space. In sub-functional block 1005, the signal input to the Mod type detection and carrier wave estimation functional block 620 is raised to an appropriate power based on the number of symbols (N) in the outer ring of the constellation (2 for BPSK, 4 for QPSK / OQPSK, 8 for 8PSK, 12 for 16APSK, etc.), and then mixed in sub-functional block 1010 by the coarse carrier frequency provided by differential conjugate calculation (if provided). The mixed signal is then decimated in sub-functional block 1015, and then in sub-functional block 1020, an FFT is performed on the signal, and in sub-functional block 1025, the peak-to-average ratio of the selected modulation type is determined. This process is then repeated for all of the desired modulation types to be detected. The result with the highest peak-to-average value is the most likely modulation type. As a method of minimizing the memory bandwidth, sub-functional block 1005, sub-functional block 1010, and sub-functional block 1015 can be combined to form one sub-functional block, thereby reducing the memory bandwidth. To further increase the data rate, in sub-functional block 1020, each modulation type trial can be executed on its own thread to further increase the throughput.

[0081] The next processing block in method 500 is Doppler compensation processing block 520, an example of which is shown schematically in FIG. 7. Processing block 520 may include multiple functional blocks, such as, but not limited to, phase pre-calculator functional block 705, and continuous phase adjustment functional block 710. Phase correction is pre-calculated in functional block 705 based on the carrier estimation from signal analyzer processing block 510. Doppler compensator processing block 520 smooths the compensation in various embodiments such that a PLL-based receiver can track the compensated signal. This pre-calculated phase information from functional block 705 is then supplied to continuous phase adjustment functional block 710 which removes most of the measured Doppler.

[0082] Referring again to FIG. 6, Es / No estimator functional block 625 measures Es / No. There are several approaches to measuring Es / No depending on the modulation type. One exemplary example for measuring Es / No is to calculate (C / N)×(B / fs), where C / N is one of the carrier-to-noise ratio or signal-to-noise ratio, B is the channel bandwidth in Hertz, and fs is the symbol rate or symbols per second. However, it will be understood that any approach for measuring Es / No is equally applicable to the embodiments disclosed herein.

[0083] FIG. 8 schematically illustrates an example of a diversity combiner processing block 530. The processing block 530 can be used as a stand-alone application when the timing and Doppler differences between a plurality of signals (e.g., two in this example) are small and the modulation type is known in advance. In various embodiments, the processing block 530 can also be used in methods 300 and / or 400. The processing block includes a plurality of functional blocks, e.g., but not limited to, a coarse timing estimator functional block 805, a timing and phase error calculator functional block 810, one or more timing adjustment functional blocks 815a-815n (collectively referred to as timing adjustment 815), one or more phase adjustment functional blocks 820a-820n (collectively referred to as phase adjustment 820), and a weighted combiner functional block 825. In the illustrated example, a plurality of timing adjustment functional blocks 815a-815n and phase adjustment functional blocks 820a-n for performing functions on a plurality of signals are shown. Any number of signals is possible, but the examples described herein are described with reference to two signals.

[0084] The coarse timing estimator functional block 805 can be used when the time delta between two arriving signals cannot be ignored. This may be required in the case of antenna handover. The estimator functional block 805 looks for correlation spikes between a plurality of signals (e.g., two in this example) to determine the time difference. The estimator functional block 805 can utilize FFT and / or IFFT to perform the correlation quickly, but any correlation technique can be utilized. If many correlations are required, methods 300 and 400 can be applied to increase the throughput.

[0085] An example of the timing and phase error calculation function block 810 is schematically shown in FIG. 11. An exemplary example of the timing and phase error calculation function block 810 includes a plurality of sub-elements or sub-function blocks, such as, but not limited to, a cross-correlator sub-function block 1105, a decimate sub-function block 1110, a timing estimation update sub-function block 1115, a phase delta generation sub-function block 1120, and a phase unwrap sub-function block 1125. The cross-correlator sub-function block 1105 can calculate the early, immediate, and late (EPL) terms between a plurality of input signals (e.g., two in this example) used in a delay lock loop (DLL). However, the SIMD technique can be used to efficiently calculate the EPL terms. When the EPL terms are calculated, the signal is decimated in the sub-function block 1110. The delta between the early term and the late term can be used for the timing update sub-function block 1115 on the cross-correlator sub-function block 1105, as well as for the feed-forward error term for subsequent timing adjustment.

[0086] Next, the phase of the immediate term is calculated in the phase delta generation sub-function block 1120 and supplied to the phase unwrap sub-function block 1125. In one example, the phase unwrap sub-function block 1125 can use a phase lock loop (PLL) in various embodiments. The phase unwrap function block 1125 includes the phase calculation of the decimated signal executed before the phase unwrap function block 1125. The phase unwrap calculation can provide continuous phase information regarding the data samples. The phase unwrap calculation stitches the phases together when the phase wraps either from π (pi) to -π radians or from -π to π radians. This unwrapping of the angle enables a curve fitting function to be executed on the phase signal without any discontinuities. Thereby, the processor 202 can reassemble the demodulated signal based on the timing and phase of the processed signal. It may be possible to replace the phase unwrap calculation with a Kalman filter to obtain phase, frequency, and Doppler rate information, or use a PLL.

[0087] Referring back to FIG. 8, this phase information from functional block 810 is then supplied to one or more timing adjustment block functional blocks 815 and phase adjustment block functional blocks 820 to adjust the phase of the signals and properly align the signals for later combination. Examples of timing adjustment and phase adjustment sub-elements are described in U.S. Patent No. 10,790,920, the disclosure of which is incorporated herein by reference as if fully set forth. For example, each timing adjustment sub-element functional block 815 can apply the timing phase information calculated by the timing and phase error calculation functional block 810. Then, functional block 815 can use a filter, such as a polyphase FIR filter, where, for example, but not limited to, an appropriate bank of filters is selected based on the provided phase information as known in the art. Then, the timing can be adjusted by functional block 815 such that the timing is efficient in both CPU usage and bandwidth usage. The filter used in functional block 815 can use SIMD techniques to further increase throughput. It is also possible to use linear, cubic, parabolic, or other interpolation formats. Each functional block 815 performs the above-described adjustments on the corresponding signal (e.g., two in the illustrated example) of the plurality of signals.

[0088] Each of the phase adjustment functional blocks 820 can apply the carrier phase information calculated by the timing and phase error calculation functional block 810. Functional block 820 can apply the phase information and use SIMD techniques to adjust the phase of the entire data block such that a plurality of signals are properly aligned for combination. Each functional block 820 performs the above-described adjustments on the corresponding signal (e.g., two in the illustrated example) of the plurality of signals.

[0089] When the signals are time and phase aligned, the weighted combiner function block 825 may apply scaling based on the Es / No estimate value and the power estimate value calculated by the signal analyzer processing block 510. For example, a signal having a better signal-to-noise ratio compared to another signal may be assigned a higher weight and scaled accordingly. Similarly, a higher Es / No estimate value and / or power estimate value may be assigned a larger weight and scaled accordingly. Using SIMD techniques, multiple signals (e.g., two signals in this example) can be efficiently scaled and combined. Exemplary embodiments of a digital signal channelizer and combiner on a general-purpose CPU that uses parallel processing on multiple cores to achieve high-throughput operation in a cloud environment

[0090] Figures 12 to 16 schematically illustrate various embodiments of a method for digital signal channelization and combination according to the embodiments disclosed herein. In various embodiments, the channelization and combination shown in Figures 12 to 16 can be implemented to manage the spectral bandwidth of one or more downlink signals. In some implementations, the methods shown in Figures 12 to 14A for digital signal channelization and / or combination may be executed using the method 300 of Figure 3 and / or the method 400 of Figure 4. The channelizer can be configured to execute a DSP algorithm that separates the spectral bandwidth from one channel into multiple channels (1-to-N), as shown in Figure 12. In another example, the combiner can be configured to execute a DSP algorithm in which the spectral bandwidth is combined from multiple channels into one channel (M-to-1), as shown in Figure 13. In some implementations, the process begins with a network device that digitizes the analog signal (e.g., one or more digitizers as described in connection with Figure 1), and then the samples are conveyed to the channelizer (1-to-N approach shown in Figure 12), or the samples from the combiner are supplied to the network device, which converts them back to an analog signal (M-to-1 approach shown in Figure 13).

[0091] In the case of a satellite as shown in FIG. 1, the digital signal channelizer and combiner according to the embodiments of this specification can be used for bandwidth compression. A digitizer (e.g., one of digitizers 124, 134, and / or 144 in FIG. 1) digitizes a downlink signal having a wide bandwidth (e.g., one of downlink signals 160 and / or 170), but not all of the bandwidth is useful for processing. In one example, a downlink signal having a bandwidth of 500 MHz can be digitized, but only a subset of the bandwidth slices (e.g., 2 slices) actually contains useful data. In this case, the channelizer creates channels from the 500 MHz bandwidth spectrum at 512, and then the combiner combines two channels together, where one channel can be, for example, 50 MHz and the other channel can be 100 MHz. These two smaller channels (50 MHz and 100 MHz) are transmitted to the processing server 150 to be appropriately processed in any case. In this example, only the sum of the subset of the smaller channels (150 MHz) needs to be transmitted to the data center instead of the entire 500 MHz. This can save the carrier cost of transmitting data from the antenna to the data center. In this case, the digitized bandwidth is compressed, and the network bandwidth transmitted via the LAN or WAN is also reduced. Although a satellite is used as an example in this specification, other wireless transmission systems such as a wireless antenna (e.g., antenna 122) or other types of transmitters can be implemented. Therefore, the use of a satellite does not limit the present disclosure.

[0092] In another example, the digital signal channelizer and combiner according to the embodiments of this specification can be used for channel splitting where many digitized bandwidths contain many independent carriers. Then, the channelizer associated with the combiner according to the embodiments of this specification creates appropriate channels to be processed by a receiver such as the receiver described in U.S. Patent No. 10,790,920, the disclosure of which is incorporated herein by reference as if fully set forth.

[0093] In another example, the digital signal channelizer and combiner according to the embodiments herein can be used for channel combining. In this case, several downlink signals received from various types of sources, which can be either other antennas or modulators, can be combined and digitally transmitted to an antenna to create a larger composite bandwidth for broadcast to a satellite.

[0094] FIG. 12 shows an exemplary channelizer processing block 1200 for separating one channel into N channels. The illustrated exemplary channelizer processing block 1200 includes a plurality of functional blocks, such as, but not limited to, an N-path filter functional block 1210, an N-point circular buffer functional block 1220, and an N-point IFFT functional block 1230. The channelizer block processing block 1200 can be configured to capture digitized samples having a first bandwidth from a network device (e.g., a digitizer as described in connection with FIG. 1) via a given network protocol such as, but not limited to, TCP / IP or UDP. The sample stream is then processed in an N-path filter functional block 1210 using a polyphase filter bank, then in an N-point circular buffer functional block 1220, and then in an N-point IFFT functional block 1230 to split the signal into N channels, each N-channel having a corresponding bandwidth smaller than the first bandwidth of the captured digitized samples.

[0095] Exemplary sources for the input shown in FIG. 12 include, but are not limited to, a modulator, a digitizer (e.g., one of digitizers 124, 134, and / or 144 of FIG. 1), the output of a diversity combiner (e.g., such as those described above in connection with FIGS. 5-11), the output of a Doppler compensator (e.g., such as those described above in connection with FIGS. 5-11), and a digital sample file player.

[0096] As shown in FIGS. 12, 14A, and 14B, the input supply to 1210 can represent an input signal that is interpolated by two (as indicated by the splitting of the input signal into two arrows), as is known in the art. As a result, since all N outputs are interpolated by two from the start of processing, the channelizer 1200 can avoid aliasing at the output.

[0097] FIG. 13 shows an exemplary combiner processing block 1300 for combining M channels into one channel. The illustrated exemplary combiner processing block 1300 includes a plurality of functional blocks, such as, but not limited to, an M-path filter and adder functional block 1310, an M-point circular buffer functional block 1320, and an M-point IFFT functional block 1330. In the M-to-1 processing block 1300, the process is reversed compared to the channelizer processing block 1200 of FIG. 12. For example, the M channels are supplied to an M-point filter functional block 1310 using an M-point IFFT functional block 1330, then an M-point circular buffer functional block 1320, and then a polyphase filter bank. The filtered channels are then combined using the adder of the functional block 1310.

[0098] Exemplary sources for the input shown in FIG. 13 include, but are not limited to, a modulator, a digitizer (e.g., one of digitizers 124, 134, and / or 144 of FIG. 1), the output of a diversity combiner (e.g., such as those described above in connection with FIGS. 5 - 11), the output of a Doppler compensator (e.g., such as those described above in connection with FIGS. 5 - 11), a channelizer (e.g., such as those described in connection with FIG. 12), and a digital sample file player. As described above in connection with FIG. 12, and as illustrated in FIGS. 14A and 14B, each input supplied to 1200a - 1200n, respectively, can represent an input signal that is interpolated by two (as shown by the splitting of the input signal into two arrows) as is known in the art. Thereby, since all N outputs are interpolated by two from the start of processing, each channelizer 1200a - 1200n can avoid aliasing in the output.

[0099] In some embodiments, one or more channelizers and one or more combiners can be used in combination to achieve any desired bandwidth from one or more channelizers as shown in FIGS. 14A and 14B. For example, as shown in FIG. 14A, the digitized downlink signal can be split into N channels (e.g., 512 channels in this example) using one or more channelizers, e.g., channelizer processing blocks 1200a - 1200n. Each of the N channels can have a bandwidth smaller than the digitized downlink signal. In some embodiments, each channelizer processing block 1200a - 1200n can receive a discrete downlink signal such that each input a - input n is not necessarily part of the same downlink signal. Thus, each input can be independent of the other inputs.

[0100] When the input signal is split into N channels each, M of those channels (e.g., 20 in this example) can be recombined to form a larger channel using a combiner, e.g., combiner processing block 1300 (e.g., only a single combiner 1300 as shown in FIG. 14A).

[0101] In some embodiments, as shown in FIG. 14B, a plurality of combiner processing blocks 1300a - 1300n can be utilized. As shown in FIG. 14B, the digitized downlink signal can be split into N channels (e.g., 512 channels in this example) using a channelizer, e.g., channelizer processing block 1200. When the input signal is split into each of the N channels, M of those channels (e.g., 20 in this example) can be recombined to form a larger channel using a combiner, e.g., a plurality of combiner processing blocks 1300a - 1300n. Each combiner processing block 1300a - n can take in N channels, combine M channels, and form a larger combined channel that is output from each processing block 1300 as output a - output n.

[0102] Channels selected for recombination may include channels over which useful data is transmitted. Which data is useful may depend on the system that processes the downlink. This method can be used to divide a digitized signal into N channels of a smaller bandwidth and then to gather any of the N channels into a wider channel of any size M. Thereby, the bandwidth is compressed, reducing the network resources required to process and transmit the downlink signal. Thus, an exemplary example illustrates that 512 elements may be channelized by a given channelizer 1200 in the above example and 20 elements may be combined by a given combiner 1300, but it will be understood that any number of elements may be channelized by one or more channelizers and any number of elements may be combined using combiners as needed. This enables the channelizer channel bandwidth to be fully programmable and to have any number of channels.

[0103] In FIG. 14A, two input channels are shown, but any number of input channels, e.g., one, two, or more input channels, 50 or more input channels, 100 or more input channels, can be implemented as needed. If one input channel is used, FIG. 14 may utilize a single channelizer processing block. Similarly, FIG. 14B shows two output channels, but any number of combiner processing blocks, e.g., one, two, or more output channels, can be implemented as needed. If one output channel is used, FIG. 14A may utilize a single combiner processing block. Thus, in some embodiments, a plurality of combiners (e.g., combiner processing blocks 1300a - n) may be arranged after one or more channelizer processing blocks to output many channels of the scaled bandwidth of M / N. Additionally, a gain stage (not shown) can be added after the channelizer and combiner to achieve either manual or automatic gain control for each channel.

[0104] Exemplary sources for each of the inputs shown in FIGS. 14A and 14B include, but are not limited to, modulators, digitizers (e.g., one of digitizers 124, 134, and / or 144 of FIG. 1), outputs of diversity combiners (e.g., those described above in connection with FIGS. 5 - 11), outputs of Doppler compensators (e.g., those described above in connection with FIGS. 5 - 11), and digital sample file players. In some examples, the channelizer and combiner can be cascaded together. For example, a first channelizer / combiner can be fed to another channelizer / combiner. In this case, the first channelizer / combiner can be configured to process any number of inputs at a low sample rate and output the combined signal to another channelizer / combiner that outputs at a higher rate. For example, 100 modulators each operating at 10 kSPS can be combined into one channel operating at 1 MSPS. This 1 MSPS channel can then be fed to another channelizer / combiner with a final output sample rate of 512 MSPS. In some embodiments, each channelizer / combiner can be implemented by one or more of blocks 315a, 315b,... 315n of FIG. 4, and / or one or more of blocks 306a, 306b,... 306n of FIG. 3, etc., by separate processing blocks. That is, for example, a first channelizer / combiner can be executed in a first one or more processing blocks, and a second channelizer / combiner can be executed in a second one or more processing blocks.

[0105] In some embodiments, methods 1400a and 1400b may include an optional combiner input control 1410 configured to time-align the inputs received from one or more channelizer processing blocks 1200 (shown in FIG. 14A but not in FIG. 14B). For example, one possible exemplary implementation of the channelizer / combiner shown in FIG. 14 is to replace an existing analog radio frequency (RF) switch matrix, so it may be desirable to maintain and align the current capabilities of the RF switch matrix. One such capability is time alignment. Since the RF switch matrix is a combination or splitting of signals with near-zero delay, time alignment is a trivial matter in the analog domain. However, once the signal is digitized and transmitted to a network such as a cloud environment carried via a LAN or WAN, time alignment may no longer be trivial. Timestamps may be applied to each input channel and maintained through all processing within one or more channelizer processing blocks 1200. Then, in the combiner input control 1410, data may be collected and buffered for a short programmable (e.g., preset) duration in order to allow the inputs to arrive within the duration. If all inputs arrive in time, they may be carefully time-aligned by the combiner input control 1410 based on the corresponding input rate and timestamp of each channel. However, if a channel does not arrive on time, the combiner input control 1410 can replace that channel with a data source of all zeros, and thus, timely channels are not further blocked or delayed. In this way, the channelizer / combiner of FIG. 14A can replicate the near-zero delay combination present in the RF switch matrix. Similarly, the combiner input control 1410 may be included in method 1400b between the channelizer 1200 and the plurality of combiners 1300a - 1300n and may be configured in a manner similar to that described herein.

[0106] In some embodiments, methods 1400a and 1400b may also include a data splitting function that splits the data of a given downlink signal into parallel data streams by a data splitter. Each of the channelizer of FIG. 12, the combiner of FIG. 13, and / or the combined channelizer / combiner of FIGS. 14A and 14B may be preceded by a data splitting function such as block 310 of FIG. 4. For example, each channelizer 1200 and / or each combiner 1300 shown in FIGS. 14A and 14B can be an example of one or more processing blocks 315 of FIG. 4, and block 310 can split data into parallel data streams as described in connection with FIG. 4. Block 310 may be referred to as a manager (or dealer) processing block in some embodiments. By splitting the data into multiple parallel streams, multiple channelizer portions 1200a - n can function in parallel threads to channelize the downlink signal, multiple combiner portions 1300a - n can function in parallel threads to combine the downlink signal, and / or both one or more channelizers 1200 can function in parallel with one or more combiners 1300. For example, in the case of the channelizer processing block 1200 of FIG. 14A, each channelizer processing block 1200 can be implemented as one or more of blocks 315a - n of FIG. 4. In the case of the combiner processing block 1300 of FIG. 14B, each combiner 1300 can be implemented as one or more of blocks 315a - n of FIG. 4. In both cases, block 310 can split the data from the downlink signal into parallel data streams that each overlap with an adjacent data stream. Each parallel data stream can then be sent to one of the processing blocks 315a - n for processing. When each processing block 315a - n completes processing of that portion of the data, the processed portion of the data is sent, for example, to block 320 of FIG. 4 that outputs the processed data. Block 320 then waits for the next processing block (e.g., the second block 315b) to complete its processing, etc., and outputs each portion of the processed data.This process is sometimes referred to as a round-robin processing scheme.

[0107] The higher the throughput desired for a given application, the more processing blocks 315 are required. For example, referring to FIG. 14B, if 512 MSPS is required as the output of a given combiner processing block 1300, it may be possible to create up to "n" processing blocks 315 each implemented as a combiner processing block 1300 (e.g., one of the combiner processing blocks 1300a - n) to transmit outputs a - n. In this case, the first portion of the input data is transmitted from block 310 to, or processed by, the first processing block 315a (e.g., the first combiner processing block 1300a implemented) to be processed. A small portion of the processed data from block 315a can be stored and prepended by block 310 as duplicate data on the next portion of the data stream. This next portion of the data stream including the duplicate data is then transmitted to the second processing block 315b (e.g., implemented as the second combiner processing block 1300a). This pattern is repeated for the number of processing blocks required to achieve the desired high throughput. When each processing block (e.g., blocks 315a - n) finishes processing its portion of the data, the processed portion of the data is transmitted to, for example, block 320 and output as combined processed data. Block 320 then waits for the next processing block (e.g., the second block 315b) to complete its processing and the like, and outputs a block of data. In some embodiments, block 310 can receive the output from the channelizer 1200 and split the output as described above to supply a plurality of combiners 1300a - 1300n. In another embodiment, block 310 may precede the channelizer 1200 to split the downlink signal. The processing blocks 315a - 315n are used in this round-robin processing scheme where block 310 as a manager or dealer cycles through all the processing blocks.

[0108] The above example has been described in relation to one or more combiners 1300, but it will be understood that the above-described round-robin processing method for splitting and processing data portions into each processing block can be utilized in any of the embodiments described herein. For example, the round-robin dealer method can also be used for the channelizer processing blocks 1200a-n of FIG. 14a to achieve high throughput. That is, each channelizer processing block 1200 can be implemented as processing blocks 315a-n, and the data splitter block 310 can split the data into parallel streams and process each parallel portion in a given processing block 315a-n as a channelizer. Each channelizer 1200a-n can process its portion and output each plurality of N channels from each channelizer 1200a-n to block 320 upon completion. Block 320 outputs the processed portion of the data as described above and waits for the next processing block 315 to complete its processing. Similarly, this round-robin processing method may be utilized in relation to the diversity combiner with the blind detection and Doppler compensation methods of FIGS. 5-11. That is, any one of the processing blocks 510-530 may be implemented as one or more processing blocks 315a-n, and the data splitter function at block 310 may split the input downlink signal into parallel data streams and supply them to each processing block 315a-n in the manner described above. Then, for each processing block 315, the resulting processed data can be transmitted to block 320 for output as described above.

[0109] In some implementations, the digital channelizer and / or combiner processing blocks (e.g., such as 1200, 1300, and / or 1400 described above) enabled by SPS150 (e.g., using a general-purpose CPU and, for example, without limitation, SIMD techniques including SSE, SSE2, SSE3, SSE4.1, SSE4.2, AVX, AVX2, and AVX512 instruction sets) can process data distributed across several cores of the CPU to improve throughput. The data processing can be managed through multiple cores of the processor (e.g., block 315) to achieve the required throughput without using dedicated signal processing hardware such as an FPGA or high-performance computing (HPC) hardware such as a graphics processing unit (GPU). An example of a representative block 315 implemented as a channelizer is shown in FIG. 15. The ability to perform this processing on a general-purpose server CPU enables the deployment of functionality within a virtualized processing architecture in a general-purpose cloud processing environment without the need for dedicated hardware such as, without limitation, the x86 architecture, Cortex-A76, NEON, and AWS Graviton, Graviton2.

[0110] FIG. 15 and FIG. 16 schematically show exemplary processing blocks implemented as a channelizer processing block 1500 or a combiner processing block 1600, respectively. The channelizer processing block 1500 may be substantially similar to the channelizer processing block 1200 described in connection with FIG. 12. For example, the channelizer processing block 1500 can separate one channel into N channels using, for example, an N-path filter function block 1510, an N-point circular buffer function block 1520, and an N-point IFFT function block 1530. The channelizer processing block 1500 is an exemplary example of a channelizer implemented as one of processing blocks 315a-315n (e.g., FIG. 4) and / or one of functions 306a-306n (e.g., FIG. 3). Multiple channelizer processing blocks 1500 may each be provided as one of processing blocks 315a-315n (e.g., FIG. 4) and / or one of functions 306a-306n (e.g., FIG. 3). In another embodiment, the functionality for executing a single channelizer processing block 1500 may be distributed among one of multiple processing blocks 315a-315n (e.g., FIG. 4) and / or functions 306a-306n (e.g., FIG. 3). In some embodiments, the round-robin processing scheme described above in connection with FIGS. 12-14B may be implemented using multiple processing blocks 1500. For example, multiple processing blocks 1500 may be implemented as processing blocks 315a-n. Thus, multiple processing blocks 1500 may be operable to operate in parallel using the round-robin processing scheme as described above.

[0111] Similarly, the combiner processing block 1600 can be substantially the same as the combiner processing block 1300 described in connection with FIG. 13. For example, the combiner processing block 1600 can combine M channels into one channel using, for example, an M-path filter and adder function block 1610, an M-point circular buffer function block 1620, and an M-point IFFT function block 1630. The combiner processing block 1600 is an exemplary example of a combiner implemented as one of processing blocks 315a - 315n (e.g., FIG. 4) and / or one of functions 306a - 306n (e.g., FIG. 3). A plurality of combiner processing blocks 1600 can each be provided as one of processing blocks 315a - 315n (e.g., FIG. 4) and / or one of functions 306a - 306n (e.g., FIG. 3). In another embodiment, the functionality for executing a single combiner processing block 1600 can be distributed among one of a plurality of processing blocks 315a - 315n (e.g., FIG. 4) and / or functions 306a - 306n (e.g., FIG. 3). In some embodiments, the round-robin processing scheme described above in connection with FIGS. 12 - 14B can be implemented using a plurality of processing blocks 1600. For example, the plurality of processing blocks 1600 may be implemented as processing blocks 315a - n. Thus, the plurality of processing blocks 1600 may be operable to operate in parallel using the round-robin processing scheme as described above.

[0112] The channelizer and / or combiner described herein are examples of the memory bandwidth optimization described above. The functionality of the channelizer can be separated into separate blocks, as shown in FIGS. 15 and 16, or combined into one functional sub-element or block where small portions of data are processed. In this exemplary example, the small portion may refer to data corresponding to one IFFT. When the functionality is separated for the channelizer, many filter calculations may be performed, then many circular buffers, and then many IFFTs may be performed. Similarly, when the functionality is separated for the combiner, many IFFTs may be performed, then many circular buffers, and then many filter calculations may be performed.

[0113] Exemplary non-limiting advantages of using a general-purpose CPU are the dynamic nature of resource allocation. For example, in the case of a channelizer / combiner as described in connection with FIGS. 14A and 14B, it may be desirable to reconfigure the system in real time such that input channels can enter and exit at any time. As shown in FIG. 14A, the processing blocks of channelizers 1200a-n can create and destroy each of these channelizers 1200a-n to achieve this goal. To accommodate the input bandwidth, any value of N can be used, and any number of channelizers 1200n can be instantiated to process an appropriate amount of input channels. In this way, as an exemplary example, at one instant, 10 channels each of 15 MHz can be combined, and then one second later, 5 channels of 45 MHz can be combined. For all techniques described in this disclosure, such as multi-core, SIMD techniques, and memory bandwidth optimization approaches, high-throughput data processing can be distributed. The examples provided herein are described in connection with a channelizer / combiner as shown in FIG. 14A, but this beneficial result can equally be achieved by using the techniques described herein with respect to, for example, diversity combiners and Doppler compensation as described in connection with FIGS. 5-11 and FIG. 14B. For example, any number of values of processing block 510, processing block 520, and / or processing block 530 of FIG. 5 can be used to achieve the desired throughput. The number of each processing block 510-530 in FIG. 5 may be dynamically created and destroyed in a similar manner as described above to achieve the desired throughput. Other aspects

[0114] The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope of the present disclosure. The various components shown in the figures may be implemented, but are not limited to, for example, software and / or firmware on a processor, or dedicated hardware. Also, the features and attributes of the specific exemplary embodiments disclosed above can be combined in different ways to form additional embodiments, all of which fall within the scope of the present disclosure.

[0115] The foregoing description of the method and process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the operations of the various embodiments must be performed in the order presented. As will be understood by those skilled in the art, the order of operations in the above embodiments can be performed in any order. Words such as "thereafter," "then," "next," etc. are not intended to limit the order of operations, and these words are merely used to guide the reader through the description of the method. Further, any reference to a claim element in the singular, using, for example, the articles "a," "an," or "the," is not to be construed as limiting the element to the singular.

[0116] The various illustrative logical blocks, modules, and algorithmic operations described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in varying ways for each particular application, but such implementation decisions are not to be construed as causing a departure from the scope of the inventive concept.

[0117] The hardware used to implement the various exemplary logics, logic blocks, and modules described in connection with the various embodiments disclosed herein can be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of receiver devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.

[0118] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The operations of the methods or algorithms disclosed herein may be implemented with processor-executable instructions that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage medium may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disks and discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically with a laser. Combinations of the above are also included within the scope of non-transitory computer-readable media and processor-readable media. Further, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable storage medium and / or a computer-readable storage medium that can be incorporated into a computer program product.

[0119] It is understood that the specific order or hierarchy of blocks in the disclosed process / flowchart is an illustration of an exemplary approach. It should be understood that, based on design preferences, the specific order or hierarchy of blocks within the process / flowchart can be reconfigured. Also, some blocks may be combined or omitted. The appended method claims present the elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.

[0120] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.

[0121] Accordingly, the claims are not intended to be limited to the aspects shown herein, but rather are to be accorded the full scope consistent with the language of the claims, and the reference to singular elements is not meant to mean "sole" unless so stated, but rather "one or more".

[0122] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects. Unless otherwise specified, the term "some" refers to one or more.

Claims

1. A method for combining a plurality of downlink signals representing communication signals, the method comprising: Receiving the plurality of downlink signals from a plurality of antenna feeds; In a first one or more processing blocks within one or more processors, Performing a first blind detection operation on a first data packet of a first one of the plurality of downlink signals; Performing a first Doppler compensation operation on a first data packet of a first one of the plurality of downlink signals; In a second one or more processing blocks within the one or more processors, parallel to the first one or more processing blocks, Performing a second blind detection operation on a second data packet of a second one of the plurality of downlink signals; Performing a second Doppler compensation operation on a second data packet of a second one of the plurality of downlink signals; Combining the first signal and the second signal based on (i) aligning the timing and phase of the first data packet with the second data packet and (ii) performing a weighted combiner operation that applies scaling to each of the first data packet and the second data packet based on corresponding signal quality; Including The first blind detection operation includes Receiving the first data packet as a sample of the first signal, the sample having an unknown symbol rate, modulation type, and frequency; Estimating a symbol rate based on a detected maximum peak frequency; Determining a timing error of the sample based on the estimated symbol rate; Synchronizing the samples of the first signal by performing a timing recovery operation based on the determined timing error; Detecting a modulation type and estimating a carrier frequency based on the synchronized samples; Estimating energy per symbol-to-noise power spectral density; A method including a first plurality of functions including

2. The method according to claim 1, wherein the first one or more processing blocks include first one or more central processing unit (CPU) cores, and the second one or more processing blocks include second one or more CPU cores.

3. The one or more processors include a plurality of processors, and in the first one or more processing blocks, are included in a first processor of the plurality, and the second one or more processing blocks are included in a second processor of the plurality. The method according to claim 1 or 2.

4. The one or more processors include a single processor including the first one or more processing blocks and the second one or more processing blocks. The method according to claim 1 or 2.

5. The first one or more processing blocks include a first processing block and a second processing block, the first blind detection operation is executed in the first processing block, and the first Doppler compensation is executed in the second processing block. The method according to any one of claims 1 to 4.

6. The second one or more processing blocks include a third processing block and a fourth processing block, the second blind detection operation is executed in the third processing block, and the second Doppler compensation is executed in the fourth processing block. The method according to claim 5.

7. The third processing block executes the second blind detection operation in parallel with the first processing block executing the first blind detection operation. The method according to claim 6.

8. The fourth processing block executes the second Doppler compensation operation in parallel with the second processing block executing the first Doppler compensation operation. The method according to claim 6.

9. At least one of the first blind detection operation and the second blind detection operation includes estimating one or more of the symbol rate, modulation type, and center frequency of the first signal or the second signal, respectively, without input from a user. The method according to any one of claims 1 to 8.

10. The first one or more processing blocks include a first plurality of processing blocks, and at least two of the first plurality of functions are executed in parallel by separate processing blocks of the first plurality of processing blocks. The method according to any one of claims 1 to 9.

11. The second blind detection operation is Receiving the second data packet as a sample of the second signal, wherein the sample has an unknown symbol rate, modulation type, and frequency; estimating a symbol rate based on the detected maximum peak frequency; determining a timing error of the sample based on the estimated symbol rate; synchronizing the sample of the second signal by performing a timing recovery operation based on the determined timing error; detecting a modulation type and estimating a carrier frequency based on the synchronized sample; estimating energy per symbol-to-noise power spectral density; The method according to any one of claims 1 to 9, comprising a second plurality of functions including the above.

12. The method according to claim 11, wherein the second one or more processing blocks include a second plurality of processing blocks, and at least two of the second plurality of functions are executed in parallel by separate processing blocks among the second plurality of processing blocks.

13. The method according to claim 11, wherein one or more of the second plurality of functions are executed in the second one or more processing blocks in parallel with one or more of the first plurality of functions executed in the first one or more processing blocks.

14. The first Doppler compensation operation includes: receiving the first data packet as a sample of the first signal; receiving an estimated value of the carrier frequency of the sample; calculating phase correction information of the sample based on the estimated carrier frequency; removing at least a part of the Doppler from the sample based on the calculated phase correction information; The method according to any one of claims 1 to 13, comprising a third plurality of functions including the above.

15. The method according to claim 14, wherein the first one or more processing blocks include a first plurality of processing blocks, and at least two of the third plurality of functions are executed in parallel by separate processing blocks among the first plurality of processing blocks.

16. The second Doppler compensation operation includes: receiving the second data packet as a sample of the second signal; receiving an estimated value of the carrier frequency of the sample; calculating phase correction information of the sample based on the estimated carrier frequency; Based on the calculated phase correction information, removing at least a part of the Doppler from the sample, The method according to claim 14, comprising a fourth plurality of functions including

17. The method according to claim 16, wherein the second one or more processing blocks include a second plurality of processing blocks, and at least two of the fourth plurality of functions are executed in parallel by separate processing blocks among the second plurality of processing blocks.

18. The method according to claim 16, wherein one or more of the fourth plurality of functions are executed in one or more of the second processing blocks in parallel with one or more of the third plurality of functions executed in the one or more of the first processing blocks.

19. A method for combining a plurality of downlink signals representing communication signals, the method comprising Receiving the plurality of downlink signals from a plurality of antenna feeds, In a first one or more processing blocks within one or more processors, Performing a first blind detection operation on a first data packet of a first signal among the plurality of downlink signals, Performing a first Doppler compensation operation on a first data packet of a first signal among the plurality of downlink signals, In a second one or more processing blocks within the one or more processors in parallel with the first one or more processing blocks, Performing a second blind detection operation on a second data packet of a second signal among the plurality of downlink signals, Performing a second Doppler compensation operation on a second data packet of a second signal among the plurality of downlink signals, Combining the first signal and the second signal based on (i) aligning the timing and phase of the first data packet with the second data packet and (ii) performing a weighted combiner operation of applying scaling to each of the first data packet and the second data packet based on corresponding signal qualities, Including Combining the first plurality of data packets and the second plurality of data packets is Calculating initial, immediate, and late (EPL) terms between the first data packet and the second data packet Determining the timing and phase error between the first data packet and the second data packet based on the EPL term; Adjusting the timing and phase of the first data packet with respect to the second data packet; A method comprising a fifth plurality of functions including the above.

20. Adjusting the timing and phase of samples of a first portion of the plurality of data packets in the third one or more processing blocks; Adjusting the timing and phase of samples of a second portion of the plurality of data packets in parallel with the third one or more processing blocks in the third one or more processing blocks; The method according to claim 19, further comprising the above.

21. Dividing a digital bit stream into a plurality of data packets in the one or more processors, wherein each of the data packets of the plurality of data packets includes an overlap of data from adjacent packets The method according to any one of claims 1 to 20, further comprising the above.

22. The method according to claim 21, wherein the adjacent packets overlap in time by 1% of the length of the packet.

23. The method according to claim 21 or 22, wherein the plurality of data packets have various lengths.

24. The method according to any one of claims 1 to 23, wherein the one or more processors are one or more general-purpose central processing units (CPUs).

25. The method according to any one of claims 1 to 23, wherein the one or more processors use single instruction multiple data (SIMD) techniques to achieve high throughput.

26. A system for combining a plurality of downlink signals representing communication signals, the system comprising: A plurality of antennas configured to receive the plurality of downlink signals; One or more processors communicatively coupled to the plurality of antennas, the one or more processors having a plurality of processing blocks and being operable to perform the method according to any one of claims 1 to 25; a processor; A system including the above.

Citation Information

Patent Citations

  • Automatic kernel migration for heterogeneous cores

    JP2014513853A

  • Per-shader preamble for graphics processing

    JP2019517078A

  • Repeater diversity spread spectrum communication system

    US5233626A