Interpolation processing apparatus and method for wave number domain imaging
Patent Information
- Application Number
- CN202611329347.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请提供了一种波数域成像插值处理装置及方法,以解决相关技术中波数域成像的插值处理效率较低的问题
[0015]本申请实施例的波数域成像插值处理装置,插值处理单元内集成有依次连接的坐标计算模块、邻域地址生成模块和加权累加模块,由此,波数域成像插值计算中所涉及的坐标计算、邻近数据点的寻址以及加权累加分别由相应的模块承担,各模块分工明确;并且,由于上述三个模块依次连接,坐标计算模块所计算出的坐标可以直接提供给邻域地址生成模块用于生成各邻近数据点的地址,并进一步用于加权累加模块确定各邻近数据点对应的插值核系数,前一模块的处理结果能够直接作为后一模块的处理依据,使得一个输出点从坐标计算、邻域寻址到加权累加的完整插值计算过程能够在插值处理单元内部连贯地完成,减少了插值计算过程中数据在处理单元之外的中转,从而提高了波数域成像中插值处理的效率。
Smart Images

Figure CN122836736A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, specifically to a wavenumber domain imaging interpolation processing apparatus and method. Background Technology
[0002] Wavenumber domain imaging algorithms are a type of frequency domain imaging algorithm with high imaging accuracy in applications such as synthetic aperture radar (SAR) imaging. During the imaging process, they require interpolation and resampling of the input frequency domain data. Specifically, for each output point, the coordinates mapped to that output point in the input frequency domain data are calculated, multiple data points near those coordinates are determined, and the data from these multiple data points are weighted and summed to obtain the value of the output point. In related technologies, this interpolation and resampling is typically implemented in software using a general-purpose processor. The coordinate calculations, determination of neighboring data points, and weighted summation operations involved in the interpolation process are all executed sequentially point-by-point as instructions. This results in numerous computational steps, a large data processing volume, and low interpolation efficiency, making it difficult to meet the processing requirements of real-time imaging scenarios.
[0003] Therefore, there is an urgent need for a wavenumber domain imaging interpolation processing device to solve the problem of low interpolation processing efficiency in wavenumber domain imaging in related technologies. Summary of the Invention
[0004] This application provides a wavenumber domain imaging interpolation processing apparatus and method to solve the problem of low interpolation processing efficiency in wavenumber domain imaging in related technologies.
[0005] In a first aspect, this application provides a wavenumber domain imaging interpolation processing device, which includes an interpolation processing unit, and the interpolation processing unit integrates a coordinate calculation module, a neighborhood address generation module and a weighted accumulation module connected in sequence. The coordinate calculation module is used to calculate the coordinates of the current output point mapped to the input frequency domain data; The neighborhood address generation module is used to determine multiple neighboring data points of the current output point based on the coordinates, and generate the addresses of multiple neighboring data points; The weighted accumulation module is used to obtain data from multiple neighboring data points based on the address, determine the interpolation kernel coefficients of each neighboring data point based on the coordinates, and weight and accumulate the data of multiple neighboring data points according to the interpolation kernel coefficients to obtain the interpolation result of the current output point.
[0006] In one alternative implementation, the coordinate calculation module includes a pipelined coordinate rotation digital computer and a multiplier array; The multiplier array is used to perform exponentiation on the frequency coordinates of the current output point, and the pipelined coordinate rotation digital computer is used to perform square root operation on the result of the exponentiation to obtain the coordinates of the current output point mapped to the input frequency domain data.
[0007] In one optional implementation, the neighborhood address generation module stores multiple sets of preset offsets, each set of preset offsets corresponding to a different number of neighboring data points. The neighborhood address generation module selects a set of preset offsets from multiple preset offsets based on the configuration signal, splits the coordinates into integer and fractional parts, and adds each offset in the selected set of preset offsets to the integer part to obtain the addresses of multiple neighboring data points.
[0008] In one optional implementation, the weighted accumulation module includes a complex multiply-accumulate tree, which includes multiple complex multiply-accumulate units. The number of complex multiply-accumulate units participating in the operation is set according to the number of multiple neighboring data points. The coordinate calculation module, the neighborhood address generation module, and the weighted accumulation module process different output points simultaneously to output the interpolation results point by point.
[0009] In one optional implementation, the device stores an interpolation kernel coefficient table. The weighted accumulation module uses the decimal part of the coordinates as an index to obtain the interpolation kernel coefficients of each neighboring data point from the interpolation kernel coefficient table. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets in a selected set of preset offsets. The interpolation kernel coefficients corresponding to negative offsets are obtained by the weighted accumulation module by mirroring the stored interpolation kernel coefficients according to the sign of the offset.
[0010] In one alternative implementation, the device further includes shared memory and a data prefetch engine; The shared memory comprises multiple independent memory banks, and the input frequency domain data is interleaved and stored in multiple independent memory banks according to frequency coordinates; The device includes multiple interpolation processing units, each of which acquires data of neighboring data points distributed in different independent storage units in parallel via arbitration logic; Each output point is interpolated in a predetermined order. The data prefetching engine performs differential prediction based on the coordinates of multiple output points with obtained coordinates to obtain the address of the neighboring data point of the subsequent output point. The subsequent output point is the output point that follows the multiple output points with obtained coordinates in a predetermined order. Before performing interpolation calculations on subsequent output points, the data prefetch engine loads the data corresponding to the address obtained from the differential prediction from off-chip memory into the on-chip cache, and each interpolation processing unit retrieves the loaded data from the on-chip cache.
[0011] In one optional implementation, the operating voltage of the multiple complex multiply-accumulate units can be adjusted independently. The multiple complex multiply-accumulate units include a first complex multiply-accumulate unit and a second complex multiply-accumulate unit. Data of neighboring data points whose interpolation kernel coefficient is greater than a preset threshold is processed by the first complex multiply-accumulate unit, and data of neighboring data points whose interpolation kernel coefficient is not greater than the preset threshold is processed by the second complex multiply-accumulate unit. The operating voltage of the first complex multiply-accumulate unit is higher than that of the second complex multiply-accumulate unit. Furthermore, the operating frequency and operating voltage of the interpolation processing unit are adjusted according to the number of neighboring data points, and the larger the number, the higher the operating frequency and operating voltage.
[0012] In one alternative implementation, the interpolation processing unit is a processing unit in a reconfigurable array; The reconfigurable array is used to sequentially perform multiple processing stages of wavenumber domain imaging. The device also includes a configuration controller and an on-chip configuration cache. During the execution of the current processing stage by the reconfigurable array, the configuration controller loads the configuration information of the next processing stage from off-chip memory into the on-chip configuration cache. After the current processing stage is completed, the reconfigurable array switches to the next processing stage according to the configuration information in the on-chip configuration cache.
[0013] Secondly, this application provides a wavenumber domain imaging interpolation processing method, which includes: Calculate the coordinates of the current output point mapped to the input frequency domain data; Determine multiple neighboring data points of the current output point based on the coordinates, and obtain the data of multiple neighboring data points; The interpolation kernel coefficients of each neighboring data point are determined based on the coordinates, and the data of multiple neighboring data points are weighted and accumulated according to the interpolation kernel coefficients to obtain the interpolation result of the current output point.
[0014] In one optional implementation, while performing weighted accumulation to obtain the interpolation result of the current output point, multiple neighboring data points of the next output point are determined, data of multiple neighboring data points of the next output point are obtained, and the coordinates of the next output point mapped to the input frequency domain data are calculated. Determine the interpolation kernel coefficients for each neighboring data point based on the coordinates, including: Using the decimal part of the coordinates as an index, the interpolation kernel coefficients of each neighboring data point are obtained from the interpolation kernel coefficient table. The offset is the offset of the neighboring data point relative to the coordinate. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets. The interpolation kernel coefficients corresponding to negative offsets are obtained by mirroring the stored interpolation kernel coefficients according to the sign of the offset.
[0015] The wavenumber domain imaging interpolation processing apparatus of this application embodiment integrates a coordinate calculation module, a neighborhood address generation module, and a weighted accumulation module connected in sequence within the interpolation processing unit. Thus, the coordinate calculation, neighbor addressing, and weighted accumulation involved in wavenumber domain imaging interpolation calculation are each handled by the corresponding modules, with clear division of labor among the modules. Furthermore, since the three modules are connected in sequence, the coordinates calculated by the coordinate calculation module can be directly provided to the neighborhood address generation module to generate the addresses of each neighboring data point, and further used by the weighted accumulation module to determine the interpolation kernel coefficients corresponding to each neighboring data point. The processing result of the previous module can be directly used as the processing basis for the next module, enabling the complete interpolation calculation process of an output point from coordinate calculation, neighborhood addressing to weighted accumulation to be completed continuously within the interpolation processing unit. This reduces the data transfer outside the processing unit during the interpolation calculation process, thereby improving the efficiency of interpolation processing in wavenumber domain imaging. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a structural block diagram of a wavenumber domain imaging interpolation processing apparatus according to an embodiment of this application; Figure 2 This is a schematic flowchart of a wavenumber domain imaging interpolation processing method according to an embodiment of this application; Figure 3 This is a flowchart illustrating a reconfigurable chip co-design method for accelerating Stolt interpolation in wavenumber domain imaging according to an embodiment of this application. Figure 4 This is a schematic diagram of the core processing stage and interpolation operation decomposition of the wavenumber domain algorithm according to an embodiment of this application; Figure 5 This is a schematic diagram of a dedicated processing unit and hierarchical mapping architecture for interpolation acceleration according to an embodiment of this application; Figure 6 This is a timing diagram illustrating configuration prefetching and pipeline switching according to an embodiment of this application; Figure 7 This is a schematic diagram of a multi-granularity parallel processing model according to an embodiment of this application; Figure 8 This is a schematic diagram of the storage system hierarchical optimization design according to an embodiment of this application; Figure 9This is a schematic diagram of the complete processing flow of the wavenumber domain algorithm according to the embodiments of this application on a dynamically reconfigurable chip. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0020] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained.
[0021] SAR (Synthetic Aperture Radar): A type of active microwave imaging radar that uses a moving platform (such as a satellite or aircraft) to form an equivalent large-aperture antenna to perform high-resolution imaging of ground scenes, providing all-day, all-weather observation capabilities.
[0022] Echo data: The signal data reflected by the scene and received by the radar receiver after the radar emits electromagnetic waves into the observed scene, is the raw input for imaging processing.
[0023] Range and azimuth: two dimensions in SAR imaging. Range refers to the radar line of sight, and azimuth refers to the platform's flight direction.
[0024] Fourier Transform and Inverse Fourier Transform: Mathematical transformations that convert signals between the time domain (or spatial domain) and the frequency domain. The fast algorithm for this transformation is called the Fast Fourier Transform (FFT), and the corresponding inverse transformation is called the Inverse Fast Fourier Transform (IFFT).
[0025] Two-dimensional frequency domain: The transform domain after performing Fourier transform on two-dimensional data along both dimensions. Each data point in the two-dimensional frequency domain corresponds to a set of range frequency and azimuth frequency.
[0026] Wavenumber domain algorithm: A type of frequency domain imaging algorithm that completes SAR imaging focusing in the two-dimensional frequency domain, also known as the ω-K algorithm. It achieves unified processing of range and azimuth directions through phase compensation and coordinate transformation, and theoretically has high imaging accuracy.
[0027] Range migration: In SAR imaging, due to the change in slant range between the platform and the target over time, the target echo energy shifts in the range direction, which needs to be corrected during the imaging process.
[0028] Stolt interpolation (interpolation resampling): A coordinate transformation operation in wavenumber domain algorithms. It decouples the range and azimuth directions by resampling non-uniformly distributed data in the two-dimensional frequency domain onto a uniform grid. It is usually implemented by interpolation.
[0029] Interpolation kernel: The weighting function used in interpolation calculation to weight neighboring data points. The length of the interpolation kernel (i.e. the number of neighboring data points involved in the weighting, such as 4 points, 8 points, or 16 points) affects the interpolation accuracy and computational cost.
[0030] Kaiser window: A commonly used window function whose shape can be adjusted by the parameter β, often used to construct interpolation kernels.
[0031] Peak Sidelobe Ratio (PSLR) and Integrated Sidelobe Ratio (ISLR): Commonly used indicators for measuring SAR imaging quality, reflecting the level of sidelobes relative to the main lobe in the point target response, usually expressed in decibels (dB). The lower the value, the better the imaging quality.
[0032] DSP (Digital Signal Processor): A programmable processor optimized for digital signal processing operations.
[0033] FPGA (Field Programmable Gate Array): An integrated circuit whose hardware logic can be configured by the user, and whose hardware structure remains fixed during operation after configuration.
[0034] GPU (Graphics Processing Unit): A processor with large-scale thread-level parallel capabilities, often used for general-purpose parallel computing.
[0035] Dynamically reconfigurable chips and reconfigurable arrays: a type of integrated circuit whose core computing resource is a reconfigurable array. The reconfigurable array consists of multiple processing units connected by an interconnection network. The functions and interconnection relationships of the processing units can be changed during operation by loading configuration information, thus balancing the energy efficiency of dedicated hardware with the flexibility of software.
[0036] Processing Element (PE): The basic unit for performing operations in a reconfigurable array, which generally includes an arithmetic logic unit, a local register file, and a routing unit.
[0037] Configuration information: This information is used to set the operating mode and interconnection relationship of each processing unit in the reconfigurable array. Loading different configuration information allows the array to perform different computing functions.
[0038] CORDIC (Coordinate Rotation Digital Computer): A hardware algorithm and corresponding circuit structure that performs operations such as square root, exponentiation, and trigonometric functions through iterative shift and addition operations. The circuit implementation is relatively simple.
[0039] Fixed-point representation: This method represents floating-point numbers as binary numbers with a fixed decimal point position and a finite bit width. It can reduce hardware computation and storage overhead. The choice of bit width requires a trade-off between precision and resources.
[0040] A bank is a partition in memory that can be read and written independently. Multiple banks can respond to different access requests at the same time.
[0041] Interleaved storage: A storage method that stores data in multiple storage banks in turn according to certain rules, so that data with adjacent addresses are distributed in different storage banks.
[0042] Data prefetching: A technique that moves data from slower-access memory to faster-access memory in advance before the processor actually uses the data, in order to hide memory access latency.
[0043] Dynamic Voltage and Frequency Scaling (DVFS): A technique that dynamically adjusts the circuit's operating voltage and frequency according to the computational load to reduce power consumption. The circuit's dynamic power consumption is approximately proportional to the square of the operating voltage.
[0044] As an optional application scenario of this application embodiment, the wavenumber domain imaging interpolation processing device provided in this application embodiment can be applied to the processing chain of wavenumber domain imaging. During wavenumber domain imaging, echo data is transformed to the frequency domain via Fourier transform to obtain frequency domain data. After completing the corresponding frequency domain processing, this frequency domain data needs to undergo interpolation resampling before the focused image can be recovered through inverse Fourier transform. The wavenumber domain imaging interpolation processing device provided in this application embodiment is used to complete the interpolation resampling: the device uses frequency domain data as input frequency domain data, performs interpolation processing on the input frequency domain data, and outputs the interpolation result to subsequent processing stages.
[0045] For example, one deployment method of the device may include: a radar receiver receiving echo data, an imaging processing device performing wavenumber domain imaging processing on the echo data, a wavenumber domain imaging interpolation processing device being disposed in the imaging processing device to perform interpolation calculations in the imaging processing, and the image output by the imaging processing device being provided to back-end devices such as image storage devices and image interpretation systems.
[0046] In practical applications, imaging processing devices can take many forms. For example, an imaging processing device can be a real-time signal processing device mounted on a platform such as a satellite, aircraft, or drone, which performs real-time imaging of received echo data while the platform is in orbit or during flight; it can also be an imaging processing server or workstation in a ground station, which performs imaging processing on the transmitted or recorded echo data. Correspondingly, wavenumber domain imaging interpolation processing devices can be integrated into the imaging processing device in various forms. For example, it can be set as a separate chip on the circuit board of the imaging processing device, it can be integrated as a functional circuit inside the chip with other processing circuits on the same chip, or it can be connected to the imaging processing device as a board or embedded module. The specific implementation form of this device is not limited in the embodiments of this application.
[0047] It should be noted that the above-mentioned optional application scenarios are merely examples of one application scenario and do not limit the scope of protection of this application. For example, the echo data may not come from the real-time reception of the radar receiver, but from a storage device that has pre-recorded the echo data; the input frequency domain data of the interpolation processing device is not limited to being provided by other circuits within the same device, and the embodiments of this application do not limit these aspects.
[0048] In related technologies, interpolation resampling in wavenumber domain imaging is usually performed sequentially point by point in software, including coordinate calculation, neighbor data point determination, and weighted summation. To address the technical problem of multiple interpolation processing steps and low processing efficiency, this application provides a wavenumber domain imaging interpolation processing device that can continuously complete the interpolation calculation of each output point within the interpolation processing unit, thereby improving the interpolation processing efficiency.
[0049] This embodiment provides a wavenumber domain imaging interpolation processing device, which can be installed in the aforementioned imaging processing equipment, such as a spaceborne real-time signal processing device, an imaging processing server in a ground station, or can be implemented as an independent chip, a functional circuit inside the chip, or a board.
[0050] Figure 1 This is a structural block diagram of a wavenumber domain imaging interpolation processing apparatus according to an embodiment of this application, as shown below. Figure 1As shown, the device includes an interpolation processing unit 110, which integrates a coordinate calculation module 111, a neighborhood address generation module 112, and a weighted accumulation module 113 connected in sequence. Each module is described below.
[0051] The coordinate calculation module 111 is used to calculate the coordinates of the current output point mapped to the input frequency domain data.
[0052] The input frequency domain data refers to the frequency domain data to be interpolated during wavenumber domain imaging. The input frequency domain data includes multiple data points, which are arranged according to a certain grid. Each data point has its position in the grid and its corresponding data.
[0053] An output point refers to the data point obtained through interpolation. The purpose of interpolation is to calculate the data of each output point using data from several data points in the input frequency domain data, thereby resampling the input frequency domain data into output data arranged according to the output grid. The current output point is the output point that the interpolation processing unit 110 is currently processing.
[0054] Since the output grid where the output point is located does not coincide with the grid where the input frequency domain data is located, the position of each output point in the input frequency domain data needs to be calculated through a mapping relationship. This calculated position is the coordinate of the current output point mapped to the input frequency domain data. It is understandable that this coordinate usually does not fall exactly on a grid point of the input frequency domain data. For example, the coordinate can be a non-integer value such as 123.3, meaning it falls between the 123rd and 124th data points of the input frequency domain data.
[0055] For example, the coordinate calculation module 111 can be pre-configured with a mapping relationship between the output point and the input frequency domain data. The coordinate calculation module 111 calculates the coordinates of the current output point in the output grid according to the mapping relationship, and provides the calculated coordinates to the neighbor address generation module 112 connected to it.
[0056] The neighborhood address generation module 112 is used to determine multiple neighboring data points of the current output point based on the coordinates, and generate the addresses of the multiple neighboring data points.
[0057] In this context, "neighboring data points" refers to data points in the input frequency domain data located near the aforementioned coordinates. Since these coordinates typically do not fall on the grid points of the input frequency domain data, the data for the current output point cannot be directly read from the input frequency domain data. Instead, it needs to be calculated from the data of multiple data points near these coordinates. These data points involved in the calculation are the neighboring data points of the current output point. The number of neighboring data points can be set according to the needs of the interpolation calculation; for example, it can be 4, 8, etc. A larger number generally results in more accurate interpolation results, but also a greater computational load.
[0058] The address of a neighboring data point refers to the storage location of the data of a neighboring data point in memory. Based on this address, the data of the corresponding neighboring data point can be read from memory. For example, the neighborhood address generation module 112 can split the coordinates output by the coordinate calculation module 111 into an integer part and a fractional part. Using the integer part as a reference, it determines multiple grid points on both sides of the coordinate as multiple neighboring data points of the current output point. Then, it generates the address of each neighboring data point based on its position in the input frequency domain data. For example, when the coordinate is 123.3 and the number of neighboring data points is 4, the four data points at positions 122, 123, 124, and 125 in the input frequency domain data can be determined as neighboring data points of the current output point, and addresses corresponding to these four data points can be generated. The neighborhood address generation module 112 provides the generated addresses to the weighted accumulation module 113 connected to it.
[0059] The weighted accumulation module 113 is used to obtain the data of the multiple neighboring data points according to the address, determine the interpolation kernel coefficient of each neighboring data point according to the coordinate, and weight and accumulate the data of the multiple neighboring data points according to the interpolation kernel coefficient to obtain the interpolation result of the current output point.
[0060] The interpolation kernel coefficient refers to the weighting coefficient used when weighting the data of each neighboring data point. It characterizes the contribution of each neighboring data point to the interpolation result of the current output point. Generally speaking, the closer the neighboring data point is to the above coordinate, the greater its influence on the interpolation result, and the larger the corresponding interpolation kernel coefficient; conversely, the farther away, the smaller the corresponding interpolation kernel coefficient. Since the distance between each neighboring data point and this coordinate is determined by this coordinate, the interpolation kernel coefficient of each neighboring data point can be determined based on this coordinate. For example, the offset of each neighboring data point relative to this coordinate can be determined based on the decimal part of the coordinate, and thus the interpolation kernel coefficient corresponding to each neighboring data point can be determined.
[0061] Weighted accumulation refers to multiplying the data of each neighboring data point by its corresponding interpolation kernel coefficient, and then summing the products. For example, the weighted accumulation module 113 reads the data of each neighboring data point from the memory based on the address provided by the neighborhood address generation module 112; determines the interpolation kernel coefficient corresponding to each neighboring data point based on the coordinates calculated by the coordinate calculation module 111; then multiplies the data of each neighboring data point by its corresponding interpolation kernel coefficient, and sums the products. The summation result is the interpolation result of the current output point. Using the previous example, the weighted accumulation module 113 reads four data points based on the addresses of four neighboring data points, determines the interpolation kernel coefficient corresponding to each of the four neighboring data points based on the coordinates 123.3, multiplies the four data points by their respective interpolation kernel coefficients, and then sums the products to obtain the interpolation result of the current output point.
[0062] It is understandable that the interpolation processing unit 110 can process each output point in the output grid as the current output point in the manner described above, thereby obtaining the interpolation result of each output point and completing the interpolation processing of the input frequency domain data in wavenumber domain imaging.
[0063] In summary, the wavenumber domain imaging interpolation processing device provided in this application integrates a coordinate calculation module, a neighborhood address generation module, and a weighted accumulation module connected in sequence within the interpolation processing unit. The coordinate calculation, determination and addressing of neighboring data points, and weighted accumulation in wavenumber domain imaging interpolation are each handled by a corresponding module, with clear division of labor among the modules. Furthermore, the coordinates calculated by the coordinate calculation module are directly provided to the neighborhood address generation module to generate addresses of neighboring data points, and further used by the weighted accumulation module to determine the interpolation kernel coefficients. The processing result of the previous module directly serves as the processing basis for the next module, enabling the complete interpolation calculation process for each output point, from coordinate calculation and neighborhood addressing to weighted accumulation, to be completed continuously within the interpolation processing unit. This reduces the data transfer outside the interpolation processing unit during the interpolation calculation process, thereby improving the efficiency of interpolation processing in wavenumber domain imaging.
[0064] This embodiment provides a wavenumber domain imaging interpolation processing device. This device can be incorporated into the aforementioned imaging processing equipment, such as a spaceborne real-time signal processing device or an imaging processing server in a ground station. Alternatively, it can be implemented as a separate chip, an internal functional circuit of the chip, or a circuit board. The structure of this device is as follows: Figure 1 As shown, the system includes an interpolation processing unit 110, which integrates a coordinate calculation module 111, a neighborhood address generation module 112, and a weighted accumulation module 113 connected in sequence. Each module will be described below.
[0065] The coordinate calculation module 111 is used to calculate the coordinates of the current output point mapped to the input frequency domain data.
[0066] The input frequency domain data refers to the frequency domain data to be interpolated during wavenumber domain imaging. This data includes multiple data points arranged according to a specific grid. The output points are the data points to be obtained through interpolation. The purpose of interpolation is to calculate the data for each output point using data from several data points in the input frequency domain data, thereby resampling the input frequency domain data into output data arranged according to the output grid.
[0067] The current output point is the output point that the interpolation processing unit 110 is currently processing. Since the output grid where the output point is located does not coincide with the grid where the input frequency domain data is located, the position of each output point in the input frequency domain data needs to be calculated according to a mapping relationship. The calculated position is the coordinate of the current output point mapped to the input frequency domain data. It is understandable that this coordinate usually does not fall exactly on a grid point of the input frequency domain data; for example, the coordinate can be a non-integer value such as 123.3.
[0068] In wavenumber domain imaging, the aforementioned mapping relationship is typically nonlinear, and coordinate calculation involves operations such as exponentiation and square root extraction. If these operations are performed line by line using a general-purpose computing circuit, coordinate calculation can easily become the bottleneck of the entire interpolation process. Therefore, in some embodiments, the coordinate calculation module 111 includes a pipelined coordinate rotating digital computer and a multiplier array; the multiplier array performs exponentiation on the frequency coordinates of the current output point, and the pipelined coordinate rotating digital computer performs square root extraction on the result of the exponentiation to obtain the coordinates of the current output point mapped to the input frequency domain data.
[0069] Here, frequency coordinates refer to the position of the output point in the frequency domain grid, which may include, for example, the range frequency and azimuth frequency of the output point. A multiplier array is an operational circuit composed of multiple multipliers. Multiple multipliers can perform multiplication operations in parallel. The exponentiation of the frequency coordinates of the output point (e.g., squaring) can be achieved by the multipliers in the multiplier array.
[0070] A coordinate rotating digital computer is an operational circuit that performs operations such as square root extraction and trigonometric function calculations through iterative shifting and addition operations. Its circuit structure is relatively simple. A pipelined coordinate rotating digital computer refers to expanding the iterative stages of a coordinate rotating digital computer into a pipeline structure, so that each stage of iteration is handled by a corresponding circuit stage. When the previous data is executing the subsequent iteration, the next data can enter the previous iteration.
[0071] For example, when calculating the coordinates of the current output point, the coordinate calculation module 111 first performs a power operation on the frequency coordinates of the current output point using a multiplier array, and then performs a square root operation on the result of the power operation using a pipelined coordinate rotation digital computer to obtain the coordinates of the current output point mapped to the input frequency domain data. In this way, the power operation and square root operation in the coordinate calculation are handled by dedicated circuits adapted to them, and the square root operation is executed in a pipelined manner. The coordinate calculation of multiple output points can be sequentially advanced in the pipeline, which helps to improve the throughput of coordinate calculation and avoids coordinate calculation becoming a bottleneck in the interpolation process.
[0072] The neighborhood address generation module 112 is used to determine multiple neighboring data points of the current output point based on the coordinates, and generate the addresses of the multiple neighboring data points.
[0073] In this context, "nearby data points" refers to data points in the input frequency domain data located near the aforementioned coordinates. Since these coordinates typically do not fall on the grid points of the input frequency domain data, the data for the current output point needs to be calculated from the data of multiple data points near these coordinates. These data points involved in the calculation are the multiple neighboring data points of the current output point. The address refers to the storage location of the neighboring data points in memory; the corresponding data can be retrieved from memory based on the address.
[0074] In practical applications, different imaging tasks may have different requirements for interpolation accuracy, and the number of neighboring data points required will also vary. For example, when high interpolation accuracy is required, 8 neighboring data points can be selected, while when high processing speed is required, 4 neighboring data points can be selected. To enable the same device to adapt to different numbers of neighboring data points, in some embodiments, the neighborhood address generation module 112 stores multiple sets of preset offsets, each set corresponding to a different number of neighboring data points. The neighborhood address generation module 112 selects a set of preset offsets from the multiple sets of preset offsets according to the configuration signal, splits the coordinates into an integer part and a fractional part, and adds each offset in the selected set of preset offsets to the integer part to obtain the addresses of multiple neighboring data points.
[0075] The offset is used to characterize the positional offset of a neighboring data point relative to the integer part of the coordinate. For example, the neighborhood address generation module 112 can store three sets of preset offsets corresponding to 4, 6, and 8 neighboring data points, respectively. The preset offset corresponding to 8 neighboring data points can be [-3,-2,-1,0,1,2,3,4].
[0076] The configuration signal is used to indicate the number of neighboring data points used in this interpolation process. It can be provided by an external control circuit or by a configuration register within the device. For example, when the coordinate is 123.3 and the configuration signal indicates that 8 neighboring data points are used, the neighborhood address generation module 112 splits the coordinate 123.3 into an integer part 123 and a fractional part 0.3, selects the offset group [-3,-2,-1,0,1,2,3,4], and adds each offset to the integer part 123 to obtain eight addresses: 120, 121, 122, 123, 124, 125, 126, and 127, which are the addresses of the eight neighboring data points. Therefore, the address of a neighboring data point can be obtained by looking up the preset offset and performing an addition operation. The circuit involved in the operation only involves simple operations such as splitting and addition, and the circuit overhead for address generation is small. Furthermore, by configuring the signal, it is possible to switch between multiple sets of preset offsets, so that the same neighborhood address generation module can adapt to different numbers of neighboring data points, thereby improving the device's adaptability to different interpolation accuracy requirements.
[0077] The weighted accumulation module 113 is used to obtain data from multiple neighboring data points based on the address, determine the interpolation kernel coefficients of each neighboring data point based on the coordinates, and weight and accumulate the data of multiple neighboring data points according to the interpolation kernel coefficients to obtain the interpolation result of the current output point.
[0078] The interpolation kernel coefficient refers to the weighting coefficient used when weighting the data of each neighboring data point. It characterizes the contribution of each neighboring data point to the interpolation result of the current output point. Generally, the closer the neighboring data point is to the aforementioned coordinates, the larger the corresponding interpolation kernel coefficient. Since the relative position of each neighboring data point to these coordinates is determined by them, the interpolation kernel coefficients of each neighboring data point can be determined based on these coordinates. Weighted summation refers to multiplying the data of each neighboring data point by its corresponding interpolation kernel coefficient, and then summing the products. The summation result is the interpolation result of the current output point.
[0079] Considering that the input frequency domain data in wavenumber domain imaging is usually complex, and weighted accumulation involves multiple complex multiplications and additions, in order to improve the processing efficiency of weighted accumulation, in some embodiments, the weighted accumulation module 113 includes a complex multiply-add tree, which includes multiple complex multiply-add units. The number of complex multiply-add units participating in the operation is set according to the number of multiple neighboring data points.
[0080] In this context, a complex multiply-add unit refers to an arithmetic circuit capable of performing complex multiplication and addition. A complex multiply-add tree refers to an arithmetic circuit formed by connecting multiple complex multiply-add units in a tree-like structure. Multiple complex multiply-add units first multiply the data of each neighboring data point with its corresponding interpolation kernel coefficient in parallel. The products are then added level by level along the tree structure, thus completing the entire weighted accumulation within a relatively short computation cycle. The number of complex multiply-add units participating in the operation is set according to the number of neighboring data points. For example, using 8 neighboring data points activates the computational paths corresponding to 8 complex multiply-add units, while using 4 neighboring data points activates only the computational paths corresponding to 4 of the complex multiply-add units. In this way, multiple multiplication operations in the weighted accumulation can be executed in parallel, and the computational scale matches the number of neighboring data points, avoiding idle computational resources.
[0081] Furthermore, in the above embodiments, the coordinate calculation module 111, the neighborhood address generation module 112, and the weighted accumulation module 113 process different output points simultaneously to output interpolation results point by point. That is, the three modules form a pipeline: when the weighted accumulation module 113 performs weighted accumulation on the first output point, the neighborhood address generation module 112 can simultaneously generate addresses of neighboring data points for the second output point, and the coordinate calculation module 111 can simultaneously calculate the coordinates of the third output point. Thus, the three modules process different output points simultaneously. After the pipeline operates stably, the interpolation processing unit 110 can continuously output interpolation results point by point, improving the throughput of interpolation processing compared to the method of processing the same output point serially by the three modules.
[0082] If the interpolation kernel coefficients are calculated in real time for each interpolation, it will introduce additional computational overhead. Therefore, in some embodiments, the device stores an interpolation kernel coefficient table. The weighted accumulation module 113 uses the decimal part of the coordinates as an index to obtain the interpolation kernel coefficients of each neighboring data point from the interpolation kernel coefficient table. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets in a selected set of preset offsets. The interpolation kernel coefficients corresponding to negative offsets are obtained by the weighted accumulation module 113 mirroring the stored interpolation kernel coefficients according to the sign of the offset.
[0083] The interpolation kernel coefficient table is a pre-established data table that records the correspondence between the decimal part of the coordinates and the interpolation kernel coefficients corresponding to each offset. Since the position of each neighboring data point relative to the coordinates is determined by both the decimal part of the coordinates and each offset, the interpolation kernel coefficients of each neighboring data point can be obtained by looking up the table using the decimal part as an index, without the need for real-time calculation. Furthermore, the interpolation kernel coefficients are usually symmetrically distributed about the coordinates, that is, two neighboring data points with symmetrical positive and negative offsets have the same interpolation kernel coefficient. For example, two neighboring data points with an offset of +1 and an offset of -1 have the same interpolation kernel coefficient at symmetrical index positions. Utilizing this symmetry, the interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets. When it is necessary to obtain the interpolation kernel coefficients corresponding to negative offsets, the weighted accumulation module 113 mirrors the stored interpolation kernel coefficients according to the sign of the offset, that is, it reads the interpolation kernel coefficient corresponding to the symmetrical positive offset as the lookup result. Therefore, the interpolation kernel coefficients are obtained by looking up a table, which reduces the real-time computation overhead of the coefficients; and the coefficient table only needs to store about half of the interpolation kernel coefficients, saving the storage space occupied by storing the interpolation kernel coefficients.
[0084] The above describes the internal structure of a single interpolation processing unit 110. When processing large-scale input frequency domain data, multiple interpolation processing units can be set up to process different output points in parallel. In this case, multiple interpolation processing units will simultaneously initiate the reading of data from neighboring data points, which places higher demands on the data supply capacity. To this end, in some embodiments, the device also includes a shared memory and a data prefetching engine; the shared memory includes multiple independent memory banks, and the input frequency domain data is interleaved and stored in multiple independent memory banks according to frequency coordinates; the device includes multiple interpolation processing units 110, and each interpolation processing unit 110 acquires data from neighboring data points distributed in different independent memory banks in parallel through arbitration logic.
[0085] Independent storage units refer to storage cells that can be read and written independently, and each independent storage unit can respond to different access requests at the same time. Interleaved storage refers to storing the input frequency domain data in each independent storage unit in turn according to the frequency coordinates, so that data points with adjacent frequency coordinates are distributed in different independent storage units.
[0086] For example, the shared memory can be divided into 16 independent memory banks, with 16 data points adjacent in frequency coordinates stored in 16 different independent memory banks. Since the multiple neighboring data points accessed by the interpolation calculation are adjacent to each other in frequency coordinates, interleaved storage distributes these neighboring data points across different independent memory banks. When multiple interpolation processing units 110 simultaneously initiate data reads, each read request is distributed to its corresponding independent memory bank via arbitration logic: read requests with target addresses in different independent memory banks can be responded to in parallel; read requests with target addresses in the same independent memory bank are responded to sequentially by the arbitration logic. Thus, the reading of neighboring data points by multiple interpolation processing units is carried out in parallel across multiple independent memory banks, alleviating access conflicts when multiple interpolation processing units read data simultaneously and improving the parallelism of data supply.
[0087] Furthermore, the amount of input frequency domain data is usually large and is often stored in off-chip memory. However, off-chip memory has a long access latency. If data from neighboring data points is only read from off-chip memory during interpolation calculation, the interpolation processing unit will pause due to waiting for data. To address this, in the above embodiment, interpolation calculations are performed on each output point in a predetermined order. The data prefetch engine performs differential prediction based on the coordinates of multiple output points with obtained coordinates to obtain the addresses of neighboring data points of the subsequent output points. The subsequent output points are those that are arranged in a predetermined order after the multiple output points with obtained coordinates. Before performing interpolation calculations on the subsequent output points, the data prefetch engine loads the data corresponding to the addresses obtained from differential prediction from off-chip memory into on-chip cache. Each interpolation processing unit 110 retrieves the loaded data from the on-chip cache.
[0088] The predetermined order refers to the order in which the output points undergo interpolation calculations; for example, they can be processed point by point according to the row and column order of the output grid. Differential prediction refers to inferring the trend of coordinate changes in subsequent output points based on the differences between the coordinates of multiple output points whose coordinates have already been obtained, thereby predicting the coordinates of subsequent output points and the addresses of their neighboring data points. In wavenumber domain imaging, the changes between the coordinates mapped by adjacent output points in a predetermined order are relatively gradual; therefore, the coordinates of subsequent output points can be predicted relatively accurately based on the differences between the coordinates of the preceding output points.
[0089] For example, the data prefetching engine can record the coordinates of several recent output points, and based on the difference between adjacent coordinates and the change in the difference, calculate the coordinates of the next output point, thereby obtaining the addresses of its neighboring data points, and pre-loading the data corresponding to these addresses from off-chip memory into the on-chip cache. When the interpolation processing unit 110 actually processes the next output point, the data of its neighboring data points is already in the on-chip cache, and the interpolation processing unit 110 can obtain the data from the on-chip cache without waiting for the off-chip memory access process. Thus, the off-chip memory access latency is hidden in the interpolation calculation process of the previous output point, reducing the pauses caused by the interpolation processing unit waiting for data and improving the continuity of the interpolation processing.
[0090] Regarding power consumption, considering that the device may be used in scenarios where power consumption is sensitive, such as spaceborne platforms, further optimization of the device's power consumption is possible. It is noted that in the weighted accumulation, the contribution of each neighboring data point to the interpolation result is not the same: neighboring data points with larger interpolation kernel coefficients have a greater impact on the interpolation result, and their computational accuracy also has a greater impact on image quality; neighboring data points with smaller interpolation kernel coefficients have a smaller impact on the interpolation result. Based on this, in some embodiments, the operating voltage of multiple complex multiply-accumulate units can be independently adjusted. These multiple complex multiply-accumulate units include a first complex multiply-accumulate unit and a second complex multiply-accumulate unit. Data from neighboring data points whose interpolation kernel coefficients are greater than a preset threshold are processed by the first complex multiply-accumulate unit, while data from neighboring data points whose interpolation kernel coefficients are not greater than the preset threshold are processed by the second complex multiply-accumulate unit. The operating voltage of the first complex multiply-accumulate unit is higher than that of the second complex multiply-accumulate unit.
[0091] The ability to independently adjust the operating voltage means that each complex multiply-accumulate unit can have its own power supply voltage set. Generally, the dynamic power consumption of the arithmetic circuit decreases as the operating voltage decreases, while the circuit's operational accuracy may be affected at lower operating voltages. A preset threshold is used to differentiate the size of the interpolation kernel coefficients; its value can be set according to imaging quality requirements, for example, it can be set based on the ratio of the interpolation kernel coefficient to the largest coefficient.
[0092] For example, for neighboring data points whose interpolation kernel coefficients are greater than a preset threshold, their data is processed by a first complex multiply-accumulate unit with a higher operating voltage to maintain the accuracy of this part of the operation; for neighboring data points whose interpolation kernel coefficients are not greater than the preset threshold, their data is processed by a second complex multiply-accumulate unit with a lower operating voltage to reduce the power consumption of this part of the operation. Since neighboring data points with smaller interpolation kernel coefficients contribute less to the interpolation result, reducing the operating voltage of this part of the operation has a smaller impact on the interpolation result, thereby reducing the power consumption of the weighted accumulation module while maintaining interpolation accuracy.
[0093] Furthermore, in the above embodiments, the operating frequency and operating voltage of the interpolation processing unit 110 are adjusted according to the number of neighboring data points; the larger the number, the higher the operating frequency and operating voltage. It is understood that the larger the number of neighboring data points, the greater the computational load required for interpolation calculation at each output point. In this case, operating at a higher frequency helps maintain the throughput of the interpolation processing, and correspondingly, a higher operating voltage is needed to support the higher operating frequency. Conversely, when the number of neighboring data points is small, the computational load is small, and operating at a lower frequency and operating voltage can meet the processing requirements while reducing power consumption. Therefore, the operating state of the interpolation processing unit matches the actual computational load, avoiding redundancy in voltage and frequency when the computational load is small, further reducing the power consumption of the device.
[0094] The above embodiments illustrate the interpolation processing unit and its associated storage and power management structure. In the overall process of wavenumber domain imaging, interpolation processing is only one of the processing stages. To enable the device to handle the overall processing of wavenumber domain imaging, in some embodiments, the interpolation processing unit 110 is a processing unit in a reconfigurable array. The reconfigurable array is used to sequentially execute multiple processing stages of wavenumber domain imaging. The device also includes a configuration controller and an on-chip configuration cache. During the execution of the current processing stage by the reconfigurable array, the configuration controller loads the configuration information of the next processing stage from the off-chip memory into the on-chip configuration cache. After the current processing stage is completed, the reconfigurable array switches to the next processing stage according to the configuration information in the on-chip configuration cache.
[0095] A reconfigurable array refers to a computational array composed of multiple processing units whose functions and interconnections can be changed through configuration information. By loading different configuration information, the same reconfigurable array can perform different computational functions at different times. The multiple processing stages of wavenumber domain imaging may include, for example, spectrum generation, phase compensation, interpolation resampling, and image restoration. Each processing stage has different computational characteristics, and the corresponding configuration information is also different.
[0096] Configuration information is used to set the operating mode of each processing unit in the reconfigurable array and the interconnection relationship between the processing units. For example, when the reconfigurable array is performing the spectrum generation phase, the configuration controller begins to load the configuration information of the phase compensation phase from off-chip memory into the on-chip configuration cache; when the spectrum generation phase is completed, the configuration information of the phase compensation phase is ready in the on-chip configuration cache, and the reconfigurable array switches to the phase compensation phase according to the configuration information in the on-chip configuration cache, and so on.
[0097] Therefore, the loading process of configuration information for the next processing stage overlaps with the execution process of the current processing stage in time. The reconfigurable array does not need to wait for the configuration information to be loaded from the off-chip memory when switching stages, which reduces the waiting time introduced by the switching of processing stages and improves the continuity of the overall wavenumber domain imaging processing. At the same time, the same reconfigurable array can sequentially undertake the operation of multiple processing stages by switching configuration information, which improves the utilization rate of computing resources in the entire imaging process.
[0098] In summary, the wavenumber domain imaging interpolation processing device provided in this application integrates a coordinate calculation module, a neighborhood address generation module, and a weighted accumulation module connected in sequence within the interpolation processing unit. Coordinate calculation, neighborhood addressing, and weighted accumulation in the interpolation calculation are handled by the corresponding modules, and the processing result of the previous module directly serves as the processing basis for the next module, enabling the complete interpolation calculation for each output point to be completed coherently within the interpolation processing unit. Furthermore, through a pipelined coordinate calculation circuit, configurable neighborhood addressing, a complex multiply-accumulate tree matching the number of neighboring data points, and pipelined processing between the three modules, the throughput of the interpolation processing is improved. The interpolation kernel coefficient table and its mirror read reduce coefficient calculation and storage overhead. Interleaved storage and data prefetching of multiple independent memory banks alleviate the data supply pressure when multiple interpolation processing units process in parallel. Voltage and frequency adjustments adapted to the interpolation kernel coefficients and the number of neighboring data points reduce the device's power consumption. The reconfigurable array and pre-loading of configuration information enable the device to continuously execute multiple processing stages of wavenumber domain imaging with high resource utilization.
[0099] This embodiment also provides a wavenumber domain imaging interpolation processing method, which can be executed by the wavenumber domain imaging interpolation processing device in the foregoing embodiment, or by other chips, processors or computing devices with corresponding processing capabilities. Figure 2 This is a flowchart of a wavenumber domain imaging interpolation processing method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Calculate the coordinates of the current output point mapped to the input frequency domain data.
[0100] The input frequency domain data refers to the frequency domain data to be interpolated during wavenumber domain imaging, which includes multiple data points arranged in a grid. The output point refers to the data point to be obtained through interpolation; the current output point is the one currently being processed. Since the output grid where the output point is located does not coincide with the grid where the input frequency domain data is located, it is necessary to convert the current output point to its corresponding position in the input frequency domain data according to a mapping relationship. The converted position is the coordinate of the current output point mapped to the input frequency domain data. This coordinate usually does not fall exactly on a grid point of the input frequency domain data; for example, it can be a non-integer value such as 123.3. For example, the coordinates of the current output point mapped to the input frequency domain data can be obtained by performing calculations according to the mapping relationship corresponding to wavenumber domain imaging, based on the frequency coordinates of the current output point.
[0101] Step S202: Determine multiple neighboring data points of the current output point based on the coordinates, and obtain the data of multiple neighboring data points.
[0102] In this context, "nearby data points" refers to data points in the input frequency domain data located near the aforementioned coordinates. Since the coordinates typically do not lie on grid points, the data for the current output point needs to be calculated from the data of multiple data points near the coordinates. The number of nearby data points can be set according to the required interpolation accuracy, for example, it can be 4 or 8. For example, the coordinates can be split into an integer part and a fractional part. Using the integer part as a reference, multiple grid points on both sides of the coordinate are selected as multiple nearby data points for the current output point. Then, based on the position of each nearby data point in the input frequency domain data, the data of each nearby data point is read from the memory storing the input frequency domain data. For example, when the coordinate is 123.3 and the number of nearby data points is 4, the four data points at positions 122, 123, 124, and 125 in the input frequency domain data can be determined as nearby data points, and the data of these four data points can be read.
[0103] Step S203: Determine the interpolation kernel coefficients of each neighboring data point based on the coordinates, and weight and accumulate the data of multiple neighboring data points according to the interpolation kernel coefficients to obtain the interpolation result of the current output point.
[0104] The interpolation kernel coefficient refers to the weighting coefficient used when weighting the data of each neighboring data point, representing the contribution of each neighboring data point to the interpolation result. Generally, the closer the neighboring data point is to the coordinate, the larger the corresponding interpolation kernel coefficient. Since the relative position between each neighboring data point and the coordinate is determined by the coordinate, the interpolation kernel coefficient of each neighboring data point can be determined based on the coordinate. Weighted accumulation refers to multiplying the data of each neighboring data point by its corresponding interpolation kernel coefficient, and then summing the products. Using the previous example, multiplying the data of the four neighboring data points by their respective interpolation kernel coefficients and summing the four products yields the interpolation result for the current output point. By sequentially performing the above steps on each output point in the output grid, the interpolation processing of the input frequency domain data in wavenumber domain imaging can be completed.
[0105] In interpolation processing, each output point undergoes three sequential steps: coordinate calculation, determination and acquisition of neighboring data points, and weighted accumulation. If all steps are executed serially for each output point before processing the next, the processing resources for each step are mostly in a waiting state. To improve the utilization rate of each step, in some embodiments, while performing weighted accumulation to obtain the interpolation result of the current output point, multiple neighboring data points of the next output point are determined, their data are acquired, and the coordinates of the next output point mapped to the input frequency domain data are calculated.
[0106] In other words, the three stages proceed simultaneously for different output points in a pipeline manner: when the first output point is in the weighted accumulation stage, the second output point is simultaneously in the neighbor data point determination and data acquisition stage, and the third output point is simultaneously in the coordinate calculation stage; after the interpolation result of the first output point is output, the second output point enters the weighted accumulation stage, the third output point enters the neighbor data point determination and data acquisition stage, the fourth output point enters the coordinate calculation stage, and so on. Thus, at any given time, different output points are in different processing stages, and the processing resources corresponding to each stage are continuously operational, allowing the interpolation results to be output continuously point by point. Compared to the method of executing all stages serially point by point, this improves the throughput of the interpolation processing.
[0107] Furthermore, calculating the interpolation kernel coefficients of each neighboring data point in real time during each interpolation would introduce additional computational overhead. Therefore, in some embodiments, the interpolation kernel coefficients of each neighboring data point are determined based on the coordinates. This includes: using the decimal part of the coordinates as an index, retrieving the interpolation kernel coefficients of each neighboring data point from the interpolation kernel coefficient table. Here, the offset is the offset of the neighboring data point relative to the coordinates. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets; the interpolation kernel coefficients corresponding to negative offsets are obtained by mirroring the stored interpolation kernel coefficients according to the sign of the offset.
[0108] The interpolation kernel coefficient table is a pre-established data table recording the correspondence between the decimal part of the coordinates and the interpolation kernel coefficients corresponding to each offset. Since the position of each neighboring data point relative to the coordinates is determined by both the decimal part of the coordinates and each offset, the interpolation kernel coefficients of each neighboring data point can be obtained by looking up the table using the decimal part as an index, eliminating the need for real-time calculation. Furthermore, the interpolation kernel coefficients are typically symmetrically distributed about the coordinates; that is, two neighboring data points with symmetrical positive and negative offsets have the same interpolation kernel coefficient. Utilizing this symmetry, the interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets. When it is necessary to obtain the interpolation kernel coefficients corresponding to negative offsets, the offset is mirrored according to its sign; that is, the interpolation kernel coefficient corresponding to the positive offset symmetrical to the negative offset is read as the lookup result. For example, to obtain the interpolation kernel coefficient of a neighboring data point with an offset of -1, the interpolation kernel coefficient at the storage location corresponding to an offset of +1 is read. Therefore, the interpolation kernel coefficients are obtained by looking up a table, which reduces the real-time computation overhead of the coefficients; and the interpolation kernel coefficient table only needs to store about half of the interpolation kernel coefficients, saving the storage space occupied by the interpolation kernel coefficients.
[0109] In summary, the wavenumber domain imaging interpolation processing method provided in this application completes the interpolation calculation for each output point through three stages: coordinate calculation, determination and acquisition of neighboring data points, and weighted accumulation. By enabling the three stages to be carried out simultaneously for different output points in a pipeline manner, the throughput of interpolation processing is improved. By using the decimal part of the coordinates as an index to look up the interpolation kernel coefficients and utilizing the symmetry of the coefficients to store only the interpolation kernel coefficients corresponding to positive offsets, the computational and storage overhead of the interpolation kernel coefficients is reduced.
[0110] To better illustrate the wavenumber domain imaging interpolation processing apparatus and method of this application, a preferred embodiment will be provided below. This embodiment is intended to describe the implementation process of this application in detail, but is not intended to limit the scope of protection of this application.
[0111] In this embodiment, a dynamically reconfigurable chip for accelerating Stolt interpolation in spaceborne SAR wavenumber domain imaging is taken as a preferred example to comprehensively describe the wavenumber domain imaging interpolation processing device and wavenumber domain imaging interpolation processing method provided in this application. This example is chosen because in real-time imaging scenarios of spaceborne SAR, the interpolation resampling operation of wavenumber domain algorithms is computationally intensive and involves irregular memory access. Furthermore, spaceborne platforms have stringent requirements for real-time processing and power consumption, thus fully demonstrating the technical details of various aspects of this application's embodiments. It should be understood that this application's embodiments are not limited to this example. For instance, this application's embodiments can also be applied to airborne or ground-based SAR imaging processing, and can also be applied to other signal processing scenarios requiring interpolation resampling in the frequency domain; the wavenumber domain imaging interpolation processing method can be executed by the dynamically reconfigurable chip in this embodiment, or by other processing devices with corresponding processing capabilities.
[0112] Wavenumber domain algorithms are among the most theoretically accurate SAR frequency domain imaging algorithms. They do not rely on approximate assumptions about the scene or platform trajectory, correcting all range migrations with a single interpolation operation, and can adapt to complex imaging scenarios such as large oblique angles, wide beamwidths, and ultra-high resolutions. For an N×N pixel image, the computational complexity of the two-dimensional Fourier transform in the wavenumber domain algorithm is approximately O(N²log₂N), and the computational complexity of interpolation resampling is approximately O(N²M). Here, N is the number of pixels in a single dimension of the image, log₂N is the logarithm of N to the base 2, M is the interpolation kernel length, i.e., the number of neighboring data points involved in the weighted summation for each output point; O(·) indicates the order of magnitude of computational complexity. It is evident that the computational complexity of interpolation resampling increases with the product of the square of the image size and the interpolation kernel length, often becoming a critical factor affecting system performance under high-precision imaging requirements.
[0113] In related technologies, the hardware implementation of wavenumber domain algorithms mainly includes three categories: software implementation schemes based on multi-core DSPs, which are flexible in development but limited in parallelism, and whose throughput and energy efficiency are difficult to meet the requirements of high-resolution real-time imaging; custom acceleration schemes based on FPGAs, which have good real-time performance, but whose hardware logic is fixed once configured, conflicting with the differences in computational characteristics at each stage of the algorithm, resulting in uneven resource utilization, and requiring redesign and verification when imaging parameters change; and parallel schemes based on GPUs, which have strong computing power, but whose power consumption is usually hundreds of watts, and whose storage hierarchy is difficult to adapt to the irregular neighborhood access of interpolation operations. Based on this, this embodiment maps wavenumber domain algorithms to a dynamically reconfigurable chip and conducts co-design around interpolation acceleration, providing a reconfigurable chip co-design method for Stolt interpolation acceleration in wavenumber domain imaging.
[0114] Figure 3 This is a schematic diagram of the overall process of this embodiment. (For example...) Figure 3As shown, the overall process of this embodiment includes four core steps: wavenumber domain core processing mechanism analysis, hardware unit design for interpolation acceleration, multi-stage streaming processing and storage optimization, and precision-configurable energy efficiency management. These steps are described in detail below.
[0115] Figure 4 This is a schematic diagram illustrating the core processing stage and interpolation operation decomposition of the wavenumber domain algorithm. For example... Figure 4 As shown, the wavenumber domain algorithm comprises four cascaded processing stages: the two-dimensional spectrum generation stage transforms the echo data to the two-dimensional frequency domain through range and azimuth Fourier transforms, exhibiting regular streaming processing characteristics; the consistent phase compensation stage constructs a phase compensation function in the two-dimensional frequency domain with the scene center slant range as a reference and multiplies it with the signal spectrum to achieve focusing at the reference distance, with a point-to-point data flow; the frequency domain interpolation and resampling stage nonlinearly stretches the original range-frequency axis through coordinate transformation to decouple the range and azimuth, requiring non-uniform to uniform resampling in the two-dimensional frequency domain, making it the most computationally intensive and memory-access-irregular stage; and the image spatial restoration stage performs a two-dimensional inverse Fourier transform on the interpolated uniform grid data to obtain the focused SAR image.
[0116] like Figure 4 As shown, interpolation resampling can be further decomposed into three sub-steps. The first is mapping coordinate calculation: calculating the original frequency coordinates corresponding to each output point based on the nonlinear mapping relationship, involving operations such as exponentiation and root extraction. The second is neighborhood addressing: determining the nearest integer coordinates in the original uniform grid based on the calculated non-integer coordinates, typically requiring 4, 8, or 16 points of neighborhood support; the memory access location varies with the output point, exhibiting strong randomness. The third is weighted accumulation: determining the weighting coefficients for each neighboring point according to the interpolation kernel and performing multiplication-accumulation. The computational characteristics of these three sub-steps differ significantly, making them suitable for collaborative completion by specially designed hardware.
[0117] Analysis of the algorithm's data flow and dependencies reveals a strict sequential dependency among the four processing stages: two-dimensional spectrum generation must be completed before consistent phase compensation can be performed, and so on. However, each stage exhibits high data parallelism; for example, phase compensation at different frequency points is independent, and interpolation operations at different output points are also independent. Quantitative analysis of computational intensity shows that interpolation resampling constitutes the main computational component of the algorithm, accounting for over 60% of the total computation. The computational load is directly proportional to the interpolation kernel length; for example, using 4-point interpolation requires 4 complex multiplications and additions per output point, while using 8-point interpolation doubles the computational load. Furthermore, each output point needs to access multiple geographically dispersed neighboring data points, resulting in high concurrency and high randomness in memory access.
[0118] In terms of accuracy, this embodiment determines differentiated fixed-point bit widths for each computational stage: the input for the Fourier transform butterfly operation uses a 16-bit fixed-point representation, with 1 bit for the sign bit, 2 bits for the integer bits, and 13 bits for the decimal bits; intermediate accumulation uses a 32-bit representation to avoid overflow during accumulation; the interpolation kernel coefficients use a 12-bit representation, with 1 bit for the sign bit, 1 bit for the integer bits, and 10 bits for the decimal bits. Furthermore, the interpolation kernel coefficients are evenly symmetrical about the center. For example, in an 8-point interpolation kernel, the coefficients corresponding to the offsets of +0.5 and -0.5 are the same. Therefore, only the four coefficients corresponding to the positive offset are stored. The coefficients corresponding to the negative offset are read from the symmetrical storage content based on the sign bit, thus halving the coefficient storage space. With the above fixed-point and compressed storage settings, imaging quality indicators such as peak sidelobe ratio and integral sidelobe ratio still meet the requirements.
[0119] Figure 5 This is a schematic diagram of a dedicated processing unit and a hierarchical mapping architecture for interpolation acceleration. For example... Figure 5 As shown, the core computing resources of the chip in this embodiment are a coarse-grained reconfigurable array, consisting of multiple processing units connected by a hierarchical interconnect network. Each basic processing unit includes an arithmetic logic unit (supporting fixed-point / floating-point addition, subtraction, multiplication, comparison, coordinate rotation, and other operations), a local register file, and a routing unit. The hierarchical interconnect network is divided into three levels: the first level is an intra-cluster fully crossover switch, where four processing units within the same cluster can be arbitrarily interconnected in pairs; the second level is a two-dimensional mesh interconnect, where 4×4 or 8×8 processing unit arrays transmit data through connections in adjacent directions; the third level is a global bus used for broadcasting instructions and synchronization signals. The chip also integrates an embedded programmable logic unit for implementing high-speed interface protocols.
[0120] For the critical step of interpolation resampling, this embodiment incorporates a dedicated interpolation processing unit within the reconfigurable array as an enhanced processing unit. For example... Figure 5As shown, the dedicated interpolation processing unit integrates three collaborative modules: the coordinate calculation module adopts a pipelined CORDIC and multiplier array, supporting fast square root and exponentiation operations, and can output a mapped coordinate every clock cycle after the pipeline is established; the neighborhood address generator supports neighborhood addressing modes of various interpolation kernels such as 4-point, 6-point, and 8-point, and can generate the address sequence of neighborhood points in real time based on non-integer coordinates. Its working process is as follows: the non-integer coordinate (e.g., 123.3) is split into an integer part (123) and a fractional part (0.3). The system retrieves a preset offset table based on the selected interpolation kernel length (e.g., offsets for 8-point interpolation [-3, -2, -1, 0, 1, 2, 3, 4]), adds each offset to its integer part to obtain the address of each neighboring point (120, 121, ..., 127), and then performs an out-of-bounds check on each address (i.e., determines whether the address is between 0 and N-1). The entire process is completed by the adder and comparator within a single clock cycle. The weighted accumulation module contains a configurable complex multiply-accumulate tree, supporting parallel weighted accumulation of interpolation kernels of different lengths. These three modules form a deep pipeline, which, once established, can output the complete interpolation result of one output point within each clock cycle. Furthermore, the dedicated interpolation processing unit is equipped with a high-speed channel directly connected to memory, eliminating the need for mesh interconnect forwarding for memory access.
[0121] To deploy wavenumber domain algorithms in an orderly manner onto the aforementioned reconfigurable array, such as Figure 5 As shown, this embodiment adopts a three-level hierarchical task mapping structure: the imaging task package is the highest level, corresponding to a complete wavenumber domain imaging application, containing the configuration information required for the entire processing flow; the processing function package is decomposed from the imaging task package, corresponding to specific processing stages, such as two-dimensional spectrum generation, consistent phase compensation, frequency domain interpolation resampling, and image spatial restoration; the hardware configuration behavior is the lowest level unit, corresponding to the specific operation mode and interconnection relationship of the basic processing unit or dedicated interpolation processing unit. Through hierarchical indexing, the upper-layer software only needs to manage a small number of imaging task packages, and the lower-layer hardware is parsed step by step, realizing fine control of the array, loading configuration information on demand, and effectively compressing the amount of configuration information.
[0122] Figure 6 This is a timing diagram for configuring prefetching and pipeline switching. For example... Figure 6As shown, in the traditional serial mode, the array must complete the current processing stage before loading the configuration information for the next processing stage, during which the array is in a waiting state. However, in the pipelined mode of this embodiment, the chip has a built-in independent configuration controller managing the on-chip configuration cache. When the reconfigurable array executes the current processing function package (e.g., two-dimensional spectrum generation), the configuration controller loads the configuration information for the next processing function package (e.g., consistent phase compensation) from off-chip memory into the on-chip configuration cache according to a predetermined processing flow. When the current function package is completed, the configuration information for the next function package is ready, and the array completes the switch within a single clock cycle by updating registers. Therefore, the configuration loading process overlaps with the calculation process in time, and the time overhead introduced by the processing stage switching is reduced to a negligible level.
[0123] Figure 7 This is a schematic diagram of a multi-granularity parallel processing model. For example... Figure 7 As shown, this embodiment organizes parallelism at three levels: stage-level parallelism maps different processing functions to different regions of the array. For example, the spectrum generation unit group and the interpolation unit group work together in a pipeline manner, with the output of the former group directly serving as the input of the latter group; data-level parallelism divides the data within the stage, for example, decomposing the two-dimensional spectrum generation into multiple parallel one-dimensional transformations, or dividing the imaging frequency domain into multiple data blocks for parallel interpolation; and operation-level parallelism is implemented within the processing unit through pipelines and vectorization. For example, the depth pipeline in the dedicated interpolation processing unit works with the multi-operand reading of the local register file, outputting one interpolation result per clock cycle.
[0124] Figure 8 This is a schematic diagram of a layered optimization design for a storage system. For example... Figure 8 As shown, to address the high-bandwidth neighborhood access requirements of interpolation operations, this embodiment designs a multi-bank interleaved shared memory structure: the shared memory is divided into multiple (e.g., 16) independent memory banks, and the two-dimensional frequency domain data is stored in an interleaved manner according to distance and azimuth frequencies, so that data points with adjacent frequency coordinates are distributed in different memory banks. When multiple dedicated interpolation processing units initiate neighborhood access simultaneously, each request is sent to the shared memory controller, and the arbitration logic distributes each request to the corresponding memory bank according to the target address: requests with target addresses belonging to different memory banks can be responded to in parallel; requests with target addresses in the same memory bank are cached by the request queue corresponding to that memory bank (e.g., a queue with a depth of 8) and processed sequentially. The compiler also pre-schedules accesses to distribute multiple requests to different memory banks as much as possible.
[0125] like Figure 8As shown, this embodiment also includes a data prefetching engine. Although the coordinate mapping of Stolt interpolation is non-linear, the coordinate changes of adjacent output points are continuous and smooth. Based on this, the data prefetching engine uses a differential prediction method to predict the subsequent access location. The prediction formula can be expressed as: Next coordinate ≈ Current coordinate + (Current coordinate) (Previous coordinate) + 0.5 × Δ²; Wherein, the next coordinate is the coordinate mapped to the output point to be predicted, which is arranged after the current output point in the processing order; the current coordinate and the previous coordinate are the coordinates mapped to the two most recent output points whose coordinate calculations have been completed, respectively; (current coordinate) The previous coordinate reflects the first-order change of the coordinate, that is, the rate of change of the coordinate between adjacent output points; Δ² represents the change of the first-order change of the coordinate (i.e., the second-order difference), which reflects the trend of the change of the coordinate rate itself; the coefficient 0.5 is the weighting coefficient of the second-order term.
[0126] The formula means: using the current coordinates as a reference, extrapolate one step according to the rate of change of the coordinates, and correct according to the trend of the rate of change, thereby estimating the position of the next coordinate. The data prefetch engine maintains the historical records of the four most recent coordinates, and loads the neighborhood data of the predicted position from off-chip memory to on-chip cache 2 to 4 output points in advance. Because the coordinates of adjacent output points change slowly, the prediction hit rate is high, and the off-chip memory access latency is hidden in the calculation process of the previous output points. The local register file, global temporary register, and shared memory constitute a hierarchical storage system, which together ensures the data supply of the computing array.
[0127] To meet the power consumption constraints of the spaceborne platform, this embodiment also integrates an energy efficiency optimization mechanism that works in conjunction with interpolation accuracy.
[0128] First, it features wide voltage dynamic adjustment. The chip supports adjusting the core operating voltage at the processing unit array level or at the individual processing unit level: for regular calculation tasks such as Fourier transform, a higher operating voltage is maintained to ensure throughput; for non-critical path operations such as interpolation coefficient calculation, the supply voltage is reduced to reduce dynamic power consumption (dynamic power consumption is approximately proportional to the square of the operating voltage).
[0129] Secondly, the voltage mapping of interpolation kernel weight perception. The weight of each coefficient in the interpolation kernel refers to the degree of influence of that neighboring point on the interpolation result. The center point (with the smallest offset) has the largest weight, and the weight decreases as it gets closer to the edge. In this embodiment, the calculation of each neighboring point is divided into three levels according to the relative weight (i.e., the ratio of the weight of that point to the maximum weight of the center point): the relative weight is not less than 0.5, which is the large weight level; the relative weight is between 0.1 and 0.5, which is the medium weight level; and the relative weight is less than 0.1, which is the small weight level. Taking the 8-point Kaiser window interpolation kernel with parameter β=6 as an example: the relative weight at offset 0.5 is 1.0, which belongs to the large weight level; the relative weight at offset 1.5 is 0.672, which belongs to the large weight level; the relative weight at offset 2.5 is 0.279, which belongs to the medium weight level; and the relative weight at offset 3.5 is 0.058, which belongs to the small weight level. During the hierarchical mapping stage, the compiler maps high-weight operations to processing units in the high-voltage domain (e.g., 1.0V) to ensure computational accuracy, medium-weight operations to processing units in the medium-voltage domain (e.g., 0.9V), and low-weight operations to processing units in the low-voltage domain (e.g., 0.7V) to save power. At the aforementioned default thresholds, the interpolation module's power consumption can be reduced by approximately 40%, while the peak-to-sidelobe ratio can still achieve a value better than -35dB. These thresholds (0.5 and 0.1) are not fixed and can be configured by the compiler according to actual imaging quality requirements.
[0130] Third, dynamic voltage and frequency adjustment. The operating frequency and voltage of the chip or local area are adjusted according to the real-time computing load: during pulse gaps or task idle periods, the frequency and voltage are reduced to enter a low-power mode, and the normal operating state is restored when new data arrives; for interpolation operations, the operating frequency of the computing unit can also be adjusted according to the length of the selected interpolation core. When the interpolation core is longer, the frequency is increased to ensure throughput, and when the interpolation core is shorter, the frequency is reduced to save power.
[0131] Figure 9 This is a schematic diagram illustrating the complete processing flow of wavenumber domain algorithms on a dynamically reconfigurable chip. For example... Figure 9 As shown, the echo data is first transformed into a two-dimensional frequency domain through two-dimensional spectrum generation, then focused at the reference distance through consistent phase compensation, then decoupled from the range and azimuth through frequency domain interpolation and resampling, and finally obtained as a focused SAR image through image space recovery; the processing stages are connected by the aforementioned configuration prefetching and pipeline switching mechanism, and the data flows between the processing array and the hierarchical storage system according to the aforementioned parallel and prefetching mechanism.
[0132] It should be noted that there are alternative solutions for some of the specific implementation methods in this embodiment. For example, in terms of the architecture of the interpolation processing unit, the coordinate calculation module can be decoupled into a shared accelerator, which can be used by multiple interpolation processing units through a competitive process, or a more general CORDIC unit can be used to implement the square root operation, trading some efficiency for flexibility. In terms of interpolation methods, in addition to weighted accumulation based on interpolation kernels, interpolation methods based on Fourier transform can also be used, which can achieve frequency domain interpolation by zero-padding Fourier transform, avoiding explicit interpolation calculation, but increasing the number of transformation points. Alternatively, interpolation based on spline functions can be used, achieving different trade-offs between smoothness and computational load. In terms of memory structure, in addition to multi-body interleaved shared memory, a temporary register structure in which data prefetching and replacement are explicitly managed by the compiler can be used to reduce hardware complexity, or a three-dimensional stacked memory can be used to alleviate memory access pressure through a wider on-chip bandwidth. In terms of parallel models, in addition to multi-granularity parallel processing models, a data flow driven model that triggers computation by data arrival can be used to achieve finer-grained dynamic parallelism, or a single instruction multiple data flow model can be used to achieve regularized parallelism through vectorized instructions to simplify control logic. All of the above alternative solutions can achieve the inventive objectives of the embodiments of this application.
[0133] Simulation results show that, under the configuration of this embodiment, the throughput of a single dedicated interpolation processing unit can reach one output point per clock cycle, and the overall interpolation processing efficiency is several times higher than that of the FPGA-based implementation when multiple units are in parallel. The configuration of prefetching and pipeline switching mechanisms reduces the time overhead of switching processing stages to a negligible level, and multi-volume interleaved storage and data prefetching can increase the effective storage bandwidth utilization to over 75%, resulting in an overall processing performance improvement of more than 100%. Combined with wide voltage regulation and weighted voltage mapping, the system energy efficiency ratio can be improved by 2 to 3 times compared to the FPGA-based implementation. At the same time, since the interpolation core configuration, algorithm core combination, and execution order can all be switched during operation through configuration information, this embodiment can flexibly adapt to different imaging quality requirements and imaging parameters without redesigning the hardware for parameter changes.
[0134] In summary, this embodiment analyzes the computational and memory access characteristics of wavenumber domain imaging algorithms and designs a dedicated interpolation processing unit that integrates coordinate calculation, neighborhood addressing, and weighted accumulation. Combined with hierarchical task mapping, configuration prefetching and pipelined switching, multi-body interleaved storage and data prefetching, and voltage and frequency adjustment adapted to the interpolation kernel weights, this enables computationally intensive and memory-undefined interpolation resampling operations in wavenumber domain imaging to be executed efficiently on a dynamically reconfigurable chip. This improves interpolation processing throughput and memory bandwidth utilization, reduces processing stage switching overhead, lowers system power consumption, and flexibly adapts to different imaging parameters and imaging quality requirements.
[0135] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A wavenumber domain imaging interpolation processing device, characterized in that, The device includes an interpolation processing unit, which integrates a coordinate calculation module, a neighborhood address generation module, and a weighted accumulation module connected in sequence. The coordinate calculation module is used to calculate the coordinates of the current output point mapped to the input frequency domain data; The neighborhood address generation module is used to determine multiple neighboring data points of the current output point based on the coordinates, and generate the addresses of the multiple neighboring data points; The weighted accumulation module is used to obtain the data of the multiple neighboring data points according to the address, determine the interpolation kernel coefficient of each neighboring data point according to the coordinate, and weight and accumulate the data of the multiple neighboring data points according to the interpolation kernel coefficient to obtain the interpolation result of the current output point.
2. The apparatus according to claim 1, characterized in that, The coordinate calculation module includes a streamlined coordinate rotation digital computer and a multiplier array; The multiplier array is used to perform exponentiation on the frequency coordinates of the current output point, and the pipelined coordinate rotation digital computer is used to perform square root operation on the result of the exponentiation to obtain the coordinates of the current output point mapped to the input frequency domain data.
3. The apparatus according to claim 1, characterized in that, The neighborhood address generation module stores multiple sets of preset offsets, each set of preset offsets corresponding to a different number of neighboring data points. The neighborhood address generation module selects a set of preset offsets from the multiple sets of preset offsets according to the configuration signal, splits the coordinates into an integer part and a decimal part, and adds each offset in the selected set of preset offsets to the integer part to obtain the addresses of the multiple neighboring data points.
4. The apparatus according to claim 1, characterized in that, The weighted accumulation module includes a complex multiply-accumulate tree, which includes multiple complex multiply-accumulate units. The number of complex multiply-accumulate units participating in the operation is set according to the number of the multiple neighboring data points. The coordinate calculation module, the neighborhood address generation module, and the weighted accumulation module process different output points simultaneously to output the interpolation results point by point.
5. The apparatus according to claim 3, characterized in that, The device stores an interpolation kernel coefficient table. The weighted accumulation module uses the decimal part of the coordinates as an index to obtain the interpolation kernel coefficients of each neighboring data point from the interpolation kernel coefficient table. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets in a selected set of preset offsets. The interpolation kernel coefficients corresponding to negative offsets are obtained by the weighted accumulation module by mirroring the stored interpolation kernel coefficients according to the sign of the offset.
6. The apparatus according to claim 1, characterized in that, The device also includes a shared memory and a data prefetch engine; The shared memory includes multiple independent storage banks, and the input frequency domain data is interleaved and stored in the multiple independent storage banks according to frequency coordinates; The device includes multiple interpolation processing units, each of which acquires data of neighboring data points distributed in different independent storage units in parallel via arbitration logic; Each output point is interpolated in a predetermined order. The data prefetching engine performs differential prediction based on the coordinates of multiple output points with obtained coordinates to obtain the address of the neighboring data point of the subsequent output point. The subsequent output point is the output point arranged after the multiple output points with obtained coordinates in the predetermined order. Before performing interpolation calculations on the later output points, the data prefetching engine loads the data corresponding to the address obtained by differential prediction from the off-chip memory into the on-chip cache, and each interpolation processing unit obtains the loaded data from the on-chip cache.
7. The apparatus according to claim 4, characterized in that, The operating voltage of the plurality of complex multiply-accumulate units can be adjusted independently. The plurality of complex multiply-accumulate units include a first complex multiply-accumulate unit and a second complex multiply-accumulate unit. Data of neighboring data points whose interpolation kernel coefficient is greater than a preset threshold is processed by the first complex multiply-accumulate unit, and data of neighboring data points whose interpolation kernel coefficient is not greater than the preset threshold is processed by the second complex multiply-accumulate unit. The operating voltage of the first complex multiply-accumulate unit is higher than that of the second complex multiply-accumulate unit. Furthermore, the operating frequency and operating voltage of the interpolation processing unit are adjusted according to the number of the plurality of neighboring data points. The larger the number, the higher the operating frequency and operating voltage.
8. The apparatus according to claim 1, characterized in that, The interpolation processing unit is a processing unit in the reconfigurable array; The reconfigurable array is used to sequentially execute multiple processing stages of wavenumber domain imaging. The device also includes a configuration controller and an on-chip configuration cache. During the execution of the current processing stage by the reconfigurable array, the configuration controller loads the configuration information of the next processing stage from off-chip memory into the on-chip configuration cache. After the current processing stage is completed, the reconfigurable array switches to the next processing stage according to the configuration information in the on-chip configuration cache.
9. A wavenumber domain imaging interpolation processing method, characterized in that, The method includes: Calculate the coordinates of the current output point mapped to the input frequency domain data; Based on the coordinates, determine multiple neighboring data points of the current output point, and acquire the data of the multiple neighboring data points; The interpolation kernel coefficients of each neighboring data point are determined based on the coordinates, and the data of the multiple neighboring data points are weighted and accumulated according to the interpolation kernel coefficients to obtain the interpolation result of the current output point.
10. The method according to claim 9, characterized in that, While performing the weighted accumulation to obtain the interpolation result of the current output point, determine multiple neighboring data points of the next output point, obtain the data of the multiple neighboring data points of the next output point, and calculate the coordinates of the next output point mapped to the input frequency domain data. The step of determining the interpolation kernel coefficients for each neighboring data point based on the coordinates includes: Using the decimal part of the coordinates as an index, the interpolation kernel coefficients of each neighboring data point are obtained from the interpolation kernel coefficient table. The offset is the offset of the neighboring data point relative to the coordinates. The interpolation kernel coefficient table only stores the interpolation kernel coefficients corresponding to positive offsets. The interpolation kernel coefficients corresponding to negative offsets are obtained by mirroring the stored interpolation kernel coefficients according to the sign of the offset.