A Hardware Acceleration Method Based on Beamforming for Underwater Detection

By designing intelligent microsystems on ARM and FPGAs, hardware acceleration processing of water acoustic signals is achieved, and the existing water acoustic signal processing system is solved, and the real-time processing rate and low-power and miniaturization characteristics of the system are improved.

CN115576230BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211045062.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-29
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The existing water acoustic signal processing system uses software to process data for a long time and high energy consumption, and has low real-time realization, which reduces the timeliness of data.

Method used

Using an intelligent micro system based on ARM and FPGA, the underwater sensor data is received through the Ethernet interface, and the format conversion and FFT calculation are performed in the FPGA's BRAM module to realize signal positioning and print the results through the serial port.

Benefits of technology

The real-time processing rate of water acoustic signals is improved, and the shortcomings of traditional water acoustic signal processing equipment are overcome, which is the disadvantages of low processing rate, high power consumption and large volume, and meets the needs of low power consumption and miniaturization of intelligent micro systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115576230B_ABST
    Figure CN115576230B_ABST
Patent Text Reader

Abstract

A hardware acceleration method based on beamforming for underwater detection provided by the present invention is applied to an intelligent microsystem. Through the PS side applied on the ARM and the PL side applied on the FPGA, the PL side sends the original format data stored in the DDR into its own BRAM module for storage; when its own custom DATA_TO_PL IP core takes out the original format data from the BRAM module and sends it to the data conversion module for format conversion, it is divided into 32 paths and sent to its own multi-channel FFT operation module for operation processing, and then sent to the positioning module to achieve signal positioning. The positioning result is sent to the PS side through the RESULT_TO_PS module. The present invention overcomes the disadvantages of low processing rate, high power consumption and large volume of traditional underwater acoustic signal processing equipment on the DSP processing platform, and improves the real-time processing rate of underwater acoustic signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of detection technology, and particularly relates to a hardware acceleration method based on beamforming for underwater detection. Background Art

[0002] Underwater acoustic signals are the most effective carriers for current underwater information transmission. They can not only reflect multiple parameters of the ocean environment but also provide comprehensive information for identifying underwater targets. The underwater acoustic processing module is a powerful payload of underwater drones. Looking ahead, the current underwater acoustic signal processing systems have tended to develop in the direction of being more miniaturized, intelligent, integrated, low-power, and capable of long-term unmanned operation.

[0003] Currently, the mainstream underwater acoustic information processing hardware platforms in domestic practical applications are DSP (He Zhijun. Research on Aviation Sonar Beamforming Algorithm and DSP Implementation [D]. Northwestern Polytechnical University), or software such as Matlab to implement preprocessing calculations such as quadrature demodulation, FIR filtering, and beamforming. Then, the data is stored in memories such as SD cards, SDRAM, and Flash. Finally, the underwater acoustic signal data is uploaded to the SD card through a network port or serial communication circuit. However, in practical applications, processing data through software first requires manually removing the SD card, then exporting the data, and then performing preprocessing calculations such as quadrature demodulation, FIR filtering, and beamforming on the data. This is time-consuming and energy-consuming, and the degree of real-time performance is not high, reducing the timeliness of the data. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a hardware acceleration method based on beamforming for underwater detection. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] A hardware acceleration method based on beamforming for underwater detection provided by the present invention is applied to an intelligent micro-system. The intelligent micro-system includes a PS side applied on an ARM and a PL side applied on an FPGA. The hardware acceleration method based on beamforming for underwater detection includes:

[0006] The PS side applied on the ARM is used to receive the sensing data of the underwater sensor transmitted by Ethernet through its own Ethernet interface, parse the sensing data according to a custom frame parsing function to extract the original format data, and store it in its own DDR.

[0007] Applied to the PL side of the FPGA, it is used to send the original format data stored in the DDR into its own BRAM module for storage; when its custom DATA_TO_PL IP core receives a request for data from the signal receiving module on the PS side, it retrieves the original format data from the BRAM module and sends it to the data conversion module for format conversion. The converted data is divided into 32 channels and sent to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts processing after receiving all frame data. The processed data is sent to multiple PE modules for covariance matrix operation and w operation, and then sent to the positioning module to achieve signal positioning. The positioning result is sent to the PS side through the RESULT_TO_PS module;

[0008] The PS side applied to the ARM is used to correct the angle of the positioning result and print the corrected result through the serial port.

[0009] Advantages of the present invention:

[0010] A hardware acceleration method based on beamforming for underwater detection provided by the present invention is applied to an intelligent microsystem. Through the PS side applied to the ARM and the PL side applied to the FPGA, the original format data stored in the DDR is sent into its own BRAM module for storage; when its custom DATA_TO_PL IP core receives a request for data from the signal receiving module on the PS side, it retrieves the original format data from the BRAM module and sends it to the data conversion module for format conversion. The converted data is divided into 32 channels and sent to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts processing after receiving all frame data. The processed data is sent to multiple PE modules for covariance matrix operation and w operation, and then sent to the positioning module to achieve signal positioning. The positioning result is sent to the PS side through the RESULT_TO_PS module, thus completing the hardware acceleration of the beamforming algorithm on the zynq7000 series development board and packaging the IP core of the beamforming algorithm. It overcomes the disadvantages of low processing rate, high power consumption and large volume of traditional underwater acoustic signal processing equipment on the DSP processing platform, and improves the real-time processing rate of underwater acoustic signals.

[0011] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0012] Figure 1 It is a schematic diagram of the process of a hardware acceleration method based on beamforming for underwater detection provided by an embodiment of the present invention applied to an intelligent microsystem;

[0013] Figure 2 It is a data transmission flow chart provided by an embodiment of the present invention;

[0014] Figure 3 is a schematic diagram of the Ethernet data receiving process provided by an embodiment of the present invention;

[0015] Figure 4 is a schematic diagram of the AXI Ethernet IP core and the DMA IP core provided by an embodiment of the present invention;

[0016] Figure 5 is a data frame processing flow chart provided by an embodiment of the present invention;

[0017] Figure 6 is a schematic diagram of the Zynq7 Processing System IP core and the AXI Interconnect IP core provided by an embodiment of the present invention;

[0018] Figure 7 is a schematic diagram of the DATA_TO_PL IP core provided by an embodiment of the present invention;

[0019] Figure 8 is a schematic diagram of the principle of the multi-channel FFT operation module provided by an embodiment of the present invention;

[0020] Figure 9 is a schematic diagram of the 32-bit fixed-point data format provided by an embodiment of the present invention;

[0021] Figure 10 is a schematic diagram of the FFT IP core data input / output frame structure provided by an embodiment of the present invention;

[0022] Figure 11 is a schematic diagram of the FFT IP core interface provided by an embodiment of the present invention;

[0023] Figure 12 is a schematic diagram of the architecture options and the corresponding resource throughput provided by an embodiment of the present invention;

[0024] Figure 13 is a schematic diagram of the RTL principle of the multi-channel FFT operation module provided by an embodiment of the present invention;

[0025] Figure 14 is a schematic diagram of the dual-port RAM interface and the covariance matrix operation module provided by an embodiment of the present invention;

[0026] Figure 15 is a STATE state transition diagram provided by an embodiment of the present invention;

[0027] Figure 16 is a schematic diagram of the spatial spectrum function calculation principle provided by an embodiment of the present invention;

[0028] Figure 17 is a schematic diagram of the positioning module provided by an embodiment of the present invention;

[0029] Figure 18 is the flowchart of transmitting back the positioning result provided by an embodiment of the present invention;

[0030] Figure 19 is the schematic diagram of the RESULT_TO_PS IP core provided by an embodiment of the present invention;

[0031] Figure 20 is the comparison diagram of the running time of the method of the present invention applied to the intelligent microsystem and the x86 PC;

[0032] Figure 21 is the schematic diagram of the FFT calculation process provided by an embodiment of the present invention;

[0033] Figure 22 is the schematic diagram of the DSP test result. Specific embodiments

[0034] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0035] As Figure 1 shown, a hardware acceleration method based on beamforming for underwater detection provided by the present invention is applied to an intelligent microsystem. The intelligent microsystem includes: a PS end applied to an ARM and a PL end applied to an FPGA. The hardware acceleration method based on beamforming for underwater detection includes:

[0036] The PS end applied to the ARM is used to receive the sensing data of the underwater sensor transmitted by Ethernet through its own Ethernet interface, parse the sensing data according to a custom frame parsing function to extract the original format data, and store it in its own DDR;

[0037] The PL end applied to the FPGA is used to send the original format data stored in the DDR into its own BRAM module for storage; when its own custom DATA_TO_PL IP core receives a signal reception module of the PS end requesting data, it takes out the original format data from the BRAM module and sends it to a data conversion module for format conversion, divides the converted data into 32 paths and sends it to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts to process after receiving all frame data, sends the processed data to multiple PE modules for covariance matrix operation and w operation, and then sends it to a positioning module to implement signal positioning, and sends the positioning result to the PS end through the RESULT_TO_PS module;

[0038] The PS end applied to the ARM is used to correct the angle of the positioning result and print the correction result through a serial port.

[0039] The present invention designs and optimizes an operation module based on the Zynq7100 chip to achieve cross-platform transplantation of Matlab algorithms. Figure 1 This is the overall schematic diagram of an application of a 32-element signal to underwater detection based on a beamforming hardware acceleration method provided by the present invention.

[0040] Reference Figure 2 , Figure 2 describes the entire data transmission process. The data is transmitted to the PS side through Ethernet, and the original data format is retrieved through a custom frame parsing function, and then stored in the DDR. The data is transmitted to the PL side through the AXI GP interface, and then the BRAM Controller IP core of the PL side is used to send the data into the BRAM module of the PL side for storage. When the signal receiving module requests data, the data is format-converted from the BRAM through the custom DATA_TO_PL IP core and divided into 32-channel signals and sent into the multi-channel FFT operation module of the PL side. After the operation module of the intelligent micro-system receives 1024 frames of complete data, it starts to process the data. After the data processing is completed, the final result is sent to the PS side through the RESULT_TO_PS module, and the final result is printed through the serial port.

[0041] It should be noted that: the present invention does not need to send the original data into the Ethernet, but directly processes and calculates on the intelligent micro-system and sends the positioning result to the Ethernet. Compared with the prior art, the present invention has a high degree of real-time and strong timeliness of data.

[0042] During function verification, the present invention uses the signal source generated by the upper computer as the input source of the system. The signal frequency of the 32-element data is 1000 Hz, the sampling frequency is 1 kHz, and the number of sampling points is 1024 points, that is, the sampling time of the signal source is The running time of the PC to implement the function of the beamforming algorithm once is 5.3 ms, and the running time of the present invention applied to the intelligent micro-system is 0.136 ms. When continuous sampling is performed at the front end, both the PC and the intelligent micro-system can perform real-time processing of the signal source. However, when the sampling frequency is increased by 100 times, the sampling time of the signal source is 1.024 ms. At this time, the running time of the PC is greater than the signal sampling time and cannot meet the signal real-time processing requirement, but the intelligent micro-system can still meet the signal real-time processing requirement.

[0043] As a specific implementation manner of the present invention, before the sensing data of underwater sensing of the present invention is sent to the PS side, it is processed by a PHY chip at the physical layer and sent to the Ethernet IP core of the PL side through the RGMII interface, and the Ethernet IP core sends the sensing data to the PS side through the Ethernet interface of the PS side in the form of data frames through DMA for waiting for processing.

[0044] Figure 3 It is a flowchart of Ethernet receiving data. After being processed by the PHY chip at the physical layer, the data is sent to the Ethernet IP core at the PL end. The Ethernet IP core sends the data frame to the PS end through DMA and waits for data frame processing.

[0045] As a specific implementation manner of the present invention, the PS end of the present invention receives a data frame transmitted in the form of a data frame through an Ethernet interface, and determines whether the data frame structure meets the frame format requirements. If the requirements are met, the low_level_input function is called to convert the data frame into 32-bit single-precision floating-point data; if the frame format requirements are not met, the data frame is discarded, and the 32-bit single-precision floating-point data is stored in the DDR.

[0046] The Ethernet IP core of the present invention is a ZYNQ7 Processing System IP core. After receiving the sensing data processed by the PHY chip, the Ethernet IP core sends it to the S_AXIS_S2MM interface of the direct memory access module DMA through the M_AXIS_RXD interface;

[0047] As Figure 4 shown in Figure a, the AXI Ethernet IP core of the present invention is responsible for parsing the data at the data link layer to ensure the correctness of the data. The direct memory access module DMA is as Figure 4 shown in Figure b. After the Ethernet data is sent out from the PHY chip, it is sent to the AXI Ethernet IP core through the RGMII interface. The data is sent to the S_AXIS_S2MM interface of the DMA through the M_AXIS_RXD interface. The DMA is connected to the AXI HP interface of the ZYNQ7 Processing System IP core through the M_AXIS_ interface to send the data to the PS end. The PS end receives and parses the data frame through the low_level_input function, extracts the original data, and sends it to the DDR.

[0048] The direct memory access module DMA is connected to the AXI HP interface of the ZYNQ7 Processing System IP core through the M_AXIS_ interface to send the sensing data to the PS end;

[0049] The PS end receives the data frame through the low_level_input function, parses the data frame according to the custom frame parsing function to extract the original format data, and extracts the original data and sends it to the DDR.

[0050] Figure 5It is a data frame processing flow chart. The PS side processes the data frame by calling the low_level_input function. If the requirements are met, the data is converted into 32-bit single-precision floating-point data and sent to the DDR, and is responsible for discarding the frame data.

[0051] The Zynq7 Processor System IP core is as Figure 6 shown in Figure a. In this IP core, the Ethernet data reception, data frame parsing, and processing result reception are completed, and the AXI Ethernet IP core, DMA IP core, AXI Interconnect IP core, AXI BRAM Controller IP core, and DATA_TO_PL IP core are controlled. On the PS side, first, the control parameters are initialized through value_init, the dma is initialized through dma_init, the ddr is initialized through ddr_init, and the interrupt is initialized through the interupt_intial function. Subsequently, after the Ethernet data processing is completed and stored in the DDR, finally, the transfer_data function is called to control the AXI BRAM Controller IP core to transfer the data in the DDR to the BRAM on the PL side.

[0052] The AXI Interconnect module is as Figure 6 shown in Figure b. The AXI GP interface is connected to the S_AXI interface of the AXI Interconnect IP core, and the control signals for each module on the PL side are transmitted through this interface. The M0_AXI to M4_AXI interfaces are respectively connected to the AXI Ethernet IP core, DMA IP core, AXI BRAM Controller IP core, DATA_TO_PL IP core, and result_to_ps core.

[0053] As a specific implementation manner of the present invention, the application of the present invention is on the PL side of the FPGA, and is used to send the original format data stored in the DDR into its own BRAM module for storage; when its own custom DATA_TO_PL IP core receives a request for data from the signal reception module on the PS side, the original format data is taken out from the BRAM module and sent to the data conversion module for format conversion, and the converted data is divided into 32 channels and sent to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts processing after receiving all frame data, and the processed data is sent to multiple PE modules for covariance matrix operation and w operation, and then sent to the positioning module to implement signal positioning, and the positioning result is sent to the PS side through the RESULT_TO_PS module, including:

[0054] Applied to the PL side of the DSP, it is used to send the original format data stored in the DDR into its own BRAM module for storage;

[0055] When the custom DATA_TO_PL IP core of itself receives the control signal generated by the PS-side receiving module's request for data, it reads the original format data from the BRAM module through the BRAM_POREB interface;

[0056] When the data_out_en of the custom DATA_TO_PL IP core is pulled high, the custom DATA_TO_PL IP core divides the original format data into 32 data signal paths through the Data interface and transmits them in parallel to the Float-point IP core for format conversion. In the Floating-point IP core, the single-precision floating-point data is fixed-pointed to 32-bit custom data;

[0057] It should be noted that: The DATA_TO_PL module is as Figure 7 shown. This IP core receives the PS-side control signal through the S_AXI_LITE interface, reads the data from the BRAM through the BRAM_POREB interface. Data is a 32-way data bus. When the data_out_en is pulled high, this IP core divides the data into 32 data signal paths and transmits them in parallel through the Data data bus to the multi-channel FFT operation module.

[0058] The Floating-point IP core divides the 32-bit custom data into 4 groups and sends them into four FFT IP cores through the interfaces of the FFT IP core for FFT operations respectively;

[0059] It should be noted that: Figure 8 For the schematic diagram of the multi-channel FFT operation module. Here, by balancing the data error precision and system resources, the present invention adopts the IEEE-754 single-precision format as the data source format of the FFT IP core. It not only ensures the data precision, but also reduces the double-precision 64-bit width to 32-bit width, saving a large amount of resources. The external interface is connected to the Data interface of the DATA_TO_PL IP core, and 32 parallel signals are sent into 32 Float-point IP cores in parallel for format conversion. In the Floating-point IP core, the single-precision floating-point data is fixed-pointed to 32-bit custom data. Figure 9 For the fixed-pointed data format, the fixed-point data format width is 32 bits, which includes 1 bit sign bit, 15 bit integer part and 16 bit fractional part. The 32 signals are divided into four groups and sent into four FFT IP cores for operations.

[0060] Four FFT IP cores perform FFT operations on their own 32-bit custom data respectively, and a total of 32 channels of data are output and sent to the Block Memory Generator IP core for storage through the DINA interfaces of 32 BRAMs;

[0061] It should be noted that: The data input / output frame structure of the FFT IP core is as Figure 10 shown. In the frame structure, 8-channel data are merged into one frame, and the input / output bit width is 512 bits. The fixed-point data format after format conversion does not contain imaginary parts, so 32-bit {1’b0} is used to fill in the input data s_axis_data_tdata. When s_axis_config_tvalid is pulled high, the configuration channel s_axis_config_tdata transmits 8’b11111111, indicating that all 8 channels are FFT operations.

[0062] Figure 11 is the interface diagram of the FFT IP core. Because the multi-channel architecture conflicts with the pipeline architecture, the Radix_4Burst I / O architecture is selected in the present invention. In addition, to facilitate the management and processing of the output data (corresponding to the operation of taking half of the points after the fft operation in the Matlab algorithm), the natural order output and index output are adopted in the present invention, and the output index (XK_INDEX) is included in the m_axis_data_tuser signal. The value of the XK_INDEX signal starts from 0. In the present invention, the transformation length of the FFT IP core is set to 1024, so the maximum value of the XK_INDEX signal is 1023. At this time, the m_axis_data_tvalid signal of the output channel is set to 0 in the next clock cycle. After the output channel stops outputting data, the XK_INDEX signal remains at 1008 until the system is reset.

[0063] In the FFT IP core configuration interface, four FFT architecture options are provided, namely Pipelined Streaming I / O, Radix_4Burst I / O, Radix_2Burst I / O, and Radix_2Lite Burst I / O. The differences between the four architectures lie in the time and resource consumption of the Fourier transform, Figure 12 is the architecture option and resource comparison.

[0064] In the design of this module, 32 Float-Point IP cores and 4 FFT IP cores are instantiated by the Verilog language in the present invention. Figure 13RTL schematic generated for Vivado. In the figure, the first column is the external interface, which obtains 32-element single-precision data format. The second column is 32 Float-Point IP cores, and the third column is 4 FFT IP cores.

[0065] Figure 14 In Figure a, it is the interface schematic diagram of Simple Dual-port RAM. The storage unit is the Block MemoryGenerator IP core (version v8.4). The depth is set to 512 and the bit width is 64 bit. The enable signal is the m_axis_data_tvalid signal of the output channel of the FFT IP core, and the address value is the EK_INDEX signal of the FFT IP core. The 4 FFT modules output a total of 32 data paths, which are respectively connected to the DINA interfaces of 32 BRAMs. When reading data, it is read from the DOUTB interface. After DOUTB, it is connected to the DATA of the PE module. The DOUTB data is transposed and then connected to DATA_CON.

[0066] The described PE ((Process Element)) module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, multiplies them to obtain the covariance matrix, and calculates multiple sets of spatial spectrum functions PBF through its own multiplexing and sends them to the positioning module;

[0067] The schematic diagram of the covariance matrix operation is as shown in Figure 14 Figure b. In order to obtain the spatial spectrum function, it is necessary to perform covariance matrix operation on the frequency domain signals of 32 elements. The present invention adopts a state machine structure. 32 state transitions can calculate the elements in one row of the covariance matrix. The 32×512-sized FFT matrix after coming out of the multi-channel FFT operation module sends the first row of 1×512 data of the matrix to each PE module in sequence through the DATA interface. Each PE module receives the same data. At the same time, each column of the conjugate transpose matrix of the FFT matrix is sent to each PE module through the DATA_CON interface in sequence, and 32 data in the first row of the result matrix are calculated. Subsequently, each row of the FFT matrix is sent to the DATA interface in sequence to obtain the covariance matrix R.

[0068] The described positioning module compares all sets of spatial spectrum functions PBF, selects the count value of the largest PBF as the positioning result, and sends it to the PS end through the RESULT_TO_PS module.

[0069] As a specific implementation manner of the present invention, the PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, and multiplying them to obtain the covariance matrix includes:

[0070] The PE module enters the initial StandBy state after reset. In this state, if it receives the enable ENA signal generated after the four FFT IP cores have stored all the data, it enters the Operation state.

[0071] In the Operation state, the PE module is controlled by an external counter to sequentially read the target data and the conjugate data of the target data from the Block MemoryGenerator IP. After each read, it multiplies the target data and the conjugate data to obtain a row of the covariance matrix and enters the Pause state, and resumes the Operation state during the next read.

[0072] If the number of reads of the PE module reaches 32 times, it enters the Stop state to obtain the covariance matrix.

[0073] Figure 15 It is the state transition diagram of the covariance matrix operation. After the system is reset, it enters the StandBy state. When the output signal enable ENA signal of the multi-channel FFT operation module is valid, the state machine enters the Operation state, and then enters 32 calculations. After each calculation, it enters the Pause state to store the data. After 32 loop calculations are completed, it enters the Stop state. The STATE signal acts on the frequency-domain data DATA and its conjugate signal DATA-CON after the FFT operation module. The PE in the figure is the multiply-accumulate operation unit. A total of 32 PE modules are defined, and 32 elements of the covariance matrix in a row are calculated during one state transition. After 32 state transitions, 1024 elements of the covariance matrix are calculated. In the pipeline structure of the FPGA, if the multiplexing method is adopted, signal storage will inevitably be involved. In the present invention, a level of RAM storage is added between the multi-channel FFT operation module and the state machine, and 32 RAMs with a depth of 512 are instantiated to store the frequency-domain signals of 32 array elements.

[0074] As a specific implementation manner of the present invention, the PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, multiplies them to obtain the covariance matrix, and obtains multiple sets of spatial spectrum functions PBF through its own multiplexing and sends them to the positioning module, including:

[0075] The PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, multiplies them to obtain the covariance matrix;

[0076] For the reuse of the PE module, the imaginary part of each group of operation factors w sent from the external ROM memory is inverted to obtain each group of conjugate factors w'. Each group of conjugate factors w' is multiplied by the sent covariance matrix to obtain the first part w'R of the spatial spectrum function. The operation factors w are multiplied by the first part w'R of the spatial spectrum function again to obtain the second part w'Rw of the spatial spectrum function, and each group of spatial spectrum functions PBF is obtained by shifting it 5 bits to the right and sent to the positioning module.

[0077] As Figure 16 shown, the numerator operation of the spatial spectrum function is w'Rw, where R is a 32×32 covariance matrix and w is a 32×1 vector. When calculating the first half w'R, the present invention uses a systolic array architecture, and the operation result of the two is a 1×32 vector. Here, it is still calculated through the PE module. w' is read from the ROM memory, and R is read from 32 BRAMs in sequence, and finally w'R is obtained. Then, w'Rw can be calculated continuously through the PE module. Since the final PBF = w'Rw / w'w and w'w is a fixed value of 32, the PBF can be obtained by shifting the value of w'Rw 5 bits to the right.

[0078] As a specific implementation manner of the present invention, the positioning module compares all groups of spatial spectrum functions PBF, and selects the count value of the largest PBF as the positioning result and sends it to the PS end through the RESULT_TO_PS module, including:

[0079] The positioning module receives each group of spatial spectrum functions PBF and counts once using a counter at the same time, compares all groups of spatial spectrum functions PBF, selects the count value of the largest spatial spectrum function PBF as the positioning result and sends it from the angle interface to the RESULT_TO_PS IP core, and generates an enable signal end_valid and sends it from the Interupt interface to the RESULT_TO_PS IP core;

[0080] The RESULT_TO_PS IP core, under the action of the enable signal, sends the positioning result to the PS end through the S_AXI_LITE interface.

[0081] Figure 17 For the schematic diagram of the final positioning module, there are 180 groups of w values in total. Therefore, w'Rw is calculated 180 times. After each calculation is completed, the calculation result of PBF is sent to this module for comparison. When the en_cnt counts to 179, the end_valid is pulled high. The flowchart for the output of the final positioning result azimuth_angle to return the positioning result is Figure 18 .

[0082] Figure 19For the RESULT_TO_PS module (v1.0), the azimuth_angle is connected to the angle interface of this IP core, and the end_valid is connected to the Interupt interface of the ZYNQ7 Processing System IP core. When the end_valid is pulled high, the CPU receives the interrupt, and the handler function on the PS side processes the interrupt. Subsequently, the final positioning angle is sent to the PS side through the S_AXI_LITE interface, and finally the final positioning angle is printed and output through the serial port.

[0083] The following is an experiment of the present invention to illustrate the efficiency and real-time performance of the present invention.

[0084] In the present invention, the installation environment of Matlab software is an Intel(R) Core(TM) i5-8300H CPU @ 2.30GHz processor, 8GB of memory, and a 64-bit operating system. Taking the preset angle of 30° as an example, the time results of running five times are shown in Table 1.1.

[0085] Table 1.1 Simulation time of azimuth angle 30° in Matlab environment

[0086]

[0087] From the results in Table 1.1, it can be seen that in the Matlab environment, the average time cost of five simulations of the signal source with an azimuth angle of 30° is 12.281 ms. The environment on the PC side is affected by the CPU threads and is limited by the serial execution structure of the algorithm program, resulting in instability and high latency in the Matlab environment simulation.

[0088] Performance comparison between Zynq and x86

[0089] The Zynq7000 series chips adopted by the intelligent microsystem are heterogeneous multi-core chips composed of an ARM CotexA9 dual-core processor and a programmable FPGA. Inside the chip, there is not only an ARM CortexA9 dual-core processor, but also other resources such as a cache, a DDR controller, a clock generator, and various I / O interfaces. Table 1.2 shows the comparison of hardware resources between the Zynq7100 chip and the x86 PC.

[0090] Table 1.2 Comparison of hardware resources between Zynq7100 and x86

[0091]

[0092]

[0093] The intelligent micro-system integrated with Zynq chips has a size of 260mm×142.75mm, a power consumption of 0.2W in the idle state, and a power consumption of 10.8W under full load. The x86 architecture is generally applied to PCs with a power consumption as high as 30W. Therefore, the Zynq-based intelligent micro-system can meet the requirements of low power consumption and small size for intelligent micro-systems.

[0094] The main frequency of the Zynq chip is 800MHz. The PL side of the chip is the Xilinx Kintex-7 series FPGA, which includes 444K logic cells, 277.4K lookup tables, 554K flip-flops, and 2020 DSP units. To calculate the maximum computing power of the FPGA, the number of single-precision data adders can be used to calculate the FLOPs of a system. The calculation method of the computing power of the PL side is as follows:

[0095] FLOPs=(Clock1×LCbasedAdder)+(Clock1×DSP48basedAdder) (2.1)

[0096] Through the manual of the Kintex-7 adder IP Core provided by Xilinx, it is known that an adder based on DSP48E is composed of 2 DSP slices and 289 LUT-FF pairs, and an adder based on Logic Cell is composed of 517 Logic Cells. According to the PL side resources of Zynq, it can be calculated that there are 1010 adders based on DSP48E and 267 adders based on LC. According to the manual of the Kintex-7 FPGA, the clock range of the adder based on DSP48 is 600MHz (slow)~891MHz (fastest), and the clock range of the adder based on Logic Cell is 667MHz (slow)~891MHz (fastest). Substituting the main frequency of 800MHz of the Zynq chip into formula (2.1), the computing power of the PL side of the chip is calculated as 800MHz×1010 + 800MHz×267 = 1.02TFLOPs.

[0097] The PS side of Zynq is composed of an ARM Cotex-A9 dual-core processor, which adopts the ARMv7 architecture, with a CPU main frequency of 1GHz, 2 cores, and the floating-point calculation value of the floating-point operation unit FPU per cycle is 32SP FLOPs / cycle. The calculation method of the computing power of the CPU is shown in formula (2.2):

[0098] FLOPs = CPU core number×CPU single-core main frequency×single-cycle floating-point calculation value (2.2)

[0099] Therefore, the computing power of the PS side of Zynq is 1 GHz × 2 × 1 SP FLOPs / cycle = 2 GFLOPs. Then the total computing power of Zynq is 1.022 TFLOPs.

[0100] The system CPU based on the x86 architecture has a main frequency of 2.3 GHz and four built-in cores. The CPU of this server is based on the AVX-256 instruction set. According to the analysis of Intel's Haswell architecture, the floating-point calculation value per cycle is 32 SP FLOPs / cycle. Substituting the number of cores and the CPU main frequency into (2.2), the computing power of the x86 PC is 4 × 2.3 GHz × 32 SP FLOPs / cycle = 294.4 GFLOPs.

[0101] By calculating the power consumption and computing power, the energy consumption ratios of Zynq and x86 are 185.82 GFLOPs / W and 9.813 GFLOPs / W respectively. The energy consumption ratio of Zynq is approximately 19 times that of x86. Under low energy consumption conditions, the Zynq chip can provide rich computing power resources for the algorithm acceleration system, enabling the underwater drone to better realize real-time signal processing.

[0102] When obtaining the running time, in order to synchronize the Matlab algorithm, the present invention enables and starts counting the clock cycles of the last group of data transmissions until the numerical output of the last spatial spectrum function stops counting. The total running time will be calculated as: 0.005 × the number of clock cycles us. The present invention compares the processing times of the algorithm in the system and PC environments respectively. As Figure 20 shown, affected by threads, the average running time of the PC is 5.3 ms, while the running time of the intelligent micro-system is stable at 0.136 ms, which can more efficiently realize the underwater acoustic signal positioning function.

[0103] The intelligent micro-system is small in size, light in weight, low in latency, and high in computing power density. Through the collaborative development and design of software and hardware, it can greatly improve the parallelization degree of the processing algorithm, thus meeting the requirements of real-time underwater signal processing under low power consumption conditions.

[0104] DSP-based algorithm implementation

[0105] DSP is called a digital signal processor, which is a kind of CPU designed by Texas Instruments (TI). It is usually used for data algorithm processing. Compared with other processors, its biggest features are powerful data processing capabilities, running speed, and pipeline structure. To further improve the execution efficiency and energy consumption ratio of the present invention, the present invention completes the DSP implementation of the beamforming algorithm based on the TMS320C6748 development board, conducts simulation tests in the CCS (Code Composer Studio) software, and completes experimental verification and performance analysis on the DSP development board.

[0106] The Matlab beamforming algorithm involves a large number of complex matrix operations and FFT calculations. In this solution, C language programming is used on CCS to complete the FFT function, covariance matrix operation, and spatial spectrum function operation respectively.

[0107] When performing FFT operation on the excitation signal X(t), the DSPF_sp_fftSPxSP function in the DSP's built-in math library is used for calculation. As Figure 21 shown, the FFT function first determines whether the number of sampling points is a power of 2 and determines the fast Fourier transform base rad (since the TI DSP library supports a maximum of 128K points calculated at one time, the number of sampling points cannot exceed 128K), then calculates the rotation factor, and finally substitutes the excitation signal, number of sampling points, transform base, rotation factor, and inverted array into the DSPF_sp_fftSPxSP function to calculate the frequency-domain signal Y0(ω).

[0108] Y0(ω) obtained through FFT calculation is a real matrix containing real and imaginary parts (in C language, there are no complex numbers i and j, which need to be artificially identified). It is first converted into a complex matrix, then the complex matrix is horizontally flipped and half of the number of rows is taken to obtain the complex matrix Y(ω). The conjugate transpose matrix Y′(ω) of Y(ω) is calculated, and according to the covariance formula R = Y′*Y, the covariance matrix R is calculated. A large number of complex multiplications are involved in the covariance operation. The function of complex multiplication is as follows.

[0109] 1. / / Complex multiplication

[0110] 2.Complex mulcomplex(complex c1,complex c2){

[0111] 3.complex temp;

[0112] 4.temp.real = c1.real*c2.real - c1.img*c2.img;

[0113] 5.temp.img = c1.real*c2.img + c1.img*c2.real;

[0114] 6.return temp;

[0115] 7.}

[0116] The calculation of the spatial spectrum function requires the covariance matrix R and the weighting vector ω, where ω can be obtained through Euler's formula. According to the formula PBF = (ω' * R * ω') / (ω' * ω), the numerator and denominator are calculated respectively. The complex division function is called to calculate the spatial spectrum values at each angle in the range of [-89°, 90°], and normalization processing is performed.

[0117] 1. / / Division of two complex numbers

[0118] 2.Complex divi(complex a,complex b) / / Division

[0119] 3. / / Removing the denominator in division can be converted to multiplication

[0120] 4.{

[0121] 5.Complex quo;

[0122] 6.Float den = b.real * b.real + b.img * b.img; / / Denominator

[0123] 7.quo.real = a.real * b.real + a.img * b.img;

[0124] 8.quo.real / = den; [[ID=X]]

[0125] 9.quo.img = a.img * b.real + a.real * b.img;

[0126] 10.quo.img / = den;

[0127] 11.Return quo;

[0128] 12.} [[ID=3X]]

[0129] Take the modulus of the spatial spectrum values at 180 angles respectively. Compare to obtain the angle at the maximum modulus value as the finally located angle. Finally, perform normalization processing on it, save the final result to a txt file, and use Matlab software for plotting.

[0130] The results are as Figure 22 shown. The excitation signal is cos(2πft), the number of array elements is 32, the array element spacing is taken as half wavelength, the target signal frequency is taken as 1000Hz, the sampling frequency is 10 4 Hz, the sound speed is taken as 1500m / s, the target azimuths are -72°, -39°, 7°, 30°, and the number of sampling points is 1024. CCS burns the program to the development board through the emulator, and the output result is saved in.txt, and the positioning result time is output as 0.31s in the Console window. Please note that there are some tags like

[0120] and

[0129] which seem to be part of a code or formatting in a specific context and are left as is without further interpretation as per the requirements. Also, the "X" marked lines are just for the purpose of highlighting where there might be some potential issues or areas for clarification in the original text if needed.

[0131] By comparing with the DSP, the computing time of Zynq to complete the beamforming algorithm is three orders of magnitude faster than that of the DSP to complete the beamforming algorithm.

[0132] Performance Comparison between Zynq and DSP

[0133] In order to further compare the execution efficiency and energy consumption ratio of the intelligent microsystem, the DSP implementation of the beamforming algorithm is completed based on the TMS320C6748 development board in the present invention, and simulation tests are carried out in the CCS 5.5 software, and experimental verification and performance analysis are completed on the DSP development board. TI TMS320C6748 is a low-power floating-point DSP processor, suitable for base station data beamforming, image processing, speech recognition, 3D graphics, etc. TMS320C6748 supports the high-performance digital signal processing and reduced instruction computer (RISC) technology of the DSP, and adopts a high-performance 32-bit processor of the 456MHz TMS320C674x series, which can fully meet the requirements of high energy efficiency, high scalability and low power consumption. Table 1.3 shows the comparison of the hardware resources between the Zynq7100 chip and the TMS320C6748.

[0134] Table 1.3 Comparison of Hardware Resources between Zynq7100 and TMS320C6748

[0135]

[0136]

[0137] It can be seen from Table 1.3 that the TMS320C6748 chip is a floating-point DSP with a main frequency of 456MHz, and its core module implements a two-level memory architecture. The first-level memory (L1) can be divided into an independent program memory (L1P) and a data memory (L1D), and L1 can be accessed by the CPU without delay. L2 is divided into L2 RAM (normally addressable on-chip memory) and L2 cache used to cache external memory locations. The algorithm implemented based on the DSP can make full use of the advantages of the DSP in data processing capabilities, and through fast instruction execution, complete the processing of multi-channel collected signal data, achieving a good processing effect, and the DSP has strong floating-point operation capabilities. The DSP is vulnerable to interference and restricted by its serial instruction stream, so to a large extent, it limits the data processing capabilities of the entire system. When multiple tasks need to be completed quickly at the same time, the system with the DSP as the core shows its deficiencies, and due to its hardware characteristics, it is very difficult to design the peripheral circuit for the DSP. Therefore, it is difficult to meet the requirements of online processing and high scalability of the intelligent microsystem. In summary, the Zynq7100 chip better meets the requirements of the intelligent microsystem.

[0138] The present invention uses a Zynq7000 series development board to complete the cross-platform transplantation of the beamforming algorithm. The intelligent micro-system is divided into the PS and PL parts. Among them, the high-speed DDR3 data transmission rate can reach up to 1600 Mb / s, and the total storage capacity is as high as 2 GB, meeting the high-speed data processing requirements. There are 755 dual-port BRAMs in the PL part, with each BRAM having a storage space of 36 KB and a total storage space of 29.2 MB. Both the ARM system and the FPGA system have the functions of independently processing and storing data, which can meet the storage requirements of the intelligent micro-system. The intelligent micro-system has a very high energy consumption ratio, which is 19 times higher than that of the traditional PC platform, and its computing power is as high as 1.022 TFLOPs, exceeding the traditional PC by more than 3.5 times.

[0139] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0140] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases.

[0141] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A hardware acceleration method based on beamforming for underwater detection, characterized in that, Applied to an intelligent microsystem, the intelligent microsystem includes: a PS side applied on an ARM and a PL side applied on an FPGA. The hardware acceleration method based on beamforming for underwater detection includes: The PS side applied on the ARM is used to receive the sensing data of the underwater sensor transmitted by Ethernet through its own Ethernet interface, parse the sensing data according to a custom frame parsing function to extract the original format data, and store it in its own DDR; The PL side applied on the FPGA is used to send the original format data stored in the DDR into its own BRAM module for storage; when its own custom DATA_TO_PL IP core receives a signal reception module request for data from the PS side, it takes out the original format data from the BRAM module and sends it to the data conversion module for format conversion, divides the converted data into 32 channels and sends it to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts processing after receiving all frame data, sends the processed data to multiple PE modules for covariance matrix operation and w operation, and then sends it to the positioning module to implement signal positioning, and sends the positioning result to the PS side through the RESULT_TO_PS module; The PS side applied on the ARM is used to correct the angle of the positioning result and print the corrected result through the serial port.

2. The acceleration method based on beamforming hardware for underwater detection according to claim 1, wherein Before the sensing data of the underwater sensor is sent to the PS side, it is processed by a PHY chip at the physical layer and sent to the Ethernet IP core of the PL side through the RGMII interface. The Ethernet IP core sends the sensing data to the PS side in the form of data frames through the Ethernet interface of the PS side through DMA for waiting for processing.

3. The hardware acceleration method based on beamforming for underwater detection according to claim 2, wherein, The PS side receives the data frame transmitted through the Ethernet interface in the form of a data frame, judges whether the data frame structure meets the frame format requirements. If it meets the requirements, it calls the low_level_input function to convert the data frame into 32-bit single-precision floating-point data; if it does not meet the frame format requirements, it discards the data frame and stores the 32-bit single-precision floating-point data in the DDR.

4. The hardware acceleration method based on beamforming for underwater detection according to claim 2, characterized in that The Ethernet IP core is a ZYNQ7 Processing System IP core. After receiving the sensing data processed by the PHY chip, the Ethernet IP core is sent to the S_AXIS_S2MM interface of the direct memory access module DMA through the M_AXIS_RXD interface; The direct memory access module DMA is connected to the AXI HP interface of the ZYNQ7 Processing System IP core through the M_AXIS_ interface to send the sensing data to the PS side; The PS side receives the data frame through the low_level_input function, parses the data frame according to a custom frame parsing function to extract the original format data, and extracts the original data and stores it in the DDR.

5. The hardware acceleration method based on beamforming for underwater detection according to claim 1, wherein The PL side applied on the FPGA is used to store the original format data stored in the DDR into its own BRAM module; when its custom DATA_TO_PL IP core receives a signal from the signal receiving module on the PS side requesting data, it retrieves the original format data from the BRAM module and sends it to the data conversion module for format conversion. The converted data is divided into 32 channels and sent to its own multi-channel FFT operation module, so that the multi-channel FFT operation module starts processing after receiving all frame data. The processed data is sent to multiple PE modules for covariance matrix operation and w operation, and then sent to the positioning module to achieve signal positioning. The positioning result is sent to the PS side through the RESULT_TO_PS module, including: The PL side applied on the FPGA is used to store the original format data stored in the DDR into its own BRAM module; When the control signal generated by the request for data from the receiving module on the PS side is received by its custom DATA_TO_PL IP core, the original format data is read from the BRAM module through the BRAM_POREB interface; When the data_out_en of the custom DATA_TO_PL IP core is pulled high, the custom DATA_TO_PL IP core parallelly transmits the original format data as 32 data signals through the Data interface to the Float-point IP core for format conversion. The floating-point IP core fixes the single-precision floating-point data to 32-bit custom data; The Floating-point IP core divides the 32-bit custom data into 4 groups and sends them to four FFT IP cores through the interfaces of the FFT IP cores for FFT operation respectively; The four FFT IP cores respectively perform FFT operations on their own 32-bit custom data, and a total of 32 channels of data are sent to the Block Memory Generator IP core for storage through the DINA interfaces of 32 BRAMs; The PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, multiplies them to obtain a covariance matrix, and calculates multiple sets of spatial spectrum functions PBF through its own multiplexing and sends them to the positioning module; The positioning module compares all sets of spatial spectrum functions PBF, selects the count value of the largest PBF as the positioning result, and sends it to the PS side through the RESULT_TO_PS module.

6. The hardware acceleration method based on beamforming applied to underwater detection according to claim 5, characterized in that, The PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, and multiplying them to obtain a covariance matrix includes: The PE module enters the initial StandBy state after reset. In this state, if it receives the enable ENA signal generated after the four FFT IP cores have stored all data, it enters the Operation state; When in the Operation state, the PE module is controlled by an external counter to sequentially read target data and the conjugate data of the target data from the Block Memory Generator IP. After each read, the target data and the conjugate data are multiplied to obtain a row of the covariance matrix, and then the PE module enters the Pause state and resumes the Operation state during the next read. If the number of reads of the PE module reaches 32 times, it enters the Stop state to obtain the covariance matrix.

7. An hardware acceleration method based on beamforming for underwater detection according to claim 5, characterized in that, The PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core, multiplies them to obtain the covariance matrix, and obtains multiple sets of spatial spectrum functions PBF through its own multiplexing and sends them to the positioning module, including: The PE module reads the target data and the corresponding conjugate data from the Block Memory Generator IP core and multiplies them to obtain the covariance matrix. The PE module is multiplexed. The imaginary part of each set of operation factors w sent from the external ROM memory is inverted to obtain each set of conjugate factors w'. Each set of conjugate factors w' is multiplied by the sent covariance matrix to obtain the first part w'R of the spatial spectrum function. The operation factors w are multiplied by the first part w'R of the spatial spectrum function again to obtain the second part w'Rw of the spatial spectrum function, and then it is right-shifted by 5 bits to obtain each set of spatial spectrum functions PBF and sent to the positioning module.

8. A hardware acceleration method based on beamforming for underwater detection according to claim 5, characterized in that The positioning module compares all sets of spatial spectrum functions PBF, selects the count value of the largest PBF as the positioning result, and sends it to the PS end through the RESULT_TO_PS module, including: The positioning module receives each set of spatial spectrum functions PBF and counts once using a counter at the same time, compares all sets of spatial spectrum functions PBF, selects the count value of the largest spatial spectrum function PBF as the positioning result, sends it from the angle interface to the RESULT_TO_PS IP core, and generates an enable signal end_valid and sends it from the Interupt interface to the RESULT_TO_PS IP core. The RESULT_TO_PS IP core, under the action of the enable signal, sends the positioning result to the PS end through the S_AXI_LITE interface.

Citation Information

Patent Citations

  • Side-scan sonar signal processing method based on ZYNQ

    CN110244304A

  • Winograd YOLOv2 target detection model method based on FPGA acceleration

    CN111459877A