Target signal and underwater acoustic channel decoupling method and system based on conditional generative adversarial network
By using a conditional generative adversarial network model, combined with underwater acoustic characteristics and environmental information, the target signal and underwater acoustic channel were effectively separated. This solved the problems of inaccurate signal decoupling and poor adaptability in existing technologies, and improved the target detection and recognition capabilities of unmanned platforms.
Patent Information
- Application Number
- CN202511286742.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-01-06
AI Technical Summary
Existing deep learning methods struggle to effectively separate target signals from underwater acoustic channels in unmanned underwater platforms. They neglect the influence of underwater acoustic characteristics and complex marine environments, resulting in poor decoupling performance and difficulty in training with large amounts of high-quality data.
A conditional generative adversarial network (GAN)-based approach is adopted. By introducing conditional information such as receiver array configuration, receiver array depth, sound source-receiver array distance, and seabed roughness, a conditional GAN model is constructed to achieve channel response estimation and decoupling of target signal.
It improves the accuracy and stability of target signal estimation, enhances adaptability in complex environments, and enables rapid signal and channel separation, making it suitable for real-time applications on underwater unmanned platforms.
Smart Images

Figure CN121283531A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of underwater unmanned platform acoustic field information perception, specifically involving a method and system for decoupling target signals and underwater acoustic channels based on Conditional Generative Adversarial Network (CGAN). Background Technology
[0002] When underwater unmanned platforms operate in marine environments, the radiated noise from targets propagates through the underwater acoustic channel, resulting in strong distortion, fluctuations, and time delays. This makes it difficult for the unmanned platform to identify targets, thus hindering target detection and tracking. The shallow sea channel is a complex, uneven, two-interface random inhomogeneous medium channel. Its complexity is mainly reflected in strong multipath and dispersion effects, high uncertainty at the sea surface and seabed, significant spatiotemporal variations, and high noise levels. These spatiotemporal variations cause time delays, fluctuations, and distortions in the acoustic signal, and the signal-to-noise ratio at the receiving array decreases due to reverberation and environmental noise. Therefore, in situations where the target is unknown and the marine environment is complex, only by accurately estimating the channel characteristics to eliminate channel influences can the unmanned platform effectively perform target detection and identification tasks.
[0003] Currently, deep learning technology is widely applied in the field of underwater acoustics both domestically and internationally, providing a new solution to the problem of separating underwater acoustic signals and channels. Through training with a large amount of raw data, the channel is indirectly identified based on the data distribution relationship between the received signal and the channel response, achieving the separation of the target and the channel. This has certain reference value for the detection and recognition of underwater acoustic signals and underwater acoustic communication. However, existing methods for decoupling target signals and underwater acoustic channels using deep learning technology have the following drawbacks: 1. Insufficient consideration of underwater acoustic characteristics: Traditional methods often rely solely on the received signal itself for target signal and channel decoupling, neglecting the decisive influence of the complex shallow sea environment on propagation characteristics. Many deep learning separation networks used for target signal and underwater acoustic channel decoupling are borrowed from speech separation applications, failing to fully consider the unique characteristics of underwater acoustics, such as the influence of different propagation media, signal frequency characteristics, and modulation characteristics, which to some extent affect the accuracy and effectiveness of decoupling. 2. Poor adaptability to complex marine environments: Marine environmental conditions are extremely complex and variable. Underwater acoustic signals will be severely distorted due to the complex marine channels before being received by sensors. Current deep learning methods are difficult to adapt well to this complex situation, resulting in the decoupling effect failing to meet actual needs. Problems such as feature mismatch and overfitting are prone to occur, causing the algorithm performance to decline. 3. Reliance on large amounts of high-quality data: Deep learning typically requires a large amount of high-quality data for training to achieve good performance. However, in the field of underwater acoustics, obtaining a large amount of accurately labeled target signals and underwater acoustic channel data covering various scenarios is very difficult. Insufficient or low-quality data will lead to inadequate model training, which will limit the accuracy and stability of decoupling. Summary of the Invention
[0004] The purpose of this invention is to use deep learning to accurately estimate channel characteristics in situations where the target is unknown and the marine environment is complex, thereby removing the influence of the channel, achieving target-channel separation, and helping unmanned platforms to effectively perform detection and identification tasks.
[0005] To achieve the above objectives, this application proposes a method for decoupling the target signal and the underwater acoustic channel based on a conditional generative adversarial network, comprising: The preprocessed received signal and conditional information are input into the trained neural network model, which outputs the estimated channel response. The estimated channel response and the received signal are deconvolved to obtain the source signal estimate, thereby decoupling the target signal and the underwater acoustic channel. The neural network model is a conditional generative adversarial network.
[0006] As an improvement to the above method, the condition information includes: receiver array configuration, receiver array depth, sound source-receiver array distance, and seabed roughness.
[0007] As an improvement to the above method, the preprocessing includes: By modifying the parameters of receiver array depth, source-to-array distance, and seabed roughness using the Bellhop model, the underwater acoustic channel response is generated and then convolved with the source signal to obtain the corresponding received signal.
[0008] As an improvement to the above method, it also includes: The Pearson correlation coefficient is used to evaluate the correlation between the estimated and true values of the source signal. The closer the Pearson correlation coefficient is to 1, the higher the similarity between the estimated and true values of the source signal.
[0009] This application also provides a target signal and underwater acoustic channel decoupling system based on a conditional generative adversarial network, implemented using the above method. The system includes: The preprocessing module is used to preprocess the received signals and condition information; The channel response estimation module is used to input the preprocessed received signal and condition information into the trained neural network model and output the estimated channel response. The signal decoupling module is used to deconvolve the estimated channel response and the received signal to obtain the source signal estimate, thereby decoupling the target signal and the underwater acoustic channel.
[0010] As an improvement to the above system, the system further includes: The quality assessment module is used to evaluate the correlation coefficient between the estimated and true values of the source signal using the Pearson correlation coefficient, and to assess the similarity between the estimated and true values of the source signal.
[0011] Compared with existing technologies, the advantages of this application are: This invention proposes a method and system for decoupling target signals and underwater acoustic channels based on conditional generative adversarial networks (GANs), achieving effective separation of target signals and channel responses in complex shallow sea environments. Compared with traditional blind deconvolution methods: it offers stronger channel decoupling capability, as this invention automatically learns the nonlinear relationship between target signals and channel effects through a data-driven approach, enabling more effective decoupling of signals and channels; it exhibits higher robustness, as the adversarial training mechanism between the generator and discriminator allows the generator to produce target signals with consistency and realism in various complex environments, while also supporting continuous model updates and optimization based on the actual deployment environment, significantly improving adaptability to environmental changes; and it offers better real-time performance, as this invention, based on a deep learning framework, enables rapid inference after training, making it suitable for real-time applications on underwater unmanned platforms. Attached Figure Description
[0012] Figure 1 The diagram shows a flowchart of a method for decoupling the target signal from the underwater acoustic channel based on a conditional generative adversarial network. Figure 2 The image shown is a visualization of the generated dataset. Figure 3 The image shows a conditional generative adversarial network model. Figure 4 The diagram shown is a schematic of the conditional generative adversarial network generator module. Figure 5 The diagram shown is a schematic of the conditional generative adversarial network discriminator module. Figure 6 The image shows a comparison between the target channel response and the predicted channel response. Figure 7(a) shows the result of source signal recovery of the received signal based on conditional generative adversarial network - the time domain waveform of the LFM signal; Figure 7(b) shows the result of source signal recovery of the received signal based on conditional generative adversarial network - LFM signal spectrum; Figure 7(c) shows the result of source signal recovery of the received signal based on conditional generative adversarial network - time-domain estimation of LFM signal; where element1, element16 and element32 represent the 1st, 16th and 32nd array elements of the receiving array, respectively; Figure 7(d) shows the result of source signal recovery of the received signal based on conditional generative adversarial network - LFM signal spectrum estimation; where element1, element16 and element32 represent the 1st, 16th and 32nd array elements of the receiving array, respectively. Detailed Implementation
[0013] The technical solution of this application will be described in detail below with reference to the accompanying drawings.
[0014] like Figure 1 As shown, this invention proposes a method for decoupling target signals and underwater acoustic channels based on conditional generative adversarial networks. This method is applicable to the effective separation of target signals and channel responses in shallow sea environments (sea areas with a depth of less than 200m are considered shallow seas). The method includes the following steps: Step 1: Use the Bellhop model to change the corresponding parameters to generate the underwater acoustic channel response and the corresponding received signal, and divide the training set and test set.
[0015] Bellhop is a model for calculating the acoustic pressure field in marine environments. By configuring the underwater acoustic channel environmental parameters and utilizing the ray arrival structure information from Bellhop, the channel response function can be obtained under different acoustic fields and channel conditions. Parameters that can be changed using the Bellhop model include: receiver array depth, source-to-receiver array distance, and root mean square roughness of the seabed. After generating the underwater acoustic channel response, it is convolved with the source signal to obtain the corresponding received signal, which is recorded as a set of data. This process is repeated to obtain a sufficient dataset, which is then divided into training and test sets at a ratio of 70% and 30%, respectively.
[0016] Step 2: Build a Conditional Generative Adversarial Network (CGAN) model as follows Figure 2 It contains a generator G and a discriminator D.
[0017] The key to CGAN is "generation" and "adversarial" approaches. Let's assume we have a set of data... according to The distribution exists, and the generator responsible for "generating" must generate [something] that is consistent with [something else]. The data is extremely similar, and the discriminator, which is responsible for "adversarial" analysis, has to identify as much as possible which data is generated by the generator and which is the real original data. The two parts compete against each other and eventually reach a balance point. When the discriminator has a 50% probability of identifying the data, the generator has achieved its goal of being indistinguishable from the real data and has obtained a generative model that is very similar to the original data distribution.
[0018] Step 3: Set network parameters, train the network using the training set data, and alternately optimize the discriminator D and the generator G.
[0019] The received signal information and condition information need to be concatenated before being input into the generator. The generator is responsible for determining the condition information based on the received signal information. and received signals Estimated value of the Green's function for generating the channel Then add the condition information. With generating channel data and real channel data The data is concatenated separately, and then the concatenated data is input into the discriminator. The discriminator determines whether the input data was generated by the generator or is real data, and then outputs the discriminant loss as feedback to the generator, which then updates its parameters. Once it approximates... If the generator achieves a "deceptive" effect, channel response estimation can be performed based on the trained network. In each outer training loop, the discriminator is updated multiple times in the inner loop, improving its ability to distinguish between real and fake data by calculating the discriminant loss between real and generated data and backpropagating it. Subsequently, the generator is updated multiple times in the inner loop, aiming to generate more realistic data to deceive the discriminator, and optimizing itself through backpropagation of the resulting loss. The entire process continues in an adversarial manner, with the ultimate goal of making the data distribution output by the generator approximate the real data distribution, and outputting the trained generator and discriminator models along with their parameters.
[0020] Unlike existing methods based on Generative Adversarial Networks (GANs) or other blind deconvolution methods, this invention introduces a conditional information vector at the network input stage, including environmental parameters such as receiver array configuration, receiver array depth, sound source-receiver array distance, and seabed roughness, which has the following advantages: 1. Explicitly embedding environmental constraints improves the accuracy of target signal estimation; Traditional methods often rely solely on the received signal itself to decouple the target signal from the channel, neglecting the decisive influence of the complex shallow sea environment on propagation characteristics. This invention, by concatenating conditional vectors at the network input layer, enables the generator to simultaneously capture the coupling relationship between the signal and the environment during the learning process. It learns the mapping relationship of "target signal—environment—received signal," which more effectively removes channel effects and recovers the target signal, thereby improving the accuracy and stability of source signal estimation. 2. Stronger adaptability to different environments; In actual underwater operations, the acoustic environment fluctuates dramatically with changes in sea state, seabed topography, and noise levels. This invention guides the generator to learn the signal-channel mapping relationship under different environmental conditions through conditional information, enabling faster adaptation to new working scenarios during model transfer or retraining, and significantly improving the adaptability of unmanned platforms in dynamic marine environments.
[0021] Step 4: On the test set, use the generator trained by CGAN to estimate the channel response based on the received signal and conditional information. Deconvolve the predicted channel response and the received signal to obtain the source signal estimate, thereby achieving separation of the target and the channel. Compare the obtained source signal estimate with the theoretical value to verify the effect of the CGAN model.
[0022] The Pearson correlation coefficient can be used to assess the correlation between the estimated and true values of the source signal. The closer the value is to 1, the higher the similarity between the estimated signal and the real signal.
[0023] Example 1 This invention provides a method for decoupling target signals and underwater acoustic channels based on conditional generative adversarial networks. The specific steps and details are as follows: Dataset Generation: Data was generated using the Bellhop model by changing the receiver array configuration, receiver array depth, source-to-receiver distance, and seabed roughness to create a dataset containing the received signal and corresponding channel response. The relevant experimental parameters are shown in Table 1. The receiver array configuration, receiver array depth, source-to-receiver distance, and seabed roughness were input as conditional vectors into the model.
[0024] Table 1 Experimental parameter settings
[0025] Receive signal Channel response matrix These can be viewed as two-channel images with dimensions M × N × 2 and M × N × 2, respectively. M represents the number of array elements, N represents the number of frequency points, and the two channels of the image represent the real and imaginary parts of the complex matrix, respectively. Then, the channel Green's function estimation problem can be viewed as a transformation problem from image to image, requiring the transformation of a low-resolution image... Quantize the image to convert it into a channel response image. Figure 2 It is receiving data and channel response data The transformed images represent the real parts of the quantized received signal and the channel response, respectively, corresponding to amplitude information. The dataset is divided into training and test sets at a ratio of 70% and 30%, respectively. To train the proposed CGAN model, the learning rates for the generator and discriminator are set to 0.0002 and 0.00002, respectively.
[0026] Setting up a network: Figure 3 It is a conditional generative adversarial network model. Figure 4 and Figure 5 This is a schematic diagram of the generator and discriminator modules of a conditional generative adversarial network.
[0027] Generator Networks: In the generator of the conditional generative adversarial network (GAN) model, the received signal vector and the conditional vector are concatenated and input into the network layer. The size of the received signal vector is set to 64×32, and the size of the conditional vector is 1×32. The first four terms correspond to the array configuration, depth, source-array distance, and seabed roughness, respectively. The latter terms are zero-padding to fill the conditional vector to a 1×32 vector. The size of the concatenated vector of the received signal and conditional vector is 65×32. During the generation process, the hidden layers are a weight layer, a ReLU activation function layer, a weight layer, and a Sigmoid activation function layer. Finally, the generated data is output as 65×32.
[0028] Discriminator Network: After obtaining a 64×32 generated data vector from the generator, it is concatenated with a 1×32 conditional vector to obtain a 65×32 input data vector, which is then fed into the discriminator. The discriminator also has a four-layer hidden layer network structure, and the final result is the discriminator's discrimination matrix. Through adversarial training, the generator's parameter updates are not directly based on the data, but rather on the discriminator's backpropagation. The discriminator attempts to learn the real distribution from the real channel without requiring a complex loss function design.
[0029] Training the network: Using CGAN to train the model from highly quantized observation data and environmental conditions The goal is to recover the channel matrix from factors such as receiver array configuration, array depth, source-receiver distance, and seabed roughness. The ultimate objective is for the generator to synthesize the most realistic channel matrix to deceive the discriminator, while simultaneously, the discriminator needs to learn to be less easily fooled. The generator and discriminator engage in a continuous game-like process until an optimal CGAN is reached. The GAN loss is expressed as:
[0030] in, Indicates by A parameterized generator that generates a matrix similar to the real channel matrix. Channel matrices that are as similar as possible, i.e. , It is by A parameterized discriminator is designed to analyze the generated channels. With real channels Let's distinguish between them. From the generator's perspective, the generator wants to generate an estimate that can fool the discriminator as much as possible, i.e., it needs to be as large as possible; from the discriminator's perspective, the discriminator wants to distinguish between real data and generated data as much as possible, so... It should be as small as possible. To make it as large as possible, the objective function can be modified as follows: At the same time, to ensure the correctness of generator parameter optimization, add... The loss is Finally, the objective function of CGAN can be expressed as:
[0031] Based on simulation generation =120 sets of channel response data and received signals, for processing =300 training iterations, with two nested loops during the training iterations, one for iterating over the discriminator and the other for further iterations. =12 iterations of the generator =12 times. After training, the generator and discriminator models are obtained.
[0032] Table 2 Pseudocode for the training process of Conditional Generative Adversarial Networks
[0033] Estimating source and target signals: After training the model on the dataset, estimate the source signals on the test set. Figure 6 This is a pseudo-color image showing the estimated and actual Green's function values output by the network. The two channels correspond to the real and imaginary parts, respectively, and the estimated and actual values show good consistency in amplitude and phase. Using the predicted channel response value and the received signal corresponding to the input network model, the estimated value of the source signal is calculated and compared with the theoretical value of the source signal. The first, last, and middle array elements of the vertical receiving array are selected for processing. Figures 7(a)-7(d) show the results of the source signal recovery.
[0034] The correlation coefficient between the estimated source signal value obtained from the channel Green's function predicted by the network and the theoretical source signal value is calculated. The correlation coefficients at the beginning, end and middle array elements of the receiving array and the average correlation coefficient of each array element are shown in Table 3.
[0035] Table 3. Correlation coefficients between estimated and actual source signals based on simulation data.
[0036] By combining the results of channel response estimation and source signal recovery, the trained conditional generative adversarial network can effectively achieve blind estimation of the channel response, and decouple the target signal and the underwater acoustic channel by deconvolving the received signal and the estimated channel response.
[0037] Example 2 This application also provides a decoupling system for target signals and underwater acoustic channels based on conditional generative adversarial networks, implemented using the above method. The system includes: The preprocessing module is used to preprocess the received signals and condition information.
[0038] The channel response estimation module is used to input the preprocessed received signal and condition information into the trained neural network model and output the estimated channel response.
[0039] The signal decoupling module is used to deconvolve the estimated channel response and the received signal to obtain the source signal estimate, thereby decoupling the target signal and the underwater acoustic channel.
[0040] The quality assessment module is used to evaluate the correlation coefficient between the estimated and true values of the source signal using the Pearson correlation coefficient, and to assess the similarity between the estimated and true values of the source signal.
[0041] This application may also provide a computer device, including: at least one processor, memory, at least one network interface, and a user interface. The various components in this device are coupled together via a bus system. It is understood that the bus system is used to implement communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0042] The user interface can include a display, keyboard, or clicking device. Examples include a mouse, trackball, touchpad, or touchscreen.
[0043] It is understood that the memory in the embodiments disclosed in this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0044] In some implementations, the memory stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof: operating systems and applications.
[0045] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application programs include various applications, such as media players and browsers, used to implement various application functions. Programs implementing the methods of the embodiments of this disclosure can be included in the application programs.
[0046] In the above embodiments, the processor can also invoke programs or instructions stored in memory, specifically programs or instructions stored in an application program, for the following purposes: Follow the steps described above.
[0047] The above methods can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic diagrams disclosed above. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the disclosed methods can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0048] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof.
[0049] For software implementation, the technology of this application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of this application. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0050] This application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, it can implement the steps in the above method embodiments.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application, and should all be covered within the scope of the claims of this application.
Claims
1. A method for decoupling target signal and underwater acoustic channel based on conditional generative adversarial network, comprising: inputting preprocessed received signal and condition information into trained neural network model to output estimated channel response; deconvolving estimated channel response and received signal to obtain source signal estimate value, and realizing decoupling of target signal and underwater acoustic channel; the neural network model is a conditional generative adversarial network.
2. The method of claim 1, wherein, the condition information includes: receiving array array pattern, receiving array depth, sound source-receiving array distance and seabed roughness.
3. The method of claim 1, wherein, the preprocessing includes: using Bellhop model to modify receiving array depth, sound source-receiving array distance and seabed roughness parameters, generate underwater acoustic channel response, and convolve with source signal to obtain corresponding received signal.
4. The method of claim 1, wherein, further comprising: using Pearson correlation coefficient to evaluate the correlation coefficient between source signal estimate value and true value, and the closer the obtained Pearson correlation coefficient is to 1, the higher the similarity between source signal estimate value and true value is.
5. A target signal and underwater acoustic channel decoupling system based on conditional generative adversarial network, implemented based on the method of any one of claims 1-4, characterized in that, the system comprises: a preprocessing module for preprocessing received signal and condition information; a channel response estimation module for inputting preprocessed received signal and condition information into trained neural network model to output estimated channel response; and a signal decoupling module for deconvolving estimated channel response and received signal to obtain source signal estimate value, and realizing decoupling of target signal and underwater acoustic channel. 6.The target signal and underwater acoustic channel decoupling system based on conditional generative adversarial network according to claim 5, wherein, the system further comprises: a quality evaluation module for using Pearson correlation coefficient to evaluate the correlation coefficient between source signal estimate value and true value, and evaluating the similarity between source signal estimate value and true value.