A weak line spectrum enhancement method based on an autoencoder-transformer network structure
Patent Information
- Application Number
- CN202610927031.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
AI Technical Summary
然而,当输入信噪比低于-20dB时,ALE算法的增益转为负值,无法从强背景噪声中有效提取线谱特征,此外,传统ALE方法对延迟参数和步长因子敏感,在复杂海洋环境中的适应性较差
[0010]进一步的,在所述最优频带处理范围内对增强信号进行精细化增强处理,包括:
Smart Images

Figure CN122779142A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater acoustic signal processing technology, specifically to a weak line spectrum enhancement method based on an autoencoder-transformer network structure. Background Technology
[0002] Underwater target detection and identification is a core task in the field of underwater acoustic engineering. Its performance largely depends on the effective detection and enhancement of the line spectrum components in the target's radiated noise. However, with the development of underwater vehicle noise reduction technology, the target's radiated noise level is constantly decreasing, and the detection environment is becoming increasingly complex. Traditional signal processing methods face severe challenges under extremely low signal-to-noise ratio conditions.
[0003] In existing technologies, adaptive line enhancement (ALE) is a commonly used method for processing line spectrum features. This method utilizes the difference in correlation characteristics between line spectrum components and broadband noise to achieve spectral line enhancement through adaptive filtering. However, when the input signal-to-noise ratio is below -20dB, the gain of the ALE algorithm becomes negative, making it unable to effectively extract line spectrum features from strong background noise. In addition, traditional ALE methods are sensitive to delay parameters and step size factors, resulting in poor adaptability in complex marine environments.
[0004] In recent years, deep learning technology has been introduced into the field of underwater acoustic signal processing. Autoencoders (AEs) achieve feature extraction through an encoder-decoder structure, but their ability to model time-series signals is limited. Recurrent neural networks (RNNs) and their variants, while capable of processing sequential data, suffer from the vanishing gradient problem, making it difficult to capture long-range dependencies. Furthermore, most existing deep learning algorithms employ time-frequency domain enhancement. While this approach has certain advantages, it also has several significant drawbacks. Firstly, time-frequency domain enhancement inevitably leads to a substantial increase in network complexity, requiring more computational resources and time, severely impacting the algorithm's efficiency. Secondly, this approach also reduces resolution, resulting in less clear and accurate signal features, failing to meet the requirements of high-precision detection, and thus limiting the practical application and value of deep learning algorithms in passive sonar systems. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a weak line spectrum enhancement method based on an autoencoder-transformer network structure. This method acquires and normalizes the sound pressure signal, then obtains a delayed signal by delaying the sampling points by a predetermined number. The sound pressure signal and the delayed signal are then framed to obtain sound pressure frame signals and delayed frame signals. These two signals are used as input data to a pre-trained autoencoder-transformer network, which outputs enhanced signal frames. The enhanced signal frames are then merged to obtain the enhanced signal. Adaptive frequency band selection processing is applied to the enhanced signal to determine its optimal frequency band processing range. Within this optimal range, the enhanced signal undergoes refined enhancement processing to obtain the final enhanced sound pressure signal. This invention, by introducing a collaborative processing mechanism between the autoencoder-transformer deep network and adaptive frequency band selection, enables multi-level enhancement of weak line spectra.
[0006] This invention employs the following technical solution: a weak line spectrum enhancement method based on an autoencoder-transformer network structure, comprising: Acquire the sound pressure signal and perform normalization processing; The normalized sound pressure signal is delayed by a set number of sampling points to obtain the delayed signal; The sound pressure signal and the delay signal are subjected to frame segmentation processing to obtain the corresponding sound pressure frame signal and delay frame signal; The sound pressure framing signal and the delay framing signal are used as input data and input into the pre-trained autoencoder-transformer network. The autoencoder-transformer network includes an autoencoder encoding module, a transformer module, and an autoencoder decoding module; The autoencoder module performs feature compression on the input data to obtain low-dimensional features corresponding to the input data. The transformer module uses a self-attention mechanism to fuse the low-dimensional features corresponding to the input data to obtain the enhanced features corresponding to the input data. The autoencoder decoding module reconstructs the enhanced features corresponding to the input data and outputs the enhanced signal frame. The enhanced signal frames are merged to obtain the complete enhanced signal; The enhanced signal is subjected to adaptive frequency band selection processing to obtain the energy distribution of the enhanced signal in different frequency bands; The main frequency band of the enhanced signal is determined based on the energy distribution, and the optimal frequency band processing range of the enhanced signal is determined based on the main frequency band. The enhanced signal is then subjected to refined enhancement processing within the optimal frequency band processing range to obtain the final enhanced sound pressure signal.
[0007] Furthermore, the autoencoder encoding module includes a fully connected layer for compressing the input data to obtain low-dimensional features; The transformer module includes a self-attention layer and a feedforward network. The self-attention layer is used to calculate the correlation weights between different time frames in the low-dimensional features and generate a low-dimensional feature sequence by weighted summation. The feedforward network is used to perform element-wise nonlinear transformation and feature mapping on the low-dimensional feature sequence to obtain enhanced features and output the enhanced signal frame. The autoencoder decoding module includes a fully connected layer for reconstructing the enhanced features.
[0008] Furthermore, the enhanced signal undergoes adaptive frequency band selection processing, specifically as follows: Perform a Fast Fourier Transform on the enhanced signal to obtain the spectrum of the enhanced signal; The spectrum is divided into multiple equally spaced frequency bands as candidate frequency bands; Calculate the energy value of each candidate frequency band to obtain the energy distribution of the enhanced signal in different frequency bands.
[0009] Furthermore, determining the optimal frequency band processing range for the enhanced signal based on the main frequency band includes: The candidate frequency band with the highest energy value is selected as the main frequency band; The energy concentration of the main frequency band is calculated based on the ratio of its energy value to the average energy of its two adjacent frequency bands. The process expands from the main frequency band and calculates the energy concentration of adjacent frequency bands in turn. The frequency bands whose energy concentration meets the set threshold constitute the optimal frequency band processing range.
[0010] Furthermore, within the optimal frequency band processing range, the enhanced signal undergoes refined enhancement processing, including: The enhanced signal is filtered using a bandpass filter, and the enhanced signal within the optimal frequency band processing range after filtering is retained; The enhanced signal within the optimal frequency band after filtering is subjected to spectral sharpening and energy normalization to obtain the final enhanced sound pressure signal.
[0011] The beneficial effects of this invention are as follows: By introducing a collaborative processing mechanism of autoencoder-transformer deep network and adaptive frequency band selection, this invention achieves multi-level enhancement of weak line spectra. First, through joint framing processing of sound pressure signals and delay signals, rich temporal features are provided to the network. Second, the autoencoder-transformer network can still achieve high line spectrum gain under extremely low signal-to-noise ratio conditions through encoder feature compression, transducer temporal dependency modeling, and decoder signal reconstruction. Finally, adaptive frequency band selection dynamically determines the optimal processing frequency band based on energy concentration threshold, and further improves the contrast and reliability of the line spectrum through refined enhancement processing. This significantly improves the detection performance and anti-false alarm capability of traditional spectral line enhancement technology in low signal-to-noise ratio, multi-spectral-line scenarios, providing effective technical support for underwater weak target detection. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of a weak line spectrum enhancement method based on an autoencoder-transformer network structure according to an embodiment of the present invention; Figure 2 This is a schematic diagram of an autoencoder-converter network structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the time-frequency analysis results of a single-line spectrum signal under an input signal-to-noise ratio of -20dB, according to an embodiment of the present invention. Figure 4 This is a schematic diagram showing the comparison results of the power spectral density of a single-line spectrum signal under an input signal-to-noise ratio of -20dB according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the time-frequency analysis results of a single-line spectrum signal with an input signal-to-noise ratio of -30dB according to an embodiment of the present invention; Figure 6 This is a schematic diagram showing the comparison results of the power spectral density of a single-line spectrum signal when the input signal-to-noise ratio is -30dB, according to an embodiment of the present invention. Figure 7 This is a schematic diagram of the time-frequency analysis results of a multi-line spectrum signal under an input signal-to-noise ratio of -20dB, according to an embodiment of the present invention. Figure 8 This is a schematic diagram of the time-frequency analysis results of a multi-line spectrum signal with an input signal-to-noise ratio of -30dB, according to an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] A schematic flowchart of a weak line spectrum enhancement method based on an autoencoder-transformer network structure according to an embodiment of the present invention is shown below. Figure 1 As shown, it includes: Acquire the sound pressure signal and perform normalization processing; In this embodiment of the invention, the normalization process can adopt the maximum absolute value normalization method, which is as follows: for the collected sound pressure signal, the maximum absolute value of the sound pressure signal within a set time period is obtained. The set time period can be 5 seconds or can be set according to the actual situation. After obtaining the maximum absolute value, the normalized sound pressure signal is obtained by the ratio of the sound pressure signal at each moment within the time period to the maximum absolute value of the sound pressure signal.
[0016] The normalized sound pressure signal is delayed by a set number of sampling points to obtain the delayed signal; In this embodiment of the invention, the delay parameter is first set. Optional setting methods include fixed sampling point delay and fixed time delay. The fixed sampling point delay can be set to 200 sampling points, while the fixed time delay can be set to a sampling delay time of 1 second. The specific setting method can be selected according to the actual application scenario. The fixed sampling point delay is suitable for scenarios that require fast response, while the fixed time delay is suitable for scenarios that require better noise decorrelation effect.
[0017] Taking a fixed number of sampling points delay as an example, after setting the number of sampling points with delay, for each sampling point of the normalized sound pressure signal, a general-purpose processor reads the signal after a delay of 200 sampling points as the output, thereby ensuring that each output sample is the sound pressure signal after 200 sampling points, thus obtaining the corresponding delayed signal.
[0018] The sound pressure signal and the delay signal are subjected to frame segmentation processing to obtain the corresponding sound pressure frame signal and delay frame signal; Since the noise in the signal due to ship radiation is correlated at various locations, this embodiment of the invention performs frame segmentation on the signal, that is, converts the signal into a sequence suitable for an autoencoder-transformer network by dividing the sequence. The specific frame segmentation method is as follows: First, the parameters for frame segmentation are set. In this embodiment of the invention, a frame length of 256 sampling points is selected for framing, and the frame shift is set to 128 sampling points to ensure the continuity of the signal. The window function is selected as a Hamming window or a rectangular window. When performing frame segmentation on the sound pressure signal, starting from the beginning position of the sound pressure signal, the sound pressure signal segments are sequentially extracted with the set frame shift as the step size. Then, for each sound pressure signal segment, a window function is used for windowing processing, and the windowed signal frame is stored as a two-dimensional array. In the two-dimensional array, the rows represent the frame number, and the columns represent the sampling points within the frame, thereby obtaining the sound pressure framed signal.
[0019] The method for framing the delayed signal is exactly the same as that for the sound pressure signal, and the parameters for framing are also the same. By maintaining the same frame length, frame shift, and window function, the sound pressure framing signal and the delayed framing signal are strictly aligned in time.
[0020] After completing the framing process, two sets of data are output: sound pressure framing signal and delay framing signal. These are used as input data for the subsequent autoencoder-transformer network, thus providing the network with signal segments with temporal locality characteristics. The entire framing process ensures the effective preservation of the signal's time-frequency characteristics, laying a good foundation for subsequent deep feature extraction.
[0021] The sound pressure frame signal and the delayed frame signal are used as input data and input into the pre-trained autoencoder-transformer network. In this embodiment of the invention, a schematic diagram of the structure of the self-encoder-transformer network (AET-DLE) is shown below. Figure 2 As shown, it includes an autoencoder encoding module, a converter module, and an autoencoder decoding module, specifically: In this embodiment of the invention, the autoencoder encoding module adopts an AE encoding layer structure. The AE encoding layer structure is usually composed of several layers of neurons, with the number of neurons in each layer gradually decreasing, thereby realizing a mapping from high dimension to low dimension. During training, the AE encoding layer learns how to transform complex original signals into concise and key information-rich low-dimensional feature representations by continuously adjusting the connection weights between neurons based on the input frame signals. Since the input data contains a large number of data points and has a high dimension, the AE encoding layer can learn a mapping relationship, thereby transforming these complex original signals into a compact low-dimensional feature. This low-dimensional feature can capture the key information in the input data, just like compressing and refining the input data.
[0022] In one specific embodiment of the present invention, the autoencoder encoding module adopts a fully connected structure composed of two layers of neural networks. The input sample of this module is defined as a joint feature vector composed of the framed sound pressure signal and the delayed frame signal after frame processing. Taking 512 dimensions as an example, each input data corresponds to the concatenation result of the 256-sample sound pressure signal and the 256-sample delayed frame signal. The input data is used as the input of the autoencoder encoding module. The first layer is the input layer, which is used to receive the input data. Then there is a first linear layer, which is used to map the 512-dimensional input data to a 256-dimensional feature space, and then perform nonlinear transformation through the ReLU activation function. The first intermediate layer is set at ReLU. After the U activation function, preliminary abstract features are extracted. Then, the second linear layer further maps the 256-dimensional features to a 128-dimensional space, and the ReLU activation function is used again for nonlinear processing. The second intermediate layer is set after the second ReLU activation function for deep feature extraction, and finally outputs 64-dimensional low-dimensional features. In this process, the autoencoder module adjusts the weight and bias parameters according to a predefined loss function, so that the encoded features can retain the key information in the input data as well as possible. This allows the subsequent transformer module and the final spectral enhancement to focus more on this key information, improving the efficiency and performance of the entire model.
[0023] In this embodiment of the invention, the transformer module includes a self-attention layer and a feedforward network; wherein, the self-attention layer is used to calculate the correlation weights between different time frames in the low-dimensional features, and generate a low-dimensional feature sequence by weighted summation; the feedforward network is used to perform element-wise nonlinear transformation and feature mapping on the low-dimensional feature sequence to obtain enhanced features.
[0024] In one specific embodiment of the present invention, the self-attention layer first passes the input 64-dimensional low-dimensional features through three different linear transformation layers to generate a query matrix, a key matrix, and a value matrix, respectively. Then, a scaling dot product attention mechanism is used to calculate the self-attention weights. At the same time, the 64-dimensional low-dimensional features are divided into 8 attention heads, each with a dimension of 8, which are responsible for extracting features from independent subspaces. The features are then concatenated after parallel computation, which significantly improves the model's representation ability.
[0025] The feedforward network consists of two linear transformation layers. Its function is to perform spatial transformation and introduce nonlinear relationships to improve the model's ability to fit data. In this embodiment of the invention, the nonlinear relationship here adopts the GeLU activation function. By mapping the 64-dimensional input to a 256-dimensional feature space, and then activating it through the GeLU activation function, the feature dimensions are remapped back to 64 dimensions, thereby outputting enhanced features.
[0026] The autoencoder decoding module includes a fully connected layer for reconstructing the enhanced features and outputting an enhanced signal frame.
[0027] In this embodiment of the invention, the autoencoder decoding module is the reverse process of the autoencoder encoding module. Its main structure consists of two neural network layers, including two linear layers and a ReLU activation function. During reconstruction, the first linear layer receives the output of the transformer module and performs a linear transformation through the weight matrix and bias vector between neurons. Subsequently, a nonlinear factor, the ReLU activation function, is introduced, enabling the model to learn more complex feature representations and gradually reconstruct reconstructed data that is similar to the original input data, thereby outputting an enhanced signal frame. This structure allows the autoencoder decoding module to effectively recover the key information of the original signal from low-dimensional features, ensuring, to a certain extent, the similarity between the reconstructed sample and the original sample in important features. Moreover, through collaborative training with the AE encoding layer, the decoding layer continuously adjusts its weight parameters to adapt to different types of input signals, improving the accuracy and stability of reconstruction.
[0028] The enhanced signal frames are merged to obtain the complete enhanced signal; adaptive frequency band selection processing is performed on the enhanced signal to obtain the energy distribution of the enhanced signal in different frequency bands; In this embodiment of the invention, the enhanced signal frames are re-merged into a complete enhanced signal using the overlap-addition method. Then, the enhanced signal is subjected to spectral analysis, and the power spectral density of the enhanced signal is calculated using Fast Fourier Transform. In order to reduce spectral leakage, the enhanced signal is processed by Hanning window before Fast Fourier Transform to obtain the power spectrum of the enhanced signal.
[0029] After obtaining the power spectrum of the enhanced signal, it is divided into a set number of equally spaced frequency bands. The specific number can be set according to the effective frequency range of the actual signal. Then, the energy value of each frequency band is calculated based on the power spectral density value in each frequency band. The specific calculation method can refer to any of the existing calculation methods, thereby obtaining the energy distribution of the enhanced signal in different frequency bands.
[0030] The main frequency band of the enhanced signal is determined based on the energy distribution, and the optimal frequency band processing range of the enhanced signal is determined based on the main frequency band. In this embodiment of the invention, the energy values of all frequency bands are compared, and the frequency band with the largest energy value is taken as the main frequency band. Then, the energy concentration of the main frequency band is calculated based on the ratio of the energy value of the main frequency band to the average energy of its two adjacent frequency bands. The energy concentration of the adjacent frequency bands is calculated sequentially with the main frequency band as the center. The frequency bands whose energy concentration meets the set threshold constitute the optimal frequency band processing range.
[0031] The enhanced signal is then refined within the optimal frequency band processing range to obtain the final enhanced sound pressure signal.
[0032] In this embodiment of the invention, a bandpass filter is used to filter the enhanced signal and retain the enhanced signal within the optimal frequency band processing range after filtering; the enhanced signal within the optimal frequency band processing range after filtering is subjected to spectral sharpening and energy normalization to obtain the final enhanced sound pressure signal.
[0033] In a specific embodiment of the present invention, it is assumed that the sound pressure signal received by the single-vector hydrophone has a frequency of 30Hz, an amplitude of 1, a sampling frequency of 2kHz, Gaussian white noise, a signal length of 6s, and a delay time of 1s. The signal-to-noise ratios are set to -20dB and -30dB, respectively. The time-frequency analysis results of the single-line spectrum signal under the signal-to-noise ratio of -20dB are as follows: Figure 3 As shown, where, Figure 3 Part (a) shows the time spectrum of the noisy input signal. At this signal-to-noise ratio, the frequency at the target spectral line is no longer visible. Figure 3 Parts (b), (c), and (d) respectively present the time-spectrum of the output signals of the traditional ALE, AE-DLE, and the AET-DLE of this invention. It can be seen that while the traditional ALE method can extract the target spectral line, the presence of interfering spectral lines limits its performance. Although the AE-DLE method is an improvement, making the target spectral line more obvious, noise interference still exists. In contrast, the AET-DLE method proposed in this invention performs exceptionally well in extracting the target spectral line; not only are the frequency components of the target spectral line clearly visible, but noise interference is also significantly suppressed. Furthermore, Figure 4 The power spectral density comparison further confirms the superiority of the proposed AET-DLE method in terms of signal-to-noise ratio gain. Its power spectral density at the target spectral line is significantly higher than that of other methods. These results show that the AET-DLE method has significant advantages in signal extraction and noise suppression in low signal-to-noise ratio environments, demonstrating its potential in complex signal processing tasks.
[0034] Figure 5 This is a schematic diagram showing the time-frequency analysis results of a single-line spectrum signal with an input signal-to-noise ratio of -30dB. Figure 5 Part (a) is the time-spectrum diagram of the signal before training. Figure 5Parts (b), (c), and (d) are the time-frequency diagrams of the output signals of the traditional ALE, AE-DLE, and the AET-DLE of this invention, respectively. It can be seen that the traditional ALE method performs poorly in this situation, with the enhanced signal almost completely ineffective, making it difficult to identify the target spectral line. While the AE-DLE method can identify the target spectral line to some extent, its effect remains blurry and is significantly affected by noise. In contrast, the AET-DLE proposed in this invention exhibits superior performance, not only clearly identifying the target spectral line but also significantly improving the output signal-to-noise ratio after filtering. Furthermore, Figure 6 The paper further presents a schematic diagram of the power spectral density comparison results, which further confirms the superiority of AET-DLE in terms of signal-to-noise ratio gain. Its power spectral density at the target spectral line is significantly higher than that of the AE-DLE method. Specifically, the AE-DLE method achieves an output gain of approximately 3.4 dB, while AET-DLE achieves a gain of approximately 13.4 dB. This indicates that AET-DLE has better signal extraction and noise suppression capabilities in low signal-to-noise ratio environments, which is significantly better than the traditional ALE algorithm and the AE-DLE method.
[0035] In another specific embodiment of the present invention, to verify the performance of the algorithm when multiple spectral lines are input, a case with 5 spectral lines is set as the input, with amplitudes of 1 and frequencies of 30Hz, 50Hz, 80Hz, 120Hz, and 170Hz respectively. The input signal-to-noise ratio is set to -20dB and -30dB. Figure 7 The figure shows the multi-line spectrum time-frequency analysis results when the input signal-to-noise ratio is -20dB. Figure 7 Part (a) in the figure is the time spectrum of the noisy input signal. Figure 7 Part (b) in the diagram represents the time spectrum of the signal after traditional ALE training. Figure 7 Part (c) in the diagram is the time spectrum of the signal after AE-DLE training. Figure 7 Part (d) in the diagram represents the time spectrum of the signal after AET-DLE training, derived from... Figure 7 The results show that the ALE algorithm is basically ineffective when strong noise is added. Although strong spectral lines can be seen to be enhanced in the power spectral density plot, many erroneous spectral lines are added. In contrast, AE-DLE can enhance strong spectral lines but cannot enhance weak spectral lines. The AET-DLE algorithm proposed in this invention effectively enhances all spectral line components.
[0036] Figure 8 Furthermore, the multi-line spectrum time-frequency analysis results are given when the input signal-to-noise ratio is -30dB, where, Figure 8 Part (a) in the figure represents the time spectrum of the noisy input signal. Figure 8 Part (b) in the diagram represents the time spectrum of the signal after traditional ALE training. Figure 8Part (c) in the diagram is the time spectrum of the signal after AE-DLE training. Figure 8 Part (d) in the figure represents the time spectrum of the signal after AET-DLE training. It can be seen that the ALE algorithm is completely ineffective with the addition of -30dB noise. AE-DLE can still faintly see the 30Hz and 170Hz spectral lines, but they are basically submerged in noise. In contrast, the AET-DLE algorithm clearly shows that all five spectral lines are enhanced, demonstrating a good enhancement effect.
[0037] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for enhancing weak line spectra based on an autoencoder-transformer network structure, characterized in that, include: Acquire the sound pressure signal and perform normalization processing; The normalized sound pressure signal is delayed by a set number of sampling points to obtain the delayed signal; The sound pressure signal and the delay signal are subjected to frame segmentation processing to obtain the corresponding sound pressure frame signal and delay frame signal; The sound pressure framing signal and the delay framing signal are used as input data and input into the pre-trained autoencoder-transformer network. The autoencoder-transformer network includes an autoencoder encoding module, a transformer module, and an autoencoder decoding module; The autoencoder module performs feature compression on the input data to obtain low-dimensional features corresponding to the input data. The transformer module uses a self-attention mechanism to fuse the low-dimensional features corresponding to the input data to obtain the enhanced features corresponding to the input data. The autoencoder decoding module reconstructs the enhanced features corresponding to the input data and outputs the enhanced signal frame. The enhanced signal frames are merged to obtain the complete enhanced signal; The enhanced signal is subjected to adaptive frequency band selection processing to obtain the energy distribution of the enhanced signal in different frequency bands; The main frequency band of the enhanced signal is determined based on the energy distribution, and the optimal frequency band processing range of the enhanced signal is determined based on the main frequency band. The enhanced signal is then subjected to refined enhancement processing within the optimal frequency band processing range to obtain the final enhanced sound pressure signal.
2. The weak line spectrum enhancement method based on an autoencoder-transformer network structure according to claim 1, characterized in that: The autoencoder-transformer network includes an autoencoder encoding module, a transformer module, and an autoencoder decoding module, specifically: The autoencoder encoding module includes a fully connected layer for compressing the input data to obtain low-dimensional features; The transformer module includes a self-attention layer and a feedforward network. The self-attention layer is used to calculate the correlation weights between different time frames in the low-dimensional features and generate a low-dimensional feature sequence by weighted summation. The feedforward network is used to perform element-wise nonlinear transformation and feature mapping on the low-dimensional feature sequence to obtain enhanced features and output the enhanced signal frame. The autoencoder decoding module includes a fully connected layer for reconstructing the enhanced features.
3. The weak line spectrum enhancement method based on an autoencoder-transformer network structure according to claim 1, characterized in that: The enhanced signal undergoes adaptive frequency band selection processing, specifically as follows: Perform a Fast Fourier Transform on the enhanced signal to obtain the spectrum of the enhanced signal; The spectrum is divided into multiple equally spaced frequency bands as candidate frequency bands; Calculate the energy value of each candidate frequency band to obtain the energy distribution of the enhanced signal in different frequency bands.
4. The weak line spectrum enhancement method based on an autoencoder-transformer network structure according to claim 3, characterized in that: Determining the optimal frequency band processing range for the enhanced signal based on the main frequency band includes: The candidate frequency band with the highest energy value is selected as the main frequency band; The energy concentration of the main frequency band is calculated based on the ratio of its energy value to the average energy of its two adjacent frequency bands. The process expands from the main frequency band and calculates the energy concentration of adjacent frequency bands in turn. The frequency bands whose energy concentration meets the set threshold constitute the optimal frequency band processing range.
5. The weak line spectrum enhancement method based on an autoencoder-transformer network structure according to claim 1, characterized in that: Refined enhancement processing of the enhanced signal within the optimal frequency band processing range includes: The enhanced signal is filtered using a bandpass filter, and the enhanced signal within the optimal frequency band processing range after filtering is retained; The enhanced signal within the optimal frequency band after filtering is subjected to spectral sharpening and energy normalization to obtain the final enhanced sound pressure signal.