A remote speech enhancement transmission method and system based on semantic communication

By extracting and transmitting the semantic features of speech signals during the transmission process, and using semantic communication technology to restore high-quality pure speech under low signal-to-noise ratio channel conditions, the problems of waste of resources and low signal quality of traditional speech enhancement systems are solved, and efficient voice signal transmission and recovery are achieved.

CN119296566BActive Publication Date: 2025-05-16NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411845049.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-16
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Traditional voice enhancement systems have large amounts of data and wasteful resources during transmission, and it is difficult to restore high-quality pure voice under low signal-to-noise ratio channel conditions.

Method used

The remote voice enhancement transmission method based on semantic communication is adopted, and the noisy voice signal is extracted through the local transmitter. The speech semantic features are encoded and dimensionally adjusted by the semantic encoder and channel encoder. After transmission to the remote receiving end, the channel decoder and semantic decoder are used for dimensional recovery and semantic decoding, and finally the reconstructed speech signal is obtained through the inverse short-term Fourier transform.

Benefits of technology

In the case of saving communication resources, the pure voice quality recovered by the remote receiver under low signal-to-noise ratio channel conditions is significantly improved, and the transmission efficiency and signal quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296566B_ABST
    Figure CN119296566B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote speech enhancement transmission method and system based on semantic communication in the field of speech signal transmission and processing technology. The system includes: a local transmitting end, which is used to: after performing short-time Fourier transform on the noisy speech signal to be enhanced, use a semantic encoder to extract semantic features from the spectrum of the noisy speech signal; then use a channel encoder to adjust the dimension of the speech semantic features; and finally transmit it to a remote receiving end through a channel; a remote receiving end, which is used to: receive the speech semantic feature signal transmitted through the channel; after using a channel decoder to restore the dimension of the received speech semantic feature signal, use a semantic decoder to perform semantic decoding to obtain the real part and imaginary part of the predicted pure speech signal, and finally perform an inverse short-time Fourier transform to obtain a reconstructed speech signal. The present invention can significantly improve the pure speech quality restored by the remote receiving end under low signal-to-noise ratio channel transmission conditions while saving communication resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a remote speech enhancement transmission method and system based on semantic communication, belonging to the technical field of speech signal processing and transmission. Background Art

[0002] In traditional speech enhancement systems, the transmitter needs to send the entire speech segment to the receiver, which results in the transmission of more data than the data required by the terminal task, restricting the transmission efficiency and causing a waste of transmission resources. On the other hand, speech signals are easily affected by channel noise during transmission, especially in the case of low signal-to-noise ratio, and the receiver can hardly recover pure speech from the received speech signal.

[0003] Semantic communication is a new communication architecture that deeply combines deep learning with wireless communication. It can greatly improve communication efficiency and significantly enhance signal transmission quality under low signal-to-noise ratio channel conditions by integrating users' information needs and semantics into the communication process. As a novel communication paradigm, semantic communication refers to the use of neural networks to extract and encode semantic information at the sending end and restore the information from a semantic perspective at the receiving end.

[0004] Thanks to the development of deep learning technology, by designing a reasonable semantic communication model, it is possible to efficiently extract and restore the semantic information of the data, while reducing the interference of channel noise on the transmitted data, and greatly improving the transmission efficiency of the communication system and saving communication resources. Summary of the invention

[0005] Purpose: In view of at least one of the above technical problems, the present application provides a remote speech enhancement transmission method and system based on semantic communication, which can significantly improve the quality of pure speech restored by the remote receiving end under low signal-to-noise ratio channel transmission conditions while saving communication resources.

[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0007] In a first aspect, the present invention provides a remote speech enhancement transmission system based on semantic communication, comprising:

[0008] Local sender, used for:

[0009] Performing short-time Fourier transform (STFT) on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal;

[0010] Extracting semantic features from the noisy speech signal spectrum using a semantic encoder to obtain speech semantic features;

[0011] Using a channel encoder to adjust the dimension of the speech semantic feature to obtain a speech semantic feature signal after the dimension adjustment;

[0012] Transmitting the dimensionally adjusted speech semantic feature signal to a remote receiving end through a channel;

[0013] Remote receiving end, used for:

[0014] Receiving a speech semantic feature signal transmitted through a channel;

[0015] Using a channel decoder to restore the dimension of the received speech semantic feature signal to obtain the speech semantic feature after the dimension restoration;

[0016] The semantic decoder is used to semantically decode the speech semantic features after dimension restoration to obtain the real part and imaginary part of the predicted pure speech signal;

[0017] The real and imaginary parts of the predicted pure speech signal are subjected to inverse short-time Fourier transform (ISTFT) to obtain the reconstructed speech signal.

[0018] In some embodiments, the semantic encoder includes a first convolutional layer, an expanded dense convolutional network DenseNet, a second convolutional layer, and a plurality of time-frequency Conformer modules connected in sequence;

[0019] The first convolution layer includes a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are connected in sequence, and are used to expand the number of channels of the noisy speech signal spectrum;

[0020] The dilated dense convolutional network DenseNet is composed of 4 dilated convolutional layers with different dilation factors, the dilation factors of the 4 dilated convolutional layers are 1, 2, 4, and 8 respectively, each dilated convolutional layer includes a one-dimensional dilated convolutional layer with a step size of 1 and a convolution kernel size of 2×3, a normalization layer with a channel number of 64, and a PReLU activation function layer connected in sequence; the dilated dense convolutional network DenseNet is used to aggregate all feature maps obtained by the first convolutional layer to extract speech features at different levels;

[0021] The second convolution layer includes a one-dimensional convolution layer with a stride of 2 and a convolution kernel size of 1×3, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are connected in sequence; the second convolution layer is used to halve the frequency dimension to reduce complexity;

[0022] The time-frequency convolution enhanced transformer Conformer module includes a first dimensional transformation layer, a time domain Conformer network, a first residual connection layer, a second dimensional transformation layer, a frequency domain Conformer network, a second residual connection layer and a third dimensional transformation layer connected in sequence. The time-frequency Conformer module is used to capture the time dependency and frequency dependency of speech features to obtain speech semantic features.

[0023] Furthermore, the time domain Conformer network has the same structure as the frequency domain Conformer network, both of which include a first feedforward neural network, a multi-head attention mechanism, a one-dimensional convolutional layer, a second feedforward neural network and a regularization layer connected in sequence.

[0024] In some embodiments, the processing of the semantic encoder includes:

[0025] The real part, imaginary part and amplitude of the spectrum of the noisy speech signal are connected in series as the input of the semantic encoder, and the three input features of the real part, imaginary part and amplitude of the spectrum of the noisy speech signal are expanded into an intermediate feature map with C channels through the first convolutional layer; the intermediate feature map is expanded by the dilated convolution layer of the residual connection in the dense convolutional network DenseNet, while retaining the number of layers, and the receptive field is increased, and all previous feature maps are aggregated to extract features of different scales; then the frequency dimension F of the features of different scales is halved to F / 2 through the second convolutional layer to reduce the complexity; finally, the time dependency and frequency dependency of the F / 2 dimensional features are extracted in turn through the time-frequency Conformer module to obtain the speech semantic features.

[0026] In some embodiments, the channel encoder includes a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1 and a dimensionality transformation layer connected in sequence, which is used to change the dimensional shape of the speech semantic features to be suitable for channel transmission.

[0027] In some embodiments, receiving a speech semantic feature signal transmitted through a channel includes:

[0028] ,

[0029] in, is the received speech semantic feature signal, is the channel parameter, is the speech semantic feature signal after dimension adjustment, represents Gaussian noise, where represents the Gaussian noise function, represents the variance of Gaussian noise in each channel, represents the identity matrix. Further, the signal-to-noise ratio of the channel is 0-10 dB.

[0030] In some embodiments, the channel decoder includes a reshaping layer and a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, which are connected in sequence, and are used to restore the dimension of the received speech semantic feature signal.

[0031] In some embodiments, the semantic decoder includes a sequentially connected dilated dense convolutional network DenseNet, an upsampling layer, and a third convolutional layer;

[0032] The dilated dense convolutional network DenseNet is composed of 4 dilated convolutional layers with different dilation factors, the dilation factors of the 4 dilated convolutional layers are 1, 2, 4, and 8 respectively, each dilated convolutional layer includes a one-dimensional dilated convolutional layer with a step size of 1 and a convolution kernel size of 2×3, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are connected in sequence. The dilated dense convolutional network DenseNet is used to aggregate semantic features; the upsampling layer is used to restore the frequency dimension; the third convolutional layer includes a one-dimensional convolutional layer and a normalization layer connected in sequence, which are used to compress the number of channels.

[0033] In some embodiments, the processing of the semantic decoder includes:

[0034] The restored dimensional speech semantic features are used as the input of the semantic decoder, and multiple feature maps are aggregated through the expanded dense convolutional network DenseNet; then the frequency dimension of the semantic features is upsampled back to the frequency F through the upsampling layer; finally, the number of channels is compressed to 1 through the third convolutional layer, and the real and imaginary parts of the predicted pure speech signal are obtained.

[0035] In some embodiments, the semantic encoder, channel encoder, channel decoder, and semantic decoder are pre-trained, and the pre-training method includes:

[0036] The clean speech signal in the clean speech data set is mixed with the noise signal in the noise set according to a certain signal-to-noise ratio, and the corresponding clean speech signal is used as the true label to obtain a labeled noisy speech data set;

[0037] Dividing the noisy speech data set into a training data set and a verification data set;

[0038] In each round of training, the training data set is used as the input of the remote speech enhancement transmission system to be trained, and the reconstructed speech signal is used as the output, and the semantic encoder, channel encoder, channel decoder, and semantic decoder are trained. During the training process, the mean square error is used as the loss function, random gradient descent is adopted, and the parameters are adjusted using the AdamW optimizer to obtain the trained remote speech enhancement transmission system;

[0039] After each round of training is completed, the speech enhancement system is tested with the verification data set as input and the reconstructed speech signal as output, and the objective speech quality evaluation PESQ and short-time objective intelligibility STOI values ​​of the clean speech signal and the reconstructed speech signal are calculated. The hyperparameters are adjusted according to the performance indicators of PESQ and STOI, and the next round of training is continued until the preset number of training times is reached to obtain the pre-trained semantic encoder, channel encoder, channel decoder and semantic decoder.

[0040] In a second aspect, the present invention provides a remote speech enhancement transmission method based on semantic communication, comprising:

[0041] The local sender performs the following steps:

[0042] Performing short-time Fourier transform (STFT) on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal;

[0043] Extracting semantic features from the noisy speech signal spectrum using a semantic encoder to obtain speech semantic features;

[0044] Using a channel encoder to adjust the dimension of the speech semantic feature to obtain a speech semantic feature signal after the dimension adjustment;

[0045] Transmitting the dimensionally adjusted speech semantic feature signal to a remote receiving end through a channel;

[0046] The remote receiving end performs the following steps:

[0047] Receiving a speech semantic feature signal transmitted through a channel;

[0048] Using a channel decoder to restore the dimension of the received speech semantic feature signal to obtain the speech semantic feature after the dimension restoration;

[0049] The semantic decoder is used to semantically decode the speech semantic features after dimension restoration to obtain the real part and imaginary part of the predicted pure speech signal;

[0050] The real and imaginary parts of the predicted pure speech signal are subjected to inverse short-time Fourier transform (ISTFT) to obtain the reconstructed speech signal.

[0051] Beneficial effects: The present invention extracts the semantic features of speech signals at the local transmitting end, and then transmits the semantic feature information through a wireless channel, which can effectively reduce the amount of data transmitted and save communication resources. At the same time, by extracting and restoring semantic information through a neural network, the influence of channel noise on the transmission signal can be greatly reduced, ensuring the quality of reconstructed speech under low signal-to-noise ratio. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of a remote speech enhancement transmission system based on semantic communication according to an embodiment of the present invention;

[0053] Figure 2 is the PESQ value of different methods in the Gaussian channel in the embodiments of the present invention;

[0054] Figure 3 is the STOI value of different methods in the Gaussian channel in the embodiments of the present invention;

[0055] Figure 4 Schematic diagram of the network structure of the time-frequency Conformer module in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0057] In the description of this application, "several" means more than one, "more" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0058] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0059] The term "and / or" is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " generally indicates that the related objects are in an "or" relationship.

[0060] Example 1: Figure 1 As shown, this embodiment provides a remote speech enhancement transmission system based on semantic communication, including two parts: local sending end encoding and remote receiving end decoding.

[0061] Local sender, used for:

[0062] Performing short-time Fourier transform (STFT) on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal;

[0063] Extracting semantic features from the noisy speech signal spectrum using a semantic encoder to obtain speech semantic features;

[0064] Using a channel encoder to adjust the dimension of the speech semantic feature to obtain a speech semantic feature signal after the dimension adjustment;

[0065] Transmitting the dimensionally adjusted speech semantic feature signal to a remote receiving end through a channel;

[0066] Remote receiving end, used for:

[0067] Receiving a speech semantic feature signal transmitted through a channel;

[0068] Using a channel decoder to restore the dimension of the received speech semantic feature signal to obtain the speech semantic feature after the dimension restoration;

[0069] The semantic decoder is used to semantically decode the speech semantic features after dimension restoration to obtain the real part and imaginary part of the predicted pure speech signal;

[0070] The real and imaginary parts of the predicted pure speech signal are subjected to inverse short-time Fourier transform (ISTFT) to obtain the reconstructed speech signal.

[0071] The semantic encoder consists of a first convolutional layer, an expanded dense convolutional network DenseNet, a second convolutional layer, and four time-frequency Conformer modules connected in sequence.

[0072] In this embodiment, the first convolution layer includes a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are used to expand the number of channels of the noisy speech signal spectrum.

[0073] In this embodiment, the dilated dense convolutional network DenseNet is composed of 4 dilated convolutional layers with different dilation factors, and the dilation factors of the 4 dilated convolutional layers are 1, 2, 4, and 8 respectively; each dilated convolutional layer includes a one-dimensional dilated convolutional layer with a stride of 1 and a convolution kernel size of 2×3, a normalization layer with a channel number of 64, and a PReLU activation function layer connected in sequence. The dilated dense convolutional network DenseNet is used to aggregate all feature maps obtained by the first convolutional layer to extract speech features at different levels.

[0074] In this embodiment, the second convolution layer includes a one-dimensional convolution layer with a stride of 2 and a convolution kernel size of 1×3, a normalization layer with a channel number of 64, and a PReLU activation function layer in sequence. The second convolution layer is used to halve the frequency dimension to reduce complexity.

[0075] In this embodiment, if Figure 4 As shown, the Convolution-augmented transformer (Conformer) module includes a first dimension transformation layer, a time domain Conformer network, a first residual connection layer, a second dimension transformation layer, a frequency domain Conformer network, a second residual connection layer and a third dimension transformation layer connected in sequence. The time-frequency Conformer module is used to capture the time dependency and frequency dependency of speech features to obtain speech semantic features.

[0076] The time domain Conformer network includes a first feedforward neural network, a multi-head attention mechanism, a one-dimensional convolutional layer, a second feedforward neural network and a regularization layer in sequence;

[0077] The frequency domain Conformer network has the same structure as the time domain Conformer network.

[0078] The processing of the semantic encoder includes:

[0079] The real part, imaginary part and amplitude of the noisy speech signal spectrum are connected in series as the input of the semantic encoder, and the three input features (real part, imaginary part and amplitude of the noisy speech signal spectrum) are expanded into an intermediate feature map with C channels through the first convolution layer; then the intermediate feature map is expanded through the residual connected hole convolution layer in the dense convolution network DenseNet, while retaining the number of layers, increasing the receptive field and effectively aggregating all previous feature maps to extract features of different scales; then the frequency dimension F of features of different scales is halved to F / 2 through the second convolution layer to reduce the complexity; finally, the time dependency and frequency dependency of the features are extracted in turn through the time-frequency Conformer module, thereby completing the semantic feature extraction task and obtaining speech semantic features.

[0080] The dimension of the speech semantic features is adjusted to make them suitable for channel transmission, which is specifically achieved through a channel encoder. The channel encoder includes a one-dimensional convolution layer with a step size of 1 and a convolution kernel size of 1 and a dimensionality transformation layer connected in sequence.

[0081] The processing of the semantic encoder and channel encoder is expressed as follows:

[0082] ,

[0083] in, is the spectrum of noisy speech signal, The parameters are The semantic encoder The parameters are The channel encoder, It is the speech semantic feature signal after dimension adjustment.

[0084] Finally, the dimensionally adjusted speech semantic feature signal is transmitted to the remote receiving end through the channel.

[0085] The remote receiving end decoding process includes the following steps:

[0086] The dimension-adjusted speech semantic feature signal transmitted by the local sending end is received, and the received speech semantic feature signal is expressed as:

[0087] ,

[0088] in, is the received speech semantic feature signal, is the channel parameter, is the speech semantic feature signal after dimension adjustment, represents Gaussian noise, where represents the Gaussian noise function, represents the variance of Gaussian noise in each channel, Represents the identity matrix.

[0089] The dimension of the speech semantic feature signal after dimension adjustment is restored by a channel decoder, wherein the channel decoder includes a reshaping layer and a one-dimensional convolution layer with a step size of 1 and a convolution kernel size of 1, which are connected in sequence.

[0090] The restored dimensional speech semantic features are received by the semantic decoder for prediction, and the real and imaginary parts of the predicted pure speech signal are obtained. The semantic decoder includes a dilated dense convolutional network DenseNet, an upsampling layer, and a third convolutional layer connected in sequence. Among them, the structure of the dilated dense convolutional network DenseNet is the same as that of the dilated dense convolutional network DenseNet in the semantic encoder. The upsampling layer is responsible for upsampling the frequency dimension back to F. The third convolutional layer includes a one-dimensional convolutional layer with a step size of 1 and a convolution kernel size of 1×2 and a normalization layer with a channel number of 2.

[0091] The processing of the semantic decoder includes:

[0092] The restored dimensional speech semantic features are used as the input of the semantic decoder, and multiple feature maps are effectively aggregated through the expanded dense convolutional network DenseNet; then the frequency dimension of the semantic features is upsampled back to F through the upsampling layer; finally, the number of channels is compressed to 1 through the third convolutional layer, and the real and imaginary parts of the predicted pure speech signal are obtained.

[0093] The processing expressions of the channel decoder and semantic decoder are as follows:

[0094] ,

[0095] in, To predict the real and imaginary parts of a clean speech signal, Indicates that the parameter is The channel decoder, Indicates that the parameter is The semantic decoder of is the received speech semantic feature signal.

[0096] According to the real and imaginary parts of the predicted pure speech signal, the inverse short-time Fourier transform (ISTFT) is performed to obtain the reconstructed speech signal to complete the speech enhancement task.

[0097] The semantic encoder, channel encoder, channel decoder, and semantic decoder are pre-trained, and the pre-training method includes:

[0098] The clean speech signal in the clean speech data set is mixed with the noise signal in the noise set according to a certain signal-to-noise ratio, and the corresponding clean speech signal is used as the true label to obtain a labeled noisy speech data set;

[0099] Dividing the noisy speech data set into a training data set and a verification data set;

[0100] In each round of training, the training data set is used as the input of the remote speech enhancement transmission system to be trained, and the reconstructed speech signal is used as the output, and the semantic encoder, channel encoder, channel decoder, and semantic decoder are trained. During the training process, the mean square error is used as the loss function, random gradient descent is adopted, and the parameters are adjusted using the AdamW optimizer to obtain the trained remote speech enhancement transmission system;

[0101] After each round of training is completed, the speech enhancement system is tested with the verification data set as input and the reconstructed speech signal as output, and the objective speech quality evaluation PESQ and short-time objective intelligibility STOI values ​​of the clean speech signal and the reconstructed speech signal are calculated. The hyperparameters are adjusted according to the performance indicators of PESQ and STOI, and the next round of training is continued until the preset number of training times is reached to obtain the pre-trained semantic encoder, channel encoder, channel decoder and semantic decoder.

[0102] The following is a description of the remote speech enhancement transmission method and system based on semantic communication in conjunction with specific application embodiments:

[0103] The performance of the present invention is compared with that of the traditional solution. The test set in the Voice Bank dataset is used as the pure speech. The test set contains 824 speech sounds from two speakers. NoiseX-92 is used as the noise set. The noise set contains 15 kinds of environmental noise. For each pure speech in the test set, a noise is randomly selected from the noise set to add noise. The signal-to-noise ratio is 5dB to obtain the noisy speech.

[0104] The traditional scheme uses 8-bit pulse code modulation (PCM) as the source coding method, and Turbo code is used for channel coding with a coding rate of 1 / 3. The modulation method uses 64-quadrature amplitude modulation (64-QAM). The logarithm maximum a posteriori probability algorithm (log-MAP) is used for decoding, and iterates 5 times to improve the decoding accuracy. Two methods are used to enhance the noisy speech, namely:

[0105] Traditional solution 1: After the noisy speech is transmitted to the receiving end using the traditional solution, the Conformer-Based Metric Generative Adversarial Network (CMGAN) is used at the receiving end to perform speech enhancement on the received speech signal.

[0106] Traditional solution 2: First, use CMGAN to enhance the noisy speech at the sending end, and then transmit the enhanced speech signal to the receiving end through the traditional solution.

[0107] The comparison results of the objective speech quality evaluation (PESQ) and short-time objective intelligibility (STOI) performance of the present invention and two traditional schemes under Gaussian channels with different signal-to-noise ratios (SNR) are shown in the figure. Figure 2 , Figure 3 shown.

[0108] from Figure 2 and Figure 3 It can be seen that under Gaussian channels with different signal-to-noise ratios (SNRs), the PESQ value and STOI value of the present invention are higher than those of the two traditional schemes. In particular, under low signal-to-noise ratio conditions, the PESQ value and STOI value of the present invention are much higher than those of the two traditional schemes, which proves that the remote speech enhancement transmission method and system based on semantic communication proposed in the present invention can still restore high-quality pure speech signals under low signal-to-noise ratio conditions. At the same time, as the channel signal-to-noise ratio increases, the results of the present invention are still higher than those of the two traditional schemes, further confirming the effectiveness and stability of the present invention. It is worth mentioning that the performance of traditional scheme 2 is higher than that of traditional scheme 1. This is because the scheme of transmitting first and then enhancing will cause the speech signal to be mixed into the channel noise, thereby affecting the performance of CMGAN.

[0109] Embodiment 2: Based on Embodiment 1, this embodiment provides a remote speech enhancement transmission method based on semantic communication, including:

[0110] The local sender performs the following steps:

[0111] Performing short-time Fourier transform (STFT) on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal;

[0112] Extracting semantic features from the noisy speech signal spectrum using a semantic encoder to obtain speech semantic features;

[0113] Using a channel encoder to adjust the dimension of the speech semantic feature to obtain a speech semantic feature signal after the dimension adjustment;

[0114] Transmitting the dimensionally adjusted speech semantic feature signal to a remote receiving end through a channel;

[0115] The remote receiving end performs the following steps:

[0116] Receiving a speech semantic feature signal transmitted through a channel;

[0117] Using a channel decoder to restore the dimension of the received speech semantic feature signal to obtain the speech semantic feature after the dimension restoration;

[0118] The semantic decoder is used to semantically decode the speech semantic features after dimension restoration to obtain the real part and imaginary part of the predicted pure speech signal;

[0119] The real and imaginary parts of the predicted pure speech signal are subjected to inverse short-time Fourier transform (ISTFT) to obtain the reconstructed speech signal.

[0120] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0121] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0124] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A remote speech enhancement transmission system based on semantic communication, characterized in that: include: The local transmitting end is used to: perform short-time Fourier transform on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal; Using a semantic encoder to extract semantic features from the spectrum of the noisy speech signal to obtain speech semantic features; using a channel encoder to adjust the dimension of the speech semantic features to obtain a speech semantic feature signal after the dimension adjustment; transmitting the speech semantic feature signal after the dimension adjustment to a remote receiving end through a channel; The remote receiving end is used to: receive the speech semantic feature signal transmitted through the channel; perform dimension restoration on the received speech semantic feature signal using a channel decoder to obtain the speech semantic feature after dimension restoration; The semantic decoder is used to semantically decode the speech semantic features after dimensionality restoration to obtain the real part and imaginary part of the predicted pure speech signal; the inverse short-time Fourier transform is performed according to the real part and imaginary part of the predicted pure speech signal to obtain the reconstructed speech signal; The semantic encoder includes a first convolutional layer, an expanded dense convolutional network, a second convolutional layer, and a plurality of time-frequency convolution enhanced transformer Conformer modules connected in sequence; the time-frequency convolution enhanced transformer Conformer module includes a first dimensional transformation layer, a time domain Conformer network, a first residual connection layer, a second dimensional transformation layer, a frequency domain Conformer network, a second residual connection layer, and a third dimensional transformation layer connected in sequence; The channel encoder comprises a one-dimensional convolution layer with a step size of 1 and a convolution kernel size of 1 and a dimensional transformation layer connected in sequence; The channel decoder comprises a reshaping layer and a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, which are connected in sequence; The semantic decoder includes a dilated dense convolutional network, an upsampling layer and a third convolutional layer connected in sequence.

2. The remote speech enhancement transmission system based on semantic communication according to claim 1 is characterized in that: The first convolution layer includes a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are connected in sequence, and are used to expand the number of channels of the noisy speech signal spectrum; The dilated dense convolutional network is composed of four dilated convolutional layers with different dilation factors, the dilation factors of the four dilated convolutional layers are 1, 2, 4, and 8 respectively, each dilated convolutional layer includes a one-dimensional dilated convolutional layer with a stride of 1 and a convolution kernel size of 2×3, a normalization layer with a channel number of 64, and a PReLU activation function layer connected in sequence; the dilated dense convolutional network is used to aggregate all feature maps obtained by the first convolutional layer to extract speech features at different levels; The second convolution layer includes a one-dimensional convolution layer with a stride of 2 and a convolution kernel size of 1×3, a normalization layer with a channel number of 64, and a PReLU activation function layer, which are connected in sequence; the second convolution layer is used to halve the frequency dimension to reduce complexity.

3. The remote speech enhancement transmission system based on semantic communication according to claim 1 is characterized in that: The time domain Conformer network has the same structure as the frequency domain Conformer network, both of which include a first feedforward neural network, a multi-head attention mechanism, a one-dimensional convolutional layer, a second feedforward neural network and a regularization layer connected in sequence.

4. The remote speech enhancement transmission system based on semantic communication according to claim 2 is characterized in that: The processing of the semantic encoder includes: The real part, imaginary part and amplitude of the spectrum of the noisy speech signal are connected in series as the input of the semantic encoder, and the three input features of the real part, imaginary part and amplitude of the spectrum of the noisy speech signal are expanded into an intermediate feature map with C channels through the first convolution layer; the intermediate feature map is expanded by the dilated convolution layer of the residual connection in the dense convolution network, the receptive field is increased while retaining the number of layers, and all previous feature maps are aggregated to extract features of different scales; then the frequency dimension F of the features of different scales is halved to F / 2 through the second convolution layer to reduce the complexity; finally, the time dependency and frequency dependency of the F / 2 dimensional features are extracted in turn through the time-frequency convolution enhanced transformer Conformer module to obtain the speech semantic features.

5. The remote speech enhancement transmission system based on semantic communication according to claim 1 is characterized in that: The dilated dense convolutional network consists of four dilated convolutional layers with different dilation factors, the dilation factors of the four dilated convolutional layers are 1, 2, 4, and 8 respectively, each dilated convolutional layer includes a one-dimensional dilated convolutional layer with a step size of 1 and a convolution kernel size of 2×3, a normalization layer with 64 channels, and a PReLU activation function layer, which are connected in sequence. The dilated dense convolutional network is used to aggregate semantic features; the upsampling layer is used to restore the frequency dimension; the third convolutional layer includes a one-dimensional convolutional layer and a normalization layer connected in sequence, which are used to compress the number of channels.

6. The remote speech enhancement transmission system based on semantic communication according to claim 5 is characterized in that: The processing of the semantic decoder includes: The restored dimensional speech semantic features are used as the input of the semantic decoder, and multiple feature maps are aggregated through the dilated dense convolutional network; then the frequency dimension of the semantic features is upsampled back to the frequency F through the upsampling layer; finally, the number of channels is compressed to 1 through the third convolutional layer, and the real and imaginary parts of the predicted pure speech signal are obtained.

7. The remote speech enhancement transmission system based on semantic communication according to claim 1 is characterized in that: Pre-training the semantic encoder, channel encoder, channel decoder, and semantic decoder, wherein the pre-training includes: The clean speech signal in the clean speech data set is mixed with the noise signal in the noise set according to a certain signal-to-noise ratio, and the corresponding clean speech signal is used as the true label to obtain a labeled noisy speech data set; Dividing the noisy speech data set into a training data set and a verification data set; In each round of training, the training data set is used as the input of the remote speech enhancement transmission system to be trained, and the reconstructed speech signal is used as the output, and the semantic encoder, channel encoder, channel decoder, and semantic decoder are trained. During the training process, the mean square error is used as the loss function, random gradient descent is adopted, and the parameters are adjusted using the AdamW optimizer to obtain the trained remote speech enhancement transmission system; After each round of training is completed, the speech enhancement system is tested with the verification data set as input and the reconstructed speech signal as output, and the objective speech quality evaluation PESQ and short-time objective intelligibility STOI values ​​of the clean speech signal and the reconstructed speech signal are calculated. The hyperparameters are adjusted according to the performance indicators of PESQ and STOI, and the next round of training is continued until the preset number of training times is reached to obtain the pre-trained semantic encoder, channel encoder, channel decoder and semantic decoder.

8. A remote speech enhancement transmission method based on semantic communication, characterized in that: include: The local transmitter performs the following steps: performing a short-time Fourier transform on the noisy speech signal to be enhanced to obtain a spectrum of the noisy speech signal; Using a semantic encoder to extract semantic features from the spectrum of the noisy speech signal to obtain speech semantic features; using a channel encoder to adjust the dimension of the speech semantic features to obtain a speech semantic feature signal after the dimension adjustment; transmitting the speech semantic feature signal after the dimension adjustment to a remote receiving end through a channel; The remote receiving end performs the following steps: receiving a speech semantic feature signal transmitted through a channel; performing dimension restoration on the received speech semantic feature signal using a channel decoder to obtain speech semantic features after dimension restoration; The semantic decoder is used to semantically decode the speech semantic features after dimensionality restoration to obtain the real part and imaginary part of the predicted pure speech signal; the inverse short-time Fourier transform is performed according to the real part and imaginary part of the predicted pure speech signal to obtain the reconstructed speech signal; The semantic encoder includes a first convolutional layer, an expanded dense convolutional network, a second convolutional layer, and a plurality of time-frequency convolution enhanced transformer Conformer modules connected in sequence; the time-frequency convolution enhanced transformer Conformer module includes a first dimensional transformation layer, a time domain Conformer network, a first residual connection layer, a second dimensional transformation layer, a frequency domain Conformer network, a second residual connection layer, and a third dimensional transformation layer connected in sequence; The channel encoder comprises a one-dimensional convolution layer with a step size of 1 and a convolution kernel size of 1 and a dimensional transformation layer connected in sequence; The channel decoder comprises a reshaping layer and a one-dimensional convolution layer with a stride of 1 and a convolution kernel size of 1, which are connected in sequence; The semantic decoder includes a dilated dense convolutional network, an upsampling layer and a third convolutional layer connected in sequence.

Citation Information

Patent Citations

  • Intelligent control system

    CN106157971A

  • Single-channel voice echo cancellation method and device based on deep neural network

    CN115565543A