A Method and System for Eliminating Impact Noise of Power Amplifiers Based on Deep Learning

By constructing a shock sound prediction model based on long and short-term memory network based on deep learning, monitoring the audio channel signals and determining the sound source switching strategy, the problem of the inability to predict the impact sound of the power amplifier in advance in the existing technology is solved, and the sound quality and user experience of the broadcast system are improved.

CN119402797BActive Publication Date: 2025-07-22GUANGZHOU BAOLUN ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411360768.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-07-22
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

In the prior art, the amplifier impact sound cancellation method cannot predict and adapt to complex and changeable practical application scenarios in advance, resulting in the impact of the sound quality and user experience of the broadcast system.

Method used

The impact sound prediction model is constructed based on deep learning long and short-term memory network, monitor the audio channel signal, predict the magnitude and time of the impact sound, and determine the sound source switching strategy based on the prediction results, including normal switching, electronic volume adjustment and switching to other channels, so as to achieve early cancellation of impact sound.

Benefits of technology

It realizes accurate prediction of the timing of sound source switching, can eliminate the impact sound before it occurs, improves the sound quality and user experience of the broadcast system, reduces maintenance costs and improves flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119402797B_ABST
    Figure CN119402797B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for eliminating impact sound of a power amplifier based on deep learning. The method includes: monitoring audio signals of each audio channel to confirm whether a sound source switch is about to occur for each audio channel; when it is confirmed that there is a first audio channel about to perform a sound source switch, obtaining target sound source data to be switched and audio data sets corresponding to each audio channel respectively; inputting each of the audio data sets and the target sound source data into a preset impact sound prediction model, so that the impact sound prediction model outputs impact sound prediction results of each audio channel according to each of the audio data sets, and the impact sound prediction model is constructed based on a long short-term memory network and trained through a historical audio data set; determining a sound source switching strategy according to the impact sound prediction results of each audio channel; and performing the sound source switch according to the sound source switching strategy, thereby improving the sound quality of the broadcast system and the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of audio data processing and deep learning, and particularly relates to a method and system for eliminating amplifier impact noise based on deep learning. Background Art

[0002] A power amplifier, abbreviated as PA, generally specifically refers to a most basic device in an audio system. Its main function is to amplify the relatively weak signal input by the sound source equipment and then generate a large enough current to drive the speaker to reproduce sound.

[0003] In large-scale audio playback scenarios such as broadcast systems, theater audio, and conference centers, harsh impact noises often occur during the source switching process of the partitioned power amplifier. These impact noises are mainly caused by the level mutation or phase difference between the old and new source signals, seriously affecting the sound quality experience and the comfort of the audience. In the prior art, the impact noise prediction method based on a fixed threshold or simple rules is usually used to eliminate the amplifier impact noise. Its accuracy is affected by various factors such as environmental noise and source characteristics, and it is difficult to adapt to complex and changeable actual application scenarios. At the same time, the existing methods for eliminating amplifier impact noise usually need to detect and process after the impact noise occurs, and cannot predict and take measures in advance. They lack intelligent learning and adaptive capabilities and cannot automatically adjust the control strategy according to environmental changes. Summary of the Invention

[0004] In view of the above technical problems, this application provides a method and system for eliminating amplifier impact noise based on deep learning, effectively eliminating the amplifier impact noise and improving the sound quality and user experience of the broadcast system.

[0005] In a first aspect, an embodiment of this application provides a method for eliminating amplifier impact noise based on deep learning, including:

[0006] Monitoring the audio signals of each audio channel to confirm whether each audio channel is about to switch the sound source;

[0007] When it is confirmed that there is a first audio channel about to switch the sound source, obtaining the target sound source data to be switched and the audio data set corresponding to each audio channel, where the audio data set includes the sound source data currently played by each audio channel and environmental noise data;

[0008] Inputting each of the audio data sets and the target sound source data into a preset impact noise prediction model, so that the impact noise prediction model outputs the impact noise prediction results of each audio channel according to each of the audio data sets. The impact noise prediction results include the magnitude of the impact noise generated by the audio channel and the moment when the impact noise is generated. The impact noise prediction model is constructed based on a long short-term memory network and trained through a historical audio data set;

[0009] Determine a sound source switching strategy according to the impact sound prediction results of the respective audio channels;

[0010] Perform the sound source switching according to the sound source switching strategy.

[0011] An embodiment of the present application provides a method for eliminating amplifier impact sound based on deep learning. When it is monitored that there is a first audio channel about to perform a sound source switch, an impact sound prediction model constructed based on a long short-term memory network is used to predict the impact sound of each audio channel, determine the magnitude of the impact sound generated by each audio channel and the moment when the impact sound is generated, achieve accurate prediction of the sound source signal switching timing, provide a decision basis for subsequent elimination of amplifier impact sound, and finally automatically select a sound source switching strategy according to the impact sound prediction results, so that the embodiment of the present application can adapt to different audio environments and requirements. Compared with the existing methods for eliminating amplifier impact sound, the method for eliminating amplifier impact sound based on deep learning provided by the embodiment of the present application can predict the impact sound in advance, and then perform relevant operations for eliminating amplifier impact sound before the impact sound occurs, effectively eliminating the amplifier impact sound in the sound source switching scenario, and improving the sound quality and user experience of the broadcast system.

[0012] Further, the determining the sound source switching strategy according to the impact sound prediction results of the respective audio channels includes:

[0013] If it is predicted that the first audio channel will not generate an impact sound, determine the sound source switching strategy as: normally perform the sound source switching in the first audio channel;

[0014] If it is predicted that the impact sound generated by the first audio channel is less than a preset threshold, determine the sound source switching strategy as: perform the sound source switching in the first audio channel by using an electronic volume adjustment method;

[0015] If it is predicted that the impact sound generated by the first audio channel is greater than a preset threshold, determine the sound source switching strategy as: according to the impact sound prediction results of the respective audio channels, select a second audio channel from each of the audio channels, and switch the target sound source to the second audio channel.

[0016] The embodiments of the present application provide three sound source switching strategies, namely, normal sound source switching, sound source switching using electronic volume adjustment, and sound source switching using other audio channels. According to the prediction result output by the impact sound prediction model, the embodiments of the present application select from the above three sound source switching strategies. When it is predicted that no impact sound will be generated in the first audio channel, no additional impact sound elimination operation is required, and the sound source can be normally switched in the first audio channel; when it is predicted that the impact sound generated in the first audio channel is less than a preset threshold, it indicates that the impact sound can be eliminated by means of electronic volume adjustment. At this time, the sound source is still switched in the first audio channel; when it is predicted that the impact sound generated in the first audio channel is greater than the preset threshold, it indicates that the impact sound is difficult to eliminate in the current audio channel, and continuing to switch the sound source will seriously affect the sound quality experience and comfort of the listener. Therefore, according to the impact sound prediction results of each audio channel, a second audio channel for sound source switching is selected, and then the audio switching of the target sound source is performed in the second audio channel. By setting three different sound source switching strategies, the power amplifier impact sound elimination method provided by the embodiments of the present application can adapt to different audio environments and requirements. At the same time, this solution has the characteristics of high automation, reduces the operation process of manual intervention, reduces the maintenance cost, and improves the flexibility and convenience of power amplifier impact sound elimination.

[0017] In a possible implementation manner, the sound source switching using the electronic volume adjustment method in the first audio channel includes:

[0018] Determine the volume adjustment time for the electronic volume adjustment according to the time when the impact sound is generated in the first audio channel;

[0019] At the volume adjustment time, use the cross-fading technology for the first audio channel to gradually reduce the volume of the current sound source in the first audio channel, and at the same time gradually increase the volume of the target sound source.

[0020] In the embodiments of the present application, when the impact sound generated in the first audio channel is less than a preset threshold, the cross-fading technology is used to eliminate the impact sound during the sound source switching process. The key to implementing the cross-fading technology lies in how to determine the volume adjustment time. If the volume adjustment time is too early, the old sound source has not finished playing yet, which is equivalent to switching to the new sound source in advance, bringing a poor experience to the listener. If the volume adjustment time is too late, there is still a risk of generating impact sound. Therefore, the embodiments of the present application predict the time when the impact sound is generated in the first audio channel through the impact sound prediction model, and then use this as a reference to determine the volume adjustment time for the electronic volume adjustment, and use the cross-fading technology before the impact sound is expected to occur, which improves the accuracy of implementing the cross-fading technology and further improves the user experience.

[0021] In a possible implementation manner, selecting a second audio channel from each of the audio channels according to the impact sound prediction results of the respective audio channels, and switching the target sound source to the second audio channel includes:

[0022] Evaluating the anti-impact sound level of each of the audio channels according to the impact sound prediction results of the respective audio channels and the audio signal characteristics of the respective audio channels;

[0023] Performing audio channel selection among the respective audio channels according to the anti-impact sound levels of the respective audio channels, the priorities of the currently playing sound sources of the respective audio channels, and the priority of the target sound source to determine the second audio channel;

[0024] Switching the target sound source to the second audio channel.

[0025] In the embodiments of the present application, first, the anti-impact sound levels of each audio channel are evaluated, and then, in combination with the priorities of the target sound source and other sound sources, the second audio channel is determined. For example, for important speeches or music performances, higher priorities can be set in advance to ensure the normal playback of the sound source. At the same time, according to the evaluation results of the anti-impact sound levels of the audio channels, the second audio channel for sound source switching is intelligently selected. When switching the sound source, it is preferentially switched to the channel with predicted no impact sound or less impact sound and the audio channel has no currently playing sound source or the sound source priority is low.

[0026] In a possible implementation manner, constructing the impact sound prediction model based on a long short-term memory network and training it through a historical audio data set includes:

[0027] Obtaining a historical audio data set in various sound source switching scenarios;

[0028] Extracting a number of Mel frequency cepstral coefficients features from the historical audio data set;

[0029] Arranging each of the Mel frequency cepstral coefficients features in chronological order to construct a number of feature data sets;

[0030] Training a preset long short-term memory network using each of the feature data sets to obtain the impact sound prediction model.

[0031] The embodiments of the present application provide a training method for an impact sound prediction model. By constructing a multi-layer LSTM (Long Short-Term Memory) network and combining MFCC (Mel Frequency Cepstral Coefficients) feature extraction, it realizes the accurate prediction of the switching timing of the sound source signal. The LSTM (Long Short-Term Memory) network structure includes an input layer, a hidden layer (which may include multiple LSTM units), and an output layer. Through its unique gating mechanism, it can effectively handle the long-term and short-term dependencies in time series data. Among them, the forget gate can determine which old audio features (which may represent background noise or previous sound source information) should be forgotten at the current time step, so that the LSTM model can focus on the features of the current impact sound. For the update of the memory cell state, through the interaction between the input gate and the memory cell, the LSTM model can selectively add new information to the memory cell state, and allows the LSTM model to learn and retain the key features of the impact sound, which are used to eliminate the impact sound in subsequent time steps. By the forget gate and updating the memory cell state, the LSTM can capture the long-term dependencies in the audio signal. The hidden layer of the LSTM not only contains the input information of the current time step, but also integrates the memory information of the previous time step. Through the hidden layer, the LSTM can capture the short-term dependencies in the audio signal, such as the association between the impact sound and its front and back audio frames. And the output gate controls which information is transferred from the memory cell state to the hidden layer and finally serves as the output of the current time step. So the output gate can selectively output the feature information related to the impact sound elimination as needed. The LSTM network can capture the unique timing pattern and frequency characteristics of the impact sound through its gating mechanism. These feature information may be manifested as sudden amplitude changes in the audio signal, an increase in high-frequency components, or a sudden increase in energy in a specific frequency band, etc. At the same time, these feature information play a key role in the process of impact sound elimination, because they are used to guide how to modify or replace the impact sound part in the original audio frame after training the LSTM model. Therefore, the finally obtained impact sound prediction model can adapt to the changes in different environmental noises and sound source characteristics and achieve more accurate impact sound prediction.

[0032] Further, extracting several Mel frequency cepstral coefficients features from the historical audio dataset includes:

[0033] Performing noise removal, echo cancellation, pre-emphasis, framing, and windowing operations on the historical audio dataset to obtain several audio frames;

[0034] Performing Fourier transform (FFT) transform, calculating the energy spectrum, and Mel filter bank conversion operations on each of the audio frames to obtain corresponding several logarithmic energies;

[0035] Extracting the first N coefficients from each of the logarithmic energies as each of the Mel frequency cepstral coefficients features through discrete cosine transform.

[0036] In a second aspect, correspondingly, the present application provides a power amplifier impact sound cancellation system based on deep learning, including a monitoring module, an acquisition module, a prediction module, a strategy determination module, and a sound source switching module;

[0037] Among them, the monitoring module is used to monitor the audio signals of each audio channel to confirm whether each audio channel is about to perform a sound source switch;

[0038] The acquisition module is used to, when it is confirmed that there is a first audio channel about to perform a sound source switch, acquire the target sound source data to be switched and the audio data sets corresponding to each audio channel respectively. The audio data set includes the sound source data currently played by each audio channel and the ambient noise data;

[0039] The prediction module is used to input each of the audio data sets and the target sound source data into a preset impact sound prediction model, so that the impact sound prediction model outputs the impact sound prediction results of each audio channel according to each of the audio data sets. The impact sound prediction results include the magnitude of the impact sound generated by the audio channel and the moment when the impact sound is generated. The impact sound prediction model is constructed based on a long short-term memory network and trained through a historical audio data set;

[0040] The strategy determination module is used to determine a sound source switching strategy according to the impact sound prediction results of each audio channel;

[0041] The sound source switching module is used to perform the sound source switch according to the sound source switching strategy.

[0042] Further, the strategy determination module includes a first strategy determination unit, a second strategy determination unit, and a third strategy determination unit;

[0043] Among them, the first strategy determination unit is used to, if it is predicted that the first audio channel will not generate an impact sound, determine the sound source switching strategy as: normally perform a sound source switch in the first audio channel;

[0044] The second strategy determination unit is used to, if it is predicted that the impact sound generated by the first audio channel is less than a preset threshold, determine the sound source switching strategy as: perform a sound source switch in the first audio channel by using an electronic volume adjustment method;

[0045] The third strategy determination unit is used to, if it is predicted that the impact sound generated by the first audio channel is greater than a preset threshold, determine the sound source switching strategy as: according to the impact sound prediction results of each audio channel, select a second audio channel among each audio channel and switch the target sound source to the second audio channel.

[0046] In a possible implementation manner, the method of performing sound source switching by using electronic volume adjustment in the first audio channel includes:

[0047] Determine the volume adjustment moment for performing the electronic volume adjustment according to the moment when an impact sound is generated in the first audio channel;

[0048] At the volume adjustment moment, adopt the cross-fading technique for the first audio channel, gradually reduce the volume of the current sound source in the first audio channel, and simultaneously gradually increase the volume of the target sound source.

[0049] In a possible implementation manner, the method of selecting a second audio channel from each of the audio channels according to the impact sound prediction results of each audio channel and switching the target sound source to the second audio channel includes:

[0050] Evaluate the anti-impact sound level of each audio channel according to the impact sound prediction results of each audio channel and the audio signal characteristics of each audio channel;

[0051] Perform audio channel selection among each of the audio channels according to the anti-impact sound level of each audio channel, the priority of the currently playing sound source of each audio channel, and the priority of the target sound source, and determine the second audio channel;

[0052] Switch the target sound source to the second audio channel. Description of the Drawings

[0053] Figure 1 : It is a schematic flowchart of a method for eliminating impact sound of a power amplifier based on deep learning provided by an embodiment of the present application.

[0054] Figure 2 : It is a schematic flowchart of determining a sound source switching strategy in a method for eliminating impact sound of a power amplifier based on deep learning provided by an embodiment of the present application.

[0055] Figure 3 : It is a schematic structural diagram of a system for eliminating impact sound of a power amplifier based on deep learning provided by an embodiment of the present application.

[0056] Figure 4 : It is a schematic structural diagram of a strategy determination module in a system for eliminating impact sound of a power amplifier based on deep learning provided by an embodiment of the present application. Detailed Embodiments

[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0058] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve to limit the execution order of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0059] Embodiment 1:

[0060] As Figure 1 shown, Embodiment 1 provides a method for eliminating amplifier impact noise based on deep learning, including steps S1 - S5:

[0061] Step S1, monitor the audio signals of each audio channel to confirm whether each audio channel is about to perform a sound source switch;

[0062] Step S2, when it is confirmed that there is a first audio channel about to perform a sound source switch, obtain the target sound source data to be switched and the audio data sets corresponding to each audio channel respectively. The audio data set includes the sound source data currently played by each audio channel and environmental noise data;

[0063] Step S3, input each of the audio data sets and the target sound source data into a preset impact noise prediction model, so that the impact noise prediction model outputs the impact noise prediction results of each audio channel according to each of the audio data sets. The impact noise prediction results include the magnitude of the impact noise generated by the audio channel and the moment when the impact noise is generated. The impact noise prediction model is constructed based on a long short-term memory network and trained through a historical audio data set;

[0064] Step S4, determine a sound source switching strategy according to the impact noise prediction results of each audio channel;

[0065] Step S5, perform the sound source switch according to the sound source switching strategy.

[0066] An embodiment of the present application provides a method for eliminating amplifier impact noise based on deep learning. When it is detected that there is a first audio channel about to switch the sound source, the impact noise prediction model constructed based on the long short-term memory network is used to predict the impact noise of each audio channel, determine the magnitude of the impact noise generated by each audio channel and the moment when the impact noise is generated, realize the accurate prediction of the sound source signal switching timing, provide a decision basis for the subsequent elimination of amplifier impact noise, and finally automatically select a sound source switching strategy according to the impact noise prediction result, so that the embodiment of the present application can adapt to different audio environments and requirements. Compared with the existing methods for eliminating amplifier impact noise, the method for eliminating amplifier impact noise based on deep learning provided by the embodiment of the present application can predict the impact noise in advance, and then perform relevant amplifier impact noise elimination operations before the impact noise occurs, effectively eliminating the amplifier impact noise in the sound source switching scenario, and improving the sound quality and user experience of the broadcast system.

[0067] Further, in step S4, determining the sound source switching strategy according to the impact noise prediction results of the respective audio channels, as Figure 2 shown, includes steps S401 - S403:

[0068] S401. If it is predicted that the first audio channel will not generate impact noise, determine the sound source switching strategy as: normally perform sound source switching in the first audio channel;

[0069] S402. If it is predicted that the impact noise generated by the first audio channel is less than a preset threshold, determine the sound source switching strategy as: perform sound source switching in the first audio channel by using the electronic volume adjustment method;

[0070] S403. If it is predicted that the impact noise generated by the first audio channel is greater than a preset threshold, determine the sound source switching strategy as: according to the impact noise prediction results of the respective audio channels, select a second audio channel from the respective audio channels, and switch the target sound source to the second audio channel.

[0071] The embodiments of the present application provide three sound source switching strategies, namely, normal sound source switching, sound source switching using electronic volume adjustment, and sound source switching using other audio channels. According to the prediction result output by the impact sound prediction model, the embodiments of the present application select from the above three sound source switching strategies. When it is predicted that no impact sound will be generated in the first audio channel, no additional impact sound elimination operation is required, and the sound source can be switched normally in the first audio channel; when it is predicted that the impact sound generated in the first audio channel is less than a preset threshold, it means that the impact sound can be eliminated by means of electronic volume adjustment. At this time, the sound source is still switched in the first audio channel; when it is predicted that the impact sound generated in the first audio channel is greater than the preset threshold, it means that the impact sound is difficult to eliminate in the current audio channel, and continuing to switch the sound source will seriously affect the sound quality experience and comfort of the listener. Therefore, according to the impact sound prediction results of each audio channel, a second audio channel for sound source switching is selected, and then the audio switching of the target sound source is performed in the second audio channel. By setting three different sound source switching strategies, the embodiments of the present application enable the power amplifier impact sound elimination method provided by the embodiments of the present application to adapt to different audio environments and requirements. At the same time, this solution has the characteristics of high automation, reduces the operation process of manual intervention, reduces the maintenance cost, and improves the flexibility and convenience of power amplifier impact sound elimination.

[0072] In a preferred embodiment, a partition output switch control strategy can also be used to eliminate impact sounds based on the impact sound prediction result. During the sound source switching process, seamless connection of the old and new sound sources in terms of time and space is ensured, and impact sounds caused by signal interruption or overlap are avoided. Specifically, the audio output device is divided into multiple regions or partitions, and each partition corresponds to a different playback environment and audience group. For example, a meeting room can be divided into a host area, an audience area, a remote video conferencing area, etc. According to the combination of audio partitions, volume groups, and audio types, an audio routing strategy is configured. Determine which channels of audio signals each partition should receive, and how to adjust parameters such as volume and sound quality. Monitor the currently playing application or sound source in real time, and dynamically adjust the audio routing strategy according to its type and importance. For example, in a video conference, the host's voice can be automatically transmitted to all partitions with priority, and the volume is adjusted to ensure clarity. According to the audio routing strategy, control the output switches of each partition. When the impact sound prediction model predicts that an impact sound may occur, the partition output switch control strategy quickly switches the output signal of the partition to achieve smooth transition and effective elimination of the impact sound.

[0073] In a possible implementation manner, in step S402, the sound source switching using the electronic volume adjustment method in the first audio channel includes:

[0074] Determine the volume adjustment moment for the electronic volume adjustment according to the moment when the impact sound is generated in the first audio channel;

[0075] At the volume adjustment moment, use the cross-fading technique for the first audio channel to gradually reduce the volume of the current sound source of the first audio channel and gradually increase the volume of the target sound source at the same time.

[0076] In the embodiment of the present application, when the impact sound generated in the first audio channel is less than a preset threshold, the cross-fading technique is used to eliminate the impact sound during the sound source switching process. The key to implementing the cross-fading technique lies in how to determine the volume adjustment moment. If the volume adjustment moment is too early, the old sound source has not finished playing yet, which is equivalent to switching to the new sound source in advance, bringing a poor experience to the listener. If the volume adjustment moment is too late, there is still a risk of generating an impact sound. Therefore, in the embodiment of the present application, the impact sound prediction model predicts the moment when the impact sound is generated in the first audio channel, and then uses this as a reference to determine the volume adjustment moment for the electronic volume adjustment. The cross-fading technique is used before the impact sound is expected to occur, improving the accuracy of implementing the cross-fading technique and further improving the user experience.

[0077] In a possible implementation manner, in step S404, the selecting the second audio channel from each of the audio channels according to the impact sound prediction results of each of the audio channels and switching the target sound source to the second audio channel includes:

[0078] Evaluate the anti-impact sound level of each of the audio channels according to the impact sound prediction results of each of the audio channels and the audio signal characteristics of each of the audio channels;

[0079] Perform audio channel selection among each of the audio channels according to the anti-impact sound level of each of the audio channels, the priority of the currently playing sound source of each of the audio channels, and the priority of the target sound source to determine the second audio channel;

[0080] Switch the target sound source to the second audio channel.

[0081] In the embodiment of the present application, first evaluate the anti-impact sound level of each audio channel, and then combine the priorities of the target sound source and other sound sources to determine the second audio channel. For example, for important speeches or music performances, a higher priority can be set in advance to ensure the normal playback of the sound source. At the same time, according to the evaluation result of the anti-impact sound level of the audio channel, intelligently select the second audio channel for which the sound source should be switched currently. When switching the sound source, give priority to switching to the channel with predicted no impact sound or less impact sound and the audio channel has no currently playing sound source or the sound source priority is relatively low.

[0082] In a possible implementation manner, the impact sound prediction model constructed based on a long short-term memory network and trained through a historical audio dataset includes:

[0083] Obtain a historical audio dataset under various sound source switching scenarios;

[0084] Extract several Mel-frequency cepstral coefficient features from the historical audio dataset;

[0085] Arrange each of the Mel-frequency cepstral coefficient features in chronological order to construct several feature datasets;

[0086] Use each of the feature datasets to train a preset long short-term memory network to obtain the impact sound prediction model.

[0087] The embodiments of the present application provide a training method for an impact sound prediction model. By constructing a multi-layer LSTM (Long Short-Term Memory) network and combining MFCC (Mel Frequency Cepstral Coefficients) feature extraction, it can accurately predict the switching timing of the sound source signal. The LSTM (Long Short-Term Memory) network structure includes an input layer, a hidden layer (which may include multiple LSTM units), and an output layer. Through its unique gating mechanism, it can effectively handle the long-term and short-term dependencies in time series data. Among them, the forget gate can determine which old audio features (which may represent background noise or previous sound source information) should be forgotten at the current time step, so that the LSTM model can focus on the features of the current impact sound. For the update of the memory cell state, through the interaction between the input gate and the memory cell, the LSTM model can selectively add new information to the memory cell state, and allows the LSTM model to learn and retain the key features of the impact sound, which are used to eliminate the impact sound in subsequent time steps. By the forget gate and updating the memory cell state, the LSTM can capture the long-term dependencies in the audio signal. The hidden layer of the LSTM not only contains the input information of the current time step, but also integrates the memory information of the previous time step. Through the hidden layer, the LSTM can capture the short-term dependencies in the audio signal, such as the association between the impact sound and the audio frames before and after it. And the output gate controls which information is transferred from the memory cell state to the hidden layer and finally serves as the output of the current time step. So the output gate can selectively output the feature information related to the impact sound elimination as needed. The LSTM network can capture the unique timing pattern and frequency characteristics of the impact sound through its gating mechanism. These feature information may be manifested as sudden amplitude changes in the audio signal, an increase in high-frequency components, or a sudden increase in energy in a specific frequency band, etc. At the same time, these feature information play a key role in the process of impact sound elimination, because they are used to guide how to modify or replace the impact sound part in the original audio frame after training the LSTM model. Therefore, the finally obtained impact sound prediction model can adapt to the changes in different environmental noises and sound source characteristics and achieve more accurate impact sound prediction.

[0088] In a preferred embodiment, acquiring a historical audio data set under various sound source switching scenarios; using each of the feature data sets to train a preset long short-term memory network to obtain the impact sound prediction model, specifically:

[0089] Collect a historical audio dataset containing sound source switching scenarios from each audio channel. The dataset includes target sound sources and environmental noise. The target sound sources include, for example, human voices, musical instrument sounds, electronic sound effects, etc. The environmental noise includes background noise, echoes, reverberations, etc. These audio datasets contain the impact sounds that may occur during sound source switching. In the historical audio dataset, the time points of sound source switching need to be clearly marked. These marks are used in the subsequent training process to help the LSTM model learn the changing characteristics of audio signals during sound source switching. At the same time, for audio segments containing impact sounds, they also need to be marked. These marks tell the model which parts of the audio are impact sounds, so that the model can learn the characteristics of impact sounds and identify impact sounds in future predictions. The audio dataset should contain data with various sound source characteristics so that the model can adapt to and accurately predict the impact sounds during different sound source switches.

[0090] Further, extracting a number of Mel-frequency cepstral coefficient features from the historical audio dataset includes:

[0091] Perform noise removal, echo cancellation, pre-emphasis, framing, and windowing operations on the historical audio dataset to obtain a number of audio frames;

[0092] Perform Fourier transform (FFT), calculate the energy spectrum, and Mel filter bank conversion operations on each of the audio frames to obtain corresponding a number of log energies;

[0093] Extract the first N coefficients from each of the log energies through discrete cosine transform as each of the Mel-frequency cepstral coefficient features.

[0094] In a preferred embodiment, arranging each of the Mel-frequency cepstral coefficient features in chronological order to construct a number of feature datasets, specifically:

[0095] Arrange the previously extracted Mel-frequency cepstral coefficient (MFCC) features in chronological order to form a feature dataset. Each feature dataset will be given a clear label indicating whether the dataset contains impact sounds. At the same time, divide the marked feature dataset into a training set, a validation set, and a test set. The training set is used to train the LSTM model, the validation set is used to adjust the model parameters and prevent overfitting, and the test set is used to evaluate the performance and generalization ability of the model. This dataset is used to train the LSTM model so that it can learn and identify the characteristics of impact sounds that may occur during sound source switching, and improve the prediction accuracy of impact sounds. And adopt optimization strategies such as backpropagation algorithm and gradient descent method to continuously adjust the model parameters to improve the prediction accuracy and generalization ability of the model.

[0096] In the embodiment of the present application, the extracted Mel Frequency Cepstral Coefficient (MFCC) features are used as the input of the LSTM (Long Short-Term Memory) model, and the sequence learning ability of the LSTM (Long Short-Term Memory) is utilized to analyze the long-term and short-term dependencies in the audio signal. The LSTM model predicts whether a shock sound will be generated during the current sound source switching based on the input MFCC features.

[0097] Embodiment 2

[0098] As Figure 3 shown, correspondingly, Embodiment 2 provides a power amplifier shock sound cancellation system based on deep learning, including a monitoring module 10, an acquisition module 20, a prediction module 30, a strategy determination module 40, and a sound source switching module 50;

[0099] Among them, the monitoring module 10 is used to monitor the audio signals of each audio channel to confirm whether each audio channel is about to perform sound source switching;

[0100] The acquisition module 20 is used to obtain the target sound source data to be switched and the audio data sets corresponding to each audio channel when it is confirmed that there is a first audio channel about to perform sound source switching. The audio data set includes the sound source data currently played by each audio channel and the ambient noise data;

[0101] The prediction module 30 is used to input each of the audio data sets and the target sound source data into a preset shock sound prediction model, so that the shock sound prediction model outputs the shock sound prediction results of each audio channel according to each of the audio data sets. The shock sound prediction results include the magnitude of the shock sound generated by the audio channel and the moment when the shock sound is generated. The shock sound prediction model is constructed based on a long short-term memory network and trained through a historical audio data set;

[0102] The strategy determination module 40 is used to determine a sound source switching strategy according to the shock sound prediction results of each audio channel;

[0103] The sound source switching module 50 is used to perform the sound source switching according to the sound source switching strategy.

[0104] Furthermore, as Figure 4 shown, the strategy determination module 40 includes a first strategy determination unit 401, a second strategy determination unit 402, and a third strategy determination unit 403;

[0105] Among them, the first strategy determination unit 401 is used to determine the sound source switching strategy as: normally performing sound source switching in the first audio channel if it is predicted that the first audio channel will not generate a shock sound;

[0106] The second policy determination unit 402 is configured to determine the sound source switching policy as: performing sound source switching in the first audio channel by using electronic volume adjustment if it is predicted that the impact sound generated by the first audio channel is less than a preset threshold.

[0107] The third policy determination unit 403 is configured to determine the sound source switching policy as: selecting a second audio channel from each of the audio channels according to the impact sound prediction results of each audio channel, and switching the target sound source to the second audio channel if it is predicted that the impact sound generated by the first audio channel is greater than a preset threshold.

[0108] In a possible implementation manner, performing sound source switching in the first audio channel by using electronic volume adjustment includes:

[0109] Determining a volume adjustment time for performing the electronic volume adjustment according to the time when the impact sound is generated in the first audio channel;

[0110] At the volume adjustment time, applying a cross-fade technique to the first audio channel to gradually reduce the volume of the current sound source in the first audio channel and gradually increase the volume of the target sound source at the same time.

[0111] In a possible implementation manner, selecting a second audio channel from each of the audio channels according to the impact sound prediction results of each audio channel and switching the target sound source to the second audio channel includes:

[0112] Evaluating the anti-impact sound level of each audio channel according to the impact sound prediction results of each audio channel and the audio signal characteristics of each audio channel;

[0113] Selecting an audio channel from each of the audio channels to determine the second audio channel according to the anti-impact sound level of each audio channel, the priority of the current playing sound source of each audio channel, and the priority of the target sound source;

[0114] Switching the target sound source to the second audio channel.

[0115] The embodiment of the present application provides a power amplifier impact sound elimination system based on deep learning. When it is monitored that there is a first audio channel about to switch the sound source, the impact sound prediction model constructed based on the long short-term memory network is used to predict the impact sound of each audio channel, determine the magnitude of the impact sound generated by each audio channel and the moment when the impact sound is generated, realize the accurate prediction of the sound source signal switching timing, provide a decision-making basis for the subsequent elimination of the power amplifier impact sound, and finally automatically select the sound source switching strategy according to the impact sound prediction result, so that the embodiment of the present application can adapt to different audio environments and requirements. Compared with the existing power amplifier impact sound elimination methods, the power amplifier impact sound elimination method based on deep learning provided by the embodiment of the present application can predict the impact sound in advance, and then perform relevant power amplifier impact sound elimination operations before the impact sound occurs, effectively eliminating the power amplifier impact sound in the sound source switching scenario, and improving the sound quality and user experience of the broadcast system.

[0116] The more detailed working principle and step flow of this embodiment can but are not limited to refer to the relevant records of Embodiment 1.

[0117] The specific embodiments described above have further elaborated on the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for eliminating the impact sound of a power amplifier based on deep learning, characterized in that, Including: Monitoring the audio signals of each audio channel to confirm whether each audio channel is about to perform a sound source switch; When it is confirmed that there is a first audio channel about to perform a sound source switch, obtaining the target sound source data to be switched and the audio data sets corresponding to each audio channel respectively, where the audio data set includes the sound source data currently played by each audio channel and the ambient noise data; Inputting each of the audio data sets and the target sound source data into a preset impact sound prediction model, so that the impact sound prediction model outputs the impact sound prediction results of each audio channel according to each of the audio data sets, where the impact sound prediction result includes the magnitude of the impact sound generated by the audio channel and the moment when the impact sound is generated, and the impact sound prediction model is constructed based on a long short-term memory network and obtained through training with a historical audio data set; Determining a sound source switching strategy according to the impact sound prediction results of each audio channel. Among them, if it is predicted that the impact sound generated by the first audio channel is greater than a preset threshold, then it is determined that the sound source switching strategy is: according to the impact sound prediction results of each audio channel, selecting a second audio channel among each audio channel and switching the target sound source to the second audio channel, including: evaluating the anti-impact sound level of each audio channel according to the impact sound prediction results of each audio channel and the audio signal characteristics of each audio channel; selecting an audio channel among each audio channel according to the anti-impact sound level of each audio channel, the priority of the currently played sound source of each audio channel and the priority of the target sound source to determine the second audio channel; switching the target sound source to the second audio channel; Performing the sound source switch according to the sound source switching strategy.

2. The method for eliminating the impact sound of a power amplifier based on deep learning according to claim 1, characterized in that, The determining the sound source switching strategy according to the impact sound prediction results of each audio channel includes: If it is predicted that the first audio channel will not generate an impact sound, then it is determined that the sound source switching strategy is: normally performing a sound source switch in the first audio channel; If it is predicted that the impact sound generated by the first audio channel is less than a preset threshold, then it is determined that the sound source switching strategy is: performing a sound source switch in the first audio channel by using an electronic volume adjustment method.

3. The method for eliminating the impact sound of a power amplifier based on deep learning according to claim 2, characterized in that, The performing a sound source switch in the first audio channel by using an electronic volume adjustment method includes: Determining the volume adjustment moment for performing the electronic volume adjustment according to the moment when the impact sound is generated in the first audio channel; At the volume adjustment moment, adopting a cross-fading technique for the first audio channel to gradually reduce the volume of the current sound source of the first audio channel and gradually increase the volume of the target sound source at the same time.

4. A method for eliminating amplifier impact noise based on deep learning according to any one of claims 1-3, characterized in that, The constructing the impact sound prediction model based on a long short-term memory network and obtained through training with a historical audio data set includes: Obtaining a historical audio data set in multiple sound source switching scenarios; Extracting several Mel-frequency cepstral coefficient features from the historical audio data set; Arranging each of the Mel-frequency cepstral coefficient features in chronological order to construct several feature data sets; Train a preset long short - term memory network using the several feature datasets to obtain the impact sound prediction model.

5. The method for eliminating the impact sound of a power amplifier based on deep learning according to claim 4, characterized in that Extract several Mel - Frequency Cepstral Coefficient (MFCC) features from the historical audio dataset, including: Perform noise removal, echo cancellation, pre - emphasis, framing, and windowing operations on the historical audio dataset to obtain several audio frames; Perform Fourier Transform (FFT), calculate the energy spectrum, and Mel filter bank conversion operations on each of the audio frames to obtain corresponding several logarithmic energies; Extract the first N coefficients from each of the logarithmic energies through discrete cosine transform as each of the Mel - Frequency Cepstral Coefficient features.

6. A power amplifier impact sound elimination system based on deep learning, characterized in that, It includes a monitoring module, an acquisition module, a prediction module, a strategy determination module, and a sound source switching module; Among them, the monitoring module is used to monitor the audio signals of each audio channel to confirm whether each audio channel is about to perform a sound source switch; The acquisition module is used to, when it is confirmed that there is a first audio channel about to perform a sound source switch, acquire the target sound source data to be switched and the audio dataset corresponding to each audio channel, and the audio dataset includes the sound source data currently played by each audio channel and the ambient noise data; The prediction module is used to input each of the audio datasets and the target sound source data into a preset impact sound prediction model, so that the impact sound prediction model outputs the impact sound prediction results of each audio channel according to each of the audio datasets. The impact sound prediction results include the magnitude of the impact sound generated by the audio channel and the moment when the impact sound is generated. The impact sound prediction model is constructed based on a long short - term memory network and trained using a historical audio dataset; The strategy determination module is used to determine a sound source switching strategy according to the impact sound prediction results of each audio channel, including a first strategy determination unit, a second strategy determination unit, and a third strategy determination unit. Among them, the third strategy determination unit is used to, if it is predicted that the impact sound generated by the first audio channel is greater than a preset threshold, determine the sound source switching strategy as: according to the impact sound prediction results of each audio channel, select a second audio channel among each audio channel and switch the target sound source to the second audio channel, including: evaluating the anti - impact sound level of each audio channel according to the impact sound prediction results of each audio channel and the audio signal characteristics of each audio channel; selecting an audio channel among each audio channel according to the anti - impact sound level of each audio channel, the priority of the currently played sound source of each audio channel, and the priority of the target sound source to determine the second audio channel; switching the target sound source to the second audio channel; The sound source switching module is used to perform the sound source switch according to the sound source switching strategy.

7. The power amplifier impact noise cancellation system based on deep learning according to claim 6, characterized in that, The strategy determination module includes a first strategy determination unit, a second strategy determination unit, and a third strategy determination unit; Among them, the first strategy determination unit is used to, if it is predicted that the first audio channel will not generate an impact sound, determine the sound source switching strategy as: perform a sound source switch normally in the first audio channel; The second strategy determination unit is configured to determine the sound source switching strategy as follows: if the predicted impact sound generated by the first audio channel is less than a preset threshold, perform sound source switching by using electronic volume adjustment in the first audio channel.

8. The power amplifier impact sound elimination system based on deep learning according to claim 7, characterized in that, The performing sound source switching by using electronic volume adjustment in the first audio channel includes: Determining a volume adjustment time for performing the electronic volume adjustment according to the time when the impact sound is generated in the first audio channel; At the volume adjustment time, adopting a cross-fading technique for the first audio channel to gradually reduce the volume of the current sound source in the first audio channel and gradually increase the volume of the target sound source at the same time.

Citation Information

Patent Citations

  • Audio processing method and device, equipment and storage medium

    CN114724581A

  • Audio enhancement and hearing protection

    GB0723908D0