Audio signal processing method and device, audio processing equipment and storage medium
By obtaining room structure information and calculating reverberation time, the microphone array filter coefficient is optimized, and the adaptive anti-reverberation problem of audio processing equipment under different room structures is solved, improving the audio signal quality and translation effect.
Patent Information
- Application Number
- CN202510405240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, it is difficult for audio processing equipment to adaptively perform anti-reverb processing under different room structures, resulting in poor audio signal quality.
By calling the radar module to obtain room structure information, calculate the reverberation time, determine the initial filter coefficient of the microphone array, and optimize the filter coefficient based on the current room's reverberation time and audio signal to achieve adaptive anti-reverberation.
It improves the anti-reverberation capability of audio signal processing equipment under different room structures, and improves the audio signal quality and translation effect.
Smart Images

Figure CN120264183A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio signal processing, and in particular, to an audio signal processing method, apparatus, audio processing device, and storage medium. Background Art
[0002] After obtaining an audio signal to be processed, an audio processing device usually uses some methods to perform anti-reverberation to improve the audio quality. Exemplarily, in a conference scenario, a voice pickup and transcription device can transcribe the audio (such as converting the audio into text to generate a conference record); in this conference scenario, for the collected audio signal, the voice pickup and transcription device needs to perform anti-reverberation processing, and then transcribe the anti-reverberation processed audio signal to improve the transcription quality.
[0003] Currently, when performing anti-reverberation on an audio signal, it is necessary to consider the room reverberation of the room where the audio processing device is located; however, the room reverberations of different room structures are different.
[0004] It can be seen that how to perform adaptive anti-reverberation on an audio signal according to the room structure of the room where the audio processing device is located is an urgent problem to be solved. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide an audio signal processing method, apparatus, audio processing device, and storage medium to perform adaptive anti-reverberation on an audio signal according to the room structure of the room where the audio processing device is located. The specific technical solutions are as follows:
[0006] In a first aspect, the embodiments of this application provide an audio signal processing method applied to an audio processing device, and the method includes:
[0007] Invoking a radar module to transmit a radar wave to the current room and receive a target reflection signal corresponding to the radar wave;
[0008] Determining the room structure information calculated based on the radar wave and the target reflection signal as the structure information of the current room;
[0009] Determining the reverberation time of the current room according to the structure information of the current room;
[0010] Determining an initial filtering coefficient of the microphone array of the audio processing device according to the reverberation time of the current room;
[0011] Transmitting a first audio signal to the current room and receiving a second audio signal reflected by the current room;
[0012] Optimizing the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient;
[0013] Set the filtering coefficients of the microphone array to the target filtering coefficients. After the setting is completed, in response to the acquisition of the original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficients of the microphone array to obtain the target audio signal.
[0014] Optionally, the determining the initial filtering coefficients of the microphone array of the audio processing device according to the reverberation time of the current room includes:
[0015] Based on the first enhancement network, determine the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room; wherein, the first enhancement network is a network for performing anti-reverberation processing on any audio signal, and the initial filtering coefficients are: the filtering coefficients required when the first enhancement network performs anti-reverberation on the audio signal generated in the current room according to the reverberation time of the current room.
[0016] Optionally, before calling the radar module to emit radar waves to the current room and receive the target reflection signal corresponding to the radar waves, the method further includes:
[0017] Detect whether a target instruction is received; wherein, the target instruction is used to represent an instruction for reusing the currently set filtering coefficients of the microphone array;
[0018] If not received, perform the step of calling the radar module to emit radar waves to the current room and receive the target reflection signal corresponding to the radar waves;
[0019] If received, in response to the acquisition of the original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficients of the microphone array to obtain the target audio signal.
[0020] Optionally, the determining the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room based on the first enhancement network includes:
[0021] Input the reverberation time of the current room into the first enhancement network to obtain the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room;
[0022] Wherein, the training method of the first enhancement network includes:
[0023] For each sample group, input the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained, so that the first enhancement network processes the sample audio signal using the current network parameters to obtain an audio signal output result corresponding to the sample audio signal; wherein, each sample group includes a reverberation time of the samples, a sample audio signal, and a true-value audio signal; the sample audio signal is an audio signal obtained by adding a target reverberation feature to the true-value audio signal in the sample group, and the target reverberation feature is: if the true-value audio signal is played in a room structure with a reverberation time of the reverberation time of the samples, the reverberation feature existing in the signal collected by collecting the played signal; the network parameters include filter coefficients required for anti-reverberation processing.
[0024] Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, determine the current network loss of the first enhancement network.
[0025] In response to the current network loss of the first enhancement network indicating that the first enhancement network has not converged, adjust the current network parameters of the first enhancement network, and return to the step of inputting the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained to obtain an audio signal output result corresponding to the sample audio signal.
[0026] In response to the current network loss of the first enhancement network indicating that the first enhancement network has converged, based on the current network parameters, determine the filter coefficient corresponding to the reverberation time of the samples in the sample group.
[0027] Optionally, the determining the current network loss of the first enhancement network based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal includes:
[0028] Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, calculate a power-law compression loss to obtain the current network loss of the first enhancement network.
[0029] Optionally, the optimizing the initial filter coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filter coefficient includes:
[0030] Determine the sum value of the specified product and the initial filter coefficient to obtain the target filter coefficient; wherein, the specified product is the product of the first audio signal, the error signal, and the specified step factor, the error signal is used to characterize the difference between the first audio signal and the second audio signal, the specified step factor is the ratio of the specified proportionality coefficient to the reverberation time of the current room, and the specified proportionality coefficient is a predetermined empirical coefficient used for anti-reverberation of any audio signal.
[0031] Optionally, the method further includes:
[0032] Based on the second enhancement network, perform audio signal enhancement on the target audio signal to obtain the enhanced target audio signal; wherein, the second enhancement network is a network used for signal enhancement processing of any audio signal.
[0033] Optionally, the determining the reverberation time of the current room according to the structural information of the current room includes:
[0034] Substitute the structural information of the current room into the predetermined Sabine formula to obtain the reverberation time of the current room.
[0035] In a second aspect, an embodiment of the present application provides an audio signal processing device, which is applied to an audio processing device, and the device includes:
[0036] A calling module, configured to call a radar module to emit radar waves to the current room and receive a target reflection signal corresponding to the radar waves;
[0037] A first determination module, configured to determine the structural information of the room calculated based on the radar waves and the target reflection signal as the structural information of the current room;
[0038] A second determination module, configured to determine the reverberation time of the current room according to the structural information of the current room;
[0039] A third determination module, configured to determine the initial filter coefficient of the microphone array of the audio processing device according to the reverberation time of the current room;
[0040] A transmitting module, configured to transmit a first audio signal to the current room and receive a second audio signal reflected by the current room;
[0041] An optimization module, configured to optimize the initial filter coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filter coefficient;
[0042] A setting module is configured to set the filtering coefficients of the microphone array to the target filtering coefficients. After the setting is completed, in response to the acquisition of the original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficients of the microphone array to obtain the target audio signal.
[0043] In a third aspect, an embodiment of the present application provides an audio processing device, including: a radar module, a microphone array, a processor, and a memory;
[0044] The radar module is configured to transmit a radar wave to the current room and receive the target reflection signal corresponding to the radar wave under the call of the processor;
[0045] The microphone array is configured to collect audio signals in the current room;
[0046] The memory is configured to store a computer program;
[0047] The processor is configured to implement any of the audio signal processing methods when executing the program stored on the memory.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the audio signal processing methods is implemented.
[0049] An embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above audio signal processing methods.
[0050] Advantageous effects of the embodiments of the present application:
[0051] The audio signal processing method provided by the embodiments of the present application enables an audio processing device to calculate the structural information of the current room by invoking a radar module, and determine the reverberation time of the current room based on the structural information of the current room. Then, based on the reverberation time of the current room, the initial filtering coefficients of the microphone array of the audio processing device are determined. Moreover, the present application can also transmit a first audio signal to the current room and receive a second audio signal reflected by the current room, that is, the second audio signal is an audio signal with the room reverberation of the current room added to the first audio signal. Thus, considering the room reverberation of the current room, the present application further optimizes the initial filtering coefficients according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain target filtering coefficients. After setting the filtering coefficients of the microphone array to the target filtering coefficients, anti-reverberation can be performed on the original audio signal generated in the current room based on the current filtering coefficients of the microphone array to obtain a target audio signal. It can be seen that the target filtering coefficients are filtering coefficients that further optimize the initial filtering coefficients obtained based on the reverberation time of the current room by combining the reverberation time of the current room and the difference between the first audio signal and the second audio signal (i.e., the room reverberation of the current room). Then, when performing anti-reverberation based on the target filtering coefficients, the adaptability to the room reverberation formed by the room structure of the current room is relatively high. It can be seen that through the present application, adaptive anti-reverberation can be performed on audio signals according to the room structure of the room where the audio processing device is located.
[0052] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0054] Figure 1 It is a schematic flowchart of an audio signal processing method provided by an embodiment of the present application;
[0055] Figure 2 It is a schematic scenario diagram of an audio signal processing method provided by an embodiment of the present application;
[0056] Figure 3 It is another schematic flowchart of an audio signal processing method provided by an embodiment of the present application;
[0057] Figure 4 It is a schematic structural diagram of an audio signal processing device provided by an embodiment of the present application;
[0058] Figure 5 Schematic diagram of an audio processing device provided by an embodiment of the present application. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the protection scope of the present application.
[0060] With the development of technologies related to mobile terminals, more and more intelligent devices have emerged in people's lives, bringing great convenience to people. Currently, there are many audio processing devices for processing audio signals. For example, a voice pickup and transcription device applied in a conference scenario for conference minutes. In this conference scenario, the quality of the audio signal plays a decisive role in the transcription effect, that is, the better the quality of the audio signal, the better the transcription effect of the audio processing device on the audio signal. Among them, the quality of the audio signal is often affected by room reverberation and echo. Usually, after receiving the audio signal (i.e., after voice pickup), the audio processing device uses some traditional or specific algorithms to perform anti-reverberation and echo cancellation. However, the room reverberation in different room structures is different, and the influence of different room reverberations on the audio signal is also different.
[0061] There is a method in the prior art that calculates the reverberation characteristics from the real-time audio signal received by the microphone array, and then matches the reverberation characteristics with the pre-trained model data to obtain the corresponding reverberation coefficient, so as to optimize the filtering coefficient of the microphone array for anti-reverberation and enhance the audio signal. There is also another method in the prior art that calculates the spatial impulse response from the real-time audio signal received by the adaptive filter and uses the exponential decay model to obtain the key factor (reverberation time) for measuring the degree of room reverberation for subsequent optimization.
[0062] The overall idea of the above two methods is to calculate the reverberation characteristics by receiving real-time audio signals, and then match the reverberation characteristics with the training data, so as to adjust the corresponding anti-reverberation filter coefficients to enhance the audio signals. The starting point is to calculate the reverberation characteristics by receiving real-time audio signals, and then through model matching, make the audio processing device (such as a sound pickup device, a device capable of collecting audio signals) have the anti-reverberation ability to adapt to the room reverberation, and avoid the cumbersome steps of manual measurement through certain means. The above two methods enhance the audio signals to a certain extent according to the room reverberation characteristics, but the room reverberation characteristics in different rooms are different; in an actual meeting scenario, the room where the audio processing device is located may change. When the room changes, the participants need to speak under different room structures, and the audio signals need to be adaptively anti-reverberated according to the room reverberation under different room structures. Obviously, the room reverberation characteristics of the room where the audio processing device is located are related to the specific structure of the room. If the prior knowledge of the room structure information of this room can be obtained, it will be beneficial to the calculation of the room reverberation.
[0063] Based on this, the embodiments of the present application provide an audio signal processing method, device, audio processing device and storage medium to adaptively anti-reverberate the audio signal according to the room structure of the audio processing device.
[0064] First, an audio processing method provided by the present application will be introduced below.
[0065] Among them, an audio processing method provided by the embodiments of the present application is applied to an audio processing device. The audio processing device can receive audio signals and perform signal enhancement, forwarding, transcription and other processing on the audio signals; the audio processing device can have a microphone array to pick up sound (that is, obtain audio signals) and perform audio processing such as audio signal enhancement through the microphone array, etc.; for example, the audio processing device can be a microphone device or an electronic device with a microphone array (such as a mobile phone, a computer, etc.). The present application does not limit the specific form of the audio processing device.
[0066] In addition, the audio processing device in the present application can also be embedded or connected with a radar module. The radar module can be a radar device, etc. When the audio processing device is embedded with a radar module, the main control module of the audio processing device can communicate with the radar module and can call the radar module; when the audio processing device is connected with an external radar module, the radar module can be connected to the audio processing device through a connection line, and the room where the audio processing device is located should be the same as the room where the radar module is located.
[0067] An audio signal processing method provided by the embodiments of the present application is applied to an audio processing device, including the following steps:
[0068] Invoke the radar module to emit radar waves to the current room and receive the target reflection signals corresponding to the radar waves;
[0069] Determine the room structure information calculated based on the radar waves and the target reflection signals as the structure information of the current room;
[0070] Determine the reverberation time of the current room according to the structure information of the current room;
[0071] Determine the initial filtering coefficient of the microphone array of the audio processing device according to the reverberation time of the current room;
[0072] Emit a first audio signal to the current room and receive a second audio signal reflected by the current room;
[0073] Optimize the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient;
[0074] Set the filtering coefficient of the microphone array to the target filtering coefficient. After the setting is completed, in response to the acquisition of the original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficient of the microphone array to obtain a target audio signal.
[0075] The audio signal processing method provided by the embodiments of the present application enables an audio processing device to calculate the structural information of the current room by invoking a radar module, and determine the reverberation time of the current room based on the structural information of the current room; then, determine the initial filtering coefficients of the microphone array of the audio processing device according to the reverberation time of the current room; moreover, the present application can also transmit a first audio signal to the current room and receive a second audio signal reflected by the current room, that is, the second audio signal is an audio signal with the room reverberation of the current room added to the first audio signal. Thus, considering the room reverberation of the current room, the present application further optimizes the initial filtering coefficients according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain target filtering coefficients; after setting the filtering coefficients of the microphone array to the target filtering coefficients, anti-reverberation can be performed on the original audio signal generated in the current room based on the current filtering coefficients of the microphone array to obtain a target audio signal. It can be seen that the target filtering coefficients are filtering coefficients that further optimize the initial filtering coefficients obtained based on the reverberation time of the current room by combining the reverberation time of the current room and the difference between the first audio signal and the second audio signal (i.e., the room reverberation of the current room); then, when performing anti-reverberation based on the target filtering coefficients, the adaptability to the room reverberation formed by the room structure of the current room is relatively high. It can be seen that through the present application, audio signals can be adaptively anti-reverberated according to the room structure of the room where the audio processing device is located.
[0076] The following introduces a kind of audio signal processing method provided by the embodiments of the present application with reference to the accompanying drawings.
[0077] As Figure 1 shown, a kind of audio signal processing method provided by the embodiments of the present application, which is applied to an audio processing device, may include the following steps:
[0078] S101: Invoke the radar module to transmit radar waves to the current room and receive the target reflection signals corresponding to the radar waves;
[0079] In the embodiments of the present application, the structural information of the room where the audio processing device is currently located is measured by the radar module. Therefore, the radar module can be invoked first to transmit radar waves to the current room (i.e., transmit radar waves in all directions) and receive the target reflection signals corresponding to the radar waves, so as to determine the subsequent room structural information based on the target reflection signals.
[0080] Exemplarily, in one implementation, before the step of invoking the radar module to transmit radar waves to the current room and receive the target reflection signals corresponding to the radar waves, the method further includes:
[0081] Detect whether a target instruction is received; wherein, the target instruction is used to represent an instruction for multiplexing the current filtering coefficients set for the microphone array;
[0082] If not received, execute the step of calling the radar module to emit radar waves to the current room and receive the target reflection signals corresponding to the radar waves;
[0083] If received, in response to the original audio signal generated in the current room collected by the microphone array, perform reverberation suppression on the received original audio signal based on the current filtering coefficients of the microphone array to obtain a target audio signal.
[0084] In this application, considering that the room where the audio processing device is located may or may not change, and when it does not change, the filtering coefficients for reverberation suppression can be directly multiplexed. Therefore, it can be first detected whether a target instruction for multiplexing the current filtering coefficients set for the microphone array is received. If not received, it indicates that the room where the audio processing device is located has changed (for example: the current room is different from the previous room where the device was located), then the step of calling the radar module in this application to emit radar waves to the current room and receive the target reflection signals corresponding to the radar waves can be executed; if received, it indicates that the room where the audio processing device is located has not changed (for example: the current room is the same as the previous room where the device was located), then directly multiplex the current filtering coefficients set for the microphone array. When the original audio signal generated in the current room is collected by the microphone array, perform reverberation suppression on the received original audio signal through the current filtering coefficients of the microphone array to obtain a target audio signal, without having to determine the filtering coefficients corresponding to the room structure of the current room, so as to improve the efficiency of audio signal processing.
[0085] Among them, the target instruction can be issued by the user operating the audio processing device. For example, if the room where the user used the audio processing device last time is the same as the room where the user is using the audio processing device currently, the user can issue this target instruction (for example: the user clicks the button on the audio processing device representing multiplexing the filtering coefficients to issue this target instruction); or, the target instruction can be issued by other devices. The other devices are connected to the audio processing device, and the other devices can detect whether the room where the audio processing device is located has changed. For example, it can be detected whether the current room is the same as the previous room where the audio processing device was located, and it can be detected by means of image similarity, etc. If there is no change, the target instruction can be issued to the audio processing device.
[0086] In addition, whenever the audio processing device is turned on, step S101 can be executed by default, or whenever the audio processing device is initialized, step S101 is executed to measure the room structure information of the room where the audio processing device is currently located through the radar module. Of course, after the audio processing device is turned on or initialized, it is reasonable that the audio processing device may receive a target instruction.
[0087] S102: Determine the room structure information calculated based on the radar wave and the target reflection signal as the structure information of the current room;
[0088] After obtaining the target reflection signal, the audio processing device can calculate the room structure information based on the radar wave and the target reflection signal. For example, based on the time difference between the radar wave and the corresponding target reflection signal in the same direction, and the propagation speeds of the radar wave and the target reflection signal, the distance in this direction is calculated. Thus, the structure information such as the length, width, and height of the current room, the inner surface area of the room, and the room volume is calculated. Of course, other methods can also be used, such as frequency analysis of the radar wave and the target reflection signal, etc., to calculate the structure information of the current room. This application does not make any limitations in this regard. Any method that can calculate the structure information of the current room based on the radar wave and the target reflection signal can be applied to this application.
[0089] It should be noted that the radar module can calculate the room structure information based on the radar wave and the target reflection signal, and then feedback the calculated room structure information to the main control module of the audio processing device to obtain the structure information of the current room; or the main control module of the audio processing device can calculate the room structure information based on the radar wave and the target reflection signal to directly obtain the structure information of the current room. This application does not make any limitations in this regard.
[0090] In addition, the sound absorption coefficient of the current room can also be determined according to the material of the inner wall of the current room. For example, according to the empirical value (the corresponding relationship between the material and the sound absorption coefficient), the sound absorption coefficient of the material of the inner wall of the current room is determined as the sound absorption coefficient of the current room (for example: the sound absorption coefficient of the sound-absorbing board is greater than or equal to 0.2, the sound absorption coefficient of the sound insulation cotton is greater than 0.4, and the sound absorption coefficient of the floating partition is greater than 0.3); and the sound absorption coefficient of the current room is used as the auxiliary information of the current room for subsequent calculation of the reverberation time.
[0091] Among them, the sound absorption coefficient of the current room can also be determined according to measurement methods such as the reverberation chamber method, the standing wave tube method, the free field method, the two-microphone method, or the local plane wave method. The specific implementation method can be similar to the prior art and will not be elaborated here.
[0092] S103: Determine the reverberation time of the current room according to the structure information of the current room;
[0093] After obtaining the structural information of the current room, the reverberation time of the current room can be further calculated.
[0094] Exemplarily, in one implementation manner, determining the reverberation time of the current room according to the structural information of the current room includes:
[0095] Substitute the structural information of the current room into the predetermined Sabine formula to obtain the reverberation time of the current room.
[0096] This application uses the Sabine formula to calculate the reverberation time of the current room, that is, substituting the structural information of the current room (such as room volume, inner surface area of the room) into the predetermined Sabine formula, and, substituting the auxiliary information of the current room (absorption coefficient of the current room), to obtain the reverberation time of the current room.
[0097] The predetermined Sabine formula in this application can be: RT60 = 0.161×(V / AS), where V is the room volume, S is the inner surface area of the room, A is the absorption coefficient, and RT60 is the reverberation time.
[0098] In addition, for different application scenarios, a simplified Sabine formula or other formulas can also be used to calculate the reverberation time of the room; for example: for the case where uniform sound-absorbing materials are applied inside the room, the simplified Sabine formula can be used: RT60≈0.049×(V / AS); for the case where non-uniform sound-absorbing materials in the room need to be considered, the Eyring formula (a formula that modifies the Sabine formula) can be used: RT60 = 0.161×(V / AS - a), where a represents the non-uniformity correction term.
[0099] It should be noted that any method that can determine the reverberation time of the current room according to the structural information of the current room is applicable to this application. The above methods are only for exemplary illustration and should not constitute a limitation to this application.
[0100] S104: Determine the initial filtering coefficient of the microphone array of the audio processing device according to the reverberation time of the current room;
[0101] In order to perform anti-reverberation on the audio signal generated in the current room according to the room structure of the current room, the initial filtering coefficient of the microphone array of the audio processing device can be determined first according to the reverberation time of the current room.
[0102] Optionally, determining the initial filtering coefficient of the microphone array of the audio processing device according to the reverberation time of the current room includes:
[0103] Based on the first enhancement network, determine the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room; wherein, the first enhancement network is a network for performing anti-reverberation processing on any audio signal, and the initial filtering coefficients are: the filtering coefficients required when the first enhancement network performs anti-reverberation on the audio signal generated in the current room according to the reverberation time of the current room.
[0104] In this application, based on the first enhancement network for performing anti-reverberation processing on any audio signal, the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room can be determined.
[0105] Among them, the filtering coefficients can be understood as the parameters used for anti-reverberation of audio signals. Networks for processing audio signals, or filters (such as filters in a microphone array), can all use the filtering coefficients to perform anti-reverberation on audio signals.
[0106] It should be noted that the first enhancement network in this application can perform anti-reverberation on audio signals. Here, only the correspondence between the reverberation time and the filtering coefficients is concerned. For example: input the reverberation time of the current room into the first enhancement network to obtain the initial filtering coefficients of the microphone array of the audio processing device. The training process of the first enhancement network will be introduced in detail in subsequent embodiments and will not be elaborated here.
[0107] In this application, through the first enhancement network, the initial filtering coefficients corresponding to the reverberation time of the current room can be determined quickly and simply, thereby improving the speed and efficiency of subsequent anti-reverberation of audio signals.
[0108] In addition, the correspondence between the reverberation time of the room and the filtering coefficients can also be established in advance. The filtering coefficients corresponding to the reverberation time of the room are: the filtering coefficients required for anti-reverberation of the audio signal containing the reverberation characteristics of the room; the filtering coefficients corresponding to the reverberation time of the current room can be found from the pre-established correspondence to obtain the initial filtering coefficients; and this correspondence can be set by experience (i.e., setting the filtering coefficients corresponding to the reverberation time through empirical values), or obtained by calibration (i.e., calibrating the filtering coefficients corresponding to the reverberation time).
[0109] S105: Transmit a first audio signal to the current room and receive a second audio signal reflected by the current room;
[0110] In this application, in order to fully consider the actual room reverberation of the current room, the initial filtering coefficient can also be optimized in combination with the real-time audio signal; a first audio signal can be transmitted to the current room, and a second audio signal reflected by the current room can be received, and the received second audio signal can be used as an audio signal with the room reverberation characteristics of the current room added to the first audio signal.
[0111] In this application, the purpose of audio processing is to process the second audio signal into the first audio signal, that is, to process the audio signal with the room reverberation of the current room added into an audio signal without the room reverberation of the current room, that is, to perform anti-reverberation on the audio signal to remove the room reverberation of the current room. Therefore, by transmitting the first audio signal and the second audio signal, it is beneficial to the subsequent optimization of the initial filtering coefficient of the microphone array.
[0112] S106: Optimize the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient;
[0113] After transmitting and receiving the real-time first audio signal and the second audio signal, the initial filtering coefficient can be optimized to obtain a target filtering coefficient finally used for adaptive anti-reverberation of the audio signal.
[0114] In one implementation manner, the optimizing the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient includes:
[0115] Determine the sum value of a specified product and the initial filtering coefficient to obtain the target filtering coefficient; wherein, the specified product is the product of the first audio signal, an error signal, and a specified step factor, the error signal is used to characterize the difference between the first audio signal and the second audio signal, the specified step factor is the ratio of a specified proportional coefficient to the reverberation time of the current room, and the specified proportional coefficient is a predetermined empirical coefficient used for anti-reverberation of any audio signal.
[0116] When optimizing the initial filtering coefficient, the sum value of the specified product and the initial filtering coefficient can be used as the target filtering coefficient; that is, calculate the product of the ratio between the specified proportional coefficient and the reverberation time of the current room, the first audio signal, and the error signal, and its sum value with the initial filtering coefficient to optimize the initial filtering coefficient to obtain the target filtering coefficient. Among them, the specified proportional coefficient is a predetermined empirical coefficient used for anti-reverberation of any audio signal, the specified proportional coefficient can be 0.161, and can be adjusted according to the quality of the actual audio signal (such as the quality of the anti-reverberated audio signal), and this application does not limit this.
[0117] Exemplarily, the following formula can be used to optimize the initial filter coefficient to obtain the target filter coefficient:
[0118] h(n + 1) = h(n) + μ × e(n) × x(n), where h(n + 1) is the target filter coefficient; h(n) is the initial filter coefficient, x(n) is the first audio signal, e(n) is the error signal; μ is a specified step factor, μ = K / RT60, K is a specified proportionality coefficient (the empirical value of K can be 0.161 and can be adjusted according to the actual situation); RT60 is the reverberation time of the current room.
[0119] Among them, the error signal of the first audio signal and the second audio signal can be extracted first. For example: first preprocess the first audio signal and the second audio signal (such as: time alignment, denoising, equalization processing, etc.), and then calculate the difference between the preprocessed second audio signal and the preprocessed first audio signal to obtain the error signal; of course, other methods can also be used to extract the error signal of the first audio signal and the second audio signal, and the present application does not limit the extraction method of the error signal.
[0120] It should be emphasized that any method that can optimize the initial filter coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain the target filter coefficient is applicable to the present application, and the present application does not limit this. And the above formula for optimizing the initial filter coefficient is only an example. In practical applications, this formula can be modified, or other formulas can be used, and the present application does not limit this.
[0121] S107: Set the filter coefficient of the microphone array to the target filter coefficient. After the setting is completed, in response to the acquisition of the original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filter coefficient of the microphone array to obtain the target audio signal.
[0122] The target filter coefficient is a filter coefficient that further optimizes the initial filter coefficient of the reverberation time of the current room by combining the reverberation time of the current room and the difference between the first audio signal and the second audio signal (i.e., the error signal, which can characterize the room reverberation of the current room). Through the target filter coefficient, anti-reverberation adapted to the room reverberation of the current room can be performed on the audio signal generated in the current room; the filter coefficient of the microphone array can be set to the target filter coefficient, and after the setting is completed, when the original audio signal generated in the current room is acquired through the microphone array, anti-reverberation is performed on the original audio signal based on the current filter coefficient of the microphone array (i.e., the set target filter coefficient) to obtain the target audio signal for anti-reverberation against the room reverberation of the current room.
[0123] In addition, to further optimize the audio signal, the method further includes:
[0124] Based on a second enhancement network, enhancing the target audio signal to obtain an enhanced target audio signal; wherein, the second enhancement network is a network for performing signal enhancement processing on any audio signal.
[0125] This application can also enhance the target audio signal based on the second enhancement network to obtain an enhanced target audio signal. Among them, the second enhancement network can be a network that meets the user's needs and is used for performing signal enhancement processing on any audio signal, that is, the second enhancement network is a personalized signal enhancement network that meets the user's needs. Thus, the obtained enhanced target audio signal is not only an audio signal that cancels reverberation for the room reverberation of the current room, but also meets the user's personalized needs.
[0126] Exemplarily, the second enhancement network can be the same as the first enhancement network, or the second enhancement network is an enhancement network for signal equalization (which has the function of adjusting the frequency response of the audio signal, allowing the user to enhance or weaken the sound of a specific frequency band. For example, improving the overall audio quality by increasing the low frequency (bass) or decreasing the high frequency (treble)), an enhancement network for signal compression (used to control the dynamic range of the audio signal, automatically reducing overly strong signals and increasing weaker signals. By suppressing the volume peak, the audio can be made more balanced, avoiding overload and distortion, making the sound clearer and more penetrating), etc., and this application does not make any limitations in this regard.
[0127] In addition, by performing the above steps, it is equivalent to being able to determine the correspondence between the structural information of the current room and the corresponding target filtering coefficient. When the amount of data is sufficient, it is also possible to directly establish the correspondence between the structural information of the room and the target filtering coefficient, so that after obtaining the structural information of any room, the target filtering coefficient corresponding to the structural information of the room can be directly determined, thereby canceling the reverberation of the audio signal generated under the structural information of the room.
[0128] In the technical solution of this application, operations such as the acquisition, storage, use, processing, transmission, provision, and disclosure of the structural information of the room, reverberation time, filtering coefficient, and various audio signals involved are all carried out under the premise of obtaining the user's authorization.
[0129] The audio signal processing method provided by the embodiment of the present application. The audio processing device can calculate the structural information of the current room by calling the radar module, and determine the reverberation time of the current room according to the structural information of the current room; then, determine the initial filtering coefficient of the microphone array of the audio processing device according to the reverberation time of the current room; and, the present application can also transmit a first audio signal to the current room and receive a second audio signal reflected by the current room, that is, the second audio signal is an audio signal with the room reverberation of the current room added to the first audio signal. Thus, considering the room reverberation of the current room, the present application further optimizes the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient; after setting the filtering coefficient of the microphone array to the target filtering coefficient, anti-reverberation can be performed on the original audio signal generated in the current room based on the current filtering coefficient of the microphone array to obtain a target audio signal. It can be seen that the target filtering coefficient is a filtering coefficient that further optimizes the initial filtering coefficient obtained based on the reverberation time of the current room by combining the reverberation time of the current room and the difference between the first audio signal and the second audio signal (i.e., the room reverberation of the current room); then, when performing anti-reverberation based on the target filtering coefficient, the adaptability to the room reverberation formed by the room structure of the current room is relatively high. It can be seen that through the present application, adaptive anti-reverberation can be performed on the audio signal according to the room structure of the room where the audio processing device is located.
[0130] Optionally, in another embodiment of the present application, determining the initial filtering coefficient of the microphone array of the audio processing device corresponding to the reverberation time of the current room based on the first enhancement network includes:
[0131] Input the reverberation time of the current room into the first enhancement network to obtain the initial filtering coefficient of the microphone array of the audio processing device corresponding to the reverberation time of the current room;
[0132] Among them, the training method of the first enhancement network includes:
[0133] For each sample group, input the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained, so that the first enhancement network processes the sample audio signal using the current network parameters to obtain an audio signal output result corresponding to the sample audio signal; wherein, each sample group includes a reverberation time of the samples, a sample audio signal, and a true-value audio signal; the sample audio signal is an audio signal obtained by adding a target reverberation feature to the true-value audio signal in the sample group, and the target reverberation feature is: the reverberation feature existing in the signal collected by playing the signal if the true-value audio signal is played in a room structure with a reverberation time of the reverberation time of the samples; the network parameters include filter coefficients required for anti-reverberation processing.
[0134] Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, determine the current network loss of the first enhancement network.
[0135] In response to the current network loss of the first enhancement network indicating that the first enhancement network has not converged, adjust the current network parameters of the first enhancement network, and return to the step of inputting the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained to obtain an audio signal output result corresponding to the sample audio signal.
[0136] In response to the current network loss of the first enhancement network indicating that the first enhancement network has converged, based on the current network parameters, determine the filter coefficient corresponding to the reverberation time of the samples in the sample group.
[0137] In this application, when determining the initial filter coefficient corresponding to the reverberation time of the current room based on the first enhancement network, the reverberation time of the current room can be input into the first enhancement network. The first enhancement network has learned the correspondence between the reverberation time and the filter coefficient, and can directly output the initial filter coefficient corresponding to the reverberation time of the current room.
[0138] Specifically, when training the first enhancement network, multiple sample groups can be constructed first. Each sample group includes a sample audio signal, a true-value audio signal, and a sample reverberation time. Any audio signal without added room reverberation can be used as the true-value audio signal. For any true-value audio signal, target reverberation features can be added to the true-value audio signal to obtain the sample audio signal. The target reverberation features can be: if the true-value audio signal is played in a room structure with a reverberation time of the sample reverberation time, the reverberation features existing in the signal collected from the played signal, that is, the target reverberation features are the reverberation features that the room structure with the sample reverberation time can add to the true-value audio signal. Among them, the sample audio signal can be obtained by playing the true-value audio signal in a room structure with a reverberation time of the sample time and then collecting the signal; or, through a pre-trained neural network for adding reverberation, taking the true-value audio signal and the target reverberation features as inputs, and fusing them through the neural network to obtain the sample audio signal, which is all reasonable.
[0139] After that, for each sample group, the sample reverberation time and the sample audio signal in the sample group can be input into the first enhancement network to be trained. The first enhancement network uses the network parameters determined by the sample reverberation time (the filter coefficients used for anti-reverberation processing included in the network parameters are not necessarily the filter coefficients used for anti-reverberation of the sample audio signal according to the sample reverberation time; after the first enhancement network is trained, the filter coefficients used for anti-reverberation processing included in the network parameters are the filter coefficients used for anti-reverberation of the sample audio signal according to the sample time) to process the sample audio signal and obtain the audio signal output result corresponding to the sample audio signal. Among them, the network parameters can include the filter coefficients required for anti-reverberation processing, and can also include other parameters such as feature extraction of the audio signal. This application does not make any limitations in this regard. And, the parameters of the last layer in the first enhancement network, for example, the fully connected layer as the output layer, can be used as the filter coefficients required for anti-reverberation processing.
[0140] The process of training the first enhancement network is to make the output result of the trained first enhancement network match or be the same as the true audio signal in the sample group. Therefore, based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, the network loss of the first enhancement network can be determined, and it can be judged whether the first enhancement network converges. If it converges, for example, the current network loss is less than the preset threshold, then based on the current network parameters, the filter coefficient corresponding to the sample reverberation time in the sample group can be determined (for example, the parameters of the last layer of the trained first enhancement network are determined as the filter coefficient corresponding to the sample reverberation time in the sample group). If it does not converge, for example, the current network loss is not less than the preset threshold, then the first enhancement network needs to be continuously trained. First, the current network parameters of the first enhancement network can be adjusted, and the sample reverberation time and the sample audio signal in the sample group are re-input into the first enhancement network to be trained, and the audio signal output result corresponding to the sample audio signal is obtained to continue training the first enhancement network.
[0141] In this application, the first enhancement network can be trained through multiple sample groups. For each sample group, the first enhancement network can be trained to converge (that is, the network loss determined by the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal represents the convergence of the first enhancement network). Thus, the filter coefficient corresponding to the sample time in each sample group can be obtained, that is, the first enhancement network can learn the corresponding relationship between the sample reverberation time in each sample group and the filter coefficient for anti-reverberation of the sample reverberation time. Moreover, each sample group can also include multiple pairs of true signals and sample audio signals. The sample audio signal in each pair of signals is obtained by adding the target reverberation feature of the sample reverberation time in the sample group to the true audio signal in the pair of signals. For each sample group, the first enhancement network can be trained through a pair of signals in the sample group and the sample reverberation time. This application does not limit this.
[0142] In addition, it should be noted that after the first enhancement network is trained, although the first enhancement network can perform anti-reverberation on the audio signal, in this application, the first enhancement network is used to determine the initial filter coefficient corresponding to the reverberation time of the current room, that is, this application only considers the corresponding relationship between the reverberation time learned by the first enhancement network and the filter coefficient; when inputting the reverberation time of the current room into the first enhancement network, an audio signal generated in the current room can also be synchronously input into the first enhancement network. At this time, although the first enhancement network can output the audio signal after anti-reverberation according to the reverberation time of the current room, this application does not focus on this anti-reverberation audio signal, and it is sufficient to obtain the initial filter coefficient corresponding to the reverberation time of the current room; alternatively, only the reverberation time of the current room can be input into the first enhancement network to obtain the initial filter coefficient, which is reasonable.
[0143] Exemplarily, in one implementation, determining the current network loss of the first enhancement network based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal includes:
[0144] Calculating the power-law compression loss based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal to obtain the current network loss of the first enhancement network.
[0145] When determining the current network loss of the first enhancement network, the power-law compression loss can be calculated based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal to obtain the current network loss of the first enhancement network. This current network loss of the first enhancement network is the loss for this sample group; each sample group can calculate the current network loss of the first enhancement network for this sample group in a similar manner, so as to determine whether to continue training the first enhancement network based on the loss, or to determine the filter coefficient corresponding to the sample reverberation time in this sample group; Exemplarily, the power-law compression loss can be calculated according to the following formula to obtain the current network loss of the first enhancement network:
[0146]
[0147] Where L is the calculated current network loss of the first enhancement network, S is the true audio signal, E is the audio signal output result corresponding to the sample audio signal, t is the time of the audio signal, f is the frequency of the audio signal, λ is a constant, and λ is usually set to 0.1 - 0.5 (the importance of amplitude and phase needs to be balanced). Here, referring to the commonly used value of 0.113 in some speech enhancement networks.
[0148] It should be noted that when calculating the current network loss of the first enhancement network, the MES (Mean Squared Error) loss can also be calculated based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal. Any method that can calculate the loss based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal is applicable to this application, and this application does not make any limitations in this regard.
[0149] In this application, through the trained first enhancement network, the corresponding initial filter coefficients can be directly obtained according to the reverberation time of the current room. Moreover, through the calculation of the power-law compression loss, the efficiency and generalization ability of the first enhancement network can be improved, the utilization efficiency of network parameters can be enhanced, and the real-time application of the first enhancement network can be supported.
[0150] Next, in combination with another embodiment, an audio signal processing method provided by this application will be introduced.
[0151] As Figure 2 shown, for a conference scenario, in a meeting room, the audio processing device (including a microphone array) is placed on the meeting room table (for example, the table is located at the center of the meeting room, and the audio processing device is placed at the center of the table), and there are multiple chairs around the table; the audio processing device calls the radar module to emit millimeter waves, calculates the overall room structure of the current room according to the millimeter wave radar (target reflection signal) returned after emission, eliminating the need for manual measurement or calculation, automatically calculates the structure information of the current room, and determines the reverberation time of the current room according to the structure information of the current room; inputs the obtained reverberation time of the current room into the first enhancement network trained with data to obtain the initial filter coefficients related to the room structure, and determines the initial filter coefficients of the microphone array that match the initial reverberation coefficients (that is, the first optimization of the filter coefficients of the microphone array. For example: the microphone array has default filter coefficients, and the process of adjusting the default filter coefficients to the initial filter coefficients can be understood as the process of the first optimization of the filter coefficients); then, this audio processing device (such as a voice pickup and transcription device) emits and receives audio signals (that is, emits a first audio signal to the current room and receives the second audio signal reflected by the current room), and then optimizes the filter coefficients of the microphone array again according to the reverberation time of the current room, the first audio signal, and the second audio signal, that is, optimizes the initial filter coefficients to obtain the target filter coefficients; thus, this application can obtain an adaptive denoising and dereverberation algorithm, realize the improvement of the audio signal quality, and improve the effect of conference transcription.
[0152] As Figure 3 shown, an audio signal processing method provided by this application may include the following steps:
[0153] S301: The microphone array is activated; that is, the voice pickup and transcription device is placed in the middle of the conference room table, and the microphone array of the voice pickup and transcription device is activated.
[0154] S302: Whether millimeter-wave ranging is performed; that is, it is determined whether to perform the step of millimeter-wave ranging to calculate the room structure of the current room; if yes, step S306 is executed, and if no, step S303 is executed.
[0155] S303: Transmit millimeter waves around for radar ranging; that is, since the step of millimeter-wave ranging is not performed, the voice pickup and transcription device transmits millimeter waves around for radar ranging, calculates the distances in several directions, and then calculates the structure information of the current room.
[0156] S304: Calculate the reverberation time; that is, according to the calculated structure information of the current room, calculate the reverberation time of the current room; for example: substitute the structure information of the current room into the Sabine formula: RT60 = 0.161×(V / AS), where V is the room volume, S is the inner surface area of the room, A is the absorption coefficient, and RT60 is the reverberation time, so as to obtain the reverberation time of the current room.
[0157] S305: Input to the first enhancement network to obtain the initial filtering coefficient; that is, input the reverberation time of the current room into the first enhancement network to obtain the initial filtering coefficient corresponding to the reverberation time of the current room.
[0158] Among them, the first enhancement network is trained according to multiple sample groups (learning the corresponding relationship between the reverberation time and the anti-reverberation coefficient), and each sample group includes a sample audio signal, a true-value audio signal, and a sample reverberation time; according to the measured distance, the structure information of the current room can be obtained. After calculating the reverberation time of the current room according to the structure information of the current room, the reverberation time of the current room can be used as a parameter to be transmitted to the trained first enhancement network to obtain the initial filtering coefficient; among them, the network loss in the training process of the first enhancement network adopts the power-law compression loss instead of the MSE loss, and the calculation formula of the power-law compression loss is as follows:
[0159]
[0160] Among them, L is the calculated current network loss of the first enhancement network, S is the true-value audio signal, E is the audio signal output result corresponding to the sample audio signal, t is the time of the audio signal, f is the frequency of the audio signal, λ is a constant, and λ is usually set to 0.1 - 0.5 (it is necessary to balance the importance of amplitude and phase). Here, referring to the common values of some speech enhancement networks, it is 0.113.
[0161] S306: Adjust the filtering coefficients of the microphone array; that is, according to the result output by the first enhancement network (i.e., the initial filtering coefficients), adjust the relevant parameters during the voice pickup and transcription of the voice pickup and transcription device, that is, adjust the filtering coefficients of the microphone array to the initial filtering coefficients.
[0162] S307: Optimize the filtering coefficients of the microphone array again according to the reverberation time, transmitted signal, and received signal; that is, after the filtering coefficients of the microphone array are adjusted to the initial filtering coefficients, transmit a first audio signal to the current room and receive a second audio signal reflected by the current room; then, optimize the initial filtering coefficients according to the reverberation time, the transmitted first audio signal, and the received second audio signal to obtain the target filtering coefficients; and, set the filtering coefficients of the microphone array to the target filtering coefficients, and, after the setting is completed, when the original audio signal generated in the current room is collected through the microphone array, perform anti-reverberation on the received original audio signal based on the current target filtering coefficients of the microphone array to obtain the target audio signal.
[0163] Among them, the following formula can be used to optimize the initial filtering coefficients to obtain the target filtering coefficients:
[0164] h(n + 1) = h(n) + μ × e(n) × x(n), where h(n + 1) is the target filtering coefficient; h(n) is the initial filtering coefficient, x(n) is the first audio signal, e(n) is the error signal; μ is the specified step factor, μ = K / RT60, K is the specified proportionality coefficient; RT60 is the reverberation time of the current room.
[0165] S308: The second enhancement network performs audio signal enhancement; that is, input the target audio signal into the second enhancement network for further audio signal enhancement.
[0166] S309: Output the personalized enhanced audio signal; that is, output the personalized enhanced audio signal obtained by the second enhancement network for audio signal enhancement, that is, output the enhanced target audio signal.
[0167] This application can measure and obtain the structural information of the current room through a millimeter-wave radar, and then transmit this information to the voice enhancement system, that is, through the structural information of the current room, finally, the target filtering coefficients for enhancing the audio signal can be determined to perform characteristic audio enhancement on the audio signal generated in the current room.
[0168] Based on the above method embodiments, this application also provides an audio signal processing device, which is applied to an audio processing device, such as Figure 4 shown, the device includes:
[0169] A calling module 410, configured to call a radar module to transmit radar waves to a current room and receive target reflection signals corresponding to the radar waves;
[0170] A first determination module 420, configured to determine room structure information calculated based on the radar waves and the target reflection signals as the structure information of the current room;
[0171] A second determination module 430, configured to determine the reverberation time of the current room according to the structure information of the current room;
[0172] A third determination module 440, configured to determine an initial filtering coefficient of a microphone array of the audio processing device according to the reverberation time of the current room;
[0173] A transmitting module 450, configured to transmit a first audio signal to the current room and receive a second audio signal reflected by the current room;
[0174] An optimization module 460, configured to optimize the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient;
[0175] A setting module 470, configured to set the filtering coefficient of the microphone array as the target filtering coefficient. After the setting is completed, in response to collecting a raw audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received raw audio signal based on the current filtering coefficient of the microphone array to obtain a target audio signal.
[0176] The audio signal processing device provided by the embodiment of the present application. The audio processing device can calculate the structural information of the current room by calling the radar module, and determine the reverberation time of the current room according to the structural information of the current room; then, determine the initial filtering coefficients of the microphone array of the audio processing device according to the reverberation time of the current room; moreover, the present application can also transmit a first audio signal to the current room and receive a second audio signal reflected by the current room, that is, the second audio signal is an audio signal with the room reverberation of the current room added to the first audio signal. Thus, considering the room reverberation of the current room, the present application further optimizes the initial filtering coefficients according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain target filtering coefficients; after setting the filtering coefficients of the microphone array to the target filtering coefficients, anti-reverberation can be performed on the original audio signal generated in the current room based on the current filtering coefficients of the microphone array to obtain a target audio signal. It can be seen that the target filtering coefficients are filtering coefficients that further optimize the initial filtering coefficients obtained based on the reverberation time of the current room by combining the reverberation time of the current room and the difference between the first audio signal and the second audio signal (i.e., the room reverberation of the current room); then, when performing anti-reverberation based on the target filtering coefficients, the adaptability to the room reverberation formed by the room structure of the current room is relatively high. It can be seen that through the present application, adaptive anti-reverberation can be performed on the audio signal according to the room structure of the room where the audio processing device is located.
[0177] Optionally, the third determination module includes:
[0178] The first determination sub-module is configured to determine, based on the first enhancement network, the initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room; wherein, the first enhancement network is a network for performing anti-reverberation processing on any audio signal, and the initial filtering coefficients are: the filtering coefficients required when the first enhancement network performs anti-reverberation on the audio signal generated in the current room according to the reverberation time of the current room.
[0179] Optionally, the device further includes
[0180] The detection module is configured to detect whether a target instruction is received; wherein, the target instruction is used to represent an instruction for reusing the currently set filtering coefficients of the microphone array;
[0181] If not received, execute the step of calling the radar module to transmit a radar wave to the current room and receive the target reflection signal corresponding to the radar wave;
[0182] If received, in response to the original audio signal generated in the current room collected by the microphone array, based on the current filtering coefficients of the microphone array, anti-reverberation is performed on the received original audio signal to obtain a target audio signal.
[0183] Optionally, the first determination sub-module is specifically configured to:
[0184] Input the reverberation time of the current room into the first enhancement network to obtain the initial filtering coefficients of the microphone array of the audio processing device;
[0185] Wherein, the training method of the first enhancement network includes:
[0186] For each sample group, input the sample reverberation time and the sample audio signal in the sample group into the first enhancement network to be trained, so that the first enhancement network processes the sample audio signal using the current network parameters to obtain the audio signal output result corresponding to the sample audio signal; wherein, each sample group includes a sample reverberation time, a sample audio signal, and a true-value audio signal; the sample audio signal is an audio signal obtained by adding a target reverberation feature to the true-value audio signal in the sample group, and the target reverberation feature is: the reverberation feature existing in the signal collected by playing the true-value audio signal in a room structure with a reverberation time equal to the sample reverberation time of the sample group; the network parameters include the filtering coefficients required for anti-reverberation processing;
[0187] Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, determine the current network loss of the first enhancement network;
[0188] In response to the current network loss of the first enhancement network indicating that the first enhancement network has not converged, adjust the current network parameters of the first enhancement network, and return to the step of inputting the sample reverberation time and the sample audio signal in the sample group into the first enhancement network to be trained to obtain the audio signal output result corresponding to the sample audio signal;
[0189] In response to the current network loss of the first enhancement network indicating that the first enhancement network has converged, based on the current network parameters, determine the filtering coefficients corresponding to the sample reverberation time in the sample group.
[0190] Optionally, the determining the current network loss of the first enhancement network based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal includes:
[0191] Based on the true audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, calculate the power-law compression loss to obtain the current network loss of the first enhancement network.
[0192] Optionally, the optimization module is specifically configured to:
[0193] Determine the sum value of the specified product and the initial filter coefficient to obtain the target filter coefficient; wherein, the specified product is the product of the first audio signal, the error signal, and the specified step factor, the error signal is used to characterize the difference between the first audio signal and the second audio signal, the specified step factor is the ratio of the specified proportional coefficient to the reverberation time of the current room, and the specified proportional coefficient is a predetermined empirical coefficient used for anti-reverberation of any audio signal.
[0194] Optionally, the apparatus further includes:
[0195] An enhancement module, configured to perform audio signal enhancement on the target audio signal based on a second enhancement network to obtain an enhanced target audio signal; wherein, the second enhancement network is a network for performing signal enhancement processing on any audio signal.
[0196] Optionally, the second determination module is specifically configured to:
[0197] Substitute the structural information of the current room into the predetermined Sabine formula to obtain the reverberation time of the current room.
[0198] An embodiment of the present application further provides an audio processing device, as Figure 5 shown, including: a radar module 501, a microphone array 502, a processor 504, and a memory 503;
[0199] The radar module 501 is configured to transmit radar waves to the current room and receive target reflection signals corresponding to the radar waves under the call of the processor 504;
[0200] The microphone array 502 is configured to collect audio signals in the current room;
[0201] The memory 503 is used to store computer programs;
[0202] The processor 504 is configured to implement any of the audio signal processing methods when executing the programs stored on the memory 503.
[0203] And the above audio processing device may further include a communication bus and / or a communication interface, and the processor 504, the communication interface, and the memory 503 complete mutual communication through the communication bus.
[0204] The communication bus mentioned in the above audio processing device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0205] The communication interface is used for communication between the above audio processing device and other devices.
[0206] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0207] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0208] In another embodiment provided by the present application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above audio signal processing methods are implemented.
[0209] In another embodiment provided by the present application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the audio signal processing methods in the above embodiments.
[0210] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid state disk (SSD), etc.
[0211] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0212] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0213] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. An audio signal processing method, characterized in that, Applied to an audio processing device, the method includes: Invoking a radar module to transmit radar waves to a current room and receive target reflection signals corresponding to the radar waves; Determining room structure information calculated based on the radar waves and the target reflection signals as the structure information of the current room; Determining the reverberation time of the current room according to the structure information of the current room; Determining initial filtering coefficients of a microphone array of the audio processing device according to the reverberation time of the current room; Transmitting a first audio signal to the current room and receiving a second audio signal reflected by the current room; Optimizing the initial filtering coefficients according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain target filtering coefficients; Setting the filtering coefficients of the microphone array to the target filtering coefficients. After the setting is completed, in response to an original audio signal generated in the current room being collected through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficients of the microphone array to obtain a target audio signal.
2. The method according to claim 1, characterized in that, The determining the initial filtering coefficients of the microphone array of the audio processing device according to the reverberation time of the current room includes: Determining, based on a first enhancement network, initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room; wherein, the first enhancement network is a network for performing anti-reverberation processing on any audio signal, and the initial filtering coefficients are the filtering coefficients required when the first enhancement network performs anti-reverberation on an audio signal generated in the current room according to the reverberation time of the current room.
3. The method according to claim 1 or 2, characterized in that, Before the invoking the radar module to transmit radar waves to the current room and receive target reflection signals corresponding to the radar waves, the method further includes: Detecting whether a target instruction is received; wherein, the target instruction is used to represent an instruction for reusing the currently set filtering coefficients of the microphone array; If not received, performing the step of invoking the radar module to transmit radar waves to the current room and receive target reflection signals corresponding to the radar waves; If received, in response to an original audio signal generated in the current room being collected through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficients of the microphone array to obtain a target audio signal.
4. The method according to claim 2, wherein The determining, based on the first enhancement network, initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room includes: Inputting the reverberation time of the current room into the first enhancement network to obtain initial filtering coefficients of the microphone array of the audio processing device corresponding to the reverberation time of the current room; Wherein, the training method of the first enhancement network includes: For each sample group, input the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained, so that the first enhancement network processes the sample audio signal using the current network parameters to obtain an audio signal output result corresponding to the sample audio signal; wherein, each sample group includes a reverberation time of the samples, a sample audio signal, and a true-value audio signal; the sample audio signal is an audio signal obtained by adding a target reverberation feature to the true-value audio signal in the sample group, and the target reverberation feature is: if the true-value audio signal is played in a room structure with a reverberation time of the reverberation time of the samples and the played signal is collected, the reverberation feature existing in the collected signal; the network parameters include filter coefficients required for anti-reverberation processing. Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, determine the current network loss of the first enhancement network. In response to the current network loss of the first enhancement network indicating that the first enhancement network has not converged, adjust the current network parameters of the first enhancement network, and return to the step of inputting the reverberation time of the samples in the sample group and the sample audio signal into the first enhancement network to be trained to obtain an audio signal output result corresponding to the sample audio signal. In response to the current network loss of the first enhancement network indicating that the first enhancement network has converged, based on the current network parameters, determine the filter coefficient corresponding to the reverberation time of the samples in the sample group.
5. The method according to claim 4, characterized in that The determining the current network loss of the first enhancement network based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal includes: Based on the true-value audio signal in the sample group and the audio signal output result corresponding to the sample audio signal, calculate a power-law compression loss to obtain the current network loss of the first enhancement network.
6. The method according to claim 1 or 2, characterized in that The optimizing the initial filter coefficient to obtain a target filter coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal includes: Determine the sum value of a specified product and the initial filter coefficient to obtain the target filter coefficient; wherein, the specified product is the product of the first audio signal, an error signal, and a specified step factor, the error signal is used to characterize the difference between the first audio signal and the second audio signal, the specified step factor is the ratio of a specified proportionality coefficient to the reverberation time of the current room, and the specified proportionality coefficient is a predetermined empirical coefficient used for anti-reverberation of any audio signal.
7. The method according to claim 1 or 2, characterized in that, The method further includes: Based on a second enhancement network, enhance the target audio signal to obtain an enhanced target audio signal; wherein, the second enhancement network is a network used for signal enhancement processing of any audio signal.
8. The method according to claim 1 or 2, characterized in that, The determining the reverberation time of the current room according to the structural information of the current room includes: Substitute the structural information of the current room into a predetermined Sabine formula to obtain the reverberation time of the current room.
9. An audio signal processing device, characterized in that, Applied to an audio processing device, the device includes: A calling module, configured to call a radar module to emit radar waves to a current room and receive target reflection signals corresponding to the radar waves; A first determination module, configured to determine room structure information calculated based on the radar waves and the target reflection signals as the structure information of the current room; A second determination module, configured to determine the reverberation time of the current room according to the structure information of the current room; A third determination module, configured to determine an initial filtering coefficient of a microphone array of the audio processing device according to the reverberation time of the current room; A transmitting module, configured to transmit a first audio signal to the current room and receive a second audio signal reflected by the current room; An optimization module, configured to optimize the initial filtering coefficient according to the reverberation time of the current room, the first audio signal, and the second audio signal to obtain a target filtering coefficient; A setting module, configured to set the filtering coefficient of the microphone array to the target filtering coefficient. After the setting is completed, in response to collecting an original audio signal generated in the current room through the microphone array, anti-reverberation is performed on the received original audio signal based on the current filtering coefficient of the microphone array to obtain a target audio signal.
10. An audio processing device, characterized in that, Including: A radar module, a microphone array, a processor, and a memory; The radar module is configured to emit radar waves to a current room and receive target reflection signals corresponding to the radar waves under the call of the processor; The microphone array is configured to collect audio signals in the current room; The memory is configured to store a computer program; The processor is configured to implement the method according to any one of claims 1-8 when executing the program stored on the memory.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program implements the method according to any one of claims 1-8 when executed by the processor.