Sound signal processing methods, devices, equipment and media

Through delay estimation and filtering processing technology, the problem of inaccurate voice recognition when external devices are connected to display devices is solved, the accuracy of voice recognition is improved, and the user experience is enhanced.

CN116312614BActive Publication Date: 2025-10-28HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310186652.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2025-10-28
Estimated Expiration
2043-03-01

AI Technical Summary

Technical Problem

When an external device is connected to a display device, the voice recognition results are inaccurate, resulting in a poor user experience.

Method used

The delay time between the external device and the display device is determined by a delay estimation method, the target audio signal is filtered to determine a residual signal, and the delay time and the residual signal are used to extract the user voice signal from the first sound signal.

Benefits of technology

The accuracy of speech recognition is improved, and the user experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312614B_ABST
    Figure CN116312614B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a sound signal processing method, apparatus, device and medium, and relates to the field of audio processing technology; wherein, the method comprises: obtaining a first sound signal collected by a sound collection module of an external device and a second sound signal that has been sent to a display device by the external device, wherein the first sound signal comprises a user voice signal and a target audio signal played by the display device; determining the delay time between the original audio signal and the target audio signal in the second sound signal by a delay estimation method; filtering the target audio signal to determine a residual signal; processing the first sound signal by the delay time, the residual signal and the original audio signal to determine the user voice signal. The embodiment of the present disclosure can obtain a more accurate user voice signal by processing the first sound signal, thereby improving the accuracy of the recognition result when recognizing the user voice signal and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, and more particularly to a sound signal processing method, apparatus, device, and medium. Background Technology

[0002] External devices (such as TV boxes, smart speakers, etc.) have been widely used as a type of home entertainment device. In recent years, with the continuous development of voice technologies such as voice recognition, it has become possible to control external devices by voice.

[0003] When an external device is connected to a display device, the external device sends signals to the display device so that the display device can display images and / or play sound. In this case, the sound signal captured by the microphone of the external device usually includes voice commands and the sound of the program played by the display device. When the external device recognizes the sound signal, it can cause inaccurate recognition results, thus affecting the user experience. Summary of the Invention

[0004] To address the aforementioned technical issues, or at least partially address them, this disclosure provides a sound signal processing method, apparatus, device, and medium capable of processing a first sound signal, filtering out the sound of a program played by a display device from the first sound signal, and obtaining a more accurate user voice signal. This improves the accuracy of the recognition results and enhances the user experience when recognizing the user voice signal.

[0005] To achieve the above objectives, the technical solutions provided by the embodiments of this disclosure are as follows:

[0006] In a first aspect, this disclosure provides a sound signal processing method, the method comprising:

[0007] The system acquires a first sound signal acquired by the sound acquisition module of an external device and a second sound signal sent by the external device to the display device. The first sound signal includes a user voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal.

[0008] The delay time between the original audio signal and the target audio signal in the second audio signal is determined by a delay estimation method.

[0009] The target audio signal is filtered to determine the residual signal;

[0010] The user's voice signal is determined by processing the first sound signal using the delay time, the residual signal, and the original audio signal.

[0011] Secondly, this disclosure provides a sound signal processing apparatus, the apparatus comprising:

[0012] The signal acquisition module is used to acquire a first sound signal acquired by the sound acquisition module of the external device and a second sound signal sent by the external device to the display device. The first sound signal includes a user voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal.

[0013] The first determining module is used to determine the time delay between the original audio signal and the target audio signal in the second sound signal by means of a delay estimation method;

[0014] The second determining module is used to filter the target audio signal and determine the residual signal;

[0015] The third determining module is used to process the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal.

[0016] Thirdly, this disclosure also provides an electronic device, including:

[0017] One or more processors;

[0018] Storage device for storing one or more programs.

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the sound signal processing methods described in the embodiments of this disclosure.

[0020] Fourthly, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the sound signal processing methods described in the embodiments of this disclosure.

[0021] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: It acquires a first audio signal collected by the sound acquisition module of an external device and a second audio signal sent by the external device to the display device. The first audio signal includes a user's voice signal and a target audio signal played by the display device. The second audio signal includes the original audio signal corresponding to the target audio signal. A delay time between the original audio signal and the target audio signal in the second audio signal is determined using a delay estimation method. The target audio signal is filtered to determine a residual signal. The first audio signal is processed using the delay time, the residual signal, and the original audio signal to determine the user's voice signal. In this technical solution, by processing the first audio signal, the sound of the program played by the display device can be filtered out, resulting in a more accurate user voice signal. This improves the accuracy of the recognition result and enhances the user experience when recognizing the user's voice signal. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1A This is a schematic diagram of the sound signal processing process when an external device is used alone in related technologies.

[0025] Figure 1B This is a schematic diagram of the structure when an external device is connected to a display device in related technologies;

[0026] Figure 1C This is a schematic diagram illustrating an applicable scenario for a sound signal processing procedure according to an embodiment of this disclosure;

[0027] Figure 2 A schematic flowchart illustrating a sound signal processing method provided in an embodiment of this disclosure;

[0028] Figure 3 This is a schematic diagram of a sound signal processing procedure provided in an embodiment of the present disclosure;

[0029] Figure 4 A schematic flowchart of another sound signal processing method provided in an embodiment of this disclosure;

[0030] Figure 5A schematic flowchart illustrating another sound signal processing method provided in this embodiment of the present disclosure;

[0031] Figure 6A This is a schematic diagram of the structure of a sound signal processing device provided in an embodiment of the present disclosure;

[0032] Figure 6B This is a schematic diagram of the structure of the third determining module in the sound signal processing apparatus of this disclosure embodiment;

[0033] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0034] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0035] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0036] It should be noted that the brief descriptions of terms in this disclosure are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0037] It should be noted that in this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. For example, a product or apparatus that comprises a list of components is not necessarily limited to all components expressly listed, but may include other components not expressly listed or inherent to such product or apparatus.

[0038] With the rapid development of the internet, display devices (devices with display functions, such as televisions and personal computers, etc., not limited here) play an important role as essential entertainment equipment in the home. More and more people want to use large-screen TVs for chatting, watching movies, listening to music, playing games, and other functions. External devices such as TV boxes and smart speakers, which can connect to display devices to enable interactive functions, have emerged. Furthermore, with the continuous development of voice recognition and related voice technologies, it is now possible to control external devices via voice, thereby enabling user interaction with these devices.

[0039] Figure 1A This is a schematic diagram illustrating the sound signal processing process when an external device is used independently in related technologies. For example... Figure 1A As shown, the sound signal processing when the external device is used alone mainly involves the following modules: external device processing system 1101, external device sound playback module 1102, signal processing module 1103, and sound acquisition module 1104. The external device sound playback module 1102 can be a module with sound playback function, such as a speaker or loudspeaker of the external device; this embodiment does not specifically limit its functionality. The sound acquisition module 1104 can be a module with sound acquisition function, such as a microphone, linear microphone array, or recording module of the external device; this embodiment does not specifically limit its functionality.

[0040] When the external device is used independently, the user wakes it up with a wake word and can control it by speaking voice commands, such as checking the weather or news. Weather and news broadcasts are played through the external device's audio playback module 1102. If the user issues a voice command while the external device's audio playback module 1102 is playing data, the data being played can interfere with the speech recognition process of the external device processing system 1101. In this case, when the external device's audio acquisition module 1104 acquires the total audio data containing both playback data and voice command data, it simulates the audio signal played by the external device's audio playback module 1102 and sends it back to the front-end signal processing module 1103. The signal processing module 1103 performs echo cancellation and noise reduction on the total audio data, suppressing the echo signal from the external device's audio playback module 1102 to the audio acquisition module 1104 in real time. This allows the external device processing system 1101 to obtain more accurate voice command data, improving its speech recognition performance.

[0041] Figure 1BThis is a schematic diagram illustrating the connection between an external device and a display device in related technologies. When the external device and the display device are connected, they can be connected via a High Definition Multimedia Interface (HDMI) cable or an optical fiber cable, etc. The specific connection method depends on the specific circumstances and is not limited here. Figure 1B As shown, Figure 1B The system includes: an external device processing system 1101, an output port 1105, a display device processing system 1201, a sound playback module 1202 for the display device, a signal processing module 1103, and a sound acquisition module 1104. The external device processing system 1101 sends sound data and / or video data to the display device processing system 1201 via the output port 1105. After processing the sound data, the display device processing system 1201 plays the processed sound data through the sound playback module 1202. The output port 1105 can be an HDMI interface or an optical fiber interface, depending on the specific type of the external device and the display device; no specific limitation is made here. In the above scenario, the external device is in OTT (Over The Top) mode, bypassing the operator to develop various video and data services based on the open Internet, and the external device's own sound playback module does not emit sound. The voice software is integrated into the signal processing module 1103 of the external device, enabling signal noise reduction, wake-up, playback, and potentially video calls in the future. It transmits voice data to the external device processing system 1101 for voice recognition. The functions of the signal processing module 1103 and the sound acquisition module 1104 are related to their respective roles in... Figure 1A The functions are the same, so to avoid repetition, they will not be repeated here.

[0042] exist Figure 1B In this situation, because it is impossible to hardware-recapture the data played by the audio playback module 1202 of the display device, the signal processing module 1103 is also unable to eliminate the sound emitted by the audio playback module 1202 of the display device through hardware recapture. In this case, if the audio signal acquired by the external device's audio acquisition module 1104 includes both voice commands and the sound emitted by the audio playback module 1202 of the display device, the external device processing system 1101 will encounter inaccurate recognition results when recognizing the audio signal, thus affecting the user experience.

[0043] To address the aforementioned issues, this disclosure provides a sound signal processing method. The method acquires a first sound signal collected by an external device's sound acquisition module and a second sound signal sent by the external device to a display device. The first sound signal includes a user's voice signal and a target audio signal played by the display device. The second sound signal includes an original audio signal corresponding to the target audio signal. A delay time is determined between the original audio signal and the target audio signal in the second sound signal using a delay estimation method. The target audio signal is filtered to determine a residual signal. The first sound signal is processed using the delay time, the residual signal, and the original audio signal to determine the user's voice signal. In this technical solution, by processing the first sound signal, the sound of the program played by the display device can be filtered out, resulting in a more accurate user voice signal. This improves the accuracy of the recognition results and enhances the user experience when recognizing the user's voice signal.

[0044] For example, Figure 1C This is a schematic diagram illustrating an applicable scenario for a sound signal processing procedure according to an embodiment of this disclosure. For example... Figure 1C As shown, this applicable scenario is when the external device 110 is connected to the display device 120.

[0045] To illustrate the sound signal processing scheme in this disclosure in more detail, the following will be combined with examples. Figure 2 To explain, it is understandable that Figure 2 The steps involved may include more or fewer steps in actual implementation, and the order of these steps may also be different, depending on whether the sound signal processing method provided in the embodiments of this application can be implemented.

[0046] Figure 2 This is a schematic flowchart illustrating a sound signal processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations where, when an external device is connected to a display device, it processes a first sound signal acquired by a sound acquisition module, which includes the user's voice signal and a target audio signal played by the display device. The method in this embodiment can be executed by a sound signal processing device, which can be implemented in hardware / software and can be configured in an electronic device.

[0047] like Figure 2 As shown, the method specifically includes the following steps:

[0048] S210, acquire the first sound signal acquired by the sound acquisition module of the external device and the second sound signal that the external device has sent to the display device.

[0049] The first audio signal includes the user's voice signal and the target audio signal played by the display device. The second audio signal includes the original audio signal corresponding to the target audio signal. The target audio signal can be understood as the audio signal being played by the display device, acquired by the sound acquisition module during signal acquisition. The second audio signal can be understood as the audio signal sent to the display device by the external device before the display device plays the target audio signal; this second audio signal includes the original audio signal corresponding to the target audio signal and the original audio signal corresponding to the audio signal not played by the display device. Since the target audio signal has been processed by the display device, while the original audio signal is the unprocessed audio signal sent to the display device by the external device, there is a one-to-one correspondence between the two; that is, each target audio signal has a corresponding original audio signal.

[0050] To address the issue that the sound signals collected by the external device's sound acquisition module include voice commands and sounds emitted by the display device's sound playback module, which can lead to inaccurate recognition results when the external device's processing system identifies the sound signals, this embodiment requires acquiring the first sound signal collected by the external device's sound acquisition module and the second sound signal that the external device has sent to the display device. This facilitates subsequent processing of the first and second sound signals.

[0051] S220, using a delay estimation method, determines the time delay between the original audio signal and the target audio signal in the second audio signal.

[0052] Since the original audio signal in the second audio signal is an unprocessed audio signal sent to the display device by an external device, and the target audio signal is the audio signal being played on the display device when the external device's sound acquisition module is acquiring the signal, there is a significant delay between these two audio signals; the target audio signal lags behind the original audio signal in the second audio signal. Therefore, in order to align the target audio signal with the original audio signal in the second audio signal, it is necessary to determine the delay time between the original audio signal and the target audio signal using a delay estimation method. Specifically, the delay time can be determined using a time delay estimation model; it can also be determined using algorithms such as the generalized correlation method, the generalized phase spectrum method, the bispectral method, or the higher-order cumulant method; other delay estimation algorithms can also be used, but this embodiment does not specifically limit them.

[0053] S230 performs filtering on the target audio signal to determine the residual signal.

[0054] The target audio signal is acquired by the sound acquisition module of an external device. After processing by the display device's processing system and power amplifier, the original audio signal is played back by the display device's sound playback module. Because the display device may have sound effects, the target audio signal acquired by the external device's sound acquisition module may differ from the original audio signal, resulting in distortion. Therefore, the target audio signal needs to be filtered. This can be done using filters, or filtering algorithms such as amplitude limiting filtering, median filtering, and recursive averaging filtering, to determine the residual signal—that is, the difference signal and noise signal between the target audio signal and the original audio signal.

[0055] S240 processes the first sound signal by using the delay time, residual signal, and original audio signal to determine the user's voice signal.

[0056] After determining the delay time and residual signal, the user's voice signal can be determined by comparing the original audio signal with the target audio signal played by the display device through the delay time and residual signal and filtering it out from the first sound signal.

[0057] The sound signal processing method provided in this embodiment acquires a first sound signal collected by the sound acquisition module of an external device and a second sound signal sent by the external device to a display device. The first sound signal includes a user's voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal. A delay estimation method is used to determine the delay time between the original audio signal and the target audio signal in the second sound signal. The target audio signal is filtered to determine a residual signal. The first sound signal is processed using the delay time, the residual signal, and the original audio signal to determine the user's voice signal. In this technical solution, by processing the first sound signal, the sound of the program played by the display device can be filtered out, resulting in a more accurate user voice signal. This improves the accuracy of the recognition result and enhances the user experience when recognizing the user's voice signal.

[0058] Figure 3 This is a schematic diagram illustrating the structure of a sound signal processing procedure provided in an embodiment of this disclosure. Figure 3 As shown: In Figure 1B Based on this, signal feedback is performed from output port 1105 to obtain the second audio signal sent by the external device to the display device. The second audio signal is then processed by signal processing module 1103, thereby realizing the audio signal processing method in this embodiment. The functions of the external device processing system 1101, output port 1105, display device processing system 1201, the audio playback module of the display device, and the audio acquisition module 1104 are as follows: Figure 1B The functions are the same, so to avoid repetition, they will not be repeated here.

[0059] In some embodiments, optionally, determining the time delay between the original audio signal and the target audio signal in the second audio signal using a delay estimation method includes:

[0060] Feature extraction is performed on the original audio signal in the second sound signal to obtain the first acoustic feature;

[0061] Feature extraction is performed on the target audio signal to obtain the second acoustic feature;

[0062] The first acoustic feature and the second acoustic feature are input into the time delay estimation model to obtain the delay time.

[0063] The time delay estimation model can be a trained deep learning filtering model or other models that can estimate the delay time. This embodiment does not make any specific limitations on this.

[0064] Specifically, by using network structures with feature extraction capabilities, such as convolutional layers, to extract features from the original audio signal in the second audio signal, the first acoustic features corresponding to the original audio signal can be obtained. Similarly, by using network structures with feature extraction capabilities, such as convolutional layers, to extract features from the target audio signal, the second acoustic features corresponding to the target audio signal can be obtained. After obtaining the first and second acoustic features, they are input into a time delay estimation model. Through calculation using this time delay estimation model, the time delay between the original audio signal and the target audio signal can be obtained.

[0065] In this embodiment, the delay time between the original audio signal and the target audio signal is determined by the above method, which is simple, fast, efficient, and highly accurate.

[0066] In some embodiments, the time delay estimation model may be specifically determined in the following ways:

[0067] Acquire training samples, wherein the training samples include first sound data collected by the sound acquisition module after the external device establishes a connection with different types of display devices and second sound data sent by the external device to the different types of display devices. Each first sound data includes user voice data and target audio data corresponding to the display device. The second sound data includes: the original audio data corresponding to the target audio data.

[0068] The time delay estimation model is trained using the training samples until the time delay estimation model converges, at which point training stops.

[0069] The preset loss function can be a connectionist temporal classification (CTC) loss function, a multi-class cross-entropy loss function, a cosine loss function, or a mean square loss function, etc. The specific loss function can be determined according to actual usage requirements, or it can be set by the user. This embodiment of the disclosure does not limit this.

[0070] A large number of training samples are obtained. These samples include first audio data collected after the external device establishes a connection with different types of display devices, as well as second audio data sent by the external device to the aforementioned display devices. Therefore, the training samples are diverse. The latency estimation model is trained using these training samples. Specifically, the training samples are sequentially input into the latency estimation model to obtain the predicted latency. Based on the predicted latency and the actual latency, a loss value is calculated using a preset loss function. The parameters of the latency estimation model are adjusted based on the loss value until the model converges, at which point training stops.

[0071] In this embodiment, by training the time delay estimation model, a well-trained time delay estimation model is obtained, which helps to improve the accuracy of the delay time obtained by the time delay estimation model, thereby improving the accuracy of the user's voice signal.

[0072] In some embodiments, optionally, after processing the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal, the method may further include:

[0073] The user's voice signal is processed by speech recognition to obtain the control command corresponding to the user's voice signal;

[0074] The external device is controlled to perform corresponding operations based on the control commands.

[0075] Specifically, after receiving the user's voice signal, the external device processing system performs speech recognition processing on the user's voice signal to obtain semantic understanding results. These semantic understanding results can then be used to determine the control commands corresponding to the user's voice signal. These control commands can then be used to control the external device to perform corresponding operations.

[0076] In this embodiment, the above method enables better voice interaction with external devices, achieving better control of external devices and enhancing the user's interactive experience.

[0077] Figure 4This is a flowchart illustrating another audio signal processing method provided in this embodiment. This embodiment is an optimization based on the above embodiment. Optionally, this embodiment provides a detailed explanation of the process of processing the first audio signal using delay time, residual signal, and original audio signal to determine the user's voice signal. Figure 4 As shown, the method specifically includes the following steps:

[0078] S210, acquire the first sound signal acquired by the sound acquisition module of the external device and the second sound signal that the external device has sent to the display device.

[0079] S220, using a delay estimation method, determines the time delay between the original audio signal and the target audio signal in the second audio signal.

[0080] S230 performs filtering on the target audio signal to determine the residual signal.

[0081] S2401 aligns the original audio signal based on the delay time to obtain a third sound signal.

[0082] Specifically, after determining the delay time between the original audio signal and the target audio signal in the second audio signal, since the target audio signal lags behind the original audio signal in the second audio signal, the original audio signal and the target audio signal are aligned by using this delay time. That is, by delaying the original audio signal by the delay time, the third audio signal can be obtained.

[0083] S2402, subtract the first sound signal from the third sound signal to obtain the reference speech signal.

[0084] Since the first audio signal contains the user's voice signal and the target audio signal played by the display device, subtracting the first audio signal from the third audio signal yields a reference audio signal, which contains a residual signal.

[0085] S2403 determines the user's voice signal based on the reference voice signal and the residual signal.

[0086] After obtaining the reference speech signal, since the reference speech signal contains a residual signal and the residual signal has been determined in S230, the user speech signal can be determined by the reference speech signal and the residual signal.

[0087] The audio signal processing method provided in this embodiment acquires a first audio signal acquired by the audio acquisition module of an external device and a second audio signal sent by the external device to a display device; determines the delay time between the original audio signal and the target audio signal in the second audio signal using a delay estimation method; filters the target audio signal to determine the residual signal; aligns the original audio signal based on the delay time to obtain a third audio signal; subtracts the first audio signal from the third audio signal to obtain a reference speech signal; and determines the user speech signal based on the reference speech signal and the residual signal. In the above technical solution, by aligning the original audio signal with the target audio signal using the delay time between the original audio signal and the target audio signal, and filtering out the third audio signal and the residual signal from the first audio signal, the sound of the program played by the display device can be filtered out from the first audio signal, resulting in a more accurate user speech signal. This improves the accuracy of the recognition result and enhances the user experience when recognizing the user speech signal.

[0088] In some embodiments, optionally, determining the user voice signal based on the reference voice signal and the residual signal may specifically include:

[0089] The residual signal is subjected to nonlinear processing to determine the power spectrum corresponding to the residual signal;

[0090] The reference speech signal is denoised based on the power spectrum and the denoising algorithm to obtain the user speech signal.

[0091] The noise reduction algorithm can be the Wiener filter algorithm, or other algorithms capable of noise reduction; no specific limitation is made here.

[0092] Specifically, after determining the residual signal, nonlinear processing is applied to it, such as multiple Fourier transforms or other nonlinear operations, to obtain the processed signal. Then, the processed signal is estimated based on correlation estimation to determine the power spectrum corresponding to the residual signal. By applying the power spectrum and a noise reduction algorithm to the reference speech signal, an accurate user speech signal can be obtained.

[0093] In this embodiment, the above method can obtain a relatively accurate user voice signal with little or no noise, which is beneficial for subsequent speech recognition processing of the user voice signal.

[0094] In some embodiments, optionally, determining the user voice signal based on the reference voice signal and the residual signal may further include:

[0095] The user's voice signal is obtained by subtracting the reference voice signal from the residual signal.

[0096] In this embodiment, the method for determining the user's voice signal is simple and efficient.

[0097] Figure 5 This is a flowchart illustrating another audio signal processing method provided in this disclosure. This embodiment is an optimization based on the above embodiments. Optionally, this embodiment mainly provides a detailed explanation of the process of filtering the target audio signal to determine the residual signal. Figure 5 As shown, the method specifically includes the following steps:

[0098] S210, acquire the first sound signal acquired by the sound acquisition module of the external device and the second sound signal that the external device has sent to the display device.

[0099] S220, using a delay estimation method, determines the time delay between the original audio signal and the target audio signal in the second audio signal.

[0100] S2301, performs filtering on the target audio signal to determine the echo signal in the target audio signal.

[0101] Specifically, by filtering the target audio signal using a linear adaptive filter or other filters, the linear echo component in the target audio signal can be obtained, that is, the echo signal in the target audio signal can be determined.

[0102] S2302, subtract the target audio signal from the echo signal to obtain the residual signal.

[0103] After obtaining the echo signal in the target audio signal, the residual signal can be obtained by subtracting the target audio signal from the echo signal.

[0104] S240 processes the first sound signal by using the delay time, residual signal, and original audio signal to determine the user's voice signal.

[0105] The sound signal processing method provided in this embodiment acquires a first sound signal collected by the sound acquisition module of an external device and a second sound signal sent by the external device to a display device; determines the delay time between the original audio signal and the target audio signal in the second sound signal using a delay estimation method; filters the target audio signal to determine the echo signal in the target audio signal; subtracts the echo signal from the target audio signal to obtain a residual signal; and processes the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user's voice signal. In the above technical solution, by filtering the target audio signal to determine the echo signal in the target audio signal, and then subtracting the echo signal from the target audio signal to obtain the residual signal, the residual signal is more accurate. This facilitates the subsequent filtering of the sound of the program played by the display device from the first sound signal to obtain a more accurate user voice signal. Therefore, when recognizing the user's voice signal, it is beneficial to improve the accuracy of the recognition result and enhance the user experience.

[0106] Figure 6A This is a schematic diagram of a sound signal processing apparatus provided in an embodiment of this disclosure. This apparatus is configured in an electronic device and can implement the sound signal processing method described in any embodiment of this application. Figure 6A As shown, the device specifically includes the following:

[0107] The signal acquisition module 601 is used to acquire a first sound signal acquired by the sound acquisition module of the external device and a second sound signal sent by the external device to the display device. The first sound signal includes a user voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal.

[0108] The first determining module 602 is used to determine the delay time between the original audio signal and the target audio signal in the second sound signal by means of a delay estimation method;

[0109] The second determining module 603 is used to filter the target audio signal and determine the residual signal;

[0110] The third determining module 604 is used to process the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal.

[0111] As an optional implementation of this disclosure, Figure 6B This is a schematic diagram of the structure of the third determining module in the sound signal processing apparatus of this disclosure embodiment, as shown below. Figure 6B As shown, the third determining module 604 includes: an alignment unit 6041, a first determining unit 6042, and a second determining unit 6043;

[0112] The alignment unit 6041 is used to align the original audio signal based on the delay time to obtain a third sound signal.

[0113] The first determining unit 6042 is used to subtract the first sound signal from the third sound signal to obtain a reference speech signal;

[0114] And a second determining unit 6043, used to determine the user voice signal based on the reference voice signal and the residual signal.

[0115] As an optional implementation of this disclosure, the second determining unit 6043 is specifically used for:

[0116] The residual signal is subjected to nonlinear processing to determine the power spectrum corresponding to the residual signal;

[0117] The reference speech signal is denoised based on the power spectrum and the denoising algorithm to obtain the user speech signal.

[0118] As an optional implementation of this disclosure, the second determining module 603 is specifically used for:

[0119] The target audio signal is filtered to determine the echo signal in the target audio signal;

[0120] The residual signal is obtained by subtracting the target audio signal from the echo signal.

[0121] As an optional implementation of this disclosure, the first determining module 602 is specifically used for:

[0122] Feature extraction is performed on the original audio signal in the second sound signal to obtain the first acoustic feature;

[0123] Feature extraction is performed on the target audio signal to obtain the second acoustic feature;

[0124] The first acoustic feature and the second acoustic feature are input into the time delay estimation model to obtain the delay time.

[0125] As an optional implementation of this disclosure, the time delay estimation model is determined in the following manner:

[0126] Acquire training samples, wherein the training samples include first sound data collected by the sound acquisition module after the external device establishes a connection with different types of display devices and second sound data sent by the external device to the different types of display devices. Each first sound data includes user voice data and target audio data corresponding to the display device. The second sound data includes: the original audio data corresponding to the target audio data.

[0127] The time delay estimation model is trained using the training samples until the time delay estimation model converges, at which point training stops.

[0128] As an optional implementation of this disclosure, the above-described apparatus further includes:

[0129] The recognition processing module is used to process the first sound signal through the delay time, the residual signal and the original audio signal to determine the user voice signal, and then perform voice recognition processing on the user voice signal to obtain the control command corresponding to the user voice signal.

[0130] The control module is used to control the external device to perform corresponding operations based on the control commands.

[0131] The sound signal processing apparatus provided in this embodiment acquires a first sound signal acquired by the sound acquisition module of an external device and a second sound signal sent by the external device to a display device. The first sound signal includes a user's voice signal and a target audio signal played by the display device. The second sound signal includes an original audio signal corresponding to the target audio signal. A delay time between the original audio signal and the target audio signal in the second sound signal is determined using a delay estimation method. The target audio signal is filtered to determine a residual signal. The first sound signal is processed using the delay time, the residual signal, and the original audio signal to determine the user's voice signal. In this technical solution, by processing the first sound signal, the sound of the program played by the display device can be filtered out from the first sound signal, resulting in a more accurate user's voice signal. This improves the accuracy of the recognition result and enhances the user experience when recognizing the user's voice signal.

[0132] The sound signal processing apparatus provided in this disclosure can execute the sound signal processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0133] This disclosure provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any of the sound signal processing methods described in this disclosure.

[0134] The electronic device may be a personal computer (PC), a server, or a mainframe computer, etc., and this disclosure does not specifically limit it.

[0135] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 7 As shown, the electronic device includes a processor 710 and a storage device 720; the number of processors 710 in the electronic device can be one or more. Figure 7 Taking a processor 710 as an example; the processor 710 and the storage device 720 in the electronic device can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.

[0136] The storage device 720, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the sound signal processing method in the embodiments of this disclosure. The processor 710 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the storage device 720, thereby implementing the sound signal processing method provided in the embodiments of this disclosure.

[0137] Storage device 720 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, storage device 720 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, storage device 720 may further include memory remotely located relative to processor 710, which can be connected to an electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0138] The electronic device provided in this embodiment can be used to execute the sound signal processing method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0139] This disclosure also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program is used to implement the sound signal processing method provided in this disclosure and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0140] Of course, the computer-executable instructions provided in the embodiments of this disclosure are not limited to the method operations described above, but can also perform related operations in the sound signal processing method provided in any embodiment of this disclosure.

[0141] Based on the above description of the implementation methods, those skilled in the art will clearly understand that this disclosure can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.

[0142] It is worth noting that in the embodiments of the above-mentioned sound signal processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this disclosure.

[0143] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The above descriptions are merely specific embodiments of this disclosure, and the selection and description of these embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations suitable for specific use considerations. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A sound signal processing method, characterized in that, The method includes: The system acquires a first sound signal acquired by the sound acquisition module of an external device and a second sound signal sent by the external device to the display device. The first sound signal includes a user voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal. The delay time between the original audio signal and the target audio signal in the second audio signal is determined by a delay estimation method. The target audio signal is filtered to determine the residual signal; The user's voice signal is determined by processing the first sound signal using the delay time, the residual signal, and the original audio signal. The step of processing the first audio signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal includes: aligning the original audio signal based on the delay time to obtain a third audio signal; subtracting the first audio signal from the third audio signal to obtain a reference audio signal; and determining the user voice signal based on the reference audio signal and the residual signal. By using the delay time between the original audio signal and the target audio signal, the original audio signal is aligned with the target audio signal, and the third audio signal and residual signal are filtered out from the first audio signal.

2. The method according to claim 1, characterized in that, Determining the user voice signal based on the reference voice signal and the residual signal includes: The residual signal is subjected to nonlinear processing to determine the power spectrum corresponding to the residual signal; The reference speech signal is denoised based on the power spectrum and the denoising algorithm to obtain the user speech signal.

3. The method according to claim 1, characterized in that, The step of filtering the target audio signal to determine the residual signal includes: The target audio signal is filtered to determine the echo signal in the target audio signal; The residual signal is obtained by subtracting the target audio signal from the echo signal.

4. The method according to claim 1, characterized in that, Determining the time delay between the original audio signal and the target audio signal in the second audio signal using a delay estimation method includes: Feature extraction is performed on the original audio signal in the second sound signal to obtain the first acoustic feature; Feature extraction is performed on the target audio signal to obtain the second acoustic feature; The first acoustic feature and the second acoustic feature are input into the time delay estimation model to obtain the delay time.

5. The method according to claim 4, characterized in that, The time delay estimation model is determined in the following way: Acquire training samples, wherein the training samples include first sound data collected by the sound acquisition module after the external device establishes a connection with different types of display devices and second sound data sent by the external device to the different types of display devices. Each first sound data includes user voice data and target audio data corresponding to the display device. The second sound data includes: the original audio data corresponding to the target audio data. The time delay estimation model is trained using the training samples until the time delay estimation model converges, at which point training stops.

6. The method according to any one of claims 1-5, characterized in that, After processing the first audio signal using the delay time, the residual signal, and the original audio signal to determine the user's voice signal, the method further includes: The user's voice signal is processed by speech recognition to obtain the control command corresponding to the user's voice signal; The external device is controlled to perform corresponding operations based on the control commands.

7. A sound signal processing device, characterized in that, The device includes: The signal acquisition module is used to acquire a first sound signal acquired by the sound acquisition module of the external device and a second sound signal sent by the external device to the display device. The first sound signal includes a user voice signal and a target audio signal played by the display device. The second sound signal includes the original audio signal corresponding to the target audio signal. The first determining module is used to determine the time delay between the original audio signal and the target audio signal in the second sound signal by means of a delay estimation method; The second determining module is used to filter the target audio signal and determine the residual signal; The third determining module is used to process the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal. The process of processing the first sound signal using the delay time, the residual signal, and the original audio signal to determine the user voice signal includes: aligning the original audio signal based on the delay time to obtain a third sound signal; subtracting the third sound signal from the first sound signal to obtain a reference voice signal; determining the user voice signal based on the reference voice signal and the residual signal; aligning the original audio signal with the target audio signal using the delay time between the original audio signal and the target audio signal; and filtering out the third sound signal and the residual signal from the first sound signal.

8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Sound processing method and device and electronic equipment

    CN112785999A

  • Method and device for audio denoising

    WO2021196042A1