Display device and echo cancellation method
By identifying and filtering the sound output by the audio output device and the sub-sound type and sound source position information in the sound collector collected by the sound collector in the display device, the problem of accidentally deleting sound during echo cancellation in the prior art is solved, and high-quality sound acquisition is achieved.
Patent Information
- Application Number
- CN202510176204.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-24
AI Technical Summary
When the echo cancellation of existing display devices, they may accidentally delete the sound collected by the sound collector, resulting in poor sound quality.
By identifying the sound output by the audio output device and the sub-sound types in the sound collected by the sound collector, the sub-sound that is most likely to be echo is filtered out using the sound source position information, and the type recognition reliability of the sub-sound type is selected for screening, and finally using the characteristics of the sub-sound to filter out the echo sub-sound sound from the sound.
Accurately eliminate echoson sounds, maintain sound fidelity, and improve the sound quality collected by sound acquisition equipment.
Smart Images

Figure CN120201220A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a display device and an echo cancellation method. Background Art
[0002] With the rapid development of the functions of display devices, the functions that display devices can provide for users are becoming increasingly rich. Currently, display devices include smart TVs, smart set-top boxes, smart boxes, and products with smart display screens, etc. Taking smart TVs as an example, smart TVs have started to support functions such as video conferencing and KTV functions. In these scenarios, the sound collected by the sound collector and the sound that needs to be output by other audio output devices are synthesized by the smart TV and then output through the audio output device and enter the sound collection device again. The sound output by the audio output device may cause echo interference to the sound collected by the sound collector.
[0003] Traditional display devices perform echo cancellation mainly by directly removing the overall sound output by the audio output device from the sound collected by the sound collector, filtering out the sound whose frequency point threshold exceeds the preset frequency point threshold, or removing the sound highly similar to the echo source based on artificial intelligence learning. However, although these methods partially eliminate echoes, they may accidentally delete the sound collected by the sound collector, resulting in poor sound quality collected by the sound collection device. Summary of the Invention
[0004] This application provides a display device and an echo cancellation method to solve the problem that only part of the echo is eliminated, and the sound collected by the sound collector is accidentally deleted, resulting in poor sound quality collected by the sound collection device.
[0005] In a first aspect, some embodiments provide a display device, including: an audio output device, a sound collector, and a controller. Among them, the audio output device is configured to output the sound of the display device;
[0006] The sound collector is configured to collect the input interactive voice and the sound output by the audio output device;
[0007] The controller is configured to execute instructions to enable the display device to:
[0008] During the process of the audio output device outputting the first sound, in response to the input interactive voice, obtain the second sound collected by the sound collector; the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector;
[0009] Identify at least one sub-sound included in the first sound and the second sound respectively, obtain the sub-sound type to which each sub-sound belongs, and obtain the sound source position information of the sub-sound of the second sound;
[0010] When the first sound and the second sound contain the same type of reference sub - sound, for the reference sub - sounds belonging to the reference sub - sound type in the second sound, the reference sub - sounds are screened according to the sound source position information of the reference sub - sounds to obtain target sub - sounds;
[0011] Based on the type recognition credibility of the sub - sound type to which the target sub - sounds belong, the target sub - sound type is screened within the sub - sound type to which the target sub - sounds belong; according to the sound characteristics of the sub - sounds belonging to the target sub - sound type in the target sub - sounds and the sound characteristics of the sub - sounds belonging to the target sub - sound type in the first sound, the target sub - sounds are screened to obtain echo sub - sounds, and the echo sub - sounds are filtered out from the second sound.
[0012] Technical effect: Some embodiments provide a display device. During the process of the audio output device outputting the first sound, in response to the input interactive voice, the sound collector collects the second sound. The second sound contains the interactive voice and the sound formed when the first sound is collected by the sound collector. The first sound may cause echo interference to the second sound; by identifying the sub - sound types to which the sub - sounds contained in the first sound and the second sound respectively belong, and the first sound and the second sound contain the same type of reference sub - sound, it indicates that the reference sub - sounds belonging to the reference sub - sound type in the second sound may be echoes. Screening the reference sub - sounds according to the sound source position information of the reference sub - sounds can determine the target sub - sounds that are most likely to be echoes; then screening the target sub - sound type with a type recognition credibility that meets the conditions within the sub - sound type to which the target sub - sounds belong to improve the accuracy of echo cancellation; finally, using the sound characteristics of the sub - sounds belonging to the target sub - sound type to screen the echo sub - sounds from the target sub - sounds. This method of screening echo sub - sounds step by step is beneficial to accurately eliminate the echo sub - sounds in the second sound, making the second sound as faithful as possible and improving the sound quality collected by the sound collection device.
[0013] In a second aspect, some embodiments further provide an echo cancellation method, which is applied to the display device provided in the first aspect. The display device includes: an audio output device, a sound collector, and a controller. The method includes:
[0014] During the process of the audio output device outputting the first sound, in response to the input interactive voice, obtain the second sound collected by the sound collector; the second sound contains the interactive voice and the sound formed when the first sound is collected by the sound collector;
[0015] Respectively identify at least one sub - sound contained in the first sound and the second sound, obtain the sub - sound type to which each sub - sound belongs, and obtain the sound source position information of the sub - sounds of the second sound;
[0016] When the first sound and the second sound contain the same type of reference sub - sound, for the reference sub - sounds belonging to the reference sub - sound type in the second sound, the reference sub - sounds are screened according to the sound source position information of the reference sub - sounds to obtain target sub - sounds;
[0017] Based on the type recognition credibility of the sub - sound type to which the target sub - sounds belong, the target sub - sound type is screened among the sub - sound types to which the target sub - sounds belong; according to the sound characteristics of the sub - sounds belonging to the target sub - sound type in the target sub - sounds and the sound characteristics of the sub - sounds belonging to the target sub - sound type in the first sound, the target sub - sounds are screened to obtain echo sub - sounds, and the echo sub - sounds are filtered out from the second sound.
[0018] Technical effect: Some embodiments provide an echo cancellation method. During the output of the first sound by the audio output device, in response to the input interactive voice, the sound collector collects the second sound. The second sound contains the interactive voice and the sound formed when the first sound is collected by the sound collector. The first sound may cause echo interference to the second sound. By identifying the sub - sound types to which the sub - sounds contained in the first sound and the second sound belong respectively, and the first sound and the second sound contain the same type of reference sub - sound, it indicates that the reference sub - sounds belonging to the reference sub - sound type in the second sound may be echoes. Screening the reference sub - sounds according to the sound source position information of the reference sub - sounds can determine the target sub - sounds that are most likely to be echoes; then screening the target sub - sound type with a type recognition credibility that meets the conditions among the sub - sound types to which the target sub - sounds belong to improve the accuracy of echo cancellation; finally, using the sound characteristics of the sub - sounds belonging to the target sub - sound type to screen the echo sub - sounds from the target sub - sounds. This method of screening echo sub - sounds step by step is conducive to accurately eliminating the echo sub - sounds in the second sound, making the second sound as faithful as possible and improving the sound quality collected by the sound collection device. Brief Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided by some embodiments of the present application;
[0021] Figure 2 It is a schematic diagram of the hardware configuration of the display device provided by some embodiments of the present application;
[0022] Figure 3Schematic diagram of the hardware configuration of the control device provided by some embodiments of the present application;
[0023] Figure 4 Schematic diagram of the software configuration of the display device provided by some embodiments of the present application;
[0024] Figure 5 Schematic flowchart of the echo cancellation method provided by some embodiments of the present application;
[0025] Figure 6 Schematic sub - flowchart of S508 provided by some embodiments of the present application;
[0026] Figure 7 Schematic flowchart of the echo cancellation method provided by some other embodiments of the present application;
[0027] Figure 8 Timing diagram of the echo cancellation method provided by some embodiments of the present application;
[0028] Figure 9 Overall schematic flowchart of the echo cancellation method provided by some embodiments of the present application;
[0029] Figure 10 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0030] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0031] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0032] The terms "first", "second", "third", etc. in the specification, claims and the above - mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0033] The terms "comprising" and "having" and any variations thereof are intended to cover inclusion without exclusion. For example, a product or device comprising a series of components need not be limited to all the components clearly listed, but may include other components not clearly listed or inherent to such products or devices.
[0034] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code capable of performing functions related to that element.
[0035] In the embodiments of the present application, the display device 200 generally refers to a device having the capabilities of displaying images and processing data. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0036] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided for some embodiments of the present application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300, and the control device 100. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.
[0037] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 to perform data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and connection communication can be achieved through network communication protocols to achieve the purpose of one-to-one control operations and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.
[0038] As Figure 1 also shown, the display device 200 also performs data communication with the server 400 through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0039] The display device 200 can provide a broadcast reception TV function, and can also additionally provide an intelligent network TV function with computer support functions, including but not limited to, network TV, smart TV, Internet Protocol TV (IPTV), etc.
[0040] Figure 2 For some embodiments of the present application Figure 1 is the hardware configuration block diagram of the display device 200.
[0041] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0042] In some embodiments, the detector 230 is configured to collect signals of the external environment or for interacting with the outside. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; alternatively, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.
[0043] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving image display. The display 260 is configured to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.
[0044] In some embodiments, the communication device 220 is a component for communicating with external devices or a server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0045] The communication device 220 can enable the display device 200 to communicate with external devices or the server 400 in a wireless or wired connection manner. Among them, the wired connection can connect the display device 200 with an external device through components such as a data cable and an interface. The wireless connection can connect the display device 200 with an external device through a wireless signal or a wireless network. The display device 200 can directly establish a connection relationship with an external device, or can indirectly establish a connection relationship through a gateway, a router, a connection device, etc.
[0046] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first interfaces to n interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0047] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different split devices, that is, the tuner-demodulator 210 may also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0048] In some embodiments, the user may input a user command in the graphical user interface (GUI) displayed on the display 260, and then the user input interface receives the user input command through the graphical user interface (GUI).
[0049] In some embodiments, the audio output device 270 may be the built-in speaker of the display device 200, or may be an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0050] In some embodiments, the user input interface 280 can be used to receive instructions from user input.
[0051] Figure 3 The hardware configuration block diagram of the control device provided by some embodiments of the present application. As Figure 1 shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply. Figure 3
[0052] The control device 100 is configured to control the display device 200, and can receive input operation instructions from the user, and convert the operation instructions into instructions recognizable and responsive by the display device 200, playing an intermediary role in the interaction between the user and the display device 200.
[0053] In some embodiments, the control device 100 may be an intelligent device. For example: the control device 100 can install various applications for controlling the display device 200 according to user needs.
[0054] Figure 1 In some embodiments, as shown, after installing the application for controlling the display device 200, the mobile terminal 300 or other intelligent electronic devices can play a similar function to the control device 100.
[0055] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and internal.
[0056] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as a WiFi chip 131, a Bluetooth module 132, an NFC module 133, etc.
[0057] The user input / output interface 140, where the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a button 144, and other input interfaces.
[0058] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The communication interface 130 configured in the control device 100, such as modules like WiFi, Bluetooth, NFC, etc., can encode user input instructions through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol and send them to the display device 200.
[0059] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0060] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0061] To perform user interaction, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program for managing and controlling the hardware resources and software resources in the display device 200. The operating system can (control the display device) provide a user interface, allowing the user to interact with the display device 200 and support the running of various application programs.
[0062] It should be noted that the operating system can be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system developed specifically for the display device.
[0063] The operating system can be divided into different modules or levels according to the functions implemented, such as Figure 4 as shown Figure 4 For some embodiments provided in this application Figure 1 is a schematic diagram of the software configuration in the display device. In some embodiments, the system of the display device 200 can be divided into three layers, from top to bottom are the application layer, the middleware layer, and the hardware layer.
[0064] The application layer mainly includes common applications on the TV and the Application Framework. Among them, the common applications are mainly applications developed based on the Browser, such as HTML5 APPs, and native applications (Native APPs). The Application Framework is a complete program model with all the basic functions required for standard application software, such as file access, data exchange, etc., and the usage interfaces for these functions (toolbars, status bars, menus, dialog boxes).
[0065] Native applications (Native APPs) can support online or offline, message push or local resource access.
[0066] The middleware layer includes various middleware such as TV protocols, multimedia protocols, and system components. The middleware can use the basic services (functions) provided by the system software to connect various parts of the application system on the network or different applications, and can achieve the purpose of resource sharing and function sharing.
[0067] The hardware layer mainly includes the Hardware Abstraction Layer (HAL) interface, hardware, and drivers. Among them, the HAL interface is the unified interface for all TV chips to connect, and the specific logic is implemented by each chip. The drivers mainly include: audio drivers, display drivers, Bluetooth drivers, camera drivers, WIFI drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.
[0068] It should be noted that the above examples are only simple divisions of the functions of the operating system and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of levels and the specific level types included in the operating system can be in other forms.
[0069] With the rapid development of the functions of display devices, the functions that display devices can provide for users are becoming increasingly rich. Currently, display devices include smart TVs, smart set-top boxes, smart boxes, and products with smart display screens, etc. Taking smart TVs as an example, smart TVs have started to support functions such as video conferencing and KTV functions. In these scenarios, the sound (Sound2_1) received by the sound collector is received by the smart TV through a receiver (such as Bluetooth, USB, or AV), and after a series of sound effects processing, such as noise reduction, gain addition, etc., the processed sound (Sound2_2) is obtained. Sound2_2 and other sounds Sound3 that need to be output by the audio output device (such as the background music played by the smart TV through the audio output device) are synthesized by the smart TV and then output as Sound1 through the audio output device. When Sound1 enters the sound collection device again, it is mixed with the newly input sound (Sound2_3) of the sound collection device, and Sound1 may cause echo interference to Sound2_3 to form (Sound2_0). Therefore, if the sound collector directly inputs Sound2_0 containing the echo source Sound1 into the smart TV without echo cancellation, the echo component Sound1 in Sound2_0 will be amplified by gain and then input into the audio output device, and the echo sound in the audio output device will be transmitted back to the sound collector again to form (Sound3_0), and after subsequent gain amplification, Sound1 will be infinitely amplified in a short time, resulting in sharp and harsh sounds such as echo howling.
[0070] For traditional display devices to perform echo cancellation, it mainly directly removes the overall sound output by the audio output device from the sound collected by the sound collector, filters out the sound whose frequency point threshold exceeds the preset frequency point threshold, or removes the sound highly similar to the echo source based on the method of artificial intelligence learning. Among them, the method of directly removing the overall sound output by the audio output device does not consider the correlation between the sound collected by the sound collector and the sound output by the audio output device, and may accidentally delete the sound collected by the sound collector; the method of filtering out the sound whose frequency point threshold exceeds the preset frequency point threshold can directly and quickly suppress the howling problem caused by echoes, but it may also remove the sound collected by the sound collector; the method of removing the sound highly similar to the echo source based on the method of artificial intelligence learning learns the echo source and the newly input sound source, compares their similarities and differences, and removes the same sound, but the sound highly similar to the echo source in the sound collected by the sound collector may also be the newly input sound source. Therefore, although the traditional echo cancellation method partially eliminates echoes, it may accidentally delete the sound collected by the sound collector, resulting in the problem of poor sound quality collected by the sound collection device.
[0071] Based on this, an embodiment of the present application provides a display device and an echo cancellation method. The display device includes: an audio output device, a sound collector, and a controller. Among them, the audio output device is configured to output the sound of the display device; the sound collector is configured to collect the input interactive voice and the sound output by the audio output device; the controller is configured to execute instructions to enable the display device, during the process of the audio output device outputting a first sound, in response to the input interactive voice, the sound collector collects a second sound, and the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector. The first sound may cause echo interference to the second sound. By identifying the sub-sound types to which the sub-sounds included in the first sound and the second sound respectively belong, and the first sound and the second sound include the same reference sub-sound type, it is characterized that the reference sub-sound belonging to the reference sub-sound type in the second sound may be an echo. By screening the reference sub-sound according to the sound source position information of the reference sub-sound, the target sub-sound most likely to be an echo can be determined; then, in the sub-sound type to which the target sub-sound belongs, the target sub-sound type with a type recognition credibility meeting the conditions is screened to improve the accuracy of echo cancellation; finally, the echo sub-sound is screened from the target sub-sounds by using the sound characteristics of the sub-sounds belonging to the target sub-sound type. This method of screening echo sub-sounds step by step is beneficial to accurately cancel the echo sub-sounds in the second sound, making the second sound as faithful as possible and improving the sound quality collected by the sound collection device.
[0072] In an exemplary embodiment, as Figure 5 shown, an echo cancellation method is provided. Taking the application of this method to the display device 200 as an example, the following steps S502 to S508 are included.
[0073] S502, during the process of the audio output device outputting a first sound, in response to the input interactive voice, obtain the second sound collected by the sound collector; the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector.
[0074] Among them, the first sound is the sound output by the audio output device. The audio output device can output different sounds according to the requirements of the application scenario. For example, in a KTV scenario, the first sound can be the background music output by the display device through the audio output device, or the sound collected by the sound collector and input to the display device, output through the audio output device, or the sound obtained after the background music and the sound collected by the sound collector and input to the display device are synthesized by the display device.
[0075] Interactive voice refers to the voice input by the user. In a voice control scenario, it can be a control instruction in the form of voice for a display device. In a karaoke scenario, it can be the singing voice of the user. In a video conferencing scenario, it can be the voice input by the user that contains the conference content.
[0076] The second sound refers to the sound collected by the sound collector. Since the first sound is distorted due to multiple reflections in a closed or semi-closed environment during the process of the audio output device outputting the first sound, the second sound contains the sound formed when the first sound is collected by the sound collector. At the same time, the interactive voice is also collected by the sound collector. Therefore, the second sound contains the interactive voice and the sound formed when the first sound is collected by the sound collector.
[0077] Since the first sound (Sound1) will cause echo interference to the interactive voice, the present application proposes that after the audio output device outputs the first sound (Sound1), the controller performs echo cancellation on the second sound (Sound2_0) to improve the sound quality collected by the sound collector. S504: respectively identify at least one sub-sound included in the first sound and the second sound, obtain the sub-sound type to which each sub-sound belongs, and obtain the sound source position information of the sub-sounds of the second sound.
[0078] Among them, the first sound and the second sound are respectively composed of at least one sub-sound. Each sub-sound may belong to different sub-sound types, or some sub-sounds belong to the same sub-sound type. The timbres corresponding to different sub-sound types may be different. For example, the sub-sound types may include child singing type (Child_Singing), guitar type (Guitar), music type (Music), drum type (Drum), bird call type (Bird), rain sound type (Rain), etc.
[0079] The controller can use a sound recognition algorithm or an artificial intelligence model to respectively identify at least one sub-sound included in the first sound and the second sound, and obtain the sub-sound type to which each sub-sound belongs.
[0080] The sound source position information refers to the position information of the sound source that emits the sub-sound. Exemplarily, the position information may include at least one of the distance between the sub-sound and the sound collector, the direction of the sub-sound relative to the sound collector, and the three-dimensional angle. For example, the second sound contains a sub-sound type, and each sub-sound type includes a sub-sound. The sound source position information LocationM2_0 of each sub-sound in the second sound can be expressed as follows:
[0081] LocationM2_0:
[0082]
[0083] For example: LocationM2_0:
[0084]
[0085] In some embodiments, the controller can obtain the sound source position information of each sub - sound by means of the data communicated between the audio output device and the sound collector, by sensor ranging, or by taking images.
[0086] S506. When the first sound and the second sound contain the same type of reference sub - sound, for the reference sub - sounds belonging to the reference sub - sound type in the second sound, according to the sound source position information of the reference sub - sounds, screen the reference sub - sounds to obtain target sub - sounds.
[0087] Among them, the reference sub - sound type refers to the sub - sound type jointly contained in the first sound and the second sound. The first sound and the second sound may or may not contain the same type of sub - sound. When the first sound and the second sound contain the same type of reference sub - sound, it indicates that the first sound causes an echo of the second sound, and echo cancellation processing needs to be performed on the second sound.
[0088] The reference sub - sound refers to the sub - sound belonging to the reference sub - sound type in the second sound. In the echo cancellation process, according to the sound source position information of the reference sub - sounds, target sub - sounds that may be echoes can be further screened from the reference sub - sounds. Exemplarily, if the sound source position information indicates that the reference sub - sound comes from the audio output device, it indicates that the reference sub - sound may be an echo, and this reference sub - sound is used as the target sub - sound. The target sub - sound is a sub - sound that is roughly screened from the reference sub - sounds and may be an echo.
[0089] S508. Based on the type recognition confidence of the sub - sound type to which the target sub - sound belongs, screen the target sub - sound type among the sub - sound types to which the target sub - sound belongs; according to the sound characteristics of the sub - sounds belonging to the target sub - sound type in the target sub - sounds and the sound characteristics of the sub - sounds belonging to the target sub - sound type in the first sound, screen the target sub - sounds to obtain echo sub - sounds, and filter out the echo sub - sounds from the second sound.
[0090] Among them, the type recognition confidence of the sub - sound type refers to the confidence level that there is a corresponding sub - sound type in the first sound or the second sound. The first sound or the second sound may contain at least one sub - sound type, and the type recognition confidences of various sub - sound types are different. The composition matrix CompositionM1 of the first sound Sound1 is composed of the sub - sound types contained in the first sound and the type recognition confidences of the sub - sound types:
[0091]
[0092] CompositionM1:
[0093]
[0094] For example, CompositionM1 (m = 6):
[0095]
[0096] Wherein, m represents the number of sub - sound types included in the first sound, Type[i] represents the i - th sub - sound type, and Prob1[i] represents the type recognition credibility of the i - th sub - sound type.
[0097] Similarly, the composition matrix CompositionM2_0 of the second sound Sound2 is composed of the sub - sound types included in the second sound and the type recognition credibility of the sub - sound types:
[0098]
[0099] CompositionM2_0:
[0100]
[0101] For example, CompositionM2_0 (n = 5)
[0102]
[0103] Wherein, n represents the number of sub - sound types included in the second sound, Type[j] represents the j - th sub - sound type, and Prob1[j] represents the type recognition credibility of the j - th sub - sound type.
[0104] Referring to the above embodiments, by taking the intersection of CompositionM1 and CompositionM2_0, the reference sub - sound types can be determined, and the reference sub - sound types include Child_Singing, Guitar, Music, Drum.
[0105] The greater the type recognition credibility of the sub - sound type, the greater the possibility that there is a corresponding sub - sound type in the first sound or the second sound. The greater the type recognition credibility of the sub - sound type, it indicates that the corresponding sub - sound type may be noise in the first sound or the second sound, or a mis - recognized sub - sound type. For example, noises such as faint or continuous low - volume background sounds that are far from the sound collector. The controller can use, as the target sub - sound type, the sub - sound types in the sub - sound type to which the target sub - sound belongs, where the type recognition credibility is greater than the preset credibility or within the preset credibility interval, so as to remove the noise in the target sub - sound.
[0106] The controller can use a feature extraction algorithm to extract the sound features of each sub - sound. The feature extraction algorithm includes but is not limited to any one of Mel - Frequency Cepstral Coefficients (MFCC), Linear Predictive Coefficients (LPC), and Short - Time Fourier Transform (STFT). The sound features are the key information for analyzing the echo source. If there is an echo source in the second sound that is the same as the echo source in the first sound, there is a certain degree of similarity between the sound features of the sub - sounds belonging to the target sub - sound type in the target sub - sound and the sound features of the sub - sounds belonging to the target sub - sound type in the first sound. Based on the sound features of the two, the echo sub - sounds can be further screened out from the target sub - sound.
[0107] The controller can use a filter to filter out the echo sub - sounds from the second sound. Exemplarily, an adaptive filter can be used, such as a filter based on the LMS (Least Mean Squares) algorithm or a filter based on the RLS (Recursive Least Squares) algorithm, etc.
[0108] The echo sub - sounds determined through multiple screenings have high accuracy. Filtering out the echo sub - sounds from the second sound accurately eliminates the echo in the second sound, making the second sound as faithful to the original as possible and improving the sound quality collected by the sound collector.
[0109] Some embodiments provide an echo cancellation method. During the process of an audio output device outputting a first sound, in response to an input interactive voice, a sound collector collects a second sound. The second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector. The first sound may cause echo interference to the second sound. By identifying the sub-sound types to which the sub-sounds included in the first sound and the second sound respectively belong, and the first sound and the second sound include the same reference sub-sound type, it indicates that the reference sub-sound belonging to the reference sub-sound type in the second sound may be an echo. By screening the reference sub-sound according to the sound source position information of the reference sub-sound, the target sub-sound that is most likely to be an echo can be determined. Then, in the sub-sound type to which the target sub-sound belongs, the target sub-sound type with a type recognition credibility meeting the conditions is screened to improve the accuracy of echo cancellation. Finally, using the sound characteristics of the sub-sounds belonging to the target sub-sound type to screen the echo sub-sounds from the target sub-sounds. This method of screening echo sub-sounds step by step is beneficial to accurately cancel the echo sub-sounds in the second sound, making the second sound as faithful as possible and improving the sound quality collected by the sound collection device.
[0110] In an exemplary embodiment, as Figure 6 described, screening the target sub-sound according to the sound characteristics of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the sound characteristics of the sub-sounds belonging to the target sub-sound type in the first sound to obtain the echo sub-sound, including:
[0111] S602, obtaining the similarity between the sound characteristics of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the sound characteristics of the sub-sounds belonging to the target sub-sound type in the first sound.
[0112] S604, when the similarity is not less than the first similarity threshold, determining the target sub-sound as the echo sub-sound.
[0113] Wherein, the controller calculates the similarity between the target sub-sound and the sound characteristics of the sub-sounds belonging to the target sub-sound type in the first sound. For example, the target sub-sound includes a' types of target sub-sound types, and the controller calculates the similarity of the sound characteristics between the target sub-sounds belonging to the same target sub-sound type and the sub-sounds of the first sound, avoiding the problem that the similarity calculated between different target sub-sound types meets the similarity conditions, resulting in incorrect echo recognition, which is beneficial to improving the accuracy of echo recognition.
[0114] A similarity not less than the first similarity threshold indicates a relatively high similarity between the acoustic features of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the acoustic features of the sub-sounds belonging to the target sub-sound type in the first sound. The sub-sound belonging to the target sub-sound type in the first sound is the sound source of the sub-sound belonging to the target sub-sound type in the target sub-sound. Therefore, the sub-sound belonging to the target sub-sound type in the target sub-sound should be regarded as an echo sub-sound, that is, this echo sub-sound comes from the audio output device, rather than an independently newly emitted sound from the display device, and this echo sub-sound will cause an echo.
[0115] Exemplarily, the acoustic feature of the sub-sound belonging to the target sub-sound type in the first sound is FeatureM1, and the acoustic feature of the sub-sound belonging to the target sub-sound type in the target sub-sound is FeatureM2_0. Taking the example that each target sub-sound type contains one sub-sound, the controller calculates the similarity between the acoustic feature FeatureM1[i] corresponding to sub-sound i in FeatureM1 and the acoustic feature FeatureM2_0[i] corresponding to sub-sound i in FeatureM2_0. When the similarity is not less than the first similarity threshold (for example, it can be 10%), the sub-sound i belonging to the target sub-sound type in the target sub-sound is regarded as an echo sub-sound and removed from the second sound.
[0116] In this embodiment, by comparing the similarity between the acoustic features of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the acoustic features of the sub-sounds belonging to the target sub-sound type in the first sound, the echo sub-sounds whose acoustic features meet the feature similarity condition are further screened out from the target sub-sounds, improving the accuracy of echo recognition.
[0117] In an exemplary embodiment, as Figure 7 described, the echo cancellation method further includes:
[0118] S702, when the similarity is less than the first similarity threshold, update the feature extraction algorithm matching the target sub-sound type, and re-extract the acoustic features of the target sub-sound and the acoustic features of the sub-sounds belonging to the target sub-sound type in the first sound based on the updated feature extraction algorithm.
[0119] S704, when the similarity between the re-extracted acoustic features of the target sub-sound and the re-extracted acoustic features of the sub-sound is not less than the second similarity threshold, determine that the target sub-sound is an echo sub-sound.
[0120] Among them, the similarity being less than the first similarity threshold indicates that the similarity between the voice features of the sub-voices belonging to the target sub-voice type in the target sub-voice and the voice features of the sub-voices belonging to the target sub-voice type in the first voice is relatively low. To further improve the accuracy of echo recognition, the controller updates the feature extraction algorithm matching the target sub-voice type, re-extracts the voice features using the updated feature extraction algorithm, and recalculates the similarity between the voice features.
[0121] When the recalculated similarity is not less than the second similarity threshold, it is determined that the target sub-voice is an echo sub-voice. The recalculated similarity not being less than the second similarity threshold indicates that the sub-voice belonging to the target sub-voice type in the first voice is not the sound source of the sub-voice belonging to the target sub-voice type in the target sub-voice, and the sub-voice belonging to the target sub-voice type in the target sub-voice is not an echo sub-voice and does not need to be eliminated. Among them, since the feature extraction algorithm has changed, the first similarity threshold and the second similarity threshold may be different.
[0122] In this embodiment, when the similarity between the voice features of the sub-voices belonging to the target sub-voice type in the target sub-voice and the voice features of the sub-voices belonging to the target sub-voice type in the first voice is less than the first similarity threshold, the controller updates the voice features by updating the feature extraction algorithm and re-determines the similarity between the updated voice features, so as to more accurately and correctly identify the echo and avoid mis-eliminating normal voices.
[0123] In an exemplary embodiment, the type recognition credibility includes the first type recognition credibility of the sub-voice type to which the target sub-voice belongs in the first voice, and the second type recognition credibility of the sub-voice type to which the target sub-voice belongs in the second voice; the controller executes type recognition credibility based on the sub-voice type to which the target sub-voice belongs to screen the target sub-voice type among the sub-voice types to which the target sub-voice belongs, and is configured to: exclude the sub-voice types in which at least one of the first type recognition credibility or the second type recognition credibility satisfies the first credibility condition from the sub-voice types to which the target sub-voice belongs to obtain the target sub-voice type.
[0124] Among them, the controller takes the type recognition credibility of the sub-voice type to which the target sub-voice belongs in the first voice as the first type recognition credibility, and takes the type recognition credibility of the sub-voice type to which the target sub-voice belongs in the second voice as the second type recognition credibility.
[0125] The first credibility condition may be that the recognition credibility of the first type is lower than the first preset credibility, or the recognition credibility of the second type is lower than the second preset credibility, or the recognition credibility of the first type is lower than the first preset credibility and the recognition credibility of the second type is lower than the second preset credibility. The sub - sound type that satisfies the first credibility condition among the sub - sound types to which the target sub - sound belongs refers to the noise or mis - recognized sub - sound type in the first sound or the second sound.
[0126] The controller eliminates the sub - sound type that satisfies the first credibility condition from the sub - sound types to which the target sub - sound belongs, and obtains the target sub - sound type. For example, both the first preset credibility and the second preset credibility are 10%. The recognition credibility of the first type corresponding to the sub - sound type to which the target sub - sound belongs is 8% (refer to Drum in CompositionM2_0), and the recognition credibility of the second type is 5% (refer to Rain in CompositionM1). Then the controller can eliminate this sub - sound type from the sub - sound types to which the target sub - sound belongs to obtain the target sub - sound type.
[0127] In this embodiment, the controller eliminates the sub - sound type that satisfies the first credibility condition from the sub - sound types to which the target sub - sound belongs, which is beneficial to eliminating the noise or mis - recognized sub - sound type in the target sub - sound and improving the accuracy of echo recognition.
[0128] In an exemplary embodiment, filtering out the echo sub - sound from the second sound includes: determining a reference sub - sound type whose type recognition credibility satisfies the second credibility condition among the sub - sound types included in the second sound; filtering out the echo sub - sound from the second sound when the similarity between the sound characteristics of the echo sub - sound and the sound characteristics of the sub - sound belonging to the reference sub - sound type in the second sound is not less than the third similarity threshold.
[0129] Among them, the second credibility condition may be that the type recognition credibility in the sub - sound types included in the second sound is higher than the preset credibility. Determining that the type recognition credibility satisfies the second credibility condition in the sub - sound types included in the second sound indicates that this sub - sound type is a sub - sound type with a higher confidence level in the second sound.
[0130] Exemplarily, the preset credibility is 80%, and the type recognition credibility of Child_Singing in CompositionM2_0 is 90%. Then Child_Singing can be used as the reference sub - sound type.
[0131] In some embodiments, other methods can also be used to determine the reference sub - sound type. For example, the preset sub - sound type can be used as the reference sub - sound type, and the sub - sound type with the highest type recognition credibility in the second sound or the combination of multiple sub - sound types in the second sound, etc., can be flexibly selected according to the actual situation.
[0132] The controller calculates the similarity between the sound characteristics of the echo sub - sound and the sound characteristics of the sub - sound belonging to the reference sub - sound type in the second sound. If the similarity is less than the third similarity threshold, it indicates that the echo sub - sound is quite different from the sub - sound belonging to the reference sub - sound type. If this echo sub - sound is eliminated, it will affect the main sound of the second sound. Therefore, when the similarity is less than the third similarity threshold, there is no need to eliminate this echo sub - sound. On the contrary, when the similarity is not less than the third similarity threshold, this echo sub - sound is eliminated from the second sound, and the main sound of the second sound can still be retained with high - fidelity.
[0133] In this embodiment, the controller selects the reference sub - sound type from the second sound and compares the similarity between the echo sub - sound and the sub - sound belonging to the reference sub - sound type, avoiding the elimination of the echo sub - sound that affects the subjectivity of the second sound, making the second sound as faithful as possible, and improving the sound quality collected by the sound acquisition device.
[0134] In some embodiments, the echo cancellation method further includes: when the first sound and the second sound do not contain the same reference sub - sound type, directly using the second sound as the sound input to the sound collector.
[0135] Among them, when the first sound and the second sound do not contain the same reference sub - sound type, it indicates that the first sound does not cause the echo of the second sound, the first sound is not beneficial to the echo cancellation of the second sound, and the current echo cancellation is completed.
[0136] The controller can directly use the second sound as the sound input to the controller. This second sound can be processed by the controller and transmitted to the audio output device, and synthesized with other sounds that need to be transmitted to the audio output device to form a new round of the first sound, and then output from the audio output device.
[0137] In this embodiment, when the first sound and the second sound do not contain the same reference sub - sound type, the controller completes the current round of echo cancellation, and directly uses the second sound as the sound input to the controller, avoiding the misdeletion of the sub - sounds in the second sound that do not belong to the echo, making the second sound as faithful as possible, and improving the sound quality collected by the sound acquisition device.
[0138] In an exemplary embodiment, filtering the reference sub-sound according to the sound source position information of the reference sub-sound to obtain the target sub-sound includes: obtaining the position information of the audio output device; determining the degree of position proximity according to the sound source position information and the position information of the audio output device; and determining the reference sub-sound as the target sub-sound when the degree of position proximity is less than a preset degree.
[0139] Among them, the position information of the audio output device indicates the distance, direction, and three-dimensional angle of the audio output device relative to the sound collector. The degree of position proximity refers to the differences in distance, direction, and three-dimensional angle in space between the sound source position information and the audio output device.
[0140] The degree of position proximity being less than the preset degree can mean that the distance in space between the sound source position information and the audio output device is less than the preset distance, the directions are the same, and the difference in each angle of each three-dimensional angle is less than the preset angle difference. The degree of position proximity being less than the preset degree indicates that the sound source position information and the audio output device are relatively close, and the possibility that the reference sub-sound corresponding to the sound source position information comes from the audio output device is relatively high. Therefore, the controller can use this reference sub-sound as the target sub-sound.
[0141] For example, the position information of the audio output device indicates that the distance of the audio output device relative to the sound collector is 15, the direction is North, and the three-dimensional angle is (15°, 16°, 20°), and the sound source position information is North and the three-dimensional angle is (15°, 18°, 27°). The controller determines that the degree of position proximity is less than the preset degree and uses the reference sub-sound as the target sub-sound.
[0142] In this embodiment, the controller filters out the target sub-sound that is relatively close to the audio output device from the reference sub-sounds by comparing the degree of position proximity between the sound source position information and the audio output device, which is beneficial to identifying the target sub-sound that may be an echo from the reference sub-sounds.
[0143] To illustrate the effects of the display device and the echo cancellation method in this solution in detail, the following uses a most detailed embodiment for illustration:
[0144] The display device includes: an audio output device, a sound collector, and a controller. Among them, the audio output device is configured to output the sound of the display device; the sound collector is configured to collect the input interactive voice and the sound output by the audio output device; and the controller is configured to execute the echo cancellation method. The echo cancellation method can be applied to various scenarios, such as video conferencing scenarios, KTV scenarios, voice control scenarios, etc.
[0145] As Figure 8The timing diagram of the echo cancellation method provided by some embodiments of the present application. During the process of the audio output device outputting the first sound Sound1, the controller, in response to the input interactive voice, acquires the second sound Sound2_0 collected by the sound collector; the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector. The controller respectively identifies at least one sub-sound included in each of the first sound Sound1 and the second sound Sound2_0, obtains the sub-sound type to which each sub-sound belongs, and acquires the sound source position information of the sub-sounds of the second sound Sound2_0. When the first sound Sound1 and the second sound Sound2_0 include the same reference sub-sound type, for the reference sub-sounds belonging to the reference sub-sound type in the second sound Sound2_0, the reference sub-sounds are screened according to the sound source position information of the reference sub-sounds to obtain the target sub-sounds. Based on the type recognition credibility of the sub-sound type to which the target sub-sounds belong, the target sub-sound type is screened in the sub-sound type to which the target sub-sounds belong; according to the sound characteristics of the sub-sounds belonging to the target sub-sound type in the target sub-sounds and the sound characteristics of the sub-sounds belonging to the target sub-sound type in the first sound Sound1, the target sub-sounds are screened to obtain the echo sub-sounds, and the echo sub-sounds are filtered out from the second sound Sound2_0.
[0146] Such as Figure 9 The overall flowchart of the echo cancellation method provided by some embodiments of the present application.
[0147] S1. Acquire at least one sub-sound included in the first sound Sound1 and the sub-sound type to which each sub-sound belongs.
[0148] Perform intelligent identification of the sound source composition type on the first sound Sound1 to obtain at least one sub-sound included in the first sound Sound1, the sub-sound type to which each sub-sound belongs, and the type recognition credibility of each sub-sound type. Sort the various sub-sound types in descending order of the type recognition credibility, and the sub-sound type with the largest type recognition credibility is the main sound type. The composition matrix CompositionM1 of the first sound Sound1 is composed of the sub-sound type included in the first sound and the type recognition credibility of the sub-sound type:
[0149]
[0150] CompositionM1:
[0151]
[0152] For example, CompositionM1(m = 6)
[0153]
[0154] Among them, m represents the number of sub - sound types included in the first sound, Type[i] represents the i - th sub - sound type, and Prob1[i] represents the type recognition confidence of the i - th sub - sound type.
[0155] S2. Obtain at least one sub - sound included in the second sound Sound2_0 and the sub - sound type to which each sub - sound belongs.
[0156] Perform intelligent sound source type recognition on the second sound Sound2_0 to obtain at least one sub - sound included in the second sound Sound2_0, the sub - sound type to which each sub - sound belongs, and the type recognition confidence of each sub - sound type. Sort each sub - sound type in descending order of the type recognition confidence, and the sub - sound type with the largest type recognition confidence is the main sound type. The composition matrix CompositionM2_0 of the second sound Sound2_0 is composed of the sub - sound types included in the second sound and the type recognition confidence of the sub - sound types:
[0157]
[0158] CompositionM2_0:
[0159]
[0160] For example, CompositionM2_0(n = 5):
[0161]
[0162] Among them, n represents the number of sub - sound types included in the second sound, Type[j] represents the j - th sub - sound type, and Prob1[j] represents the type recognition confidence of the j - th sub - sound type.
[0163] S3. Determine the difference in sub - sound types included in the first sound Sound1 and the second sound Sound2_0 respectively.
[0164] Determine the intersection of the sound types between the composition matrix CompositionM2_0 of sound 2_0 and the composition matrix CompositionM1 of sound 1.
[0165] If there is no intersection, it indicates that the first sound Sound1 has not caused an echo of the second sound Sound2_0, and the first sound Sound1 is not beneficial to the echo cancellation of the second sound Sound2_0, and the current round of echo cancellation is completed.
[0166] If there is an intersection, then perform subsequent echo cancellation.
[0167] For example, when taking the intersection of CompositionM2_0 (n = 5) and CompositionM1 (m = 4), the resulting reference sub - sound types are as follows:
[0168]
[0169] If there is an intersection and the number of intersection types is a, then delete the information rows of the corresponding types of the sub - sound types not in the intersection from CompositionM1 of the first sound Sound1. The corresponding matrix after cleaning is CompositionM1_2, and the number of remaining member rows of it is all a. Synchronously, delete the information rows of the corresponding types of the sub - sound types not in the intersection from CompositionM2_0 of the second sound Sound2_0. The corresponding matrix after cleaning is CompositionM2_1, and the number of remaining member rows of it is all a.
[0170] For example, delete the sub - sound types not in the intersection from CompositionM1 (n = 6) to obtain CompositionM1_2 (a = 4):
[0171]
[0172] Delete the sub - sound types not in the intersection from CompositionM2_0 (n = 5) to obtain CompositionM2_1 (a = 4):
[0173]
[0174] For the latest remaining a sub - sound types in the second sound Sound2_0, collect the sound source position information of each sub - sound in the sub - sound types, determine the distance, direction, three - dimensional angle and other position information between each sub - sound of the second sound Sound2_0 and the sound collector, and form the position matrix LocationM2_0 of the second sound Sound2_0. The number of row members of LocationM2_0 is a, and the number of its columns is the position information dimension of each member (where one column is the distance between the sound source of the sub - sound and the sound collector).
[0175] LocationM2_0:
[0176]
[0177] For example, LocationM2_0:
[0178]
[0179] Specifically, if a certain sub - sound type Type[j] is measured to have more than one distance information in the second sound Sound2_0, that is, for the same type of sound, K[j] sound sources at multiple different positions are found (the specific details such as the position / sound intensity / energy / phase of each sound source may be different), it indicates that the sound sources of the same sub - sound type arriving at the sound collector simultaneously may have multiple sources at different heights, or come from the echo of the same sound source in the subsequent sound and the previous time period. Therefore, in this case, (K[j] - 1) new members of the same sub - sound type Type[j] with different sound source position information need to be added correspondingly in LocationM2_0. After adding, the number of rows of the original a - row members in LocationM2_0 becomes a’:
[0180] a’ = K[0]+K[1]+…+K[j]+K[a - 1]; (j ∈ [0, (a’ - 1)])
[0181] This can ensure that each sub - sound in each row of the matrix corresponds uniquely to its sound source position information.
[0182] For example, for LocationM2_0 (a = 4, a’ = 5), if the Child_Singing in it comes from two different positions, its position matrix is as follows:
[0183]
[0184] The corresponding CompositionM2_1 (a’ = 5) is:
[0185]
[0186] After the above - mentioned LocationM2_0 is completed, add "four columns" in front of it to form a new EchoM2_0:
[0187] (1) The first column: Add the sub - sound type Type[j] according to CompositionM2_1; (j ∈ [0, (a’ - 1)]);
[0188] (2) The second column: Add the recognition credibility Prob1[i] of the first type corresponding to this sub - sound according to CompositionM1_2; (i ∈ [0, (a’ - 1)]);
[0189] (3) The third column: Add the recognition credibility Prob2[j] of the second type corresponding to this sub - sound according to CompositionM2_1; (j ∈ [0, (a’ - 1)]);
[0190] (4) Fourth column: Add a flag bit isEcho[j] indicating whether the sub - sound is judged as an echo, with a default value of 0 (j ∈ [0, (a’ - 1)]);
[0191] Following these four columns is the position information column of the corresponding sub - sound of the original LocationM2_0. In this way, the matrices CompositionM1_2 and CompositionM2_1 are combined with LocationM2_0 and saved as EchoM2_0. The sub - sound type, type recognition credibility, and sound source position information jointly constitute a new information database EchoM2_0.
[0192] For example, after merging the above - mentioned LocationM2_0 (a’ = 5) and CompositionM2_1 (a’ = 5), the following is the EchoM2_0 matrix (a’ = 5):
[0193]
[0194] S4. Screen out the target sub - sounds that may be echoes based on the sound source position information of the sub - sounds.
[0195] Inspect each sub - sound in EchoM2_0 one by one, and compare the sound source position information therein with the distance and azimuth in the above - mentioned SpeakerLocation.
[0196] (1) If the proximity degree between the sub - sound Type[j] and the position of SpeakerLocation is less than the preset degree, it indicates that the sub - sound Type[j] is very likely to be the echo source from the audio output device, and isEcho[j] is assigned a value of 1.
[0197] (2) If the proximity degree between the sub - sound Type[j] and the position of SpeakerLocation is not less than the preset degree, it means that the sub - sound Type[j] is very likely not to be the echo source from the audio output device, but just a newly input sound source. For this newly input sound source, when processing echo cancellation, it can be directly input to the sound collector, and isEcho[j] remains the default value of 0 unchanged.
[0198] For example, after the isEcho[j] value of EchoM2_0 is updated, it is as follows:
[0199]
[0200] S5. Screen out the target sub - sound type based on the type recognition credibility of the sub - sound type.
[0201] Among the types corresponding to all sub - sounds where isEcho[j] == 1,
[0202] If Prob1[i] < B, isEcho[j] should be assigned 0;
[0203] If Prob2[j] < B, isEcho[j] should be assigned 0;
[0204] For all other sub - sounds where isEcho[j] == 1, isEcho[j] remains 1.
[0205] According to the above rules, if the threshold B = 10%, then for EchoM2_0 with different Drum position values:
[0206]
[0207] The latest value of isEcho[j] should be updated as follows:
[0208]
[0209] After the above steps, retain all sub - sounds in EchoM2_0 where isEcho[j] == 1, and delete all other items. The resulting EchoM2_0 is as follows:
[0210]
[0211] The refined latest EchoM2_0 is the smallest echo set that most likely contains all the true sub - sounds with echoes.
[0212] S6. Extract the acoustic features of the sub - sounds.
[0213] Based on the a' target sub - sound types included in EchoM2_0, perform targeted feature extraction on the sub - sounds in the first sound that belong to the target sub - sound types, and the sub - sounds in the target sub - sounds that belong to the target sub - sound types. The feature extraction algorithms include but are not limited to Mel - Frequency Cepstral Coefficients (MFCC), Linear Predictive Coefficients (LPC), Short - Time Fourier Transform (STFT), etc. These acoustic feature information form the echo feature matrices FeatureM1 and FeatureM2_0. In particular, in FeatureM1 and FeatureM2_0, when the row member serial number index is the same, the sub - sound types should be the same, so as to ensure that the rows of the two feature matrices correspond to the echo sources and the second sound of the same sub - sound type.
[0214] S7. Filter out the echo sub - sounds from the second sound according to the acoustic features.
[0215] 1. Feature calculation. Calculate the similarity between the sub - sound feature FeatureM1[i] in FeatureM1 and FeatureM2_0[i] in FeatureM2_0.
[0216] 2. The voice features match. When the similarity is not less than the first similarity threshold, determine the target sub-voice as the echo sub-voice.
[0217] 3. Echo sub-voice filtering. Use an adaptive filter to filter out the echo sub-voice from the second voice, which can be expressed as follows:
[0218] LeftSound2_0[j] = Sound2_0 - Sound2_0[j]
[0219] Where, after eliminating the sub-voice i in this round, the obtained LeftSound2_0[j] is the remaining voice of Sound2_0 after eliminating the sub-voice i.
[0220] 4. Fidelity check. Determine the reference sub-voice type whose type recognition credibility meets the second credibility condition among the sub-voice types included in the second voice; when the similarity between the voice features of the echo sub-voice and the voice features of the sub-voice belonging to the reference sub-voice type in the second voice is not less than the third similarity threshold, filter out the echo sub-voice from the second voice.
[0221] 5. Take FeatureM1[j + 1] and perform the above steps 1 to 3 in the next round.
[0222] 6. The voice features do not match. When the similarity in step 2 is less than the first similarity threshold, update the feature extraction algorithm matching the target sub-voice type, and re-extract the voice features of the target sub-voice and the voice features of the sub-voice belonging to the target sub-voice type in the first voice based on the updated feature extraction algorithm; when the similarity between the re-extracted voice features of the target sub-voice and the re-extracted voice features of the sub-voice is not less than the second similarity threshold, determine the target sub-voice as the echo sub-voice.
[0223] After performing the above steps for a' rounds, all echo sub-voices are cleared, and LeftSound2_0[a' - 1] is the finally remaining voice:
[0224]
[0225] Sound2_1 refers to the next round of new voice that has had all echoes cleared and is simply collected by the sound collector.
[0226] Some embodiments provide a display device and an echo cancellation method. During the output of a first sound by an audio output device, in response to an input interactive voice, a sound collector collects a second sound. The second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector. The first sound may cause echo interference to the second sound. By identifying the sub-sound types to which the sub-sounds included in the first sound and the second sound respectively belong, and the first sound and the second sound include the same reference sub-sound type, it is indicated that the reference sub-sound belonging to the reference sub-sound type in the second sound may be an echo. By screening the reference sub-sound according to the sound source position information of the reference sub-sound, the target sub-sound that is most likely to be an echo can be determined. Then, in the sub-sound type to which the target sub-sound belongs, the target sub-sound type whose type recognition credibility meets the conditions is screened to improve the accuracy of echo cancellation. Finally, the echo sub-sound is screened from the target sub-sounds by using the sound characteristics of the sub-sounds belonging to the target sub-sound type. This method of screening echo sub-sounds step by step is beneficial to accurately cancel the echo sub-sounds in the second sound, making the second sound as faithful as possible and improving the sound quality collected by the sound collection device.
[0227] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0228] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an echo cancellation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0229] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0230] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0231] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0232] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0233] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0234] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processing modules, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0235] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0236] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several variations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A display device, characterized in that: include: an audio output device configured to output sound of the display device; A sound collector configured to collect the input interactive voice and the sound output by the audio output device; A controller configured to execute instructions to enable the display device to: In the process of the audio output device outputting the first sound, in response to the input interactive voice, obtaining a second sound collected by the sound collector; the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector; Respectively identifying at least one sub-sound contained in the first sound and the second sound, obtaining a sub-sound type to which each sub-sound belongs, and obtaining sound source position information of the sub-sounds of the second sound; In a case where the first sound and the second sound include the same reference sub-sound type, for reference sub-sounds in the second sound that belong to the reference sub-sound type, the reference sub-sounds are screened according to sound source position information of the reference sub-sounds to obtain a target sub-sound; Based on the type recognition credibility of the sub-sound type to which the target sub-sound belongs, screening a target sub-sound type from the sub-sound types to which the target sub-sound belongs; The target sub-sound is screened according to the sound features of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the sound features of the sub-sounds belonging to the target sub-sound type in the first sound to obtain an echo sub-sound, and the echo sub-sound is filtered out from the second sound.
2. The display device according to claim 1, characterized in that The controller performs screening of the target sub-sound according to the sound features of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the sound features of the sub-sounds belonging to the target sub-sound type in the first sound to obtain the echo sub-sound, and is configured as follows: Acquire a similarity between a sound feature of a sub-sound in the target sub-sound that belongs to the target sub-sound type and a sound feature of a sub-sound in the first sound that belongs to the target sub-sound type; In a case where the similarity is not less than a first similarity threshold, the target sub-sound is determined to be an echo sub-sound.
3. The display device according to claim 2, characterized in that The controller is also configured to: When the similarity is less than a first similarity threshold, updating a feature extraction algorithm that matches the target sub-sound type, and re-extracting sound features of the target sub-sound and sound features of sub-sounds in the first sound that belong to the target sub-sound type based on the updated feature extraction algorithm; In a case where the similarity between the re-extracted sound feature of the target sub-sound and the re-extracted sound feature of the sub-sound is not less than a second similarity threshold, the target sub-sound is determined to be an echo sub-sound.
4. The display device according to claim 1, characterized in that The type recognition credibility includes a first type recognition credibility of the sub-sound type to which the target sub-sound belongs in the first sound, and a second type recognition credibility of the sub-sound type to which the target sub-sound belongs in the second sound; the controller performs type recognition credibility based on the sub-sound type to which the target sub-sound belongs, and filters the target sub-sound type in the sub-sound type to which the target sub-sound belongs, and is configured to: The sub-sound types to which the target sub-sound belongs are eliminated, the sub-sound types that satisfy at least one of the first type recognition credibility or the second type recognition credibility satisfying the first credibility condition, to obtain the target sub-sound type.
5. The display device according to claim 1, characterized in that The controller performs filtering of the echo sub-sound from the second sound, and is configured to: Determining, among the sub-sound types included in the second sound, a reference sub-sound type whose type recognition credibility satisfies a second credibility condition; When the similarity between the sound feature of the echo sub-sound and the sound feature of the sub-sound belonging to the reference sub-sound type in the second sound is not less than a third similarity threshold, the echo sub-sound is filtered out from the second sound.
6. The display device according to claim 1, characterized in that The controller is also configured to: When the first sound and the second sound do not include the same reference sub-sound type, the second sound is directly used as the sound input by the sound collector.
7. The display device according to claim 1, characterized in that The controller performs screening of the reference sub-sounds according to the sound source position information of the reference sub-sounds to obtain the target sub-sounds, and is configured to: Acquire location information of the audio output device; Determining a degree of position proximity based on the sound source position information and the position information of the audio output device; When the position proximity is less than a preset degree, the reference sub-sound is determined to be a target sub-sound.
8. An echo cancellation method, characterized in that: The method comprises: In the process of outputting the first sound by the audio output device, in response to the input interactive voice, obtaining a second sound collected by the sound collector; the second sound includes the interactive voice and the sound formed when the first sound is collected by the sound collector; Respectively identifying at least one sub-sound contained in the first sound and the second sound, obtaining a sub-sound type to which each sub-sound belongs, and obtaining sound source position information of the sub-sounds of the second sound; In a case where the first sound and the second sound include the same reference sub-sound type, for reference sub-sounds in the second sound that belong to the reference sub-sound type, the reference sub-sounds are screened according to sound source position information of the reference sub-sounds to obtain a target sub-sound; Based on the type recognition credibility of the sub-sound type to which the target sub-sound belongs, a target sub-sound type is screened from the sub-sound types to which the target sub-sound belongs; based on the sound features of the sub-sounds belonging to the target sub-sound type in the target sub-sound and the sound features of the sub-sounds belonging to the target sub-sound type in the first sound, the target sub-sound is screened to obtain an echo sub-sound, and the echo sub-sound is filtered out from the second sound.
9. The method according to claim 8, characterized in that The step of screening the target sub-sound according to the sound feature of the target sub-sound and the sound feature of the sub-sound of the first sound belonging to the target sub-sound type to obtain the echo sub-sound includes: Acquire a similarity between a sound feature of the target sub-sound and a sound feature of a sub-sound in the first sound that belongs to the target sub-sound type; In a case where the similarity is not less than a first similarity threshold, the target sub-sound is determined to be an echo sub-sound.
10. The method according to claim 9, characterized in that The method further comprises: When the similarity is less than a first similarity threshold, updating a feature extraction algorithm that matches the target sub-sound type, and re-extracting sound features of the target sub-sound and sound features of sub-sounds in the first sound that belong to the target sub-sound type based on the updated feature extraction algorithm; In a case where the similarity between the re-extracted sound feature of the target sub-sound and the re-extracted sound feature of the sub-sound is not less than a second similarity threshold, the target sub-sound is determined to be an echo sub-sound.