Speech processing method and apparatus, and storage medium and electronic device

By acquiring a reference signal for echo cancellation processing, the problem of poor echo cancellation effect when external devices play sound is solved, improving the device's performance and user experience.

WO2025223029A1PCT designated stage Publication Date: 2025-10-30SHENZHEN TCL DIGITAL TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079140
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-02-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

In scenarios where the local device plays sound through an external device, the echo cancellation effect is poor, affecting the normal use of the device and the user experience.

Method used

By acquiring the reference signal generated by the reference subject based on the transmission delay value, the microphone signal is subjected to echo cancellation processing. The reference subject is either the first device or the second device. The delay value is calculated or weighted averaged using the interactive delay value and the calibration delay value to achieve effective echo cancellation.

Benefits of technology

It improves echo cancellation, enhances device performance and user experience, and is particularly effective at accurately recognizing user voice commands when the TV plays sound through the speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025079140_30102025_PF_FP_ABST
    Figure CN2025079140_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech processing. Disclosed are a speech processing method and apparatus, and a storage medium and an electronic device. The speech processing method is applied to a first device, and comprises: sending a speech signal to a second device to perform sound playing; a microphone acquiring a microphone signal; acquiring a reference signal, wherein the reference signal is generated by a reference main body on the basis of a transmission delay value; and performing echo cancellation processing on the microphone signal on the basis of the reference signal. The present application improves an echo cancellation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Voice processing methods, devices, storage media and electronic devices

[0001] This application claims priority to Chinese Patent Application No. 202410518896.6, filed on April 26, 2024, entitled “Speech Processing Method, Apparatus, Storage Medium and Electronic Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of speech processing technology, specifically to a speech processing method, apparatus, storage medium, and electronic device. Background Technology

[0003] The scenario where the local device (first device) plays voice remotely through an external device (second device) is commonly used in various situations. For example, in order to pursue high-definition sound quality on their TVs, users often connect external speakers to their TVs to play the voice messages from the TVs.

[0004] In scenarios where voice is played remotely via an external device, the sound received by the microphone on the local device is often mixed with the sound played from the external device. For example, if a user controls the local device via voice while sound is playing from an external device, the sound received by the local device will be a mixture of the user's control voice and the sound played from the external device. In this case, echo cancellation is usually required on the local device. However, in scenarios where sound is played remotely via an external device, it is difficult to obtain an effective reference signal for echo cancellation on the local device, or the obtained reference signal may be ineffective for echo cancellation. Technical issues

[0005] Currently, in scenarios where the local device plays sound through an external device, there is a problem with poor echo cancellation, which affects the normal use of the device and the user experience. Technical solutions

[0006] This application provides a voice processing solution that can effectively improve echo cancellation in scenarios where the local device plays sound through an external device, thereby enhancing device performance and user experience.

[0007] The embodiments of this application provide the following technical solutions:

[0008] According to one embodiment of this application, a voice processing method is applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a far-end speaker, the method comprising: sending a voice signal to the second device so that the second device plays the voice signal through the far-end speaker; acquiring a microphone signal through the microphone, wherein the microphone signal is infused with a sound signal corresponding to the far-end speaker playing the voice signal; acquiring a reference signal generated by a reference subject based on a transmission delay value, wherein the reference subject is either the first device or the second device; and performing echo cancellation processing on the microphone signal based on the reference signal to obtain a processed signal with the sound signal canceled.

[0009] In some embodiments of this application, when the reference subject is the first device, the transmission delay value includes a first delay value, and the step of obtaining the reference signal in the reference subject includes: obtaining at least one of an interaction delay value and a calibration delay value, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the sound reception of the microphone; obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value; and obtaining the reference signal by locally caching the voice signal for the duration corresponding to the first delay value.

[0010] In some embodiments of this application, obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value includes one of the following methods: using one of the interaction delay value and the calibration delay value as the first delay value; or calculating the first delay value by performing a weighted average of the interaction delay value and the calibration delay value.

[0011] In some embodiments of this application, the step of calculating the first delay value by weighted averaging the interaction delay value and the calibration delay value includes: obtaining device information of the first device and the second device; obtaining voice playback environment information of the first device and the second device; determining corresponding weighting coefficients based on the device information and the voice playback environment information; and calculating the first delay value by weighted averaging the interaction delay value and the calibration delay value based on the weighting coefficients.

[0012] In some embodiments of this application, the first device and the second device are connected via a multimedia interface cable, and the interaction delay value is obtained by: receiving delay information returned by the second device via the multimedia interface cable, the delay information being the delay information of the second device playing the voice signal transmitted by the first device; and calculating the interaction delay value based on the delay information.

[0013] In some embodiments of this application, the first device includes a near-end speaker, and the calibration delay value is obtained by: acquiring the near-end reception time and the far-end reception time, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and obtaining the calibration delay value based on the time difference between the near-end reception time and the far-end reception time.

[0014] In some embodiments of this application, when the reference subject is the second device, the transmission delay value includes a second delay value, and the step of obtaining the reference signal in the reference subject includes: receiving the reference signal transmitted by the second device, wherein the reference signal is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to the second delay value, the second delay value is obtained by the second device based on the audio transmission delay value, and the audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

[0015] In some embodiments of this application, the second delay value is obtained in one of the following ways: using the audio transmission delay value as the second delay value; obtaining a preset adjustment coefficient corresponding to the first device, and adjusting the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

[0016] In some embodiments of this application, the first device is a television and the second device is a speaker.

[0017] According to one embodiment of this application, a voice processing apparatus is applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a far-end speaker, the voice processing apparatus comprising: a transmitting module for transmitting a voice signal to the second device, so that the second device plays the voice signal through the far-end speaker; an acquiring module for acquiring a microphone signal through the microphone, wherein the microphone signal is infused with a corresponding sound signal from the far-end speaker playing the voice signal; a reference module for acquiring a reference signal from a reference body, wherein the reference signal is generated by the reference body based on a transmission delay value, the reference body being either the first device or the second device; and an echo cancellation module for performing echo cancellation processing on the microphone signal based on the reference signal to obtain a processed signal with the sound signal canceled.

[0018] In some embodiments of this application, when the reference subject is the first device, the transmission delay value includes a first delay value. The reference module is configured to: obtain at least one of an interaction delay value and a calibration delay value, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the microphone's sound reception; obtain the first delay value based on at least one of the interaction delay value and the calibration delay value; and obtain the reference signal by locally caching the voice signal for the duration corresponding to the first delay value.

[0019] In some embodiments of this application, the reference module is configured to perform one of the following: taking one of the interaction delay value and the calibration delay value as the first delay value; and performing a weighted average calculation of the interaction delay value and the calibration delay value to obtain the first delay value.

[0020] In some embodiments of this application, the reference module is configured to: acquire device information of the first device and the second device; acquire voice playback environment information of the first device and the second device; determine corresponding weighting coefficients based on the device information and the voice playback environment information; and calculate the first delay value by performing a weighted average calculation on the interaction delay value and the calibration delay value based on the weighting coefficients.

[0021] In some embodiments of this application, the first device and the second device are connected via a multimedia interface cable. The first delay determination module is configured to: receive delay information returned by the second device via the multimedia interface cable, wherein the delay information is the delay information of the audio signal transmitted by the first device being played in the second device; and calculate the interaction delay value based on the delay information.

[0022] In some embodiments of this application, the first device includes a near-end speaker and a first delay determination module, configured to: acquire near-end reception time and far-end reception time, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and obtain the calibration delay value based on the time difference between the near-end reception time and the far-end reception time.

[0023] In some embodiments of this application, when the reference subject is the second device, the transmission delay value includes a second delay value. The reference module is configured to: receive the reference signal transmitted by the second device, wherein the reference signal is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to the second delay value, wherein the second delay value is obtained by the second device based on the audio transmission delay value, and wherein the audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

[0024] In some embodiments of this application, the second delay value is obtained in one of the following ways: using the audio transmission delay value as the second delay value; obtaining a preset adjustment coefficient corresponding to the first device, and adjusting the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

[0025] In some embodiments of this application, the first device is a television and the second device is a speaker.

[0026] According to another embodiment of this application, a storage medium stores a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the methods described in the embodiments of this application.

[0027] According to another embodiment of this application, an electronic device may include: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the methods described in the embodiments of this application.

[0028] According to another embodiment of this application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of this application. Beneficial effects

[0029] In this embodiment, the voice processing method is applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a remote speaker, the voice processing method including: sending a voice signal to the second device so that the second device plays the voice signal through the remote speaker; acquiring a microphone signal through the microphone, wherein the microphone signal is combined with a sound signal corresponding to the remote speaker playing the voice signal; acquiring a reference signal transmitted by a reference subject according to a transmission delay value, wherein the reference subject is the first device or the second device; and performing echo cancellation processing on the microphone signal according to the reference signal to obtain a processed signal with the sound signal canceled.

[0030] In this way, when the first device plays voice remotely through the second device, the first device can obtain a reference signal transmitted by a reference subject based on the transmission delay value. This reference subject can be either the second device or the first device. Consequently, the first device can obtain an effective reference signal for echo cancellation, resulting in a processed signal that eliminates the sound played by the remote speaker in the second device. In scenarios where the first device plays sound through an external device, the echo cancellation effect can be effectively improved, enhancing the device's performance and user experience. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 shows a flowchart of a speech processing method according to an embodiment of this application.

[0033] Figure 2 shows a framework diagram of a speech processing system according to an embodiment of this application.

[0034] Figure 3 shows a framework diagram of a speech processing system according to an embodiment of this application.

[0035] Figure 4 shows a block diagram of a speech processing apparatus according to an embodiment of the present application.

[0036] Figure 5 shows a block diagram of an electronic device according to an embodiment of this application.

[0037] Implementation methods of this application

[0038] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments provided below are some embodiments for implementing the present disclosure, and not all embodiments for implementing the present disclosure. Unless otherwise specified, the technical solutions described in the embodiments of the present disclosure can be implemented in any combination.

[0039] It should be noted that, in the embodiments of this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the apparatus, such as portions of circuitry, processors, programs, or software, etc.) in the method or apparatus that includes that element.

[0040] For example, the speech processing method provided in this disclosure includes a series of steps, but the speech processing method provided in this disclosure is not limited to the steps described. Similarly, the speech processing apparatus provided in this disclosure includes a series of units, but the apparatus provided in this disclosure is not limited to the units explicitly described, but may also include units that need to be set up for obtaining relevant information or processing based on information.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.

[0042] Figure 1 schematically illustrates a flowchart of a voice processing method according to an embodiment of this application. The entity executing the voice processing method can be any first device with processing capabilities. The first device includes a microphone, and the first device is electrically connected to a second device. The second device includes a far-end speaker. Examples of the first and second devices include televisions, speakers, computers, mobile phones, smartwatches, and home appliances.

[0043] As shown in Figure 1, the first device can execute the voice processing method shown in Figure 1, which may include steps S110 to S140.

[0044] Step S110: Send the voice signal to the second device so that the second device can play the voice signal through a remote speaker; Step S120: Obtain a microphone signal through a microphone, wherein the microphone signal is combined with the sound signal corresponding to the remote speaker playing the voice signal; Step S130: Obtain a reference signal transmitted by a reference subject according to a transmission delay value, wherein the reference subject is the first device or the second device; Step S140: Perform echo cancellation processing on the microphone signal according to the reference signal to obtain a processed signal with the sound signal canceled.

[0045] The first device can send voice signals to the second device, which then transmits the voice signals to its far-end speaker for sound playback. The far-end speaker will then play the sound signal. The near-end speaker in the first device does not produce sound (e.g., the Amp (amplifier) ​​in the first device can be turned off, preventing the voice signal from being transmitted to the near-end speaker).

[0046] When the remote speaker of the second device plays a sound signal, the microphone in the first device can record sound (i.e., pick up sound), so that the first device can obtain the microphone signal recorded by the microphone. The microphone signal will be combined with the sound played by the remote speaker and other sounds (such as user control voice).

[0047] The first device can obtain a reference signal generated by the reference subject based on the transmission delay value. The reference subject is either the second device or the first device. The microphone signal received in the first device and the signal from which the sound played in the second device originates (this signal can be used to generate the reference signal) usually have a delay. The reference subject (i.e., the first device or the second device) generates the reference signal based on the transmission delay value using the signal from which the sound played in the second device originates. The first device can then effectively perform echo cancellation on the microphone signal based on this reference signal.

[0048] In one approach, the reference subject is a second device, and the signal from which the sound played in the second device originates is "the signal transmitted by the second device to the remote microphone for playback (this signal is the signal obtained by the second device after processing the aforementioned voice signal transmitted by the first device according to the playback settings in the second device)". Specifically, the reference signal can be a copy of "the signal transmitted by the second device to the remote microphone for playback" in the second device and cached for the duration corresponding to the second delay value, and then used as the reference signal.

[0049] In another approach, the reference subject is the first device, and the signal from which the sound played in the second device originates is the "voice signal sent from the first device to the second device." Specifically, this reference signal can be the duration corresponding to a first delay value that the first device buffers in its cache. Specifically, the first device can process the "voice signal sent from the first device to the second device" using a hardware codec (e.g., analog-to-digital conversion) before buffering.

[0050] Furthermore, by performing echo cancellation processing on the microphone signal based on the reference signal in the first device, an echo-cancelling signal can be obtained that effectively eliminates the sound played from the far-end speaker. For example, in some scenarios, the echo-cancelling signal can be effectively used for speech recognition, improving the accuracy of speech recognition when the first device plays sound through an external second device.

[0051] In this way, based on steps S110 to S140, in the scenario where the first device plays voice remotely through the second device, the first device can obtain the reference signal transmitted by the reference subject according to the transmission delay value. The reference subject is either the second device or the first device. Then, the first device can obtain an effective reference signal to perform echo cancellation and obtain a processed signal that eliminates the sound played by the remote speaker in the second device. In the scenario where the local device plays sound through an external device, the echo cancellation effect can be effectively improved, thereby improving the device's performance and user experience.

[0052] In one specific embodiment of this application, the first device is a television, and the second device is a speaker. When a user uses the speaker to play sound from the television, if the user controls the television via voice, after the user outputs "control voice", the microphone signal received by the microphone in the television will at least combine the sound played by the speaker and the control voice. If the echo cannot be effectively eliminated, the television cannot effectively recognize the user control instruction in the microphone signal. However, based on the embodiments of this application, the television can effectively eliminate the processed signal of the sound played by the speaker in the microphone signal, and accurately perform voice recognition based on the processed signal to obtain the user control instruction and execute the response operation.

[0053] The following describes further optional embodiments of the steps performed during speech processing in the embodiment shown in Figure 1.

[0054] In one embodiment, when the reference subject is the first device, the reference signal is obtained by the first device based on a signal transmitted to the second device, the transmission delay value includes a first delay value, and obtaining the reference signal in the reference subject may include:

[0055] At least one of an interaction delay value and a calibration delay value is obtained, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the sound reception of the microphone; and, the first delay value is obtained based on at least one of the interaction delay value and the calibration delay value; the reference signal is obtained by caching the voice signal locally for the duration corresponding to the first delay value.

[0056] The interaction delay value is calculated based on the delay information returned by the second device. In other words, the interaction delay value can reflect the signal delay from the perspective of the actual delay in the second device.

[0057] The calibration delay value is obtained by calibrating based on the microphone's sound reception, meaning that the calibration delay value can reflect the signal delay from the perspective of the sound reception in the first device.

[0058] A first delay value is obtained based on at least one of the interaction delay value and the calibration delay value. When the reference signal is obtained by the first device based on the signal transmitted to the second device, the first device, as the reference subject, can use the voice signal transmitted to the second device to buffer the duration corresponding to the first delay value locally as the obtained reference signal. This reference signal is used by the first device to perform echo cancellation, effectively ensuring the echo cancellation effect.

[0059] Specifically, the reference subject is the first device, and the signal from which the sound played in the second device originates is the "voice signal sent from the first device to the second device." This reference signal can be obtained by the first device caching the "voice signal sent from the first device to the second device" for a duration corresponding to a first delay value within the first device. Specifically, the first device can process the "voice signal sent from the first device to the second device" using a hardware codec (e.g., analog-to-digital conversion) before caching.

[0060] For example, Figure 2 shows a framework diagram of a voice processing system according to an embodiment of this application. The voice processing system includes a first device and a second device. In an optional scenario, the first device is a television 210, and the second device is a speaker 220.

[0061] When the TV 210 is performing "audio stream playback": In the first mode, the TV 210 can send the voice signal to the speaker 220 through the passthrough path 230. In this mode, the TV 210 and the speaker 220 are in a passthrough scenario (i.e., the TV 210 sends the undecoded source data (voice signal) to the speaker 220, and the speaker 220 either "decodes" itself or "does not support decoding and is silent"); In the second mode, the TV 210 can send the voice signal to the speaker 220 through the decoding path 240. In this mode, the TV 210 sends the "voice signal" decoded by the decoder 2101 to the speaker 220.

[0062] Furthermore, as shown in Figure 2, the power amplifier (Amp) 2103 in the television 210 is turned off, preventing the voice signal from being transmitted to the near-end speaker (Spk) 2104 in the first device, thus preventing the near-end speaker (Spk) 2104 from producing sound. However, the television 210 locally buffers the voice signal sent to the speaker 220 for the duration corresponding to the first delay value as a reference signal. Then, the reference signal (Ref) is transmitted to the echo cancellation module (AEC) 2106 in the first device (i.e., the television) through the reference path 2105. The echo cancellation module (AEC) 2106 performs echo cancellation on the microphone signal received by the microphone 2107. Specifically, the reference signal (Ref) can be the voice signal sent to the speaker 220, processed by the hardware codec (including ADC / DAC functions) 2102 in the television 210 (e.g., analog-to-digital conversion), and buffered for the duration corresponding to the first delay value as a reference signal.

[0063] Specifically, by applying a delay of a "first delay value" to the reference path 2105, the voice signal input to the speaker 220 is locally buffered for the duration corresponding to the first delay value and used as the obtained reference signal. Then, the reference signal (Ref) is transmitted to the echo cancellation module (AEC) 2106 in the first device through the reference path 2105.

[0064] In a further embodiment, the first device and the second device are connected via a multimedia interface cable, and the interaction delay value can be obtained in the following manner: the first device receives delay information returned by the second device through the multimedia interface cable, the delay information being the delay information of the audio signal transmitted by the first device being played in the second device; and the interaction delay value is obtained by calculation based on the delay information.

[0065] When the first device and the second device are connected via a multimedia interface cable (HDMI cable), some private commands can be customized to allow the second device to transmit corresponding latency information to the first device. The first device can then use this latency information to calculate the interaction latency value.

[0066] The voice signal sent from the first device to the second device can be marked with a corresponding transmission time. The second device can calculate the time difference between the "time of receiving the voice signal or the time of transmitting the voice signal to the remote speaker" and the "transmission time" as latency information. The first device can use the latency information (time difference) as the interaction latency value, or the first device can increase the latency information (time difference) by a predetermined first increment value and use it as the interaction latency value.

[0067] Furthermore, in one embodiment, the first device includes a near-end speaker, and the calibration delay value can be obtained as follows: the first device acquires a near-end reception time and a far-end reception time, the near-end reception time being the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time being the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and the calibration delay value is obtained based on the time difference between the near-end reception time and the far-end reception time.

[0068] The first and second devices can also be electrically connected via other methods such as BT and SPDIF, and the calibration delay value can be obtained through active calibration. Referring to Figure 3, the near-end speaker 3101 in the first device 310 and the far-end speaker 3201 in the second device 320 play the first sound and the second sound respectively. The first sound and the second sound are played according to the signals simultaneously transmitted in the first device. By having the microphone 3102 in the first device 310 pick up the sound, the near-end sound pickup time and the far-end sound pickup time can be obtained. Furthermore, the time difference between the near-end sound pickup time and the far-end sound pickup time can be calculated.

[0069] The calibration delay value can be obtained based on the time difference between the near-end reception time and the far-end reception time. Alternatively, the calibration delay value can be obtained by increasing the time difference between the near-end reception time and the far-end reception time by a predetermined second increment.

[0070] Furthermore, in one embodiment, obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value includes one of the following methods:

[0071] The first method involves using either the interaction delay value or the calibration delay value as the first delay value.

[0072] The second method involves calculating the first delay value by performing a weighted average of the interaction delay value and the calibration delay value.

[0073] In the first approach, one of the interaction delay value and the calibration delay value can be selected as the first delay value. Transmitting a reference signal based on this first delay value can ensure the echo cancellation effect to a certain extent.

[0074] In the second approach, the interaction delay value and the calibration delay value are further weighted and averaged. The calculated weighted average value is used as the first delay value. The echo cancellation effect can be further improved by transmitting the reference signal based on the first delay value.

[0075] In one embodiment, the step of calculating the first delay value by weighting the interaction delay value and the calibration delay value may include: calculating the first delay value by weighting the interaction delay value and the calibration delay value according to a predetermined weighting coefficient.

[0076] In another embodiment, the step of calculating the first delay value by weighting the interaction delay value and the calibration delay value may include:

[0077] Obtain device information for the first device and the second device; obtain voice playback environment information for the first device and the second device; determine corresponding weighting coefficients based on the device information and the voice playback environment information; calculate the first delay value by weighting the interaction delay value and the calibration delay value based on the weighting coefficients.

[0078] In this embodiment, the device information of the first device and the second device, as well as the voice playback environment information of the first device and the second device, are dynamically acquired. Based on the device information and the voice playback environment information, the corresponding weighting coefficients are dynamically determined to perform a weighted average calculation on the interaction delay value and the calibration delay value to obtain a first delay value. Based on the first delay value, the transmission of the reference signal can further improve the echo cancellation effect.

[0079] The device information for the first and second devices can include device models, while the audio playback environment information for the first and second devices can include network strength, device distance, location information, etc. The corresponding weighting coefficients can be determined by querying the preset coefficient table to obtain the device information and audio playback environment information.

[0080] In one embodiment, when the reference subject is the second device, the reference signal is obtained by the second device based on a signal transmitted to the remote microphone for playback, the transmission delay value includes a second delay value, and obtaining the reference signal in the reference subject includes:

[0081] The reference signal transmitted by the second device is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to a second delay value. The second delay value is obtained by the second device based on the audio transmission delay value. The audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

[0082] In this embodiment, the reference subject is the second device, and the signal from which the sound played in the second device originates is "the signal sent by the second device to the remote microphone for playback (this signal is the signal obtained by the second device after processing the aforementioned voice signal transmitted by the first device according to the playback settings in the second device. For example, after the second device receives the voice signal transmitted by the first device, it trims the voice signal to obtain a trimmed signal, and then the trimmed signal is transmitted to the remote microphone for playback. This trimmed signal is the signal obtained after processing according to the playback settings in the second device)". Specifically, the reference signal can be a copy of "the signal sent by the second device to the remote microphone for playback" in the second device and cached for the duration corresponding to the second delay value, and then used as the reference signal.

[0083] In the second device, the difference between the time it takes for the microphone of the first device to receive sound from the remote speaker and the time it takes for the first device to receive the signal transmitted by the second device is used as the audio transmission delay value. Based on this audio transmission delay value, a second delay value is obtained. When the reference subject is the second device, the second device obtains a reference signal based on the second delay value and sends it to the first device. The first device can then transmit the reference signal to the echo cancellation module for effective echo cancellation.

[0084] In one embodiment, the second delay value may be obtained in one of the following ways:

[0085] The first method is to use the radio transmission delay value as the second delay value;

[0086] The second method involves obtaining a preset adjustment coefficient corresponding to the first device and adjusting the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

[0087] In the first method, the audio transmission delay value is used as the second delay value. The second device sends the reference signal to the first device with a delay based on the second delay value. The first device can then send the reference signal to the echo cancellation module for effective echo cancellation.

[0088] In the second method, a preset adjustment coefficient corresponding to the first device is obtained, and the audio transmission delay value is adjusted according to the preset adjustment coefficient to obtain a second delay value. The second device sends the reference signal to the first device with a delay based on the second delay value. The first device can then send the reference signal to the echo cancellation module for further effective echo cancellation, thereby further improving the echo cancellation effect.

[0089] To facilitate better implementation of the speech processing method provided in the embodiments of this application, the embodiments of this application also provide a speech processing apparatus based on the above-described speech processing method. The meanings of the terms used are the same as in the speech processing method described above, and specific implementation details can be found in the descriptions in the method embodiments. Figure 4 shows a block diagram of a speech processing apparatus according to an embodiment of this application.

[0090] As shown in Figure 4, the voice processing device 400 can be applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a remote speaker, and the device 400 including: a transmitting module 410 for transmitting a voice signal to the second device so that the second device plays the voice signal through the remote speaker; an acquiring module 420 for acquiring a microphone signal through the microphone, wherein the microphone signal is fused with the corresponding sound signal from the remote speaker playing the voice signal; a reference module 430 for acquiring a reference signal in a reference body, wherein the reference signal is generated by the reference body based on a transmission delay value, and the reference body is either the first device or the second device; and an echo cancellation module 440 for performing echo cancellation processing on the microphone signal based on the reference signal to obtain a processed signal with the sound signal canceled.

[0091] In some embodiments of this application, when the reference subject is the first device, the transmission delay value includes a first delay value. The reference module is configured to: obtain at least one of an interaction delay value and a calibration delay value, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the microphone's sound reception; obtain the first delay value based on at least one of the interaction delay value and the calibration delay value; and obtain the reference signal by locally caching the voice signal for the duration corresponding to the first delay value.

[0092] In some embodiments of this application, the reference module is configured to perform one of the following: taking one of the interaction delay value and the calibration delay value as the first delay value; and performing a weighted average calculation of the interaction delay value and the calibration delay value to obtain the first delay value.

[0093] In some embodiments of this application, the reference module is configured to: acquire device information of the first device and the second device; acquire voice playback environment information of the first device and the second device; determine corresponding weighting coefficients based on the device information and the voice playback environment information; and calculate the first delay value by performing a weighted average calculation on the interaction delay value and the calibration delay value based on the weighting coefficients.

[0094] In some embodiments of this application, the first device and the second device are connected via a multimedia interface cable. The first delay determination module is configured to: receive delay information returned by the second device via the multimedia interface cable, wherein the delay information is the delay information of the audio signal transmitted by the first device being played in the second device; and calculate the interaction delay value based on the delay information.

[0095] In some embodiments of this application, the first device includes a near-end speaker and a first delay determination module, configured to: acquire near-end reception time and far-end reception time, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and obtain the calibration delay value based on the time difference between the near-end reception time and the far-end reception time.

[0096] In some embodiments of this application, when the reference subject is the second device, the transmission delay value includes a second delay value. The reference module is configured to: receive the reference signal transmitted by the second device, wherein the reference signal is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to the second delay value, wherein the second delay value is obtained by the second device based on the audio transmission delay value, and wherein the audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

[0097] In some embodiments of this application, the second delay value is obtained in one of the following ways: using the audio transmission delay value as the second delay value; obtaining a preset adjustment coefficient corresponding to the first device, and adjusting the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

[0098] In some embodiments of this application, the first device is a television and the second device is a speaker.

[0099] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0100] Furthermore, this application also provides an electronic device, as shown in FIG5. FIG5 shows a block diagram of an electronic device according to an embodiment of this application, specifically:

[0101] The electronic device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that the electronic device structure shown in FIG. 5 does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0102] The processor 501 is the control center of the electronic device. It connects to various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and application programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.

[0103] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0104] The electronic device also includes a power supply 503 that supplies power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0105] The electronic device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0106] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 502 according to the following instructions, and the processor 501 runs the computer programs stored in the memory 502, thereby realizing the various functions in the foregoing embodiments of this application.

[0107] If an electronic device is used as the first device, the processor 501 may perform the following steps:

[0108] The voice signal is sent to the second device so that the second device can play the voice signal through a remote speaker; a microphone signal is acquired through a microphone, wherein the microphone signal is combined with the sound signal corresponding to the sound signal played by the remote speaker; a reference signal is acquired based on the transmission delay value by a reference subject, wherein the reference subject is the first device or the second device; echo cancellation processing is performed on the microphone signal based on the reference signal to obtain a processed signal with the sound signal canceled.

[0109] In some embodiments of this application, when the reference subject is the first device, the transmission delay value includes a first delay value, and the step of obtaining the reference signal in the reference subject includes: obtaining at least one of an interaction delay value and a calibration delay value, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the sound reception of the microphone; obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value; and obtaining the reference signal by locally caching the voice signal for the duration corresponding to the first delay value.

[0110] In some embodiments of this application, obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value includes one of the following methods: using one of the interaction delay value and the calibration delay value as the first delay value; or calculating the first delay value by performing a weighted average of the interaction delay value and the calibration delay value.

[0111] In some embodiments of this application, the step of calculating the first delay value by weighted averaging the interaction delay value and the calibration delay value includes: obtaining device information of the first device and the second device; obtaining voice playback environment information of the first device and the second device; determining corresponding weighting coefficients based on the device information and the voice playback environment information; and calculating the first delay value by weighted averaging the interaction delay value and the calibration delay value based on the weighting coefficients.

[0112] In some embodiments of this application, the first device and the second device are connected via a multimedia interface cable, and the interaction delay value is obtained by: receiving delay information returned by the second device via the multimedia interface cable, the delay information being the delay information of the second device playing the voice signal transmitted by the first device; and calculating the interaction delay value based on the delay information.

[0113] In some embodiments of this application, the first device includes a near-end speaker, and the calibration delay value is obtained by: acquiring the near-end reception time and the far-end reception time, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and obtaining the calibration delay value based on the time difference between the near-end reception time and the far-end reception time.

[0114] In some embodiments of this application, when the reference subject is the second device, the transmission delay value includes a second delay value, and the step of obtaining the reference signal in the reference subject includes: receiving the reference signal transmitted by the second device, wherein the reference signal is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to the second delay value, the second delay value is obtained by the second device based on the audio transmission delay value, and the audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

[0115] In some embodiments of this application, the second delay value is obtained in one of the following ways: using the audio transmission delay value as the second delay value; obtaining a preset adjustment coefficient corresponding to the first device, and adjusting the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

[0116] In some embodiments of this application, the first device is a television and the second device is a speaker.

[0117] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0118] Therefore, embodiments of this application also provide a storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the methods provided in embodiments of this application.

[0119] The storage medium can be a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0120] Since the computer program stored in the storage medium can execute the steps of any of the methods provided in the embodiments of this application, the beneficial effects that the methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0121] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0122] It should be understood that this application is not limited to the embodiments described above and shown in the accompanying drawings, but various modifications and changes can be made without departing from its scope.

Claims

1. A speech processing method, wherein, Applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a far-end speaker, the method includes: The voice signal is sent to the second device so that the second device can play the voice signal through a remote speaker; The microphone signal is acquired through the microphone, and the microphone signal is combined with the sound signal corresponding to the sound signal played by the far-end speaker to the voice signal. Acquire a reference signal from a reference body, wherein the reference signal is generated by the reference body based on a transmission delay value, and the reference body is either the first device or the second device; The microphone signal is subjected to echo cancellation processing based on the reference signal to obtain a processed signal in which the sound signal has been eliminated.

2. The method according to claim 1, wherein, When the reference subject is the first device, the transmission delay value includes a first delay value, and the acquisition of the reference signal in the reference subject includes: At least one of an interaction delay value and a calibration delay value is obtained, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the sound reception of the microphone; The first delay value is obtained based on at least one of the interaction delay value and the calibration delay value; The reference signal is obtained by caching the voice signal locally for the duration corresponding to the first delay value.

3. The method according to claim 2, wherein, The step of obtaining the first delay value based on at least one of the interaction delay value and the calibration delay value includes one of the following methods: One of the interaction delay value and the calibration delay value is used as the first delay value; The first delay value is obtained by weighting the interaction delay value and the calibration delay value.

4. The method according to claim 3, wherein, The step of calculating the first delay value by weighting the interaction delay value and the calibration delay value includes: Obtain device information for the first device and the second device; Obtain the voice playback environment information of the first device and the second device; The corresponding weighting coefficients are determined based on the device information and the voice playback environment information; The first delay value is obtained by performing a weighted average of the interaction delay value and the calibration delay value based on the weighting coefficient.

5. The method according to claim 3, wherein, The first device and the second device are connected via a multimedia interface cable, and the interaction delay value is obtained in the following manner: Receive delay information returned by the second device via the multimedia interface line, the delay information being the delay information for playing the voice signal transmitted by the first device in the second device; and, The interaction delay value is obtained by calculating based on the delay information.

6. The method according to claim 3, wherein, The first device includes a near-end speaker, and the calibration delay value is obtained in the following manner: Acquire near-end and far-end reception times, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first and second sounds are played based on signals simultaneously transmitted in the first device; and, The calibration delay value is obtained based on the time difference between the near-end reception time and the far-end reception time.

7. The method according to claim 1, wherein, When the reference subject is the second device, the transmission delay value includes a second delay value, and the acquisition of the reference signal in the reference subject includes: The reference signal transmitted by the second device is obtained by the second device after buffering the signal transmitted to the remote microphone for playback for a duration corresponding to a second delay value. The second delay value is obtained by the second device based on the audio transmission delay value. The audio transmission delay value is the difference between the time when the microphone of the first device receives the sound played by the remote speaker and the time when the first device receives the signal transmitted by the second device.

8. The method according to claim 5, wherein, The step of calculating the interaction delay value based on the delay information includes: The delay information can be used as the interaction delay value, or the delay information can be increased by a predetermined first increment and then used as the interaction delay value.

9. The method according to claim 6, wherein, The step of obtaining the calibration delay value based on the time difference between the near-end reception time and the far-end reception time includes: The time difference between the near-end reception time and the far-end reception time is used as the calibration delay value, or the time difference between the near-end reception time and the far-end reception time is increased by a predetermined second increment and then used as the calibration delay value.

10. The method according to claim 3, wherein, The step of calculating the first delay value by weighting the interaction delay value and the calibration delay value includes: The interaction delay value and the calibration delay value are weighted and averaged according to a predetermined weighting coefficient to obtain the first delay value.

11. The method according to claim 7, wherein, The second delay value is obtained in one of the following ways: The radio transmission delay value is used as the second delay value; Obtain the preset adjustment coefficient corresponding to the first device, and adjust the audio transmission delay value according to the preset adjustment coefficient to obtain the second delay value.

12. The method according to claim 1, wherein, The first device is a television, and the second device is a speaker.

13. A voice processing device, wherein, Applied to a first device, the first device including a microphone, the first device being electrically connected to a second device, the second device including a far-end speaker, the voice processing device comprising: The transmitting module is used to transmit the voice signal to the second device so that the second device can play the voice signal through a remote speaker; The acquisition module is used to acquire a microphone signal through the microphone, wherein the microphone signal is fused with the corresponding sound signal of the far-end speaker playing the voice signal; A reference module is used to acquire a reference signal in a reference body, wherein the reference signal is generated by the reference body based on a transmission delay value, and the reference body is either the first device or the second device; The echo cancellation module is used to perform echo cancellation processing on the microphone signal according to the reference signal to obtain a processed signal with the sound signal canceled.

14. The apparatus according to claim 13, wherein, When the reference subject is the first device, the transmission delay value includes a first delay value. The reference module is configured to: obtain at least one of an interaction delay value and a calibration delay value, wherein the interaction delay value is calculated based on the delay information returned by the second device, and the calibration delay value is obtained by calibration based on the microphone's sound reception; obtain the first delay value based on at least one of the interaction delay value and the calibration delay value; and obtain the reference signal by locally caching the voice signal for the duration corresponding to the first delay value.

15. The apparatus according to claim 14, wherein, The reference module is configured to implement one of the following methods: taking one of the interaction delay value and the calibration delay value as the first delay value; and performing a weighted average calculation on the interaction delay value and the calibration delay value to obtain the first delay value.

16. The apparatus according to claim 15, wherein, The reference module is configured to: acquire device information of the first device and the second device; acquire voice playback environment information of the first device and the second device; determine corresponding weighting coefficients based on the device information and the voice playback environment information; and calculate the first delay value by performing a weighted average calculation on the interaction delay value and the calibration delay value based on the weighting coefficients.

17. The apparatus according to claim 15, wherein, The first device and the second device are connected via a multimedia interface cable. The first delay determination module is configured to: receive delay information returned by the second device via the multimedia interface cable, wherein the delay information is the delay information of the second device playing the voice signal transmitted by the first device; and calculate the interaction delay value based on the delay information.

18. The apparatus according to claim 15, wherein, The first device includes a near-end speaker and a first delay determination module, configured to: acquire near-end reception time and far-end reception time, wherein the near-end reception time is the time when the microphone receives a first sound played by the near-end speaker, and the far-end reception time is the time when the microphone receives a second sound played by the far-end speaker, wherein the first sound and the second sound are played based on signals simultaneously transmitted in the first device; and obtain the calibration delay value based on the time difference between the near-end reception time and the far-end reception time.

19. A storage medium, wherein, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the method described in any one of claims 1 to 12.

20. An electronic device, wherein, include: Memory, which stores computer programs; A processor reads a computer program stored in memory to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Voice processing method and device, storage medium and electronic equipment

    CN118197339A

  • Signal processing method and device and computer storage medium

    CN110769352A

  • Method for improving far-field voice activation rate of television, computer equipment and storage medium

    CN113014978A

  • Echo cancellation method and device

    CN113689871A

  • Deterministic characterization and reduction of acoustic echo

    US20090247239A1