Acoustic echo cancellation method and bluetooth intercom system
By caching voice identifiers locally within the Bluetooth-connected device and merging voice data, the problem of Bluetooth communication latency jitter affecting acoustic echo cancellation is solved, improving call quality and avoiding howling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING BIG FISH SEMICON CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-06-23
AI Technical Summary
In Bluetooth communication devices, latency jitter caused by changes in Bluetooth communication distance and wireless interference affects the acoustic echo cancellation effect, leading to decreased call quality and even howling.
The device sends voice data containing voice identifiers and device identifiers to the master device from the current voice frame it collects, and caches it locally. The master device generates and sends merged voice data associated with the device identifier, and performs acoustic echo cancellation processing on the local data from the device.
It improves the time delay stability between the reference signal and the signal to be processed, eliminates the time delay jitter of the Bluetooth transmission link, improves the acoustic echo cancellation effect, and avoids howling caused by sound backhaul.
Smart Images

Figure CN122266385A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and more specifically, to an acoustic echo cancellation method and a Bluetooth communication system. Background Technology
[0002] Bluetooth intercom devices feature full-duplex, low-latency, and multi-person real-time calling capabilities, making them widely applicable in real-time voice interaction scenarios involving multiple roles, such as fleet cycling, film and television production scheduling, and industrial remote collaboration. A typical Bluetooth intercom device set consists of one master device and several slave devices. The master device receives voice data from each slave device, decodes and mixes it, and plays it locally. Simultaneously, it merges the merged voice data with the data captured by its own microphone, encodes it, and sends it back to all slave devices. The slave devices receive the merged voice data from the master device, decode it, and play it back, thus achieving full-duplex intercom communication.
[0003] The audio decoded and played by the device contains its own collected sound. This collected sound is then picked up again by the device's microphone after being played through the speaker, creating an acoustic echo. In severe cases, this can lead to a decrease in call quality or even howling. The device needs to use Acoustic Echo Cancellation (AEC) to filter out the sound collected by the device itself before playing it through the speaker.
[0004] Currently, slave devices typically use the sound captured by their local microphone as a reference signal and the voice transmitted back from the master device as the signal to be processed, performing AEC on the voice transmitted back from the master device. However, Bluetooth communication is prone to packet loss and transmission delay fluctuations due to factors such as distance changes and wireless interference, resulting in delay jitter between the reference signal and the signal to be processed. Since AEC is extremely sensitive to delay, delay jitter will significantly degrade the echo cancellation effect. Summary of the Invention
[0005] The purpose of this application is to address the shortcomings of the prior art by providing an acoustic echo cancellation method and a Bluetooth communication system to solve the practical problem of latency jitter degrading AEC performance in the prior art.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide an acoustic echo cancellation method applied to a slave device in a Bluetooth communication system; the method includes: The first voice data is sent to the master device based on the current voice frame collected. The first voice data includes the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame. The current voice frame and the corresponding voice identifier are stored as local data in the local cache area. The system receives second voice data sent by the master device. The second voice data includes voice identifiers associated with the device identifiers of each slave device and encoded voice data that has been merged from multiple sources. Based on the second voice data and the local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded voice in the second voice data to obtain processed voice, and the processed voice is played.
[0007] As an optional implementation, sending the first voice data to the master device based on the acquired current voice frame includes: Generate a voice identifier corresponding to the current voice frame based on the current voice frame; The current speech frame is encoded to obtain the encoded speech of the current speech frame; Obtain the device identifier, integrate the voice identifier corresponding to the current voice frame, the device identifier, and the encoded speech of the current voice frame into the first voice data of the slave device, and send the first voice data to the master device through the Bluetooth transmission module deployed on the slave device.
[0008] As an optional implementation, the step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data based on the second speech data and the local data to obtain processed speech includes: The second voice data is parsed to obtain the voice identifier associated with the device identifier of each slave device and the encoded voice after multi-channel merging; The target voice identifier is determined based on the voice identifier associated with the device identifier of each slave device; Based on the target voice identifier and the local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0009] As an optional implementation, the step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the target speech identifier and the local data to obtain the processed speech includes: Based on the target voice identifier, the current voice frame corresponding to the target voice identifier is obtained from the local data; Based on the current voice frame corresponding to the target voice identifier, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0010] As an optional implementation, the step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the current speech frame corresponding to the target speech identifier to obtain the processed speech includes: Use the current voice frame corresponding to the target voice identifier as a reference signal; The encoded speech obtained by multi-channel merging is decoded to obtain decoded speech obtained by multi-channel merging. The decoded speech obtained by multi-channel merging is used as the signal to be processed. The reference signal is filtered out from the signal to be processed to obtain the processed speech.
[0011] As an optional implementation, the method further includes: If the second voice data sent by the master device is lost, the master device performs packet loss compensation based on the previously processed voice data, generates compensation voice data, and plays the compensation voice data.
[0012] Secondly, embodiments of this application provide an acoustic echo cancellation method applied to a master device in a Bluetooth communication system; the method includes: Receive first voice data sent by each slave device, the first voice data including the voice identifier corresponding to the current voice frame collected by the slave device, the device identifier of the slave device, and the encoded voice of the current voice frame; Based on each of the first voice data, second voice data is generated and sent to each slave device. The second voice data includes a voice identifier associated with the device identifier of each slave device and encoded voice data that has been multi-channel merged.
[0013] As an optional implementation, the step of generating second voice data based on each of the first voice data, and sending the second voice data to each slave device respectively, includes: If the first voice data of each slave device is successfully received, the first voice data is parsed to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain the voice frames of each slave device. The voice frames from each slave device and the voice frames currently acquired by the master device are merged to obtain multi-channel merged speech, and the multi-channel merged speech is encoded to obtain the multi-channel merged encoded speech. The device identifier of each slave device and the voice identifier corresponding to the current voice frame collected by each slave device are associated to obtain the voice identifier associated with the device identifier of each slave device. The second voice data is generated based on the voice identifier associated with the device identifier of each slave device and the multi-channel merged encoded voice, and the second voice data is sent to each slave device through the Bluetooth transmission module deployed on the master device.
[0014] As an optional implementation, the step of generating second voice data based on each of the first voice data, and sending the second voice data to each slave device respectively, includes: If the first voice data sent by the target device is lost, then according to the target device voice frame obtained from the first voice data previously sent by the target device, packet loss compensation is performed on the target device to generate a new target device voice frame. Based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the first voice data previously sent by the target slave device and the device identifier of the target slave device, a new voice identifier associated with the device identifier of the target slave device is generated. Based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully transmitted by each non-target slave device, the second voice data is generated, and the second voice data is transmitted to the target slave device and each non-target slave device respectively through the Bluetooth transmission module deployed on the master device.
[0015] Thirdly, embodiments of this application provide an acoustic echo cancellation device corresponding to an acoustic echo cancellation method performed by a slave device, the acoustic echo cancellation device comprising: The first sending module is used to send first voice data to the master device according to the collected current voice frame. The first voice data includes the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame. The storage module is used to store the current voice frame and the voice identifier corresponding to the current voice frame as local data in the local cache area; The first receiving module is used to receive the second voice data sent by the master device. The second voice data includes voice identifiers associated with the device identifiers of each slave device and encoded voice data that has been merged from multiple sources. The processing module is configured to perform acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the second speech data and the local data, to obtain processed speech, and to play the processed speech.
[0016] As an optional implementation, the first sending module is specifically used for: Generate a voice identifier corresponding to the current voice frame based on the current voice frame; The current speech frame is encoded to obtain the encoded speech of the current speech frame; Obtain the device identifier, integrate the voice identifier corresponding to the current voice frame, the device identifier, and the encoded speech of the current voice frame into the first voice data of the slave device, and send the first voice data to the master device through the Bluetooth transmission module deployed on the slave device.
[0017] As an optional implementation, the processing module is specifically used for: The second voice data is parsed to obtain the voice identifier associated with the device identifier of each slave device and the encoded voice after multi-channel merging; The target voice identifier is determined based on the voice identifier associated with the device identifier of each slave device; Based on the target voice identifier and the local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0018] As an optional implementation, the processing module is specifically used for: Based on the target voice identifier, the current voice frame corresponding to the target voice identifier is obtained from the local data; Based on the current voice frame corresponding to the target voice identifier, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0019] As an optional implementation, the processing module is specifically used for: Use the current voice frame corresponding to the target voice identifier as a reference signal; The encoded speech obtained by multi-channel merging is decoded to obtain decoded speech obtained by multi-channel merging. The decoded speech obtained by multi-channel merging is used as the signal to be processed. The reference signal is filtered out from the signal to be processed to obtain the processed speech.
[0020] As an optional implementation, the processing module is further configured to: If the second voice data sent by the master device is lost, the master device performs packet loss compensation based on the previously processed voice data, generates compensation voice data, and plays the compensation voice data.
[0021] Fourthly, embodiments of this application provide an acoustic echo cancellation device corresponding to the acoustic echo cancellation method performed by the main device, the acoustic echo cancellation device comprising: The second receiving module is used to receive first voice data sent by each slave device. The first voice data includes the voice identifier corresponding to the current voice frame collected by the slave device, the device identifier of the slave device, and the encoded voice of the current voice frame. The second sending module is used to generate second voice data based on each of the first voice data, and send the second voice data to each slave device respectively. The second voice data includes a voice identifier associated with the device identifier of each slave device and multi-channel merged encoded voice.
[0022] As an optional implementation, the second sending module is specifically used for: If the first voice data of each slave device is successfully received, the first voice data is parsed to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain the voice frames of each slave device. The voice frames from each slave device and the voice frames currently acquired by the master device are merged to obtain multi-channel merged speech, and the multi-channel merged speech is encoded to obtain the multi-channel merged encoded speech. The device identifier of each slave device and the voice identifier corresponding to the current voice frame collected by each slave device are associated to obtain the voice identifier associated with the device identifier of each slave device. The second voice data is generated based on the voice identifier associated with the device identifier of each slave device and the multi-channel merged encoded voice, and the second voice data is sent to each slave device through the Bluetooth transmission module deployed on the master device.
[0023] As an optional implementation, the second sending module is specifically used for: If the first voice data sent by the target device is lost, then according to the target device voice frame obtained from the first voice data previously sent by the target device, packet loss compensation is performed on the target device to generate a new target device voice frame. Based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the first voice data previously sent by the target slave device and the device identifier of the target slave device, a new voice identifier associated with the device identifier of the target slave device is generated. Based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully transmitted by each non-target slave device, the second voice data is generated, and the second voice data is transmitted to the target slave device and each non-target slave device respectively through the Bluetooth transmission module deployed on the master device.
[0024] Fifthly, embodiments of this application provide an electronic device, which is either the slave device or the master device. The electronic device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the method steps as described in the first aspect or the second aspect of the method executed by the slave device or the master device.
[0025] Sixthly, embodiments of this application provide a Bluetooth intercom system, the Bluetooth intercom system including a master device and multiple slave devices that communicate with the master device via Bluetooth; Each of the aforementioned slave devices is used to perform the steps of the acoustic echo cancellation method described in the first aspect above; The main device is used to perform the steps of the acoustic echo cancellation method described in the second aspect above.
[0026] The beneficial effects of this application are: This application provides an acoustic echo cancellation method and a Bluetooth intercom system. A slave device sends first voice data, containing a voice identifier corresponding to the current voice frame, a device identifier of the slave device, and encoded speech of the current voice frame, to a master device based on the acquired current voice frame. The master device stores the current voice frame and its corresponding voice identifier as local data in a local cache area. The master device generates second voice data, containing voice identifiers associated with the device identifiers of each slave device and encoded speech from multiple channels, based on each first voice data, and sends this data to each slave device. The slave devices perform acoustic echo cancellation processing on the encoded speech from the multiple channels in the second voice data, based on the second voice data and the local data, to obtain processed speech. This ensures the stability of the time delay between the reference signal and the signal to be processed during AEC processing, eliminates time delay jitter caused by the Bluetooth transmission link, improves the AEC processing effect, and thus improves the quality of the processed speech, avoiding howling caused by sound backhaul. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a schematic diagram of the architecture of a Bluetooth intranet communication system provided in an embodiment of this application; Figure 2 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 1 ; Figure 3 This is a schematic diagram of the frame structure of the first voice data provided in an embodiment of this application; Figure 4 This is a schematic diagram of the frame structure of the second voice data provided in an embodiment of this application; Figure 5 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 2 ; Figure 6 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 3 ; Figure 7 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 4 ; Figure 8 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 5 ; Figure 9 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 6 ; Figure 10 A schematic diagram illustrating the determination of the current voice frame corresponding to the target voice identifier provided in an embodiment of this application; Figure 11 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 7 ; Figure 12 A modular structure diagram of an acoustic echo cancellation device provided in an embodiment of this application; Figure 13 A module structure diagram of another acoustic echo cancellation device provided in the embodiments of this application; Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0030] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0031] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0032] In multi-role real-time voice interaction scenarios, the master device receives and decodes the voice from each slave device, plays it locally, and simultaneously merges it with the voice captured by its own microphone. After encoding, the merged voice is sent back to all slave devices. The slave devices receive the merged voice from the master device, decode it, and play it, achieving full-duplex intra-device communication. Because the audio decoded and played by the slave devices contains the sound they themselves captured, this sound, after being played through the speaker, will be picked up again by their own microphone, creating an acoustic echo. In severe cases, this can lead to a decrease in call quality or even feedback. Therefore, the slave devices need to use AEC (Aspect-Oriented Echo Control) to filter out the sound they themselves captured before playing it through the speaker.
[0033] Currently, slave devices typically use the sound captured by their local microphone as a reference signal and the voice transmitted back from the master device as the signal to be processed, performing AEC on the voice transmitted back from the master device. However, Bluetooth communication is prone to packet loss and transmission delay fluctuations due to factors such as distance changes and wireless interference, resulting in delay jitter between the reference signal and the signal to be processed. Since AEC is extremely sensitive to delay, delay jitter will significantly degrade the echo cancellation effect.
[0034] Based on the above-mentioned problems, this application provides an acoustic echo cancellation method to improve the time delay stability between the reference signal and the signal to be processed, thereby improving the AEC effect and ensuring the voice playback quality of the device.
[0035] Figure 1 This is a schematic diagram of the architecture of the Bluetooth intranet communication system provided in the embodiments of this application, as shown below. Figure 1 As shown, a Bluetooth intranet system includes a master device and multiple slave devices that communicate with the master device via Bluetooth. Figure 1 For example, the main device and Bluetooth communication connection between the slave device and the device.
[0036] Specifically, each slave device sends first voice data to the master device. The first voice data includes the voice identifier corresponding to the current voice frame collected by the slave device, the device identifier of the slave device, and the encoded voice of the current voice frame. The current voice frame and the voice identifier corresponding to the current voice frame are stored in the local cache area.
[0037] The master device receives and parses the first voice data sent by each slave device, generates second voice data containing voice identifiers associated with the device identifiers of each slave device and encoded voice data that has been multiplied, and sends the second voice data to each slave device respectively.
[0038] The device receives second voice data, determines the target voice identifier of the slave device based on the voice identifier associated with the device identifier of each slave device in the second voice data, and searches for the current voice frame corresponding to the target voice identifier in the local buffer area of the slave device based on the target voice identifier of the slave device, using it as the reference signal for AEC. The multi-channel merged encoded voice in the second voice data is decoded to obtain the multi-channel merged decoded voice, which is used as the signal to be processed by AEC. The reference signal and the voice collected by the slave device itself are filtered out from the signal to be processed to obtain the processed voice, and the processed voice is played through the local speaker to avoid howling caused by sound feedback.
[0039] Since the device uses the current voice frame corresponding to the target voice identifier stored locally as the reference signal for AEC, the time delay between the reference signal and the signal to be processed changes from the dynamically changing network time delay to the fixed voice codec processing time delay, which greatly improves the time delay stability between the reference signal and the signal to be processed, thereby improving the AEC effect and ensuring the voice playback quality of the device.
[0040] Figure 2 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 1 This paper uses the interaction between a master device and a slave device in a Bluetooth communication system as an example to explain the process of acoustic echo cancellation. The slave device is... Any one of the slave devices in the set. For example... Figure 2 As shown, the method includes: S101. The slave device sends the first voice data to the master device based on the collected current voice frame.
[0041] The first voice data includes the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame.
[0042] Optionally, the device can acquire local voice data via a locally deployed microphone to obtain the current voice frame. Generate the voice identifier corresponding to the current voice frame. and the current voice frame Encoded speech as the current speech frame .
[0043] The equipment is pre-assigned a corresponding equipment identifier. ,by Figure 1 Main equipment and Taking a slave device Bluetooth communication connection as an example The device identifiers of each slave device are as follows: , … .
[0044] From the device according to the device identification The voice identifier corresponding to the current voice frame and the encoded speech of the current speech frame The system generates the first voice data and sends it to the master device, which then receives the first voice data sent by each slave device.
[0045] Figure 3 This is a schematic diagram of the frame structure of the first voice data provided in an embodiment of this application, as shown below. Figure 2 As shown, the frame structure of the first voice data includes the device identifier of the slave device. The voice identifier corresponding to the current voice frame and the encoded speech of the current speech frame .
[0046] S102. The slave device stores the current voice frame and the corresponding voice identifier as local data in the local cache area.
[0047] Optionally, the device will send the unencoded current voice frame. and the voice identifier corresponding to the current voice frame As local data, the pair is stored in the local cache area of the device.
[0048] When the second voice data is subsequently received from the master device, the local data stored in the local cache area can be used as the basis for finding the AEC reference signal, so that the time delay between the reference signal and the signal to be processed is no longer affected by network transmission jitter.
[0049] S103. The main device generates second voice data based on each first voice data.
[0050] The second voice data includes voice identifiers associated with the device identifiers of each slave device and encoded voice data that has been merged into multiple channels.
[0051] Optionally, the master device receives the first voice data sent by each slave device, and determines the voice data based on the master device's collected voice frames and the voice identifiers corresponding to the current voice frames collected by the slave devices in each of the first voice data. From the equipment identification and the encoded speech of the current speech frame This generates the second voice data.
[0052] Figure 4 This is a schematic diagram of the frame structure of the second voice data provided in an embodiment of this application, as shown below. Figure 4 As shown, the frame structure of the second voice data includes voice identifiers associated with the device identifiers of each slave device. and multi-channel merged encoded speech .
[0053] Among them, reference Figure 4 Voice identifier associated with the device identifier of each slave device include , … .
[0054] S104. The master device sends the second voice data to each slave device.
[0055] Optionally, the master device sends the second voice data to each slave device, so that each slave device can parse the voice identifier associated with its device identifier from the second voice data sent by the master device. and multi-channel merged encoded speech .
[0056] S105. The device performs acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data based on the second speech data and local data to obtain the processed speech.
[0057] Optionally, the slave device uses the voice identifier associated with the device identifier of each slave device in the second voice data. And from the local data stored in the device's local cache area, determine the reference signal and the signal to be processed during AEC processing, in order to process the multi-channel combined encoded speech in the second speech data. AEC processing is performed to obtain the processed speech.
[0058] The local data stored in the local cache area provides a basis for finding the AEC reference signal, ensuring the stability of the time delay between the reference signal and the signal to be processed during AEC processing. The processed speech has had echoes from the device itself filtered out, thus improving the quality of the processed speech.
[0059] S106. Play the processed audio from the device.
[0060] Alternatively, the device can play processed speech through a locally deployed speaker. Since the processed speech is high-quality speech with the echo from the device itself filtered out, feedback caused by sound backlash can be avoided.
[0061] In this embodiment, the slave device sends first voice data, containing the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame, to the master device based on the acquired current voice frame. The master device stores the current voice frame and its corresponding voice identifier as local data in a local cache area. The master device generates second voice data, containing voice identifiers associated with the device identifiers of each slave device and multi-channel merged encoded voice, based on each first voice data, and sends this data to each slave device. The slave devices perform acoustic echo cancellation processing on the multi-channel merged encoded voice in the second voice data based on the second voice data and the local data, obtaining processed voice. This ensures the stability of the time delay between the reference signal and the signal to be processed during AEC processing, eliminates time delay jitter caused by the Bluetooth transmission link, improves the AEC processing effect, and thus improves the quality of the processed voice, avoiding howling caused by sound backhaul.
[0062] Figure 5 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 2 ,like Figure 5 As shown, in step S101 above, the slave device sends the first voice data to the master device based on the collected current voice frame, including: S201. Generate the voice identifier corresponding to the current voice frame based on the current voice frame.
[0063] Optionally, the device uses a cyclic increment method to represent the current voice frame. Assign a unique voice identifier The voice identifier corresponding to the current voice frame. Used to uniquely identify the current speech frame .
[0064] For example, voice identifier It can be a cyclic code from 0 to 255, if the current voice frame The voice identifier corresponding to the previous audio frame is Then the device determines the voice identifier corresponding to the current voice frame as... This ensures the uniqueness of the voice identifier corresponding to the current voice frame.
[0065] S202. Encode the current speech frame to obtain the encoded speech of the current speech frame.
[0066] Optionally, from the device the current voice frame Perform encoding and compression to generate the encoded speech of the current speech frame. .
[0067] By analyzing the current speech frame Encoding and compression are performed to reduce the amount of data transmitted between master and slave devices via Bluetooth, ensuring correct voice transmission while meeting the real-time communication requirements of low bandwidth and low latency.
[0068] S203. Obtain the device identifier, integrate the voice identifier corresponding to the current voice frame, the device identifier, and the encoded voice of the current voice frame into the first voice data of the slave device, and send the first voice data to the master device through the Bluetooth transmission module deployed on the slave device.
[0069] Optionally, the device reads the device identifier assigned to it. ,according to Figure 3 The frame structure shown will identify the device. The voice identifier corresponding to the current voice frame and the encoded speech of the current speech frame The data is packaged into first voice data and sent to the master device via a locally deployed Bluetooth transmission module. The slave device's device identifier is included. It makes it easier for the master device to distinguish between different slave devices.
[0070] For example, if the device identifier of the device is The voice identifier corresponding to the current voice frame is The first voice data from the device includes the device identifier. The voice identifier corresponding to the current voice frame and the encoded speech of the current speech frame .
[0071] In this embodiment, the slave device generates a voice identifier corresponding to the current voice frame based on the current voice frame, and encodes the current voice frame to obtain the encoded speech of the current voice frame. The slave device obtains a device identifier, integrates the voice identifier corresponding to the current voice frame, the device identifier, and the encoded speech of the current voice frame into first speech data of the slave device, and sends the first speech data to the master device through the Bluetooth transmission module deployed on the slave device. This ensures that the first speech data carries the device identifier of the slave device, the voice identifier corresponding to the current voice frame, and the encoded speech of the current voice frame.
[0072] Figure 6 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 3 ,like Figure 6 As shown, in the above steps, the master device generates second voice data based on each first voice data, and sends the second voice data to each slave device respectively, including: S301. If the first voice data of each slave device is successfully received, the first voice data is parsed to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain the voice frames of each slave device.
[0073] Optionally, after successfully receiving the first voice data from each slave device (i.e., no packet loss in the first voice data from each slave device), the master device parses each first voice data to obtain the device identifier of each slave device. The voice identifier corresponding to the current voice frame collected by each device. and the encoded speech of each current speech frame .
[0074] The master device encodes the speech of each current voice frame. The decoding process restores the audio frames from each slave device, where each slave device audio frame represents the current audio frame captured by that slave device. .
[0075] S302. Merge the voice frames from each slave device and the voice frames currently acquired by the master device to obtain multi-channel merged speech, and encode the multi-channel merged speech to obtain multi-channel merged encoded speech.
[0076] Optionally, the master device acquires local voice data via a locally deployed microphone to obtain master device voice frames. It then mixes the voice frames from each slave device with the master device voice frames to obtain a single merged voice stream containing the voices of all participants (i.e., users holding the master device and users holding each slave device), and uses this merged voice stream as the multi-stream merged voice stream. .
[0077] The main device performs multi-channel voice merging. Encode and compress to generate multi-channel merged encoded speech. This serves as the audio payload to be sent to each slave device. Through multi-channel merging of the audio... Encoding and compression are performed to further reduce the amount of data transmitted between master and slave devices via Bluetooth, ensuring correct voice transmission while meeting the real-time communication requirements of low bandwidth and low latency.
[0078] S303. Perform association processing on the device identifier of each slave device and the voice identifier corresponding to the current voice frame collected by each slave device to obtain the voice identifier associated with the device identifier of each slave device.
[0079] Optionally, the master device may use the voice identifiers corresponding to the current voice frames collected by each slave device participating in the mixing process. Its equipment identification Perform one-to-one binding and generate voice identifiers associated with the device identifiers of each slave device. .
[0080] For example, if the device identifier of the device is The voice identifier corresponding to the current voice frame is The master device then determines the voice identifier associated with the device identifier of the slave device as... .
[0081] Voice identifier associated with the device identifier of each slave device To facilitate precise location of the current voice frame of each slave device participating in the mixing process. This improves the accuracy of the reference signal when each slave device performs AEC processing.
[0082] S304. Based on the voice identifier associated with the device identifier of each slave device and the encoded voice data after multi-channel merging, generate second voice data and send the second voice data to each slave device through the Bluetooth transmission module deployed on the master device.
[0083] Optionally, the main equipment is configured according to Figure 4 The frame structure shown associates the voice identifiers with the device identifiers of each slave device. and multi-channel merged encoded speech The data is packaged into second voice data and sent to each slave device via a locally deployed Bluetooth transmission module.
[0084] In this embodiment, if the master device successfully receives the first voice data from each slave device, it parses each first voice data to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain each slave device voice frame. The slave device voice frames and the master device's currently collected master device voice frames are merged to obtain multi-channel merged voice. The multi-channel merged voice is then encoded to obtain multi-channel merged encoded voice. The device identifiers of each slave device and the voice identifiers corresponding to the current voice frames collected by each slave device are associated to obtain voice identifiers associated with each slave device's device identifier. Based on the voice identifiers associated with each slave device's device identifier and the multi-channel merged encoded voice, second voice data is generated and sent to each slave device via the Bluetooth transmission module deployed on the master device. The voice identifiers associated with the device identifiers of each slave device in the second voice data facilitate accurate location of the current voice frame of the slave device participating in the mixing, thereby improving the accuracy of the reference signal when each slave device performs AEC processing.
[0085] Figure 7 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 4 ,like Figure 7 As shown, in the above steps, the master device generates second voice data based on each first voice data, and sends the second voice data to each slave device respectively, including: S401. If the first voice data sent by the target slave device is lost, then based on the target slave device voice frame obtained from the first voice data previously sent by the target slave device, packet loss compensation is performed on the target slave device to generate a new target slave device voice frame.
[0086] Optionally, if the master device does not receive the first voice data from a slave device, and takes that slave device as the target slave device (i.e., the first voice data of the target slave device is lost), then the master device queries the first voice data of the target slave device that was previously successfully received locally.
[0087] The master device retrieves the target voice frame from the first voice data of the device based on the previously successfully received target. A compensated voice frame is generated using the Packet Loss Compensation (PLC) algorithm, which serves as the new target slave device voice frame. By using packet loss compensation to fill in the audio gaps caused by the loss of the first voice data packets from the target device, the continuity of sound from the target device after mixing is ensured.
[0088] S402. Based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the first voice data previously sent by the target slave device and the device identifier of the target slave device, generate a new voice identifier associated with the device identifier of the target slave device.
[0089] Optionally, the master device uses the voice identifier corresponding to the current voice frame collected by the device from the first voice data of the previously successfully received target to determine the target. Calculate the voice identifier corresponding to the lost current voice frame. .
[0090] Specifically, because the device uses a cyclically incrementing method to generate the current voice frame... Assign a unique voice identifier Then, the voice identifier corresponding to the current voice frame collected by the target slave device in the first voice data of the target slave device that the master device successfully received in the previous instance. Adding 1 to the base value will give you the voice identifier corresponding to the lost current voice frame. .
[0091] For example, if the master device previously successfully received the voice identifier corresponding to the current voice frame collected by the target device in the first voice data of the target device... for The master device then determines the voice identifier corresponding to the lost current voice frame as follows: .
[0092] The master device will retrieve the voice identifier corresponding to the current voice frame that the device lost. With the target device's device identifier Bind, generate a new voice identifier associated with the device identifier of the target slave device.
[0093] S403. Based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully sent by each non-target slave device, generate second voice data, and send the second voice data to the target slave device and each non-target slave device respectively through the Bluetooth transmission module deployed on the master device.
[0094] Optionally, the master device parses the first voice data successfully transmitted by each non-target slave device (i.e., each slave device that did not lose packets) to obtain the device identifier of each non-target slave device and the device identifier of each non-target slave device. Voice identifier corresponding to the current voice frame collected by the device and the encoded speech of the current voice frames collected from each non-target device. And the encoded speech of the current voice frames collected from each non-target device. Decoding is performed to obtain the voice frames of each non-target slave device.
[0095] The master device acquires local voice data via a locally deployed microphone, obtaining master device voice frames. It then mixes the non-target slave device voice frames, the master device voice frames, and the new target slave device voice frames obtained through packet loss compensation to produce multi-channel merged audio. .
[0096] The main device will use the voice identifiers corresponding to the current voice frames collected by the device from each non-target device participating in the mixing process. Its equipment identification Perform one-to-one binding and generate voice identifiers associated with the device identifiers of each non-target slave device.
[0097] The master device will integrate the voice identifiers associated with the device identifiers of each non-target slave device and the new voice identifiers associated with the device identifiers of the target slave devices into a single voice identifier associated with the device identifiers of each slave device. and in accordance with Figure 4 The frame structure shown associates the voice identifiers with the device identifiers of each slave device. and multi-channel merged encoded speech The data is packaged into second voice data and sent to the target slave device and each non-target slave device via a locally deployed Bluetooth transmission module.
[0098] In this embodiment, if the first voice data sent by the target slave device is lost, the master device performs packet loss compensation on the target slave device based on the target slave device voice frame obtained from the previously sent first voice data, generating a new target slave device voice frame. Furthermore, based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the previously sent first voice data and the device identifier of the target slave device, a new voice identifier associated with the device identifier of the target slave device is generated. Second voice data is generated based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully sent by each non-target slave device. This second voice data is then sent to the target slave device and each non-target slave device via the Bluetooth transmission module deployed on the master device. By performing packet loss compensation in the event of packet loss on the target slave device and generating a voice identifier associated with the device identifier of the target slave device, the AEC (Advanced Capability Equipping) effect of the target slave device in the event of packet loss is ensured.
[0099] Figure 8 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 5 ,like Figure 8As shown, in step S105 above, the slave device performs acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data based on the second speech data and local data to obtain processed speech, and then plays the processed speech, including: S501. The second voice data is parsed to obtain the voice identifier associated with the device identifier of each slave device and the encoded voice after multi-channel merging.
[0100] Optionally, the slave device parses the second voice data sent by the master device to obtain the voice identifier associated with the device identifier of each slave device. and multi-channel merged encoded speech .
[0101] S502. Determine the target voice identifier based on the voice identifier associated with the device identifier of each slave device.
[0102] Optionally, the device can identify itself. As an index, the voice identifier associated with the device identifier of each slave device is used. Find the device identifier that matches its own The associated voice identifier serves as the target voice identifier for the slave device.
[0103] The target voice identifier of the slave device corresponds to the voice identifier of the current voice frame carried in the first voice data sent by the slave device. Exactly the same.
[0104] S503. Based on the target speech identifier and local data, perform acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0105] Optionally, the device uses the target voice identifier as an index to query the local data stored in the local cache area of the device to obtain the reference signal for AEC processing.
[0106] The device encodes the multi-channel combined speech from the second speech data based on the reference signal during AEC processing. AEC processing is performed to obtain high-quality processed speech.
[0107] In this embodiment, the slave device parses the second voice data to obtain voice identifiers associated with the device identifiers of each slave device and multi-channel merged encoded voice. Based on the voice identifiers associated with the device identifiers of each slave device, a target voice identifier is determined. Then, based on the target voice identifier and local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded voice in the second voice data to obtain the processed voice. This ensures the accuracy of the reference signal during AEC processing.
[0108] Figure 9 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 6 ,like Figure 9 As shown, in step S503 above, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data based on the target speech identifier and local data to obtain the processed speech, including: S601. Based on the target voice identifier, retrieve the current voice frame corresponding to the target voice identifier from the local data.
[0109] Optionally, the local data stored in the device's local cache area includes each current voice frame acquired from the device. and the voice identifier corresponding to the current voice frame The device uses the target voice identifier as an index to query the local data stored in the device's local cache area to obtain the current voice frame corresponding to the target voice identifier.
[0110] Figure 10 This is a schematic diagram illustrating the determination of the current voice frame corresponding to the target voice identifier provided in an embodiment of this application, with reference to... Figure 10 If the equipment identifier is From the device according to the device identifier Voice identifier associated with the device identifier of each slave device The target voice identifier was found in the middle. (like Figure 10 (As shown in the orange highlighted portion of the second speech data), then based on the target speech identifier... The target voice identifier is obtained by querying local data stored in the device's local cache area. The corresponding current voice frame, such as Figure 10 The orange highlighted portion is shown in the local data.
[0111] S602. Based on the current speech frame corresponding to the target speech identifier, perform acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0112] Optionally, the device performs multi-channel merging encoded speech processing on the second speech data based on the current speech frame corresponding to the target speech identifier. AEC processing is performed to obtain high-quality processed speech.
[0113] In this embodiment, the device retrieves the current voice frame corresponding to the target voice identifier from local data based on the target voice identifier, and then performs acoustic echo cancellation (AEC) processing on the multi-channel merged encoded speech in the second speech data based on the current voice frame corresponding to the target voice identifier, obtaining the processed speech. This improves the AEC effect and ensures the quality of the processed speech.
[0114] Figure 11 Flowchart of the acoustic echo cancellation method provided in the embodiments of this application Figure 7 ,like Figure 11 As shown, in step S602 above, acoustic echo cancellation is performed on the multi-channel merged encoded speech in the second speech data according to the current speech frame corresponding to the target speech identifier, resulting in processed speech. This includes: S701, Use the current voice frame corresponding to the target voice identifier as a reference signal.
[0115] Optionally, the slave device uses the current voice frame corresponding to the target voice identifier as a reference signal, and the current voice frame corresponding to the target voice identifier is the voice frame collected by the slave device and sent to the master device.
[0116] Using the current voice frame corresponding to the target voice identifier as a reference signal can eliminate the latency of the Bluetooth transmission link during AEC processing, so that the latency between the reference signal and the signal to be processed only includes a fixed encoding and decoding latency.
[0117] S702. Decode the multi-channel merged encoded speech to obtain the multi-channel merged decoded speech, and use the multi-channel merged decoded speech as the signal to be processed. Filter out the reference signal from the signal to be processed to obtain the processed speech.
[0118] Optionally, the device can combine multiple encoded voice streams. The decoded audio is restored to a multi-channel merged decoded audio, which contains the voices of all participants.
[0119] The slave device uses the decoded audio from multiple channels as the signal to be processed and filters out the reference signal from the signal to be processed. That is, it filters out the current voice frame collected by the slave device and sent to the master device from the decoded audio from multiple channels, so that the processed audio only contains the voices of all users except the user holding the slave device, thus avoiding feedback caused by sound transmission.
[0120] In this embodiment, the slave device uses the current voice frame corresponding to the target voice identifier as a reference signal to decode the multi-channel merged encoded speech, obtaining multi-channel merged decoded speech. This decoded speech is then used as the signal to be processed, from which the reference signal is filtered out to obtain the processed speech. AEC processing eliminates the latency of the Bluetooth transmission link, ensuring that the latency between the reference signal and the signal to be processed only includes a fixed encoding / decoding latency, thus guaranteeing latency stability and ensuring the playback quality of the processed speech.
[0121] It is worth noting that the device can also first process the encoded speech data from the multi-channel merged second voice data. Decoding is performed to obtain multi-channel merged decoded speech, which is then used as the signal to be processed. Next, based on the voice identifier associated with the device identifier of each slave device, the target voice identifier is determined. Then, based on the target voice identifier and local data, the current voice frame corresponding to the target voice identifier is determined, and this current voice frame is used as the reference signal. In other words, this embodiment does not specifically limit the order in which the slave devices determine the reference signal and the signal to be processed.
[0122] As an optional implementation, the method further includes: If the second voice data sent by the master device is lost, the master device will perform packet loss compensation based on the previously processed voice data, generate compensation voice, and play the compensation voice.
[0123] Optionally, if the slave device fails to receive the second voice data sent by the master device, i.e. the second voice data sent by the master device is lost, the slave device reads the processed voice obtained from the previous AEC processing, uses the PLC algorithm to smooth and extrapolate or waveform extend the previous processed voice, generates a frame of compensated voice, and plays the compensated voice through a speaker deployed locally on the slave device.
[0124] In this embodiment, if the second voice data sent by the master device is lost, packet loss compensation is performed on the master device based on the previously processed voice data to generate and play the compensation voice. This ensures the continuity of voice playback on the slave device.
[0125] Based on the same inventive concept, this application also provides an acoustic echo cancellation device corresponding to the acoustic echo cancellation method performed by the slave device. Since the principle of the device in this application is similar to the acoustic echo cancellation method performed by the slave device in the above-mentioned application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0126] Figure 12 This application provides a modular structure diagram of an acoustic echo cancellation device, as shown in the embodiments below. Figure 12 As shown, the device includes: The first sending module 1201 is used to send first voice data to the master device according to the collected current voice frame. The first voice data includes the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame. Storage module 1202 is used to store the current voice frame and the corresponding voice identifier as local data in the local cache area; The first receiving module 1203 is used to receive the second voice data sent by the master device. The second voice data includes voice identifiers associated with the device identifiers of each slave device and encoded voice data after multiplexing. The processing module 1204 is used to perform acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the second speech data and local data, to obtain the processed speech, and to play the processed speech.
[0127] As an optional implementation, the first sending module 1201 is specifically used for: Generate the voice identifier corresponding to the current voice frame based on the current voice frame; The current speech frame is encoded to obtain the encoded speech of the current speech frame; Obtain the device identifier, integrate the voice identifier corresponding to the current voice frame, the device identifier, and the encoded voice of the current voice frame into the first voice data of the slave device, and send the first voice data to the master device through the Bluetooth transmission module deployed on the slave device.
[0128] As an optional implementation, the processing module 1204 is specifically used for: The second voice data is parsed to obtain the voice identifier associated with the device identifier of each slave device and the encoded voice after multi-channel merging; The target voice identifier is determined based on the voice identifier associated with the device identifier of each slave device; Based on the target speech identifier and local data, acoustic echo cancellation is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0129] As an optional implementation, the processing module 1204 is specifically used for: Based on the target voice identifier, retrieve the current voice frame corresponding to the target voice identifier from the local data; Based on the current speech frame corresponding to the target speech identifier, acoustic echo cancellation is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
[0130] As an optional implementation, the processing module 1204 is specifically used for: Use the current voice frame corresponding to the target voice identifier as the reference signal; The encoded speech from multiple channels is decoded to obtain decoded speech from multiple channels. The decoded speech from multiple channels is then used as the signal to be processed. The reference signal is filtered out from the signal to be processed to obtain the processed speech.
[0131] As an optional implementation, the processing module 1204 is further configured to: If the second voice data sent by the master device is lost, the master device will perform packet loss compensation based on the previously processed voice data, generate compensation voice, and play the compensation voice.
[0132] Based on the same inventive concept, this application also provides an acoustic echo cancellation device corresponding to the acoustic echo cancellation method performed by the main device. Since the principle of the device in this application is similar to the acoustic echo cancellation method performed by the main device in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0133] Figure 13 A module structure diagram of another acoustic echo cancellation device provided in the embodiments of this application is shown below. Figure 13 As shown, the device includes: The second receiving module 1301 is used to receive first voice data sent by each slave device. The first voice data includes the voice identifier corresponding to the current voice frame collected by the slave device, the device identifier of the slave device, and the encoded voice of the current voice frame. The second sending module 1302 is used to generate second voice data based on each first voice data, and send the second voice data to each slave device respectively. The second voice data includes a voice identifier associated with the device identifier of each slave device and multi-channel merged encoded voice.
[0134] As an optional implementation, the second transmitting module 1302 is specifically used for: If the first voice data of each slave device is successfully received, the first voice data is parsed to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain the voice frames of each slave device. The voice frames from each slave device and the voice frames currently acquired by the master device are merged to obtain multi-channel merged speech. The multi-channel merged speech is then encoded to obtain multi-channel merged encoded speech. The device identifier of each slave device and the voice identifier corresponding to the current voice frame collected by each slave device are associated to obtain the voice identifier associated with the device identifier of each slave device. Based on the voice identifier associated with the device identifier of each slave device and the encoded voice data after multi-channel merging, the second voice data is generated and sent to each slave device through the Bluetooth transmission module deployed on the master device.
[0135] As an optional implementation, the second transmitting module 1302 is specifically used for: If the first voice data sent by the target device is lost, the target device will be compensated for the packet loss based on the target device voice frame obtained from the first voice data sent by the target device in the previous transmission, and a new target device voice frame will be generated. Based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the first voice data previously sent by the target slave device, and the device identifier of the target slave device, a new voice identifier associated with the device identifier of the target slave device is generated. Based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully transmitted by each non-target slave device, second voice data is generated, and the second voice data is transmitted to the target slave device and each non-target slave device respectively through the Bluetooth transmission module deployed on the master device.
[0136] This application also provides an electronic device, which can be the slave device or master device described above. For example... Figure 14 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application, including a processor 141, a memory 142, and a bus 143. The memory 142 stores machine-readable instructions executable by the processor 141. When the electronic device is running, the processor 141 communicates with the memory 142 via the bus 143, and the processor 141 executes the machine-readable instructions to perform the method steps performed by the slave device or the master device described above.
[0137] This application provides a Bluetooth intranet communication system, such as... Figure 1 As shown, the Bluetooth intranet system includes a master device and multiple slave devices that communicate with the master device via Bluetooth.
[0138] Each slave device and master device is used to perform the steps of the acoustic echo cancellation method described in the foregoing embodiments.
[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0140] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0141] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. An acoustic echo cancellation method, characterized in that, A slave device applied in a Bluetooth communication system; the method includes: The first voice data is sent to the master device based on the current voice frame collected. The first voice data includes the voice identifier corresponding to the current voice frame, the device identifier of the slave device, and the encoded voice of the current voice frame. The current voice frame and the corresponding voice identifier are stored as local data in the local cache area. The system receives second voice data sent by the master device. The second voice data includes voice identifiers associated with the device identifiers of each slave device and encoded voice data that has been merged from multiple sources. Based on the second voice data and the local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded voice in the second voice data to obtain processed voice, and the processed voice is played.
2. The method according to claim 1, characterized in that, The step of sending the first voice data to the master device based on the collected current voice frame includes: Generate a voice identifier corresponding to the current voice frame based on the current voice frame; The current speech frame is encoded to obtain the encoded speech of the current speech frame; Obtain the device identifier, integrate the voice identifier corresponding to the current voice frame, the device identifier, and the encoded speech of the current voice frame into the first voice data of the slave device, and send the first voice data to the master device through the Bluetooth transmission module deployed on the slave device.
3. The method according to claim 1, characterized in that, The step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data based on the second speech data and the local data to obtain processed speech includes: The second voice data is parsed to obtain the voice identifier associated with the device identifier of each slave device and the encoded voice after multi-channel merging; The target voice identifier is determined based on the voice identifier associated with the device identifier of each slave device; Based on the target voice identifier and the local data, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
4. The method according to claim 3, characterized in that, The step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the target speech identifier and the local data to obtain the processed speech includes: Based on the target voice identifier, the current voice frame corresponding to the target voice identifier is obtained from the local data; Based on the current voice frame corresponding to the target voice identifier, acoustic echo cancellation processing is performed on the multi-channel merged encoded speech in the second speech data to obtain the processed speech.
5. The method according to claim 4, characterized in that, The step of performing acoustic echo cancellation processing on the multi-channel merged encoded speech in the second speech data according to the current speech frame corresponding to the target speech identifier to obtain the processed speech includes: Use the current voice frame corresponding to the target voice identifier as a reference signal; The encoded speech obtained by multi-channel merging is decoded to obtain decoded speech obtained by multi-channel merging. The decoded speech obtained by multi-channel merging is used as the signal to be processed. The reference signal is filtered out from the signal to be processed to obtain the processed speech.
6. The method according to claim 1, characterized in that, The method further includes: If the second voice data sent by the master device is lost, the master device performs packet loss compensation based on the previously processed voice data, generates compensation voice data, and plays the compensation voice data.
7. An acoustic echo cancellation method, characterized in that, The method is applied to a master device in a Bluetooth communication system; the method includes: Receive first voice data sent by each slave device, the first voice data including the voice identifier corresponding to the current voice frame collected by the slave device, the device identifier of the slave device, and the encoded voice of the current voice frame; Based on each of the first voice data, second voice data is generated and sent to each slave device. The second voice data includes a voice identifier associated with the device identifier of each slave device and encoded voice data that has been multi-channel merged.
8. The method according to claim 7, characterized in that, The step of generating second voice data based on each of the first voice data, and sending the second voice data to each slave device respectively, includes: If the first voice data of each slave device is successfully received, the first voice data is parsed to obtain the device identifier of each slave device, the voice identifier corresponding to the current voice frame collected by each slave device, and the encoded voice of each current voice frame. The encoded voice of each current voice frame is then decoded to obtain the voice frames of each slave device. The voice frames from each slave device and the voice frames currently acquired by the master device are merged to obtain multi-channel merged speech, and the multi-channel merged speech is encoded to obtain the multi-channel merged encoded speech. The device identifier of each slave device and the voice identifier corresponding to the current voice frame collected by each slave device are associated to obtain the voice identifier associated with the device identifier of each slave device. The second voice data is generated based on the voice identifier associated with the device identifier of each slave device and the multi-channel merged encoded voice, and the second voice data is sent to each slave device through the Bluetooth transmission module deployed on the master device.
9. The method according to claim 7, characterized in that, The step of generating second voice data based on each of the first voice data, and sending the second voice data to each slave device respectively, includes: If the first voice data sent by the target device is lost, then according to the target device voice frame obtained from the first voice data previously sent by the target device, packet loss compensation is performed on the target device to generate a new target device voice frame. Based on the voice identifier corresponding to the current voice frame collected by the target slave device obtained from the first voice data previously sent by the target slave device and the device identifier of the target slave device, a new voice identifier associated with the device identifier of the target slave device is generated. Based on the new target slave device voice frame, the new voice identifier associated with the device identifier of the target slave device, and the first voice data successfully transmitted by each non-target slave device, the second voice data is generated, and the second voice data is transmitted to the target slave device and each non-target slave device respectively through the Bluetooth transmission module deployed on the master device.
10. A Bluetooth communication system, characterized in that, The Bluetooth communication system includes a master device and multiple slave devices that communicate with the master device via Bluetooth. Each of the aforementioned slave devices is used to perform the steps of the acoustic echo cancellation method as described in any one of claims 1 to 6; The main device is used to perform the steps of the acoustic echo cancellation method as described in any one of claims 7 to 9.