Array microphone device cascading system and method

By cascading master and slave devices, multiple array microphone devices are connected and microphone signals are processed independently to extend the coverage area, solving the problem of limited coverage of a single device and achieving high-quality sound reinforcement and audio transmission over a wider range.

WO2026085941A1PCT designated stage Publication Date: 2026-04-30AISPEECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AISPEECH CO LTD
Filing Date
2024-11-11
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

The limited coverage of a single array microphone device restricts its widespread application in large spaces such as ultra-large conference rooms, lecture halls, and exhibition halls.

Method used

By cascading master and slave devices, multiple array microphone devices are connected. The master and slave devices independently process the acquired multi-channel microphone signals based on the reference speaker signal to generate playback audio signals and uplink audio signals, which expands the coverage and ensures mutual isolation and audio quality between each array microphone device.

Benefits of technology

It expands the coverage of the array microphone equipment, ensuring sound reinforcement and remote audio quality at the meeting venue, while also guaranteeing the independence of each device and the isolation of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131240_30042026_PF_FP_ABST
    Figure CN2024131240_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are an array microphone device cascading system and method. The system comprises a master device and at least one slave device, wherein a first slave output end and a second slave output end of the slave device are connected to a slave device signal input end group of the master device; a first master output end of the master device is connected to a loudspeaker, and a second master output end of the master device is connected to a far end; a first master input end of the master device and a first slave input end of the slave device are used for inputting a loudspeaker reference signal; and the master device and the slave device are configured to perform preset processing on multiple slave microphone signals at least on the basis of the loudspeaker reference signal to generate a playback audio signal and an uplink audio signal. The array microphone device cascading system using the cascading method expands the coverage range, and ensures mutual isolation between array microphone devices, and also ensures the sound reinforcement at a conference site and the audio quality at a far end.
Need to check novelty before this filing date? Find Prior Art

Description

Array microphone device cascading system and method

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 202411487429.8, filed on October 23, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of array microphone technology, and in particular to an array microphone device cascading system and method. Background Technology

[0004] With the development of array microphone technology, array microphone equipment used in conference rooms can not only meet the local amplification needs of the conference room, but also meet the needs of remote conferencing. However, the coverage of a single array microphone device is limited, which restricts the widespread application of array microphone equipment in large spaces such as ultra-large conference rooms, lecture halls, and exhibition halls.

[0005] Summary of the Invention

[0006] This application provides an array microphone device cascading system and method to at least solve one of the above-mentioned technical problems.

[0007] In a first aspect, embodiments of this application provide an array microphone device cascade system, including a master device and at least one slave device; the master device includes a first master input terminal, a slave device signal input terminal group, a first master output terminal, and a second master output terminal; the slave device includes a first slave input terminal, a first slave output terminal, and a second slave output terminal; wherein...

[0008] The first slave output terminal and the second slave output terminal are connected to the slave device signal input terminal group, the first main output terminal is connected to the speaker, and the second main output terminal is connected to the remote end;

[0009] The first main input terminal and the first slave input terminal are used to input the speaker reference signal;

[0010] The slave device is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device; the multiple slave microphone signals are acquired by a microphone array inside the slave device.

[0011] The main device is configured to: perform preset processing on multiple main microphone signals based on the speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device;

[0012] The master device is also configured to generate playback audio signals and uplink audio signals based at least on the first master output signal, the second master output signal, the first slave output signal, and the second slave output signal.

[0013] In some embodiments, the master device further includes a second master input terminal, and the slave device further includes a second slave input terminal; the second master input terminal and the second slave input terminal are used to receive downlink audio signals from the remote end;

[0014] Generating playback audio signals and uplink audio signals based at least on the first main output signal, the second main output signal, the first slave output signal, and the second slave output signal includes:

[0015] A playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal;

[0016] An uplink audio signal is generated based on the second master output signal and the second slave output signal.

[0017] In some embodiments, the master device further includes a third master output terminal, a third master input terminal, a fourth master output terminal, and a fourth master input terminal;

[0018] Generate a playback audio signal based on the first main output signal, the first slave output signal, and the downlink audio signal, including:

[0019] The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal.

[0020] The uplink audio signal is generated based on the second main output signal and the second slave output signal, including:

[0021] The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

[0022] In some embodiments, at least based on the speaker reference signal, the multi-channel main microphone signals are pre-processed to obtain a first main output signal and a second main output signal, including:

[0023] At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result;

[0024] At least based on the multi-microphone signals, the sound source is located;

[0025] At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result;

[0026] The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal.

[0027] The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

[0028] In some embodiments, sound source localization is performed at least based on the multiple microphone signals, including:

[0029] Sound source localization is performed based on the multi-microphone signals and the downlink audio signal.

[0030] In some embodiments, the step of performing beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performing echo suppression processing on the beamformed signal, includes:

[0031] When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process.

[0032] When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a first echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

[0033] Secondly, this application also provides a method for cascading array microphone devices, applied to an array microphone cascading system, the array microphone cascading system including a master device and at least one slave device; the master device includes a first master input terminal, a slave device signal input terminal group, a first master output terminal, and a second master output terminal; the slave device includes a first slave input terminal, a first slave output terminal, and a second slave output terminal; the method includes:

[0034] The first slave output terminal and the second slave output terminal are connected to the slave device signal input terminal group, the first master output terminal is connected to the speaker, and the second master output terminal is connected to the remote end; wherein, the first master input terminal and the first slave input terminal are used to input the speaker reference signal;

[0035] The slave device is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device, wherein the multiple slave microphone signals are acquired by a microphone array inside the slave device;

[0036] The main device is configured to: perform preset processing on multiple main microphone signals based on the speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device;

[0037] The master device is also configured to generate playback audio signals and uplink audio signals based at least on the first master output signal, the second master output signal, the first slave output signal, and the second slave output signal.

[0038] In some embodiments, the master device further includes a second master input terminal, and the slave device further includes a second slave input terminal; the second master input terminal and the second slave input terminal are used to receive downlink audio signals from the remote end;

[0039] Generating playback audio signals and uplink audio signals based at least on the first main output signal, the second main output signal, the first slave output signal, and the second slave output signal includes:

[0040] A playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal;

[0041] An uplink audio signal is generated based on the second master output signal and the second slave output signal.

[0042] In some embodiments, the master device further includes a third master output terminal, a third master input terminal, a fourth master output terminal, and a fourth master input terminal;

[0043] Generate a playback audio signal based on the first main output signal, the first slave output signal, and the downlink audio signal, including:

[0044] The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal.

[0045] The uplink audio signal is generated based on the second main output signal and the second slave output signal, including:

[0046] The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

[0047] In some embodiments, at least based on the speaker reference signal, the multi-channel main microphone signals are pre-processed to obtain a first main output signal and a second main output signal, including:

[0048] At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result;

[0049] At least based on the multi-microphone signals, the sound source is located;

[0050] At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result;

[0051] The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal.

[0052] The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

[0053] In some embodiments, sound source localization is performed at least based on the multiple microphone signals, including:

[0054] Sound source localization is performed based on the multi-microphone signals and the downlink audio signal.

[0055] In some embodiments, the step of performing beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performing echo suppression processing on the beamformed signal, includes:

[0056] When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process.

[0057] When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a first echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

[0058] This application connects multiple array microphone devices via a master-slave cascading method. Each master and slave device independently performs preset processing (e.g., echo cancellation and / or noise reduction) on the acquired multi-channel microphone signals based on a reference speaker signal, resulting in independent first master output signal, second master output signal, first slave output signal, and second slave output signal. Finally, the master device generates playback audio signals and uplink audio signals based on these signals. This cascading array microphone system extends the coverage area while ensuring isolation between the array microphone devices and maintaining sound reinforcement at the conference venue and high-quality audio at distant locations. Attached Figure Description

[0059] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 is a schematic block diagram of an embodiment of the array microphone device cascade system of this application;

[0061] Figure 2 is a schematic block diagram of another embodiment of the array microphone device cascade system of this application;

[0062] Figure 3 is a schematic block diagram of another embodiment of the array microphone device cascade system of this application;

[0063] Figure 4 is a schematic block diagram of another embodiment of the array microphone device cascade system of this application;

[0064] Figure 5 is a schematic diagram of an embodiment of the array microphone device of this application;

[0065] Figure 6 is a schematic diagram of another embodiment of the array microphone device of this application;

[0066] Figure 7 is a schematic block diagram of another embodiment of the array microphone device of this application;

[0067] Figure 8 is a schematic block diagram of another embodiment of the array microphone device of this application;

[0068] Figure 9 is a schematic block diagram of another embodiment of the array microphone device of this application;

[0069] Figure 10 is a schematic block diagram of another embodiment of the array microphone device of this application;

[0070] Figure 11 is a flowchart illustrating another embodiment of the signal processing method of this application;

[0071] Figure 12 is a schematic block diagram of another embodiment of the array microphone device cascade system of this application. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0073] It should also be noted that, in this document, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0074] This application proposes a cascaded array microphone system that can be used for local or remote conferencing. For example, when used for remote conferencing, a cascaded array microphone system of this application is set up in the local conferencing location, while a cascaded array microphone system of this application or a single array microphone device is set up in the remote conferencing location. The cascaded array microphone system in the local conferencing location sends uplink audio signals to the cascaded array microphone system or the single array microphone device in the remote conferencing location via an uplink; the cascaded array microphone system or the single array microphone device in the remote conferencing location sends downlink audio signals to the cascaded array microphone system in the local conferencing location via a downlink.

[0075] As shown in Figure 1, an embodiment of this application provides a cascaded array microphone device system, including a master device 01 and at least one slave device 02; the master device 01 includes a first master input terminal Zin1, a slave device signal input terminal group Inn, a first master output terminal Zout1, and a second master output terminal Zout2; the slave device includes a first slave input terminal Cin1, a first slave output terminal Cout1, and a second slave output terminal Cout2; wherein, both the master device 01 and the slave device 02 are array microphone devices;

[0076] The first slave output terminal Cout1 and the second slave output terminal Cout2 are connected to the slave device signal input terminal group Inn, the first master output terminal Zout1 is connected to the speaker, and the second master output terminal Zout2 is connected to the remote end; wherein, the first master input terminal Zin1 and the first slave input terminal Cin1 are used to input the speaker reference signal;

[0077] The slave device 02 is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device through the link between the first slave output terminal Cout1 and the second slave output terminal Cout2 of the slave device 02 and the slave device signal input terminal group Inn of the master device 01. The multiple slave microphone signals are acquired by the microphone array inside the slave device. For example, the slave device signal input terminal group Inn may include two slave device signal input terminals, which are respectively connected to the first slave output terminal Cout1 and the second slave output terminal Cout2 of the slave device 02.

[0078] The main device 01 is configured to: perform preset processing on multiple main microphone signals based on a speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device;

[0079] The master device 01 is also configured to generate a playback audio signal and an uplink audio signal based at least on a first master output signal, a second master output signal, a first slave output signal, and a second slave output signal. Exemplarily, the playback audio signal is used for local amplified playback, such as through a speaker; the uplink audio signal is used to transmit to an array microphone device in a remote conference venue.

[0080] This application connects multiple array microphone devices via a master-slave cascading method. Each master and slave device independently performs preset processing (e.g., echo cancellation and / or noise reduction) on the acquired multi-channel microphone signals based on a reference speaker signal, resulting in independent first master output signal, second master output signal, first slave output signal, and second slave output signal. Finally, the master device generates playback audio signals and uplink audio signals based on these signals. This cascading array microphone system extends the coverage area while ensuring isolation between the array microphone devices and maintaining sound reinforcement at the conference venue and high-quality audio at distant locations.

[0081] Figure 2 shows a schematic block diagram of another embodiment of the array microphone device cascade system of this application. The difference from the embodiment shown in Figure 1 is that the master device 01 further includes a second master input terminal Zin2, and the slave device 02 further includes a second slave input terminal Cin2; the second master input terminal Zin2 and the second slave input terminal Cin2 are used to receive downlink audio signals from the remote end;

[0082] Generating a playback audio signal and an uplink audio signal based on at least a first master output signal, a second master output signal, a first slave output signal, and a second slave output signal includes: generating a playback audio signal based on the first master output signal, the first slave output signal, and the downlink audio signal; and generating an uplink audio signal based on the second master output signal and the second slave output signal.

[0083] For example, the first main output signal, the first slave output signal, and the downlink audio signal are mixed to obtain the playback audio signal; the second main output signal and the second slave output signal are mixed to obtain the uplink audio signal.

[0084] Figure 3 shows a schematic block diagram of another embodiment of the array microphone device cascade system of this application. In this embodiment, the master device 01 includes input terminals RX1 to RX6 and output terminals TX1 to TX4. The input terminal RX1 corresponds to the first master input terminal Zin1 in the aforementioned embodiment, RX2 corresponds to the second master input terminal Zin2 in the aforementioned embodiment, the input terminals RX5 and RX6 correspond to the slave device signal input terminal group Inn in the aforementioned embodiment, and the output terminal TX1 corresponds to the first master output terminal Zout1 in the aforementioned embodiment, and the output terminal TX2 corresponds to the second master output terminal Zout2 in the aforementioned embodiment.

[0085] In this embodiment, the slave device 02 includes input terminals RX1' to RX2' and output terminals TX1' to TX2'. The input terminal RX1' corresponds to the first slave input terminal Cin1 in the aforementioned embodiment, the output terminal TX1' corresponds to the first slave output terminal Cout1 in the aforementioned embodiment, and the output terminal TX2' corresponds to the second slave output terminal Cout2 in the aforementioned embodiment.

[0086] Compared with the previous embodiments, the main device 01 in the embodiment shown in FIG3 has additional input terminals RX3 and RX4 and output terminals TX3 and TX4. In some embodiments, the main device 01 further includes a third main output terminal, a third main input terminal, a fourth main output terminal, and a fourth main input terminal; wherein the third main output terminal and the fourth main output terminal correspond to output terminals TX3 and TX4, respectively, and the third main input terminal and the fourth main input terminal correspond to input terminals RX3 and RX4, respectively.

[0087] In some embodiments, generating a playback audio signal based on a first master output signal, a first slave output signal, and a downlink audio signal includes:

[0088] The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal.

[0089] The uplink audio signal is generated based on the second main output signal and the second slave output signal, including:

[0090] The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

[0091] The inventors discovered that the audio data (first slave output signal and second slave output signal) and downlink audio signal sent by device 02 are timestamped, while the audio data (first master output signal and second master output signal) of master device 01 are not transmitted over the network and therefore cannot obtain network timestamps, making data alignment and mixing impossible. To solve this problem, this embodiment performs a loopback between the amplified audio and the uplink audio of the call from master device 01 (i.e., the first master output signal and the second master output signal). The audio from both master and slave devices is synchronously received according to timestamps using Dante, thus achieving alignment. The amplified audio from master device 01 is sent via TX3 and received by RX3, while the uplink audio is sent via TX4 and received by RX4.

[0092] Furthermore, Figure 3 also includes a power amplifier unit, a converter, and a PC. The main device communicates with a remote location via the PC. The power amplifier unit is located between the first main output of the main device 01 and the speaker, pre-amplifying the audio signal input to the speaker. The converter is located between the second main output of the main device 01 and the PC, performing signal conversion to convert Dante to USB signals.

[0093] In the embodiment shown in Figure 3, the master device 01 obtains the local amplified signal 1 and the uplink audio signal 1 based on the downlink audio signal and the speaker reference signal; the slave device 02 receives the downlink audio signal and the speaker reference signal, and outputs the local amplified signal 2 and the uplink audio signal 2. Finally, the master device outputs a mixed signal 1 of the downlink audio signal, the local amplified signal 1, and the local amplified signal 2 through the first main output terminal TX1, and amplifies the sound through the power amplifier unit and the speaker; the master device 01 outputs the mixed signal 2 of the uplink audio signal 1 and the uplink audio signal 2 through the second main output terminal TX2.

[0094] For example, the data flow of the two devices in the embodiment shown in Figure 3 is described as follows:

[0095] 1) Define the master device in the cascaded network by configuring it.

[0096] 2) Downlink audio during the call is input from the PC, forwarded to the master and slave devices via a converter. The ceiling microphone completes this through receiving channel 1.

[0097] 3) Amplified audio: The amplified audio recorded by the master and slave devices is processed by algorithms; the amplified audio from the slave device is sent to the master device via the network.

[0098] 4) Uplink audio: The master and slave devices record the call audio after algorithm processing; the slave device's uplink audio is sent to the master device via the network.

[0099] 5) Reference tone, used for echo cancellation and feedback suppression. The reference tone is the locally played audio transmitted through the speaker. The locally played audio consists of several parts: first, the amplified audio recorded by the master and slave devices and processed by algorithms; second, the downlink audio of the call. These two audio parts are mixed by the master device using a strategy, EQ processing, gain processing, etc., and output to the speaker, the master device, and each slave device.

[0100] 6) The main device can output to digital speakers via Dante network, or it can output analog signals to analog speakers via the DAC on the device.

[0101] 7) For ease of configuration, the following constraints are applied: RX1 is the input for reference audio, and RX2 is the input for downlink audio. TX1 and TX2 are the outputs for amplified sound and uplink audio, respectively. RX5 and RX6 of the master device are connected to the amplified sound and uplink audio of slave device 1, respectively; RX7 and RX8 are connected to the amplified sound and uplink audio of slave device 2, respectively, and so on. Up to 63 ceiling microphones can work collaboratively in scenarios without an audio processor.

[0102] Figure 4 shows a schematic block diagram of another embodiment of the array microphone device cascading system of this application. This embodiment differs from the embodiment shown in Figure 3 in that the embodiment shown in Figure 3 cascades a single slave device with a master device, while the embodiment shown in Figure 4 cascades two slave devices with a master device. In this embodiment, the slave device signal input terminal group Inn of the master device 01 further includes input terminals RX7 and RX8, used to connect the two output signals of the newly added slave device 03. The slave device 03 receives the downlink audio signal and the speaker reference signal, and outputs the local amplification signal 3 and the uplink audio signal 3. Other signal inputs, signal outputs, and signal processing of the slave device 03 are described in the slave device 02 diagram and will not be repeated here. Furthermore, the master device 01 ultimately outputs a mixed signal 1 of the downlink audio signal, local amplification signal 1, local amplification signal 2, and local amplification signal 3 through the first main output terminal TX1, and amplifies the sound through a power amplifier unit and a speaker; the master device 01 outputs a mixed signal 2 of the uplink audio signal 1, uplink audio signal 2, and uplink audio signal 3 through the second main output terminal TX2.

[0103] In some embodiments, the main device 01 performs preset processing on the multiple main microphone signals based at least on the speaker reference signal to obtain a first main output signal and a second main output signal, including:

[0104] At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result;

[0105] At least based on the multi-microphone signals, the sound source is located;

[0106] At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result;

[0107] The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal.

[0108] The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

[0109] Similarly, device 02 can perform signal processing according to the above steps to obtain the first slave output signal and the second slave output signal, which will not be repeated here.

[0110] In some embodiments, sound source localization is performed at least based on the multiple microphone signals, including:

[0111] Sound source localization is performed based on the multi-microphone signals and the downlink audio signal. Instead of directly relying on the multi-microphone signals, sound source localization is performed based on both the multi-microphone signals and the downlink audio signal, thus avoiding the influence of potentially included far-end audio—downlink audio signal—on the accuracy of sound source localization, and consequently ensuring the accuracy of the BF (Best Before Frequency).

[0112] The internal structure of the master and slave devices and the signal processing method of this application are further illustrated below with reference to the relevant embodiments of the array microphone device shown in Figures 5 to 9.

[0113] Figure 5 shows a schematic diagram of an embodiment of the array microphone device of this application. The array microphone device includes multiple microphones and a processor. The processor includes a first input terminal In1, a first output terminal Out1, and a second output terminal Out2. The first output terminal outputs a signal to a local speaker, and the second output terminal outputs a signal to a remote location. The multiple microphones are used to acquire multiple microphone signals. The processor is exemplarily a signal processor. The array microphone device may also include signal acquisition and peripheral circuitry that cooperates with the first input terminal, the first output terminal, and the second output terminal. This application does not limit the specific circuit structure. The signal processing method for the array microphone device provided in this application includes the following steps:

[0114] S10. The processor acquires the speaker reference signal through the communication path between the first output terminal and the first input terminal;

[0115] S20. The processor performs a first preset processing on the multi-microphone signal based at least on the speaker reference signal to obtain a first output signal, and outputs it to the local speaker through the first output terminal.

[0116] S30. The processor performs a second preset processing on the multi-microphone signal based at least on the speaker reference signal to obtain a second output signal, and outputs it to the remote end through the second output terminal.

[0117] This embodiment adopts a single-input channel and dual-output channel design. It uses only the reference input channel formed by the first input terminal to receive the speaker reference signal, thereby realizing independent processing of the local sound reinforcement signal (first output signal) and the far-end signal (second output signal). It can independently realize noise reduction and echo cancellation processing of the two signals. The system hardware design is simple and has good compatibility.

[0118] Figure 6 shows a schematic diagram of another embodiment of the array microphone device of this application. The difference from Figure 5 is that the array microphone device in this embodiment further includes a second input port. Exemplarily, when the array microphone device includes a processor, the processor is provided with a second input port, through which the processor can acquire downlink audio signals from the remote end.

[0119] Further, the processor performs a first preset processing on the multi-microphone signals based at least on the speaker reference signal to obtain a first output signal, including: the processor performs a first preset processing on the multi-microphone signals based at least on the speaker reference signal, the downlink audio signal, and the multi-microphone signals to obtain a first output signal; the processor performs a second preset processing on the multi-microphone signals based at least on the speaker reference signal to obtain a second output signal, including: the processor performs a second preset processing on the multi-microphone signals based at least on the speaker reference signal, the downlink audio signal, and the multi-microphone signals to obtain a second output signal.

[0120] This embodiment implements a dual-channel input and dual-channel output design. Specifically, the single microphone array system has two independent reference input channels (In1 and In2), which respectively receive the speaker reference signal and the remote downlink audio signal, effectively distinguishing between remote speech and local sound reinforcement. Furthermore, this design provides an independent signal processing path for mixed scenarios, eliminating the need for an external DSP and achieving efficient signal processing. Moreover, the two independent output channels process the remote call and local sound reinforcement signals separately, ensuring that the remote call quality and local sound reinforcement effect do not interfere with each other in complex scenarios.

[0121] Figure 7 shows a schematic block diagram of another embodiment of the array microphone device of this application. The portion within the dashed box in Figure 7 corresponds to the portion within the dashed boxes in Figures 5 and 6. Furthermore, in the embodiment shown in Figure 7, the audio signal (i.e., the downlink audio signal) from the remote conferencing device is received through the downlink of the second input port In2, and the audio signal is transmitted to the remote conferencing device through the uplink of the second output port Out2. The embodiment shown in Figure 7 also includes an AEC module, a BF module, an NN_AFC module, a DOA module, and an AES module.

[0122] The AEC (Acoustic Echo Cancellation) module is used to eliminate echoes in multi-microphone signals. When sound from the local speaker is picked up by the local microphone and sent back to the remote end, an echo may occur. The AEC module dynamically estimates the echo path using an adaptive filter and eliminates the speaker echo from the local microphone signal. The AEC module ensures that the local speaker sound does not interfere with the remote call, providing a better call experience.

[0123] The DOA (Direction of Arrival) module is used to calculate the angle of arrival of sound signals. For example, in a local sound reinforcement channel, the DOA module calculates the azimuth of local sound sources (such as speakers), thus providing input for beamforming (BF) algorithms. The DOA module determines the direction of the sound source by analyzing the time and phase differences between different channels of the microphone array.

[0124] For example, by utilizing the signal from sampling channel 2 (corresponding to the downlink), the remote signal and the local sound source can be accurately distinguished, avoiding misjudging the speaker's playback signal as the local speaker's signal, thus enabling more accurate tracking of the actual local speaker. For example, as shown in Figure 7, the DOA module performs sound source localization based on the multi-microphone signals and the downlink audio signal, rather than directly based on the multi-microphone signals. This avoids the influence of the remote audio-downlink audio signal that may be included in the multi-microphone signals on the accuracy of sound source localization, thereby ensuring the accuracy of BF.

[0125] In some embodiments, the second input port In2 can accurately distinguish between the real speaker sound source and the speaker echo signal in a mixed conference scenario. Because in a local sound reinforcement environment, the sound played by the speaker may be picked up by the microphone array, creating false sound source directions, the DOA module may incorrectly locate the echo signal direction if not handled properly. To address this issue, the second input port In2 provides a reference signal (i.e., a downlink audio signal) to indicate whether the speaker is playing distant speech, helping the microphone array determine whether the sound currently received by the microphone array contains echo interference from the speaker. It has at least the following functions:

[0126] • Prevent DOA from misinterpreting echo signals: By introducing the In2 signal, the microphone array can effectively distinguish between the real sound source and the speaker echo, avoiding incorrect positioning of the DOA module due to echo, thereby improving the accuracy of sound source positioning.

[0127] • Improved positioning accuracy: When In2 indicates the presence of a remote downlink signal, the microphone array can automatically eliminate interference from the speaker, ensuring that the DOA algorithm always focuses on the actual local sound source, effectively improving the accuracy of beamforming (BF) and the reliability of subsequent signal processing.

[0128] The BF (Beamforming) module is a beamforming algorithm that uses the azimuth angle estimated by DOA to weight and combine the microphone array signals to form a "beam" that is focused on a sound source in a specific direction. This process can enhance the signal of the target sound source while suppressing noise from other directions, thereby improving the signal quality of local sound reinforcement.

[0129] The NN_AFC (Neural Network Adaptive Feedback Cancellation) module is a feedback cancellation algorithm that incorporates deep learning technology. Compared to traditional AFC algorithms, NN_AFC, through its neural network model, can better capture the complex characteristics of feedback paths and adapt to different environments and device configurations. This model can automatically learn feedback features and adaptively adjust, thus providing more efficient feedback cancellation in complex scenarios.

[0130] The AES (Acoustic Echo Suppression) module further processes the echoes remaining after AEC (Acoustic Echo Cancellation). Although AEC removes most of the echo, some residual echoes may still exist. AES further suppresses these residual echoes through advanced filtering and processing. AES ensures that there is almost no echo interference in long-distance calls, providing a more natural call experience.

[0131] In some embodiments, the processor performs a first preset processing on the multi-microphone signals based at least on the speaker reference signal to obtain a first output signal, including:

[0132] The processor performs a first echo cancellation process on the multi-microphone signals based at least on the speaker reference signal to obtain a first echo cancellation result; wherein the first echo cancellation process can be implemented using an AEC module;

[0133] The processor performs sound source localization based at least on the multi-microphone signals;

[0134] The processor performs beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performs second echo cancellation processing on the beamformed signal to obtain a first output signal. The second echo cancellation processing is (NN_AFC, Neural Network Adaptive Feedback Cancellation), a feedback cancellation algorithm that combines deep learning technology. Compared to the traditional AFC algorithm, NN_AFC, through a neural network model, can better capture the complex characteristics of the feedback path and adapt to different environments and device configurations.

[0135] In some embodiments, the processor performs a second preset processing on the multi-microphone signals based at least on the speaker reference signal to obtain a second output signal, including:

[0136] The processor performs a first echo cancellation process on the multi-microphone signals based at least on the speaker reference signal to obtain a first echo cancellation result; wherein the first echo cancellation process can be implemented using an AEC module;

[0137] The processor performs sound source localization based at least on the multi-microphone signals;

[0138] The processor performs beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performs echo suppression processing on the beamformed signal to obtain the second output signal.

[0139] In some embodiments, the processor performs beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performs echo suppression processing on the beamformed signal, including:

[0140] When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process.

[0141] When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a second echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

[0142] For example, when there is a distant downlink voice signal (i.e., the downlink audio signal indicates the presence of distant voice activity), the suppression intensity of the first echo suppression process (AES) is around 30-40dB to ensure no echo leakage. In amplified mode, if there is no distant voice speaking (i.e., the downlink audio signal indicates no distant voice activity), the suppression intensity of the second echo suppression process should be reduced to a maximum suppression intensity of no more than 5dB, or completely turned off, to ensure the voice quality of the near-end voice in the uplink channel.

[0143] In this embodiment, through a dual-return signal input design, the AES module can distinguish between the far-end speech played by the speaker and the local amplified sound signal in real time. When far-end speech activity is detected in return channel 2 (corresponding to the downlink), the AES module will enhance the echo suppression and reduce residual echo; while when there is no activity in return channel 2 but there is activity in return channel 1, the system recognizes it as a local amplified sound signal, and the AES module will reduce the suppression intensity to ensure the clarity of the local amplified speech.

[0144] Figure 8 shows a schematic block diagram of another embodiment of the array microphone device of this application. The difference from the embodiment in Figure 7 is that the array microphone device of this application is provided with a third output port (Out_DOA_info) for outputting the coordinate information of the sound source. Exemplarily, when the array microphone device includes a processor, the processor is configured with the third output port. Furthermore, the processor performs sound source localization at least based on the multiple microphone signals to determine the coordinate information of the sound source and outputs it to an external device through the third output port.

[0145] By adding a DOA coordinate output channel in this embodiment, the coordinate information of the current speaker can be output in real time, which can help external devices (such as cameras) to automatically and quickly track the speaker in the current meeting, thereby improving the accuracy and efficiency of meeting recording.

[0146] Figure 9 shows a schematic block diagram of another embodiment of the array microphone device of this application. In this embodiment, the device further includes: performing local sound reinforcement preprocessing on the first output signal before outputting the first output signal, the local sound reinforcement preprocessing including at least one of the following: noise suppression (NR, NN_NR), neural network reverberation suppression (NN_DER), automatic gain control (AGC), and frequency equalization processing (EQ); and / or, performing remote signal preprocessing on the second output signal before outputting the second output signal, the remote signal preprocessing including at least one of the following: noise suppression (NR, NN_NR), neural network dynamic reverberation suppression (NN_DER), automatic gain control (AGC), and frequency equalization processing (EQ).

[0147] This embodiment employs end-to-end audio signal processing, achieving comprehensive optimization from noise suppression to automatic gain control. For example, this application integrates a complete audio processing chain, including modules for noise suppression (NR), dereverberation (NN_DER), equalizer (EQ), and automatic gain control (AGC). This end-to-end processing system ensures that the audio signal undergoes multi-level optimization throughout the entire process from input to output.

[0148] The effects include: significantly improved audio clarity: through NR and dereverberation processing, the system can effectively improve audio clarity, especially in noisy or reverberant environments; smooth volume adjustment: the AGC module automatically adjusts the gain to ensure that the audio output is not too high or too low, providing a more natural listening experience.

[0149] Figure 10 shows a schematic block diagram of another embodiment of the array microphone device of this application. In this embodiment, independent signal processing modules are configured for the local sound reinforcement link and the uplink of the remote signal, respectively. Compared with the embodiment in Figure 9, Figure 10 splits the AEC in Figure 9 into an AFC for the local sound reinforcement link and an AEC for the uplink of the remote signal, splits the BF into two independent BFs, and splits the DOA into two independent DOAs.

[0150] It should be noted that the system block diagrams of the array microphone devices shown in Figures 7-10 illustrate an array microphone unit that integrates complete mixed-scene audio processing capabilities. The entire unit within the dashed box not only supports audio acquisition from multi-channel microphones but also includes dual-channel reference audio acquisition, DSP signal processing, sound reinforcement channel output, and uplink call link signal output. The entire system block diagram exhibits high integration, capable of independently handling complex audio scenarios such as remote calls and local sound reinforcement, reducing reliance on external DSP processors.

[0151] The array microphone system of this application supports multiple signal transmission formats. External signal transmission can use either analog audio signals or digital audio protocols (such as Dante and AES67). If analog audio signals are used, the array microphone unit also needs to be equipped with corresponding ADC and DAC modules to complete the analog-to-digital conversion and ensure accurate signal processing.

[0152] Figure 11 shows a flowchart of another embodiment of the signal processing method of this application.

[0153] Step 1: Audio Acquisition

[0154] Array microphone devices acquire local audio signals through microphone arrays. These arrays typically consist of multiple microphones, arranged to capture sound source information from different directions. Each microphone picks up a different audio signal, which is then transmitted to subsequent signal processing modules. With the help of the microphone array, the system can achieve spatial audio perception for subsequent beamforming (BF) and azimuth estimation (DOA).

[0155] Step 1.2: Execute the AFC algorithm

[0156] Adaptive Feedback Cancellation (AFC) algorithms are used to detect and eliminate feedback howling in local sound reinforcement systems. This feedback typically occurs when a microphone picks up sound emitted by a speaker, causing the sound to continuously cycle between the microphone and the speaker, creating howling.

[0157] AFC dynamically adjusts the signal through an adaptive filter to reduce or eliminate feedback signals, so that the system can operate normally under high gain conditions.

[0158] Step 1.3: Perform local sound reinforcement channel DOA1 calculation.

[0159] The DOA (Direction of Arrival) algorithm is used to calculate the angle of arrival of a sound signal. In the local sound reinforcement channel, the DOA1 module calculates the azimuth of local sound sources (such as speakers), thus providing input for the beamforming (BF) algorithm.

[0160] DOA1 determines the direction of the sound source by analyzing the time and phase differences between different channels of the microphone array.

[0161] Step 1.4: Execute the beamforming (BF) algorithm

[0162] Beamforming (BF) algorithms use the azimuth angle estimated by DOA1 to weight and combine the microphone array signals to form a "beam" that is focused on a sound source in a specific direction.

[0163] This process can enhance the signal of the target sound source while suppressing noise from other directions, thereby improving the signal quality of local sound reinforcement.

[0164] Step 1.5: Execute the NN_AFC algorithm

[0165] NN_AFC (Neural Network Adaptive Feedback Cancellation) is a feedback cancellation algorithm that incorporates deep learning techniques. Compared to traditional AFC algorithms, NN_AFC, through its neural network model, can better capture the complex characteristics of feedback paths and adapt to different environments and device configurations.

[0166] This model can automatically learn the features of feedback and make adaptive adjustments, thereby providing a more efficient feedback elimination effect in complex scenarios.

[0167] Step 1.6: Execute the local sound reinforcement processing link (NR, NN_NR, NN_DER, AGC, EQ)

[0168] After beamforming, the local sound reinforcement signal needs to go through a series of processing links to ensure that the output audio signal is of good quality and suitable for use with local sound reinforcement equipment.

[0169] NR (Noise Reduction): Noise reduction is applied to the signal captured by the microphone to reduce interference from ambient noise.

[0170] NN_NR (Neural Network Noise Suppression): A noise suppression algorithm based on deep learning that can better separate speech signals from complex noise.

[0171] NN_DER (Neural Network Reverb Suppression): Dynamically suppresses reverberation tails through a neural network model to ensure clear sound reinforcement signals.

[0172] AGC (Automatic Gain Control): Automatically adjusts the gain of the audio signal to ensure volume balance and prevent the signal from being too loud or too soft.

[0173] EQ (Equalization): Equalizes the frequency of an audio signal, enhancing certain frequency bands or attenuating others to ensure the sound quality of the amplified signal.

[0174] Step 1.7, Local Amplification Output

[0175] The local sound reinforcement signal, after passing through all the processing links, is finally transmitted to the local sound reinforcement speaker for playback.

[0176] The signal should have high-quality sound clarity, low noise, and no feedback to ensure that users hear clear and natural sound in their local environment.

[0177] Step 2.2: Execute the AEC algorithm

[0178] AEC (Acoustic Echo Cancellation) algorithm is used to eliminate echoes in far-end calls. Echoes can occur when sound emitted from the local speaker is picked up by the local microphone and sent back to the far end. AEC dynamically estimates the echo path using an adaptive filter and removes the speaker echo from the local microphone signal.

[0179] AEC ensures that local speakerphone sound does not interfere with remote calls, providing a better call experience.

[0180] Step 2.3: Perform remote call channel DOA2 calculation

[0181] DOA2 (Distant Azimuth Estimation) is used to calculate the azimuth of the sound source in a distant call. By analyzing the angle of arrival of the distant signal, the system can help beamforming algorithms enhance the sound in the direction of the distant call.

[0182] DOA2 provides an important input for beamforming in remote calls, ensuring clear audio transmission from a distance.

[0183] Step 2.4: Execute the beamforming (BF) algorithm

[0184] Based on the far-end azimuth angle estimated by DOA2, a beamforming (BF) algorithm is used to directionally enhance the far-end call signal. By focusing on the far-end sound source, the system can suppress other interference signals and ensure that the far-end voice signal is clearly audible.

[0185] This beamforming technology can significantly improve the signal-to-noise ratio (SNR) of remote calls, ensuring the quality of voice transmission.

[0186] Step 2.5: Execute the AES algorithm and perform residual echo post-processing.

[0187] The AES (Acoustic Echo Suppression) algorithm further processes the echoes remaining after AEC. Although AEC has removed most of the echoes, some residual echoes may still exist. AES further suppresses these residual echoes through advanced filtering and processing.

[0188] AES ensures virtually no echo interference during remote calls, providing a more natural call experience.

[0189] Step 2.6: Execute the remote call processing link (NR, NN_NR, NN_DER, AGC, EQ)

[0190] The remote call signal, after beamforming and AES processing, still needs to go through multiple processing modules to ensure the voice quality of the remote output.

[0191] NR (Noise Reduction): Reduces background noise in remote call signals.

[0192] NN_NR (Neural Network Noise Suppression): Uses deep learning algorithms to separate speech signals from complex noise, further improving speech clarity.

[0193] NN_DER (Neural Network Dynamic Echo Suppression): Through a neural network model, it dynamically suppresses reverberation tails to ensure that the call signal sent to the remote end is clear and free of reverberation tails.

[0194] AGC (Automatic Gain Control): Automatically adjusts the gain of the remote call signal to ensure balanced voice volume.

[0195] EQ (Equalization): Performs frequency equalization processing on the remote call signal to ensure that the output signal's sound quality is suitable for the remote device.

[0196] Step 2.7, Remote signal output

[0197] The processed signal from the remote end is finally sent to the remote communication device (such as a telephone or video conferencing system). This signal has undergone echo cancellation, beamforming, and noise suppression to ensure clear voice signals free from echo and noise interference. The remote user will receive high-quality audio, resulting in a better calling experience.

[0198] Figure 12 shows a schematic block diagram of another embodiment of the array microphone device cascade system of this application. The master and slave devices in this embodiment are based on the array microphone device shown in Figure 10.

[0199] Secondly, this application also provides a method for cascading array microphone devices, applied to an array microphone cascading system, the array microphone cascading system including a master device and at least one slave device; the master device includes a first master input terminal, a slave device signal input terminal group, a first master output terminal, and a second master output terminal; the slave device includes a first slave input terminal, a first slave output terminal, and a second slave output terminal; the method includes:

[0200] The first slave output terminal and the second slave output terminal are connected to the slave device signal input terminal group, the first master output terminal is connected to the speaker, and the second master output terminal is connected to the remote end; wherein, the first master input terminal and the first slave input terminal are used to input the speaker reference signal;

[0201] The slave device is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device, wherein the multiple slave microphone signals are acquired by a microphone array inside the slave device;

[0202] The main device is configured to: perform preset processing on multiple main microphone signals based on the speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device;

[0203] The master device is also configured to generate playback audio signals and uplink audio signals based at least on the first master output signal, the second master output signal, the first slave output signal, and the second slave output signal.

[0204] In some embodiments, the master device further includes a second master input terminal, and the slave device further includes a second slave input terminal; the second master input terminal and the second slave input terminal are used to receive downlink audio signals from the remote end;

[0205] Generating playback audio signals and uplink audio signals based at least on the first main output signal, the second main output signal, the first slave output signal, and the second slave output signal includes:

[0206] A playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal;

[0207] An uplink audio signal is generated based on the second master output signal and the second slave output signal.

[0208] In some embodiments, the master device further includes a third master output terminal, a third master input terminal, a fourth master output terminal, and a fourth master input terminal;

[0209] Generate a playback audio signal based on the first main output signal, the first slave output signal, and the downlink audio signal, including:

[0210] The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal.

[0211] The uplink audio signal is generated based on the second main output signal and the second slave output signal, including:

[0212] The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

[0213] In some embodiments, at least based on the speaker reference signal, the multi-channel main microphone signals are pre-processed to obtain a first main output signal and a second main output signal, including:

[0214] At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result;

[0215] At least based on the multi-microphone signals, the sound source is located;

[0216] At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result;

[0217] The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal.

[0218] The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

[0219] In some embodiments, sound source localization is performed at least based on the multiple microphone signals, including:

[0220] Sound source localization is performed based on the multi-microphone signals and the downlink audio signal.

[0221] In some embodiments, the step of performing beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performing echo suppression processing on the beamformed signal, includes:

[0222] When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process.

[0223] When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a first echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

[0224] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0225] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0226] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A cascaded array microphone system, characterized in that, It includes a master device and at least one slave device; the master device includes a first master input terminal, a slave device signal input terminal group, a first master output terminal, and a second master output terminal; the slave device includes a first slave input terminal, a first slave output terminal, and a second slave output terminal; wherein, The first slave output terminal and the second slave output terminal are connected to the slave device signal input terminal group, the first main output terminal is connected to the speaker, and the second main output terminal is connected to the remote end; The first main input terminal and the first slave input terminal are used to input the speaker reference signal; The slave device is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device; the multiple slave microphone signals are acquired by a microphone array inside the slave device. The main device is configured to: perform preset processing on multiple main microphone signals based on the speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device; The master device is also configured to generate playback audio signals and uplink audio signals based at least on the first master output signal, the second master output signal, the first slave output signal, and the second slave output signal.

2. The system according to claim 1, characterized in that, The master device further includes a second master input terminal, and the slave device further includes a second slave input terminal; the second master input terminal and the second slave input terminal are used to receive downlink audio signals from the remote end; Generating playback audio signals and uplink audio signals based at least on the first main output signal, the second main output signal, the first slave output signal, and the second slave output signal includes: A playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal; An uplink audio signal is generated based on the second master output signal and the second slave output signal.

3. The system according to claim 2, characterized in that, The main device also includes a third main output terminal, a third main input terminal, a fourth main output terminal, and a fourth main input terminal; Generate a playback audio signal based on the first main output signal, the first slave output signal, and the downlink audio signal, including: The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal. The uplink audio signal is generated based on the second main output signal and the second slave output signal, including: The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

4. The system according to any one of claims 1-3, characterized in that, At least based on the speaker reference signal, the multi-channel main microphone signals are pre-processed to obtain a first main output signal and a second main output signal, including: At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result; At least based on the multi-microphone signals, the sound source is located; At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result; The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal. The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

5. The system according to claim 4, characterized in that, The sound source localization based at least on the multi-microphone signals includes: Sound source localization is performed based on the multi-microphone signals and the downlink audio signal.

6. The system according to claim 4, characterized in that, The step of performing beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performing echo suppression processing on the beamformed signal, includes: When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process. When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a second echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

7. A method for cascading array microphone devices, applied to an array microphone cascading system, characterized in that, The array microphone cascade system includes a master device and at least one slave device; the master device includes a first master input terminal, a slave device signal input terminal group, a first master output terminal, and a second master output terminal; the slave device includes a first slave input terminal, a first slave output terminal, and a second slave output terminal; the method includes: The first slave output terminal and the second slave output terminal are connected to the slave device signal input terminal group, the first master output terminal is connected to the speaker, and the second master output terminal is connected to the remote end; wherein, the first master input terminal and the first slave input terminal are used to input the speaker reference signal; The slave device is configured to: perform preset processing on multiple slave microphone signals based at least on the speaker reference signal to obtain a first slave output signal and a second slave output signal, and send them to the master device, wherein the multiple slave microphone signals are acquired by a microphone array inside the slave device; The main device is configured to: perform preset processing on multiple main microphone signals based on the speaker reference signal to obtain a first main output signal and a second main output signal, wherein the multiple main microphone signals are acquired by a microphone array inside the main device; The master device is also configured to generate playback audio signals and uplink audio signals based at least on the first master output signal, the second master output signal, the first slave output signal, and the second slave output signal.

8. The method according to claim 7, characterized in that, The master device further includes a second master input terminal, and the slave device further includes a second slave input terminal; the second master input terminal and the second slave input terminal are used to receive downlink audio signals from the remote end; Generating playback audio signals and uplink audio signals based at least on the first main output signal, the second main output signal, the first slave output signal, and the second slave output signal includes: A playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal; An uplink audio signal is generated based on the second master output signal and the second slave output signal.

9. The method according to claim 8, characterized in that, The main device also includes a third main output terminal, a third main input terminal, a fourth main output terminal, and a fourth main input terminal; Generate a playback audio signal based on the first main output signal, the first slave output signal, and the downlink audio signal, including: The first main output signal is output from the third main output terminal to the third main input terminal, and a playback audio signal is generated based on the first main output signal, the first slave output signal, and the downlink audio signal obtained from the third main input terminal. The uplink audio signal is generated based on the second main output signal and the second slave output signal, including: The second main output signal is output from the fourth main output terminal to the fourth main input terminal, and an uplink audio signal is generated based on the second main output signal and the second slave output signal obtained from the fourth main input terminal.

10. The method according to any one of claims 7-9, characterized in that, At least based on the speaker reference signal, the multi-channel main microphone signals are pre-processed to obtain a first main output signal and a second main output signal, including: At least the multi-microphone signals are subjected to a first echo cancellation process based on the loudspeaker reference signal to obtain a first echo cancellation result; At least based on the multi-microphone signals, the sound source is located; At least based on the sound source localization result, beamforming processing is performed on the first echo cancellation result; The first main output signal is obtained by performing a second echo cancellation process on the beamformed signal. The second main output signal is obtained by performing echo suppression processing on the signal after beamforming.

11. The method according to claim 10, characterized in that, The sound source localization based at least on the multi-microphone signals includes: Sound source localization is performed based on the multi-microphone signals and the downlink audio signal.

12. The method according to claim 10, characterized in that, The step of performing beamforming processing on the first echo cancellation result based at least on the sound source localization result, and performing echo suppression processing on the beamformed signal, includes: When the downlink audio signal indicates the presence of remote voice activity, the signal after beamforming is subjected to a first echo suppression process. When the downlink audio signal indicates that there is no far-end voice activity, the beamformed signal is subjected to a first echo suppression process; wherein the intensity of the first echo suppression is greater than the intensity of the second echo suppression.

Citation Information

Patent Citations

  • Audio signal processing method and device, reverberation detection method and device, conference method and device and storage medium

    CN114143668A

  • Audio signal processing method and system for suppressing echoes

    CN114697785A

  • Echo cancellation method, device and equipment

    CN115512713A

  • Recovery method and system for omnidirectional array microphone and electronic equipment

    CN118540616A

  • Gain and spectral shape adjustment in audio signal processing

    US20090225980A1