Audio signal processing method for distributed multi-microphone sound reinforcement system multi-person conference scenario

By analyzing audio parameters of a distributed multi-microphone amplification system, the system automatically identifies sound status and implements processing strategies, solving the noise problem caused by improper microphone management in multi-person conferences. This achieves clear amplification and whisper suppression, adapting to the dynamic needs of complex conference scenarios.

CN120980430BActive Publication Date: 2026-02-10SHENZHEN POROS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511517991.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-10
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In traditional multi-person meeting scenarios, improper microphone management can lead to noise accumulation or inability to capture sound, frequently interrupting the meeting discussion and reducing communication efficiency.

Method used

By analyzing audio parameters (signal energy intensity and human voice characteristics), the system automatically identifies the sound state and determines the processing strategy for different audio signals. This ensures clear amplification of effective speech and suppresses or weakens invalid or whispered signals. The system employs a distributed multi-microphone amplification system, including a pickup module, amplification module, networking module, and audio processor, to achieve audio signal processing without human intervention.

Benefits of technology

It enables rapid and accurate dynamic optimization of audio signals when multiple people are speaking alternately, adapting to the needs of complex meeting scenarios, ensuring clear amplification of effective speech, suppressing invalid or whispered signals, and avoiding the superposition of noise from multiple microphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980430B_ABST
    Figure CN120980430B_ABST
Patent Text Reader

Abstract

The application provides an audio signal processing method for a multi-person conference scene of a distributed multi-microphone sound amplification system. The distributed multi-microphone sound amplification system comprises a sound pickup module, a sound amplification module, a networking module and an audio processor. The sound pickup module comprises at least two microphones arranged independently, and the microphones correspond to speakers one by one. The sound amplification module comprises at least one loudspeaker. The at least two microphones and the at least one loudspeaker are in communication connection with the networking module and the audio processor. The distributed multi-microphone sound amplification system automatically identifies a sound state through audio parameter analysis (signal energy intensity + human voice feature information), determines a processing strategy for different audio signals according to the sound state, ensures clear sound amplification of effective speech and suppression and weakening of invalid or whisper signals, and avoids multi-microphone noise superposition without manual intervention. Even if multiple people speak alternately, the dynamic optimization of the audio signal can be quickly and accurately completed, and the complex conference scene requirements can be adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of amplification equipment technology, and in particular relates to an audio signal processing method for a distributed multi-microphone amplification system in a multi-person conference scenario. Background Technology

[0002] In traditional multi-person meeting scenarios, each person needs to have a microphone placed in front of them. When one microphone is speaking, the others need to be manually muted. If the speaker forgets to do so, problems such as "multiple microphones being on at the same time causing noise accumulation" or "no microphone being on causing sound not to be captured" can easily occur, frequently interrupting the meeting discussion and reducing communication efficiency. Summary of the Invention

[0003] This application provides an audio signal processing method for multi-person conference scenarios in a distributed multi-microphone amplification system. It automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice characteristic information) and determines different audio signal processing strategies based on the sound state. This ensures clear amplification of valid speech while suppressing or weakening invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people speak alternately, it can quickly and accurately complete the dynamic optimization of audio signals, adapting to the needs of complex conference scenarios.

[0004] In a first aspect, embodiments of this application provide a distributed multi-microphone amplification system, including a microphone pickup module, an amplification module, a networking module, and an audio processor. The microphone pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected to the audio processor via the networking module. A single microphone is used to collect human voice audio signals and send these signals to the audio processor. The audio processor is used to determine the sound state of each human voice audio signal based on the audio parameters of each signal within the at least one human voice audio signal. The audio parameters include signal energy intensity and human voice feature information, and the sound states include speaking state, whispering state, and invalid state; audio processing operations are performed on each of the at least one human voice audio signals according to the sound state of each human voice audio signal to obtain the weight coefficient and signal gain of each human voice audio signal, and the audio signal processing operations include weight coefficient generation operation and signal gain adjustment operation; a mixed audio signal is generated according to the speaking type, weight coefficient, and signal gain of each human voice audio signal in the at least one human voice audio signal, and the mixed audio signal is sent to each of the at least one loudspeakers; a single loudspeaker is used to play the mixed audio signal.

[0005] Secondly, embodiments of this application provide an audio signal processing method for a multi-person conference scenario, applied to a distributed multi-microphone amplification system as described in any one of the first aspects. The distributed multi-microphone amplification system includes a microphone pickup module, an amplification module, a networking module, and an audio processor. The microphone pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected through the networking module and the audio processor. The method includes: receiving at least one human voice audio signal from the microphones; processing the audio signal according to each human voice audio signal in the at least one human voice audio signal... The audio parameters determine the sound state of each human voice audio signal, the audio parameters including signal energy intensity and human voice feature information, and the sound state including speaking state, whispering state, and invalid state; audio processing operations are performed based on the sound state of each human voice audio signal in the at least one human voice audio signal to obtain the weight coefficient and signal gain of each human voice audio signal, the audio signal processing operations including weight coefficient generation operation and signal gain adjustment operation; a mixed audio signal is generated based on the speaking type, weight coefficient, and signal gain of each human voice audio signal in the at least one human voice audio signal, and the mixed audio signal is sent to each of the at least one loudspeakers.

[0006] Thirdly, embodiments of this application provide an audio signal processing device for a multi-person conference scenario, applied to a distributed multi-microphone amplification system as described in any one of the first aspects. The distributed multi-microphone amplification system includes a microphone module, an amplification module, a networking module, and an audio processor. The microphone module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected to the audio processor via the networking module. The device includes: a receiving unit for receiving at least one human voice audio signal from the microphones; and a processing unit for processing each human voice audio signal in the at least one human voice audio signal. The audio parameters of the audio signal determine the sound state of each human voice audio signal. The audio parameters include signal energy intensity and human voice characteristic information. The sound state includes speaking state, whispering state, and inactive state. Audio processing operations are performed based on the sound state of each human voice audio signal in the at least one human voice audio signal to obtain the weighting coefficient and signal gain of each human voice audio signal. The audio signal processing operations include weighting coefficient generation and signal gain adjustment. A mixed audio signal is generated based on the speaking type, weighting coefficient, and signal gain of each human voice audio signal in the at least one human voice audio signal. A transmitting unit is used to transmit the mixed audio signal to each of the at least one loudspeakers.

[0007] Fourthly, embodiments of this application provide a processing chip including a processor and a memory, the memory including one or more programs, the one or more programs being invoked by the processor to execute the step instructions as described in the second aspect.

[0008] Fifthly, embodiments of this application provide an audio processor, including a processor and a memory, the memory including one or more programs, the one or more programs being invoked by the processor to execute the step instructions as described in the second aspect.

[0009] As can be seen from the embodiments of this application, the distributed multi-microphone amplification system includes a pickup module, an amplification module, a networking module, and an audio processor. The pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are all connected to the audio processor via the networking module. A single microphone is used to collect human voice audio signals and send them to the audio processor. The audio processor is used to determine the sound state of each human voice audio signal based on the audio parameters of each signal in the at least one human voice audio signal. The audio parameters include signal energy intensity and human voice characteristic information. The sound state includes speaking state, whispering state, and invalid state. Audio processing operations are performed based on the sound state of each human voice audio signal in the at least one human voice audio signal to obtain the weighting coefficient and signal gain of each signal. The audio signal processing operations include weighting coefficient generation and signal gain adjustment. A mixed audio signal is generated based on the speaking type, weighting coefficient, and signal gain of each human voice audio signal in the at least one human voice audio signal, and the mixed audio signal is sent to each of the at least one megaphone. A single megaphone is used to play the mixed audio signal. As can be seen, in this embodiment, the distributed multi-microphone amplification system automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice feature information), and determines the processing strategy for different audio signals based on the sound state. This ensures clear amplification of effective speech and suppresses or weakens invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people speak alternately, it can quickly and accurately complete the dynamic optimization of audio signals, adapting to the needs of complex meeting scenarios. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1This is a schematic diagram of the structure of a distributed multi-microphone amplification system provided in an embodiment of this application;

[0012] Figure 2 This is a schematic diagram of the microphone structure provided in an embodiment of this application;

[0013] Figure 3 A schematic diagram of a meeting scenario provided in an embodiment of this application;

[0014] Figure 4 A flowchart illustrating an audio signal processing method for a multi-person conferencing scenario provided in an embodiment of this application;

[0015] Figure 5 A functional unit structure block diagram of an audio signal processing device for a multi-person conference scenario provided in an embodiment of this application;

[0016] Figure 6 This is a schematic diagram of the structure of the audio processor provided in an embodiment of this application;

[0017] Figure 7 This is a schematic diagram of the structure of the processing chip provided in an embodiment of this application. Detailed Implementation

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0020] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but in some embodiments includes steps or units not listed, or in some embodiments includes other steps or units inherent to these processes, methods, products, or apparatuses.

[0021] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0022] In the embodiments of this application, "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone; A and B exist simultaneously; B exists alone. Among them, A and B can be singular or plural.

[0023] In this embodiment, the symbol " / " can indicate that the preceding and following objects are in an "or" relationship. Alternatively, the symbol " / " can also represent a division sign, i.e., performing a division operation. For example, A / B can mean A divided by B.

[0024] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.

[0025] In the embodiments of this application, "equal to" can be used with "greater than" and is applicable to technical solutions used when "greater than" is used; it can also be used with "less than" and is applicable to technical solutions used when "less than" is used. When "equal to" is used with "greater than", it is not used with "less than"; when "equal to" is used with "less than", it is not used with "greater than".

[0026] To address the aforementioned technical issues, this application provides an audio signal processing method for multi-person conference scenarios in a distributed multi-microphone amplification system. This method automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice characteristic information) and determines different audio signal processing strategies based on the sound state. This ensures clear amplification of valid speech while suppressing or weakening invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people are speaking alternately, it can quickly and accurately perform dynamic optimization of the audio signal, adapting to the needs of complex conference scenarios.

[0027] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0028] Please see Figure 1 , Figure 2 , Figure 1 This is a schematic diagram of the structure of a distributed multi-microphone amplification system provided in an embodiment of this application. Figure 2 This is a schematic diagram of the microphone structure provided in an embodiment of this application. Figure 1 As shown, the distributed multi-microphone amplification system includes a pickup module, an amplification module, a networking module, and an audio processor. The pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected to the audio processor via the networking module. Each microphone is used to collect human voice audio signals and send them to the audio processor. The audio processor is used to determine the sound state of each human voice audio signal based on audio parameters of each of the at least one human voice audio signals. The data includes signal energy intensity and human voice feature information, and the sound state includes speaking state, whispering state, and invalid state; audio processing operations are performed on each of the at least one human voice audio signals according to the sound state of each human voice audio signal to obtain the weight coefficient and signal gain of each human voice audio signal, and the audio signal processing operations include weight coefficient generation operation and signal gain adjustment operation; a mixed audio signal is generated according to the speaking type, weight coefficient, and signal gain of each of the at least one human voice audio signals, and the mixed audio signal is sent to each of the at least one loudspeaker; a single loudspeaker is used to play the mixed audio signal.

[0029] Specifically, the sound pickup module includes at least two independently configured microphones, each corresponding to a speaker, ensuring that the voices of different speakers can be captured independently. For example... Figure 2 As shown, a single microphone integrates at least two micro-microphones spaced apart and a processing chip. The micro-microphones are responsible for acquiring a single audio sub-signal, while the processing chip performs beam calibration, filtering, and merging on multiple audio sub-signals to ultimately generate a high-quality human voice audio signal and transmit it to the audio processor, effectively eliminating sound interference from non-target areas and improving sound pickup accuracy.

[0030] The amplification module receives the mixed audio signal transmitted by the audio processor and converts it into a sound signal that can be clearly perceived by the audience. This module includes at least one amplifier, which must have stable signal reception and linear playback characteristics to ensure no sound distortion during playback. It should also be able to adjust the playback volume according to the needs of the scene, ensuring that listeners in different locations can obtain clear audio information.

[0031] The networking module can connect dispersed microphones and amplifiers to a unified network using WiFi, Ethernet, or dedicated audio transmission protocols, enabling communication between the pickup module, amplifier module, and audio processor. By employing an adapted transmission protocol, the networking module ensures low latency and high stability of audio signals during transmission, avoiding synchronization issues caused by signal interruptions or delays. It also supports multi-device expansion to meet the needs of various scale scenarios. Furthermore, the networking module has a built-in network clock synchronization unit, controlling the clock deviation of each microphone, amplifier, and audio processor to within 1 microsecond. It also features a jitter buffer, the buffer depth of which can be dynamically adjusted according to network latency (range 50-200ms). When network jitter exceeds a threshold, the buffer depth is automatically increased to prevent signal transmission disorder.

[0032] The audio processor, as the core control unit of the system, has comprehensive functions of audio signal analysis, processing, and mixing. First, it receives the human voice audio signal transmitted from the audio module, extracts the signal energy intensity and human voice characteristic information, and determines the sound state. Second, based on the sound state, it performs weight coefficient generation and signal gain adjustment operations to assign appropriate weights and gains to audio signals in different states. Finally, it combines the speech type, weight coefficients, and signal gain to generate a mixed audio signal and transmits it to the amplification module to ensure that the mixed signal highlights effective speech and suppresses invalid interference.

[0033] The aforementioned distributed multi-microphone amplification system ensures clear amplification of effective speech while suppressing or weakening invalid or whispering signals. Please refer to [link / reference]. Figure 3 , Figure 3 A schematic diagram of a meeting scenario provided in an embodiment of this application, such as... Figure 3 As shown, taking a roundtable meeting scenario with 3 participants as an example, each participant corresponds to a microphone of the sound pickup module, the sound amplification module includes a loudspeaker set in the middle of the meeting table, the networking module uses the Ethernet protocol to connect the components, and the audio processor pre-configures the parameters.

[0034] Before the meeting begins, each microphone in the sound pickup module is powered on and enters its first phase. The loudspeaker prompts the speaker to change their speaking angle. The processing chip generates the initial calibration beam range for each participant based on the initial audio sub-signal collected by the miniature microphone and the microphone distance, locking the effective sound pickup area for each speaker and eliminating signal interference outside the effective sound pickup area. During the meeting, when a participant begins to speak, the corresponding microphone collects the human voice audio signal and transmits it to the audio processor. The audio processor detects that the signal energy intensity reaches the speaking energy threshold, and the human voice characteristics match the speaker's voice characteristics, determining it as a speaking state and assigning a first weighting coefficient and a first signal gain. At the same time, if other participants occasionally whisper, the signal collected by the corresponding microphone is determined to be a whispering state, and a second weighting coefficient and a second signal gain are assigned. Slight air conditioning noise signals in the environment are determined to be invalid, and the weighting coefficient is set to 0.

[0035] The audio processor generates a mixed audio signal based on the weighting coefficients and signal gain of each signal, and transmits it to the amplification module through the networking module. After receiving the signal, the amplification module outputs the sound in a linear playback mode, ensuring that all participants can clearly hear the speaker's remarks, while effectively suppressing whispers to avoid interfering with the main speech, thus achieving orderly amplification of the conference audio.

[0036] When no one is speaking, all audio signals are deemed invalid, the audio processor generates an empty mixed audio signal, and the amplification module outputs no sound to avoid noise affecting the meeting order.

[0037] As can be seen from the embodiments of this application, the distributed multi-microphone amplification system includes a pickup module, an amplification module, a networking module, and an audio processor. The pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are all connected to the audio processor via the networking module. A single microphone is used to collect human voice audio signals and send them to the audio processor. The audio processor is used to determine the sound state of each human voice audio signal based on the audio parameters of each signal in the at least one human voice audio signal. The audio parameters include signal energy intensity and human voice characteristic information. The sound state includes speaking state, whispering state, and invalid state. Audio processing operations are performed based on the sound state of each human voice audio signal in the at least one human voice audio signal to obtain the weighting coefficient and signal gain of each signal. The audio signal processing operations include weighting coefficient generation and signal gain adjustment. A mixed audio signal is generated based on the speaking type, weighting coefficient, and signal gain of each human voice audio signal in the at least one human voice audio signal, and the mixed audio signal is sent to each of the at least one megaphone. A single megaphone is used to play the mixed audio signal. As can be seen, in this embodiment, the distributed multi-microphone amplification system automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice feature information), and determines the processing strategy for different audio signals based on the sound state. This ensures clear amplification of effective speech and suppresses or weakens invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people speak alternately, it can quickly and accurately complete the dynamic optimization of audio signals, adapting to the needs of complex meeting scenarios.

[0038] In some embodiments, the single microphone includes at least two integrated miniature microphones and a processing chip, with a spacing between the at least two miniature microphones. Each miniature microphone is used to acquire a single audio sub-signal and send the audio sub-signal to the processing chip. Before sending the human voice audio signal to the audio processor, the processing chip is specifically configured to: generate an initial calibration beam range for at least two first audio sub-signals within a first time period, based on the at least two first audio sub-signals and the microphone spacing between the at least two miniature microphones; the first time period is a preset time period after the microphone is powered on, and the initial calibration beam range is the effective pickup range of the current speaker. For at least two second audio sub-signals in other time periods after the first time period, a current beam range is generated based on the at least two audio sub-signals and the microphone spacing; the current beam range is filtered based on the initial calibration beam range; if the current beam range falls entirely within the initial calibration beam range, the at least two second audio sub-signals are determined to be the target human voice signal of the current speaker; if the current beam range exceeds the initial calibration beam range, the audio signal within the initial calibration beam range of the at least two second audio sub-signals is determined to be the target human voice signal; the target human voice signal is then merged to obtain the human voice audio signal.

[0039] After the microphone is powered on, it automatically enters a preset "initialization window" (i.e., the first period, usually 3-5 seconds, adjustable depending on the scenario), during which the microphone is in a "pending sound pickup calibration" state. Once a valid initial human voice is detected, the processing chip simultaneously extracts the "first audio sub-signals" from at least two microphones, focusing on analyzing the "time difference" and "phase difference" of the signals. Combined with the known "microphone spacing" (fixed in the hardware design, such as 1.5-2cm), the sound source location is calculated using acoustic localization logic.

[0040] By using the "time difference" when the same voice is received by different microphones, the distance from the speaker to each microphone can be estimated;

[0041] By combining the signal "phase difference" (phase difference caused by different directions of sound propagation), the speaker's azimuth angle relative to the microphone is determined.

[0042] Based on the calculated sound source distance and azimuth angle, the initial calibration beam range is defined: usually with the sound source as the center, a horizontal coverage angle of ±15° and a vertical coverage angle of ±10° are set (which can be finely adjusted according to the scenario). This range is the "effective pickup area of ​​the current speaker", and the parameters of this range (such as angle boundaries and signal gain thresholds) are stored in the chip cache as the basis for subsequent filtering.

[0043] After initialization, the microphone enters "normal pickup mode", with at least two miniature microphones continuously collecting "second audio sub-signals" for subsequent time periods and transmitting them to the processing chip in real time.

[0044] For each newly acquired "second audio sub-signal," the processing chip repeats the initial calibration beam range calculation logic: combining the microphone spacing, it generates the current beam range for the corresponding time period. The current beam range is compared with the initial calibration beam range. If the current beam range falls entirely within the initial calibration beam range, at least two second audio word signals are determined to be the target human voice signal of the current speaker, and the audio word signals corresponding to that range are retained. If the current beam range exceeds the initial calibration beam range, such as due to noise from other directions, the excess signal is attenuated, ultimately filtering out the target human voice signal located within the initial calibration beam range.

[0045] The processing chip synchronously merges at least two target human voice signals after screening. By adjusting the gain and delay of each signal, the effective human voice signals are superimposed and enhanced, while irrelevant noise is canceled out, ultimately generating a clear human voice audio signal, which is then transmitted to the audio processor.

[0046] In addition, the microphone can initiate beam validity detection every preset time interval: analyzing the signal-to-noise ratio (SNR) of the current merged human voice audio signal; if the SNR is consistently higher than a threshold (e.g., 15dB), it indicates that the current beam range still matches the speaker's position and no adjustment is needed; if the SNR is lower than the threshold (e.g., below 10dB for one consecutive second), or a significant drop in effective human voice energy is detected, it is determined that "the speaker's position may have shifted," triggering "recalibration." During recalibration, the microphone temporarily expands the sound pickup detection range (e.g., from ±15° to ±30°) to capture new effective human voices; repeating the logic of the above embodiment, the new sound source position is calculated, the "effective pickup area" parameter is updated, and the original buffered beam range is replaced to ensure that subsequent sound pickup remains focused on the speaker.

[0047] As can be seen, in this embodiment, by detecting voice activity during the initialization phase, effective initial human voice can be accurately captured. Combined with the signal time difference and phase difference of multiple miniature microphones, the specific location of the sound source is calculated, thereby defining a dedicated initial directional beam range. This range focuses only on the current speaker. During subsequent regular sound pickup, all signals outside this range (such as ambient noise or irrelevant sounds from other directions) are attenuated, ensuring that sound pickup always revolves around the target sound source. This significantly reduces interference from irrelevant sounds on the target human voice, guaranteeing its purity. Simultaneously, the dynamic adjustment mechanism can handle minor changes in the speaker's position. When insufficient beam effectiveness is detected, the range is automatically expanded for recalibration, or manual adjustment is supported, preventing sound pickup interruptions or quality degradation due to slight changes in the speaker's position, thus ensuring continuous sound pickup.

[0048] In some embodiments, in determining the sound state of each voice audio signal based on audio parameters of each voice audio signal in at least one voice audio signal, the audio processor is specifically configured to: detect that the signal energy intensity is greater than or equal to a whisper energy threshold and less than a speech energy threshold, and that the voice feature information matches the whisper voice feature information, and determine that the sound state is the whisper state; detect that the signal energy intensity is greater than or equal to the speech energy threshold, and that the voice feature information matches the speaker voice feature information, and determine that the sound state is the speech state; detect that the signal energy intensity is less than the whisper energy threshold, or detect that the voice feature information does not match the whisper voice feature information and the speaker voice feature information, and determine that the sound state is an invalid state; wherein the voice feature information includes a fundamental frequency, harmonic richness, and high-frequency proportion, the whisper voice feature information includes a first range of values ​​for the fundamental frequency, harmonic richness, and high-frequency proportion in the whisper state, and the speaker voice feature information includes a second range of values ​​for the fundamental frequency, harmonic richness, and high-frequency proportion in the speech state.

[0049] The audio parameters include signal energy intensity and vocal characteristic information. Signal energy intensity is calculated by the energy detection module built into the audio processor. This module samples the audio signal in real time, converting the sampled values ​​into energy values ​​to quantify the loudness of the sound. Vocal characteristic information includes the fundamental frequency, harmonic richness, and high-frequency proportion. The fundamental frequency is extracted by the frequency analysis module, which uses a Fourier transform algorithm to decompose the audio signal and find the fundamental frequency. Harmonic richness is obtained by calculating the number and amplitude of frequency components that are integer multiples of the fundamental frequency. The high-frequency proportion is obtained by statistically analyzing the proportion of high-frequency signal energy in the total signal energy.

[0050] The whispering energy threshold and the speech energy threshold were determined through extensive experimental data, referencing the sound intensity range of human whispering and normal speech in different scenarios. A certain interval is maintained between the two thresholds to avoid misjudgment caused by fluctuations in signal energy intensity, ensuring that whispering and speech can still be accurately distinguished even when the signal intensity is at the critical value.

[0051] The physical characteristics of the human voice differ significantly under different sound conditions. When whispering, the vocal cords vibrate with small and irregular amplitude, resulting in a fundamental frequency slightly higher than normal speech. The low harmonic content and high frequency composition, coupled with the high frequency components generated by airflow friction, contribute to the high proportion of high frequencies. During normal speech, the vocal cords vibrate regularly with large amplitude, the fundamental frequency is within a stable normal range, the harmonic content is rich, and the high frequency proportion is low. Environmental noise lacks a fixed fundamental frequency and irregular harmonic structure, making it distinctly different from human voice characteristics and allowing for effective identification through feature matching.

[0052] Relying solely on signal energy intensity for judgment can lead to misjudgments. For example, the energy intensity of a loud cough may reach the speech energy threshold, but the cough lacks the regular characteristics of a human voice. By combining this with a dual judgment based on vocal characteristic information, such non-human voice signals or abnormal vocal signals can be excluded, ensuring the accuracy of the judgment result.

[0053] Based on the above principles, when determining the sound state, the audio processor compares the signal energy intensity with preset whisper energy thresholds and speech energy thresholds. If the signal energy intensity reaches the whisper energy threshold but does not reach the speech energy threshold, the extracted voice feature information is matched with the whisper voice feature information. The whisper voice feature information includes the value ranges of the fundamental frequency, harmonic richness, and high-frequency proportion in the whisper state. If all three fall within their respective value ranges, the sound state is determined to be a whisper state.

[0054] If the signal energy intensity reaches the speaking energy threshold, the human voice feature information is matched with the speaker's voice feature information. The speaker's voice feature information includes the range of values ​​for the fundamental frequency, harmonic richness, and high-frequency proportion during the speaking state. If all three meet the value requirements, the sound state is determined to be a speaking state.

[0055] If the signal energy intensity does not reach the whisper energy threshold, it is directly determined to be invalid. If the signal energy intensity reaches the whisper energy threshold or the speech energy threshold, but the feature matching module detects that the voice feature information does not match the whisper voice feature information or the speaker voice feature information, it is also determined to be invalid.

[0056] As can be seen, in this embodiment, the dual judgment criteria can effectively distinguish between speaking, whispering and invalid signals, avoiding misjudging non-speaking sounds as valid speaking or valid whispering as invalid signals, thus ensuring the accuracy of the processed objects; the clear sound state classification can accurately guide the subsequent weight coefficient generation and signal gain adjustment operations, allowing audio signals in different states to be processed appropriately, ensuring the final amplification effect.

[0057] In some embodiments, the weighting coefficient generation operation includes: obtaining a first quantity of first voice audio signals in the speaking state and a second quantity of second voice audio signals in the whispering state; obtaining a first preset value and a second preset value, wherein the first preset value is the sum of first weighting coefficients of the first voice audio signals and the second preset value is the sum of second weighting coefficients of the second voice audio signals, and the sum of the first preset value and the second preset value is 1; determining a first weighting coefficient of the first voice audio signal based on the first quantity and the first preset value; if the first quantity is 1, the first weighting coefficient is the first preset value; if the first quantity is at least two, obtaining at least two... The first weighting coefficient ratio of the first human voice audio signal is used to calculate the first weighting coefficient of each first human voice audio signal based on the first weighting coefficient ratio and the first preset value; the second weighting coefficient of the second human voice audio signal is determined based on the second quantity and the second preset value; if the second quantity is 1, the second weighting coefficient is the second preset value; if the second quantity is at least two, the second weighting coefficient ratio of at least two second human voice audio signals is obtained, and the second weighting coefficient of each second human voice audio signal is calculated based on the second weighting coefficient ratio and the second preset value; the third weighting coefficient of the invalid third human voice audio signal is determined to be 0.

[0058] The system pre-sets the sum of weight coefficients for all speaking state signals. This value is set according to the importance of each speaking state in the actual scenario, ensuring that speaking signals occupy the majority of the mixed signal. Similarly, it pre-sets the sum of weight coefficients for all whispering state signals. The second preset value is lower than the first preset value, and the third preset value for invalid states is 0. This is because speaking signals are core information in the scenario and need to be prioritized for transmission; therefore, the first preset value is set to a higher value to ensure its largest proportion in the mixed signal. Whispering signals are auxiliary information and do not need to be emphasized, so the second preset value is set to a lower value. Invalid signals have no informational value, and setting the weight coefficient to 0 prevents them from occupying mixed resources and affecting the propagation of valid signals. The sum of the first and second preset values ​​is 1, ensuring a reasonable proportion of each component in the mixed signal, without signal overflow or insufficient proportion.

[0059] In this case, with a total weight set, the weight of each audio signal can be further determined based on the number of audio signals. If there is only one audio signal, the weight coefficient of that audio signal is directly determined as the weight coefficient summation; for example, if there is only one first voice audio signal in the speaking state, the first weight coefficient of the first voice audio signal is determined as the first preset value; if there is only one second voice audio signal in the whispering state, the second weight coefficient of the second voice audio signal is determined as the second preset value.

[0060] When multiple signals in the same state exist simultaneously, weighting coefficients can be assigned according to a preset ratio to prevent any one signal from monopolizing the mixed resources. For example, in a multi-person debate scenario, when multiple speaking signals exist simultaneously, assigning weights according to a ratio ensures that the sound of each speaking signal can be perceived by the audience, ensuring the fairness of information transmission; if only a single signal exists, assigning it all preset values ​​can highlight the sound of that signal, allowing the audience to obtain information more clearly.

[0061] In specific implementation, if the first voice audio signal in the speaking state includes at least two, the audio processor first obtains the weighting coefficient ratio of the at least two first voice audio signals, then establishes an equation based on this ratio and a first preset value, and calculates the first weighting coefficient of each speaking state signal by solving the equation. After the calculation is completed, all first weighting coefficients are summed and verified to ensure that the sum equals the first preset value. Similarly, if the second voice audio signal in the whispering state includes at least two, the second weighting coefficient of each second voice audio signal is obtained based on the weighting coefficient ratio between the at least two second voice audio signals and the second weighting coefficient.

[0062] As can be seen, in this embodiment, by setting the weight relationship between the speaking signal and the whispering signal, multiple speakers can speak simultaneously. The speaking signal has a high proportion, which ensures that the core information is highlighted during the amplification process, and the audience can quickly grasp the key content. The whispering signal has a low proportion, which can retain some background sound, play a certain prompting effect, and avoid the abruptness caused by an overly quiet scene, without interfering with the transmission of core information. The weight coefficients of each speaking signal are reasonably distributed, so that no one speaking signal is suppressed by other signals, ensuring normal communication in multi-person interactive scenarios.

[0063] In some embodiments, regarding the acquisition of the first weighting coefficient ratio relationship of at least two first human voice audio signals, the audio processor is specifically configured to: acquire the permission priority of the microphones corresponding to at least two first human voice audio signals, and determine the first weighting coefficient ratio relationship based on the permission priority; or, acquire the signal energy intensity of at least two first human voice audio signals, and determine the first weighting coefficient ratio relationship based on the ratio relationship of the signal energy intensity.

[0064] When multiple signals have the same audio state, the weighting ratio between them can be determined based on permission priority or signal energy intensity. The permission priority method is designed based on the hierarchical requirements of information transmission in the scenario, with the core logic being "priority transmission of information from important roles." In most multi-speaker scenarios, there is a clear information leader (such as a meeting moderator or speaker), whose information plays a crucial role in advancing the scenario and should occupy a higher proportion in the mixed signal. By presetting permission levels and benchmark values, the importance of roles is transformed into quantifiable weighting ratios, ensuring that the audio signals of key roles are perceived by the audience first, consistent with the information transmission logic in the scenario.

[0065] Signal energy intensity directly reflects sound loudness. In scenarios without clear role hierarchy (such as roundtable discussions or brainstorming), louder sounds usually indicate a stronger willingness to express themselves and a greater need for attention (such as statements of opinion during heated debates). By collecting and smoothing energy intensity in real time, the differences in speakers' willingness to express themselves are converted into weighted ratios, enabling dynamic adjustments to "who speaks louder, whose voice stands out more," adapting to the needs of equal communication scenarios. Simultaneously, the energy threshold limiting mechanism is set based on human auditory comfort and the amplifier's processing capacity, avoiding volume imbalances and equipment overload caused by abnormally high energy signals.

[0066] Method 1: Determine based on permission priority:

[0067] During system initialization, staff log into the audio processor's control interface using the accompanying management software and access the permission configuration module. This module supports custom permission levels, with common levels divided into administrator, speaker, and attendee levels. Different levels have preset priority weight baseline values ​​(e.g., administrator level is 3, speaker level is 2, attendee level is 1). Staff assign a unique permission level to each microphone based on its usage role in the scenario (e.g., meeting host corresponds to administrator level, speaker corresponds to speaker level, and general audience corresponds to attendee level). After assignment, the permission level is bound to the microphone device identifier and stored in the audio processor's local database, forming a permission-weight mapping table.

[0068] When at least two first-person audio signals exist simultaneously, the audio processor activates the permission reading module to retrieve the corresponding permission level and baseline value from the database using the device identifier carried by each signal. Then, it activates the ratio calculation module to compare the baseline values ​​of each signal and generate a first weighting coefficient ratio relationship. For example, if the two first-person audio signals correspond to an administrator-level microphone and a presenter-level microphone, respectively, with baseline values ​​of 3 and 2, the ratio is determined to be 3:2. If there are three first-person audio signals, corresponding to administrator-level, presenter-level, and participant-level, respectively, with baseline values ​​of 3, 2, and 1, the ratio is determined to be 3:2:1. After generating the ratio, the audio processor combines it with a first preset value (e.g., 0.8) and allocates weighting coefficients proportionally: the total number of weighting coefficients is the sum of all baseline values, and the weighting coefficient of a single signal = (the baseline value of that signal / the total number of weighting coefficients) × the first preset value. After allocation, the processor automatically verifies whether the sum of all weighting coefficients equals the first preset value to ensure calculation accuracy.

[0069] Method 2: Determined based on the energy intensity of the first human voice audio signal:

[0070] The audio processor acquires the signal energy intensity value of each first human voice audio signal. To avoid energy intensity fluctuations caused by transient noise (such as sudden coughing or desktop collision), the system is equipped with an energy smoothing module that performs a moving average calculation on the energy intensity values ​​of five consecutive sampling periods to obtain a stable average energy intensity value, which serves as the basis for subsequent ratio calculations.

[0071] When at least two first-person voice audio signals are present simultaneously, the audio processor retrieves the average energy intensity of each signal and directly compares these averages to generate a first weighting coefficient ratio (automatically simplified if the values ​​have a common divisor). For example, if the average energy intensity of the two first-person voice audio signals are 65 and 50 respectively, the ratio simplifies to 13:10; if the average values ​​of the three signals are 70, 60, and 50, the ratio simplifies to 7:6:5. To prevent imbalances caused by abnormally high energy intensity of individual signals (e.g., exceeding twice the normal speaking intensity), the system sets an energy threshold limit module, preset an upper limit threshold for energy intensity (referring to the maximum intensity of normal human speech). If the average energy intensity of a signal exceeds the threshold, it is automatically adjusted to the upper limit before participating in the ratio calculation. After generating the ratio, weighting coefficients are allocated proportionally based on the first preset value, calculated in the same way as the permission priority method. After allocation, the sum is checked to ensure it equals the first preset value.

[0072] Similarly, when at least two second human voice audio signals exist, the weighting coefficient of each second human voice audio signal can also be obtained through method 1 or method 2 above.

[0073] The two methods can be switched with one click through the management interface. The permission priority method is suitable for scenarios with clear role hierarchies (such as formal meetings and speeches), while the energy intensity method is suitable for scenarios of equal communication (such as group discussions). Different scenario requirements can be met without modifying the hardware.

[0074] As can be seen, in this embodiment, whether based on preset permissions or real-time energy, the ratio relationship is generated through quantitative calculation, avoiding the subjectivity of manual adjustment and ensuring that the weight allocation of each signal is fair and meets the needs of the scenario.

[0075] In some embodiments, the signal gain adjustment operation includes: amplifying a first signal gain of the first human voice audio signal, reducing a second signal gain of the second human voice audio signal, and reducing a third signal gain of the third human voice audio signal to a silence level gain; wherein the first signal gain is less than the distortion level gain.

[0076] During system initialization, staff access the gain configuration module through the management interface and preset gain parameters for three types of signals: the first signal gain (corresponding to the first human voice audio signal) is set to amplification gain (e.g., +6dB), and this value is strictly less than the distortion level gain (the distortion level gain is determined through amplifier hardware parameter testing, typically 90% of the amplifier's maximum undistorted gain; for example, if the amplifier's maximum undistorted gain is +10dB, then the distortion level gain is set to +9dB, and the first signal gain ≤ +8dB); the second signal gain (corresponding to the second human voice audio signal) is set to attenuation gain (e.g., -10dB); and the third signal gain (corresponding to the third human voice audio signal) is set to silence level gain (e.g., -40dB, a value below the minimum volume threshold perceptible to the human ear). These preset parameters are stored in the audio processor's gain parameter library as a reference for real-time adjustment.

[0077] When adjusting the gain, the audio processor determines that a signal is in a speaking state and then jumps to the preset first signal gain value from the gain parameter library to amplify the signal. During amplification, the gain detection module collects the peak amplitude of the signal in real time. If the detected peak amplitude is close to the amplitude threshold corresponding to the distortion level gain (such as the 90% threshold), it automatically reduces the first signal gain (such as from +6dB to +5dB) to ensure that the signal is always in the linear amplification range without the risk of distortion.

[0078] After determining that the signal is a whisper, a preset second signal gain value (e.g., -10dB) is retrieved to attenuate the signal. After attenuation, the volume monitoring module confirms whether the signal volume is lower than the first human voice audio signal (usually 1 / 3 of the first human voice audio signal volume). If it is not lower, the attenuation is further increased (e.g., adjusted to -12dB) to prevent the whisper signal from interfering with the speaking signal.

[0079] After determining that the signal is invalid, a preset silence level gain value (e.g., -40dB) is retrieved, and the signal is deeply attenuated. After attenuation, the silence detection module confirms whether the signal has reached a state imperceptible to the human ear. If it has not reached this state, the attenuation amplitude is further increased (e.g., adjusted to -45dB) to ensure that invalid signals do not participate in the generation of mixed audio.

[0080] As can be seen, in this embodiment, by using three gain adjustment methods—amplifying the spoken signal, attenuating the whisper signal, and silencing the invalid signal—signaling a significant difference in the volume of the three types of signals, listeners can quickly distinguish between core information and interference information, thus improving information acquisition efficiency.

[0081] In some embodiments, in generating a mixed audio signal based on the speech type, weighting coefficient, and signal gain of each of the at least one human voice audio signals, the audio processor is specifically configured to:

[0082] If all of the at least one human voice audio signals are the third human voice audio signals, then the mixed audio signal is determined to be empty;

[0083] If the at least one human voice audio signal includes at least one first human voice audio signal or a second human voice audio signal, then the mixed audio signal is generated according to the weighting coefficient and signal gain of each human voice audio signal.

[0084] In this process, after the audio processor receives the human voice audio signal, judges the sound state, determines the weighting coefficients and signal gain, the next step is to generate a mixed audio signal.

[0085] If at least one human voice audio signal is in an invalid state, the audio processor directly generates an empty mixed audio signal. This empty signal is transmitted to the loudspeaker, which recognizes it as an empty data packet and does not perform playback, thus avoiding noise output. It should be noted that if an empty signal is not generated when all signals are invalid, the loudspeaker may output background noise (such as a hum) due to the lack of valid data, affecting the quietness of the environment. Generating an empty signal triggers the loudspeaker's mute protection mechanism, preventing background noise output and improving scene comfort.

[0086] If at least one human voice audio signal is in a speaking state or a whispering state, the audio processor generates a mixed audio signal according to the following formula (1). The mixed audio signal adopts a linear superposition algorithm to superimpose each effective signal according to the ratio of "weight × gain" to ensure that the signal with high weight and high gain (such as the speaking signal) dominates the mixed signal, while the signal with low weight and low gain (such as the whispering signal) is in an auxiliary position, which meets the requirements of sound hierarchy.

[0087] Mixed audio signal = Σ (single effective audio signal × signal weighting coefficient × signal gain) Formula (1)

[0088] For example, if a scene contains one first human voice audio signal (weighting coefficient 0.8, gain +6dB) and one second human voice audio signal (weighting coefficient 0.2, gain -10dB), then the mixed signal = (first human voice audio signal × 0.8 × (+6dB)) + (second human voice audio signal × 0.2 × (-10dB)). After generating the mixed signal, noise reduction processing (such as eliminating high-frequency background noise) and peak limiting (ensuring that the signal peak does not exceed the maximum input threshold of the amplifier) ​​are performed on the mixed signal. Finally, the optimized mixed signal is encapsulated into a standard playback data packet and transmitted to the amplifier module.

[0089] As can be seen, in the embodiments of this application, an empty signal is generated in the case of invalid signal to completely eliminate background noise and ensure a quiet scene; when there is a speaking or whispering signal, the effective signal is superimposed according to the "weight × gain" ratio to highlight the core information and suppress auxiliary signals, so that the audience can clearly obtain the key content.

[0090] Please see Figure 4 , Figure 4 A flowchart illustrating an audio signal processing method for a multi-person conferencing scenario provided in this application embodiment is shown below. Figure 4 As shown, the method is applied to, for example Figure 1 The distributed multi-microphone amplification system shown includes a microphone pickup module, an amplification module, a networking module, and an audio processor. The microphone pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected through the networking module and the audio processor. The method includes the following steps S401-S404:

[0091] Step S401: Receive at least one human voice audio signal from the microphone.

[0092] Step S402: Determine the sound state of each human voice audio signal based on the audio parameters of each human voice audio signal in the at least one human voice audio signal.

[0093] The audio parameters include signal energy intensity and human voice feature information, and the sound state includes speaking state, whispering state, and invalid state.

[0094] Step S403: Perform audio processing operation based on the sound state of each of the at least one human voice audio signals to obtain the weighting coefficient and signal gain of each human voice audio signal.

[0095] The audio signal processing operations include weight coefficient generation and signal gain adjustment.

[0096] Step S404: Generate a mixed audio signal based on the speaking type, weighting coefficient, and signal gain of each of the at least one human voice audio signals, and send the mixed audio signal to each of the at least one loudspeaker.

[0097] The specific implementation method in this method embodiment is consistent with the above-described distributed multi-microphone amplification system embodiment. For details, please refer to the above embodiments, and this application will not repeat them here.

[0098] As can be seen from the embodiments of this application, the distributed multi-microphone amplification system includes a microphone module, an amplification module, a networking module, and an audio processor. The microphone module includes at least two independently configured microphones, with each microphone corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and at least one megaphone are all connected to the audio processor via the networking module. The distributed multi-microphone amplification system automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice characteristic information) and determines different audio signal processing strategies based on the sound state. This ensures clear amplification of effective speech and suppresses or weakens invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people speak alternately, it can quickly and accurately complete the dynamic optimization of audio signals, adapting to the needs of complex meeting scenarios.

[0099] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the server includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] This application embodiment can divide the server into functional units according to the above method example. For example, each function can be divided into different functional units, or two or more functions can be integrated into one processing module. The integrated unit can be implemented in hardware or as a software program module. It should be noted that the unit division in this application embodiment is illustrative and only represents a logical functional division, while other division methods may be used in actual implementation.

[0101] In the case of using integrated units, please refer to Figure 5 , Figure 5 This application provides a functional unit structure block diagram of an audio signal processing device for a multi-person conferencing scenario, applicable to, for example... Figure 1 The distributed multi-microphone amplification system shown includes a microphone pickup module, an amplification module, a networking module, and an audio processor. The microphone pickup module includes at least two independently configured microphones, each corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and the at least one megaphone are communicatively connected to the audio processor via the networking module. The audio signal processing device 5 includes:

[0102] The receiving unit 501 is configured to receive at least one human voice audio signal from the microphone;

[0103] Processing unit 502 is configured to determine the sound state of each voice audio signal based on audio parameters of each voice audio signal in the at least one voice audio signal, wherein the audio parameters include signal energy intensity and voice feature information, and the sound state includes speaking state, whispering state, and invalid state; perform audio processing operations based on the sound state of each voice audio signal in the at least one voice audio signal to obtain the weighting coefficient and signal gain of each voice audio signal, wherein the audio signal processing operations include weighting coefficient generation operation and signal gain adjustment operation; and generate a mixed audio signal based on the speaking type, weighting coefficient, and signal gain of each voice audio signal in the at least one voice audio signal.

[0104] The transmitting unit 503 is used to transmit the mixed audio signal to each of the at least one loudspeaker.

[0105] As can be seen from the embodiments of this application, the distributed multi-microphone amplification system includes a microphone module, an amplification module, a networking module, and an audio processor. The microphone module includes at least two independently configured microphones, with each microphone corresponding to a speaker. The amplification module includes at least one megaphone. The at least two microphones and at least one megaphone are all connected to the audio processor via the networking module. The distributed multi-microphone amplification system automatically identifies the sound state through audio parameter analysis (signal energy intensity + human voice characteristic information) and determines different audio signal processing strategies based on the sound state. This ensures clear amplification of effective speech and suppresses or weakens invalid or whispered signals, avoiding the superposition of noise from multiple microphones without manual intervention. Even when multiple people speak alternately, it can quickly and accurately complete the dynamic optimization of audio signals, adapting to the needs of complex meeting scenarios.

[0106] This application provides an audio processor, see [link / reference] Figure 6 , Figure 6This is a schematic diagram of the structure of the audio processor provided in the embodiments of this application, as shown below. Figure 6 As shown, the audio processor 6 includes a first processor 61, a first memory 63, a first communication interface 62, and one or more first programs 631. The one or more first programs 631 are stored in the first memory 63 and configured to be executed by the first processor 61. Each of the one or more first programs 631 includes instructions for performing any step in the above method embodiments.

[0107] This application provides a processing chip, see... Figure 7 , Figure 7 This is a schematic diagram of the structure of the processing chip provided in the embodiments of this application, such as... Figure 7 As shown, the processing chip 7 includes a second processor 71, a second memory 73, a second communication interface 72, and one or more second programs 731. The one or more second programs 731 are stored in the second memory 73 and configured to be executed by the second processor 71. The one or more second programs 731 include instructions for performing any step in the above method embodiments.

[0108] This application provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any possible embodiment.

[0109] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0110] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0112] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0114] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0115] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0116] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A distributed multi-microphone amplification system, characterized in that, The system includes a microphone module, a speaker amplifier module, a networking module, and an audio processor. The microphone module includes at least two independently configured microphones, each corresponding to a speaker. The speaker amplifier module includes at least one megaphone. All at least two microphones and the at least one megaphone are communicatively connected to the audio processor via the networking module. A single microphone is used to capture human voice audio signals and send the human voice audio signals to the audio processor; The audio processor is configured to determine the sound state of each voice audio signal based on audio parameters of each voice audio signal in at least one human voice audio signal, wherein the audio parameters include signal energy intensity and human voice feature information, and the sound state includes speaking state, whispering state, and invalid state; perform audio processing operations based on the sound state of each voice audio signal in at least one human voice audio signal to obtain the weighting coefficient and signal gain of each voice audio signal, wherein the audio signal processing operations include weighting coefficient generation operation and signal gain adjustment operation; generate a mixed audio signal based on the speaking type, weighting coefficient, and signal gain of each voice audio signal in at least one human voice audio signal, and send the mixed audio signal to each of the at least one loudspeakers; A single loudspeaker is used to play the mixed audio signal.

2. The system according to claim 1, characterized in that, The single microphone includes at least two integrated miniature microphones and a processing chip, with a spacing between the at least two miniature microphones. The single miniature microphone is used to collect a single audio sub-signal and send the audio sub-signal to the processing chip. Before sending the human voice audio signal to the audio processor, the processing chip is specifically used for: For at least two first audio sub-signals within a first time period, an initial calibration beam range is generated based on the at least two first audio sub-signals and the microphone spacing of the at least two miniature microphones; the first time period is a preset time period after the microphones are powered on, and the initial calibration beam range is the effective pickup range of the current speaker; For at least two second audio sub-signals in other time periods after the first time period, the current beam range is generated based on the at least two audio sub-signals and the microphone spacing; The current beam range is filtered and selected based on the initial calibration beam range; If the current beam range falls entirely within the initial calibration beam range, then the at least two second audio sub-signals are determined to be the target human voice signal of the current speaker; If the current beam range exceeds the initial calibration beam range, then the audio signal within the initial calibration beam range of the at least two second audio sub-signals is determined to be the target human voice signal; The target human voice signal is merged to obtain the human voice audio signal.

3. The system according to claim 1, characterized in that, In determining the sound state of each human voice audio signal based on audio parameters of each human voice audio signal in at least one human voice audio signal, the audio processor is specifically configured to: If the signal energy intensity is detected to be greater than or equal to the whisper energy threshold and less than the speech energy threshold, and the voice feature information matches the whisper voice feature information, then the sound state is determined to be the whisper state. If the signal energy intensity is detected to be greater than or equal to the speech energy threshold, and the human voice feature information matches the speaker's voice feature information, then the sound state is determined to be a speaking state. If the signal energy intensity is detected to be less than the whisper energy threshold, or if the voice feature information does not match the whisper voice feature information and the speaker voice feature information, the sound state is determined to be invalid. The human voice feature information includes fundamental frequency, harmonic richness, and high frequency proportion; the whispering human voice feature information includes a first range of values ​​for the fundamental frequency, harmonic richness, and high frequency proportion in the whispering state; and the speaker's voice feature information includes a second range of values ​​for the fundamental frequency, harmonic richness, and high frequency proportion in the speaking state.

4. The system according to claim 1, characterized in that, The weight coefficient generation operation includes: Obtain a first number of first voice audio signals in the speaking state and a second number of second voice audio signals in the whispering state; Obtain a first preset value and a second preset value, wherein the first preset value is the sum of the first weighting coefficients of the first human voice audio signal, and the second preset value is the sum of the second weighting coefficients of the second human voice audio signal, and the sum of the first preset value and the second preset value is 1. The first weighting coefficient of the first human voice audio signal is determined based on the first quantity and the first preset value. If the first quantity is 1, the first weighting coefficient is the first preset value; If the first quantity is at least two, obtain the first weight coefficient ratio relationship of at least two first human voice audio signals, and calculate the first weight coefficient of each first human voice audio signal according to the first weight coefficient ratio relationship and the first preset value; The second weighting coefficient of the second human voice audio signal is determined based on the second quantity and the second preset value; If the second quantity is 1, the second weighting coefficient is the second preset value; If the second quantity is at least two, obtain the second weighting coefficient ratio relationship of at least two second human voice audio signals, and calculate the second weighting coefficient of each second human voice audio signal according to the second weighting coefficient ratio relationship and the second preset value; The third weighting coefficient of the invalid third human voice audio signal is determined to be 0.

5. The system according to claim 4, characterized in that, In acquiring the first weighting coefficient ratio relationship of at least two first human voice audio signals, the audio processor is specifically configured to: Obtain the permission priority of at least two microphones corresponding to the first human voice audio signals, and determine the first weight coefficient ratio relationship based on the permission priority; or, The signal energy intensities of at least two of the first human voice audio signals are acquired, and the ratio of the first weighting coefficients is determined based on the ratio of the signal energy intensities.

6. The system according to claim 4, characterized in that, The signal gain adjustment operation includes: Amplify the first signal gain of the first human voice audio signal, reduce the second signal gain of the second human voice audio signal, and reduce the third signal gain of the third human voice audio signal to a silence level gain. Wherein, the gain of the first signal is less than the gain of the distortion stage.

7. The system according to claim 6, characterized in that, In generating a mixed audio signal based on the speech type, weighting coefficient, and signal gain of each of the at least one human voice audio signals, the audio processor is specifically configured to: If all of the at least one human voice audio signals are the third human voice audio signals, then the mixed audio signal is determined to be empty; If the at least one human voice audio signal includes at least one first human voice audio signal or a second human voice audio signal, then the mixed audio signal is generated according to the weighting coefficient and signal gain of each human voice audio signal.

8. An audio signal processing method for a multi-person conference scenario, characterized in that, An application is made to a distributed multi-microphone amplification system as described in any one of claims 1-7, the distributed multi-microphone amplification system comprising a microphone pickup module, an amplification module, a networking module, and an audio processor, wherein the microphone pickup module comprises at least two independently configured microphones, each microphone corresponding to a speaker, the amplification module comprises at least one megaphone, and the at least two microphones and the at least one megaphone are communicatively connected through the networking module and the audio processor, the method comprising: Receive at least one human voice audio signal from the microphone; The sound state of each voice audio signal is determined based on the audio parameters of each voice audio signal in the at least one human voice audio signal. The audio parameters include signal energy intensity and human voice feature information. The sound state includes speaking state, whispering state and invalid state. Based on the sound state of each of the at least one human voice audio signals, an audio processing operation is performed to obtain the weighting coefficient and signal gain of each human voice audio signal. The audio signal processing operation includes a weighting coefficient generation operation and a signal gain adjustment operation. A mixed audio signal is generated based on the speaking type, weighting coefficient, and signal gain of each of the at least one human voice audio signals, and the mixed audio signal is sent to each of the at least one loudspeaker.

9. An audio signal processing device for a multi-person conference scenario, characterized in that, An application is made to a distributed multi-microphone amplification system as described in any one of claims 1-7, the distributed multi-microphone amplification system comprising a microphone pickup module, an amplification module, a networking module, and an audio processor, wherein the microphone pickup module comprises at least two independently configured microphones, each microphone corresponding to a speaker, the amplification module comprises at least one megaphone, and the at least two microphones and the at least one megaphone are communicatively connected through the networking module and the audio processor, the device comprising: A receiving unit is configured to receive at least one human voice audio signal from the microphone; The processing unit is configured to determine the sound state of each voice audio signal based on audio parameters of each voice audio signal in the at least one voice audio signal, wherein the audio parameters include signal energy intensity and voice feature information, and the sound state includes speaking state, whispering state, and invalid state; perform audio processing operations based on the sound state of each voice audio signal in the at least one voice audio signal to obtain the weighting coefficient and signal gain of each voice audio signal, wherein the audio signal processing operations include weighting coefficient generation operation and signal gain adjustment operation; and generate a mixed audio signal based on the speaking type, weighting coefficient, and signal gain of each voice audio signal in the at least one voice audio signal. A transmitting unit is configured to transmit the mixed audio signal to each of the at least one loudspeaker.

10. An audio processor, characterized in that, It includes a processor and a memory, the memory including one or more programs that are invoked by the processor to execute the step instructions in the method of claim 8.

Citation Information

Patent Citations

  • Sound mixing processing method of sound amplifying system, sound amplifying system and storage medium

    CN112601158A

  • Adaptive noise reduction conference system based on multi-collar microphone

    CN117479063A