Digital conference audio processing method and system

By identifying and analyzing audio scenarios, content relevance, and speaking duration, combined with device parameters and personnel distribution, the problems of audio chaos and incomplete coverage in digital meetings were solved, achieving orderly audio output and adaptive adjustment, thus improving the auditory experience of participants and meeting efficiency.

CN121692018APending Publication Date: 2026-03-17GUANGZHOU YINQIAO ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address audio chaos and incomplete local audio coverage caused by multiple speakers in digital meetings, and the volume adjustment is not adaptive enough, affecting the auditory experience and efficiency of participants.

Method used

The system identifies valid audio signals using audio acquisition equipment, determines the audio scene, and analyzes the relevance of content and speaking duration when one or more people are speaking, thus filtering out priority audio signals. It also calculates the root mean square amplitude of the audio signals and equipment parameters to determine the coverage area and makes adaptive adjustments based on the distribution characteristics of the participants.

Benefits of technology

It enables orderly audio output in multi-person speaking scenarios, avoids auditory confusion, ensures that all participants are within the audio coverage area, and improves the auditory experience and meeting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121692018A_ABST
    Figure CN121692018A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of conference audio processing, and relates to a digital conference audio processing method and system. According to the invention, the effective audio signal is identified, and the audio scene is determined; recording the effective audio signals of the single-person speech scene as to-be-output audio signals, analyzing the content association degree between the content corresponding to each path of effective audio signals of the multi-person speech scene and the content of the conference theme, obtaining the priority in combination with the speech duration, screening the to-be-output audio signals, calculating the root-mean-square amplitude of the to-be-output audio signals, and outputting the to-be-output audio signals according to the root-mean-square amplitude of the to-be-output audio signals. Determining an overall audio coverage area; judging whether the overall audio coverage area needs to be adjusted according to the distribution characteristics of participants; if so, obtaining to-be-adjusted equipment according to the distribution characteristics of the participants; calculating a pre-adjustment coverage distance of each to-be-adjusted equipment; and the to-be-output audio signal of the to-be-adjusted device is adjusted, so that the voice definition and the conference efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of conference audio processing technology, and specifically to a digital conference audio processing method and system. Background Technology

[0002] In offline digital conference venues, a clear and organized audio experience is crucial for ensuring meeting efficiency. To address the issue of chaotic audio output caused by multiple speakers speaking simultaneously, and the problem of audio gaps for certain individuals due to complex personnel and device distribution scenarios, audio processing is necessary to improve the auditory experience and communication efficiency of participants.

[0003] However, existing technologies have the following problems: 1. When faced with multiple audio streams, traditional audio processing methods often simply reduce noise and output the audio without considering prioritizing each audio stream based on the meeting topic and speaking time, outputting the highest priority audio, avoiding mutual interference from multiple audio streams outputting simultaneously, and reducing the auditory experience of the participants.

[0004] 2. Existing technologies often rely on manual adjustment of the output device to regulate audio output volume. They do not take into account the distribution characteristics of people and the coverage of the device in the audio output scenario to adaptively adjust the audio volume. This avoids the traditional manual adjustment method, making the meeting process more convenient and improving the meeting experience. Summary of the Invention

[0005] The present invention aims to address the shortcomings of the prior art by providing a digital conference audio processing method and system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: one aspect: a digital conference audio processing method, comprising: S1, acquiring conference audio signals through an audio acquisition device, identifying valid audio signals based on the conference audio signals, and determining the audio scene.

[0007] S2. When the audio scenario is a single person speaking, the valid audio signal is recorded as the audio signal to be output.

[0008] S3. When the audio scenario involves multiple speakers, analyze the relevance of the content corresponding to each valid audio signal to the meeting topic, and combine the speaking duration to obtain the priority of each valid audio signal, and then select the audio signal to be output.

[0009] S4. Calculate the root mean square amplitude of the audio signal to be output, and combine it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output by each output device.

[0010] S5. Determine the overall audio coverage range based on the effective coverage range of each output device, and determine whether the overall audio coverage range needs to be adjusted based on the distribution characteristics of the participants.

[0011] S6. If the overall audio coverage needs to be adjusted, the devices to be adjusted are obtained based on the distribution characteristics of the participants, the pre-adjustment coverage distance of each device to be adjusted is calculated, and the output audio signal of each device to be adjusted is adjusted.

[0012] On the other hand, a digital conference audio processing system includes: an audio scene recognition module, a single-person speaking scene analysis module, a multi-person speaking scene analysis module, a coverage analysis module, a coverage adjustment judgment module, and a coverage adjustment module. The connections between these modules are as follows: the audio scene recognition module is connected to both the single-person speaking scene analysis module and the multi-person speaking scene analysis module; the coverage analysis module is connected to both the single-person speaking scene analysis module and the multi-person speaking scene analysis module, as well as the coverage adjustment judgment module; and the coverage adjustment module is connected to the coverage adjustment judgment module.

[0013] The audio scene recognition module collects conference audio signals through audio acquisition devices, identifies valid audio signals based on the conference audio signals, and determines the audio scene.

[0014] The single-person speaking scenario analysis module records the valid audio signal as the audio signal to be output when the audio scenario is a single person speaking.

[0015] The multi-person speaking scenario analysis module analyzes the relevance of the content of each valid audio signal to the meeting topic when the audio scenario involves multiple speakers. It also obtains the priority of each valid audio signal based on the speaking duration and selects the audio signal to be output.

[0016] The coverage analysis module calculates the root mean square amplitude of the audio signal to be output, and combines it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output by each output device.

[0017] The coverage adjustment judgment module determines the overall audio coverage based on the effective coverage of each output device, and judges whether the overall audio coverage needs to be adjusted based on the distribution characteristics of the participants.

[0018] The coverage adjustment module, if the overall audio coverage needs to be adjusted, obtains the devices to be adjusted based on the distribution characteristics of the participants, calculates the pre-adjustment coverage distance of each device, and adjusts the output audio signal of each device.

[0019] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention analyzes the correlation between the content corresponding to each effective audio signal and the content of the meeting theme, and obtains the priority of each effective audio signal by combining the speaking time, and selects the audio signal to be output from it, so as to avoid auditory confusion caused by multiple audio signals being output at the same time, ensure that the meeting proceeds efficiently around the theme, and improve the auditory experience of the participants.

[0020] (2) The present invention calculates the root mean square amplitude of the audio signal to be output and combines it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output of each output device; providing data support for subsequent adjustment and judgment of the audio output device.

[0021] (3) The present invention determines the overall audio coverage range based on the effective coverage range of each output device, and judges whether the overall audio coverage range needs to be adjusted based on the distribution characteristics of the participants, so as to prevent the participants from being in the blind spot of the overall audio coverage range and affecting their listening experience, thereby improving the participants' listening experience and efficiency.

[0022] (4) Based on the distribution characteristics of the participants, obtain the devices to be adjusted, calculate the pre-adjustment coverage distance of each device to be adjusted, adjust the output audio signal of the device to be adjusted, achieve precise adjustment, and ensure that the participants are all within the overall audio coverage range after adjustment, taking into account both meeting efficiency and meeting experience. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the method steps of the present invention.

[0025] Figure 2 This is a schematic diagram illustrating the specific steps involved in establishing content relevance in this invention.

[0026] Figure 3 This is a schematic diagram illustrating the specific steps involved in obtaining the device to be adjusted in this invention.

[0027] Figure 4 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0028] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention. Furthermore, it should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale.

[0029] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use. Techniques, methods, and apparatus known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and apparatus should be considered part of the specification.

[0030] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0031] An embodiment of a digital conference audio processing method: The specific scenario addressed in this embodiment is: a digital conference offline venue, where each participant is equipped with an audio acquisition device, and multiple audio output devices are deployed on-site.

[0032] Please see Figure 1 As shown, the present invention provides a digital conference audio processing method, including: S1, acquiring conference audio signals through an audio acquisition device, identifying valid audio signals based on the conference audio signals, and determining the audio scene.

[0033] Considering that in multi-venue remote collaborative meetings, the audio collected by each acquisition device may contain device noise or non-speaking audio segments, directly outputting the audio without processing will lead to audio noise and affect the meeting effect. Scene recognition can first determine whether each audio signal is a valid human voice signal.

[0034] Furthermore, considering that if multiple people are speaking at the same time, and the output device directly outputs multiple audio channels simultaneously, it will cause the multiple audio channels to be mixed, affecting the accurate reception of information. Therefore, it is necessary to determine the number of audio channels through scene recognition, so as to facilitate further processing to obtain the audio to be output.

[0035] In a specific embodiment of the present invention, the method for determining the audio scene is as follows: First, extract the conference audio signals of each audio acquisition device in the current set time window, perform frame-by-frame processing on each signal to obtain each frame signal, and perform noise reduction processing on each frame signal.

[0036] Specifically, in the embodiment, the frame length is set to 20ms and the frame shift is 10ms during frame processing. Considering the need to meet the real-time requirements of audio processing, setting the processing time window within a very short time can reduce the latency of audio processing. In this invention, the current time window is set to 1s, but the implementer can also set other specific values.

[0037] It should also be noted that the noise reduction process for each frame of signal uses spectral subtraction, which is a fundamental technique in real-time speech processing and a well-known existing technology in the field of speech signal processing.

[0038] Secondly, based on the noise-reduced signals of each frame, the short-time average energy and the frequency band of the spectrum centroid of each frame are obtained, and the short-time average energy and the frequency band of the spectrum centroid are compared with the standard human voice energy threshold and the standard human voice frequency band range, respectively.

[0039] Then, if the short-time average energy of a certain frame exceeds the standard human voice energy threshold and the frequency band of the spectrum centroid is within the standard human voice frequency band, then the frame is recorded as a valid frame.

[0040] The short-time average energy is used to reflect the short-time average intensity of each frame of signal, and the spectral centroid band is used to reflect the dominant frequency of each frame of signal. Both the short-time average energy and the spectral centroid band are commonly used features in existing technologies for human voice recognition, and their acquisition methods are well-known existing technologies in the field of speech signal processing.

[0041] In a preferred embodiment of the invention, the standard human voice energy threshold is a value 10-15 dB higher than the background noise level. The standard human voice energy threshold can be obtained by collecting the short-time average energy of the background noise before the meeting begins. The standard human voice frequency range is 200 Hz-3000 Hz.

[0042] Next, the number of valid frames for each audio acquisition device within the current set time window is counted, and the ratio of this number to the total number of frames within the current set time window is recorded as the valid frame percentage. If the valid frame percentage of the conference audio signal of a certain audio acquisition device is greater than the set percentage threshold, then the conference audio signal of that audio acquisition device is determined to be a valid audio signal.

[0043] The threshold for the percentage is set to 90%. When the percentage of valid frames of an audio signal exceeds 90% within a set time window, the audio signal can be considered a valid audio signal. The implementer can also set other specific values.

[0044] Finally, count the valid audio signals within the current set time window. If there is not a unique valid audio signal, the audio scene is determined to be a multi-person speech; otherwise, the audio scene is determined to be a single-person speech.

[0045] This invention reduces the processing load by filtering valid audio signals and retaining only the speaking signals of participants; by determining the audio scene, it provides a processing basis for the selection of output audio signals for subsequent differentiated processing of audio signals.

[0046] S2. When the audio scenario is a single person speaking, the valid audio signal is recorded as the audio signal to be output.

[0047] S3. When the audio scenario involves multiple speakers, analyze the relevance of the content corresponding to each valid audio signal to the meeting topic, and combine the speaking duration to obtain the priority of each valid audio signal, and then select the audio signal to be output.

[0048] Considering that when multiple people are speaking in the audio scenario, the more relevant the speaker's content is to the meeting topic, the more important the speech is; the earlier the speech starts, the higher the priority of the speech. Therefore, the audio to be output can be selected by analyzing the relevance of the content of each valid audio signal to the meeting topic and the speech duration.

[0049] like Figure 2 As shown, the specific analysis method for the content correlation degree includes: S31, performing speech recognition on each valid audio signal within the current set time window to obtain the text content of each valid audio signal.

[0050] S32. Record the audio acquisition devices of each valid audio signal within the current set time window as speaking devices, trace the historical signals of each speaking device from the current set time window to the historical time, and determine whether there are pauses in the historical signals of each speaking device.

[0051] It should be noted that the pause points can be directly obtained using existing voice endpoint detection technology.

[0052] S33. If there are pauses in the historical signal of a certain speaking device, then the audio signal between the most recent pause and the current time point is used for speech recognition to obtain the historical text content of the speaking device and supplement it into the text content of the valid audio signal of the speaking device.

[0053] S34. Obtain the preset topic content of the meeting from the background meeting database in real time, and retrieve the topic-related words related to the preset topic content from the vocabulary database according to the preset topic content, and form them into a topic-related word set.

[0054] In a preferred embodiment of the present invention, for example, the preset topic of a conference is the planning of intelligent electric vehicle battery technology, and its related keywords include core technical terms: cycle life, safety, energy recovery, solid-state battery, energy density, fast charging, supercharging, thermal management, and related domain terms: electric vehicle, new energy vehicle, driving range, charging pile, carbon neutrality, environmental recycling, etc.

[0055] S35. Count the number of related words in the text content of each valid audio signal that are the same as those in the topic related word set, and record the ratio of the number of related words to the total number of words in the text content as the content relevance.

[0056] In a preferred embodiment of the present invention, the specific method for filtering the audio signals to be output includes: obtaining the time difference between the most recent pause point of each valid audio signal and the end of the current set time window, and recording it as the speaking duration of each valid audio signal.

[0057] The speaking duration of each valid audio signal is normalized, and the weighted sum of the duration and the corresponding content relevance is used to obtain the priority score of each valid audio signal.

[0058] Specifically, the min-max normalization method is used to normalize the speaking duration of each valid audio signal.

[0059] Considering that when multiple people speak at the same time, priority should be given to maintaining the continuity of the speaker's speech, which is in line with the basic etiquette of human dialogue and the principle of meeting efficiency, while the consideration of content relevance is mainly used to suppress speeches that are obviously unrelated to the meeting topic. Therefore, in this invention, the weights of content relevance and speaking time are set to 0.3 and 0.7, respectively. Implementers may also set other weights, but should ensure that the proportion of speaking time is higher than that of content relevance, and that the sum of the two is 1.

[0060] The valid audio signals are sorted by priority score from high to low, and the valid audio signal with the highest priority is recorded as the audio signal to be output.

[0061] This invention identifies audio scenarios. When multiple people are speaking, it analyzes the relevance of the content of each valid audio signal to the meeting topic, and obtains the priority of each valid audio signal based on the speaking duration. The output audio signal is then selected according to the priority of each valid audio signal, avoiding auditory confusion caused by simultaneous output of multiple audio signals, ensuring that the meeting proceeds efficiently around the topic, and improving the auditory experience of the participants.

[0062] S4. Calculate the root mean square amplitude of the audio signal to be output, and combine it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output by each output device.

[0063] In a specific embodiment of the present invention, step S4 includes the following: First, the frame signal amplitude sequence of the audio signal to be output is obtained, and the root mean square amplitude of the audio signal to be output is calculated by performing root mean square calculation on it.

[0064] Secondly, the speaker sensitivity is extracted from the device parameters of each output device, and the pre-calibrated standard reference root-mean-square amplitude of each output device is retrieved from the background conference database. Combined with the root-mean-square amplitude of the audio signal to be output, the initial sound pressure level of the audio output by each output device is obtained through the sound pressure level calculation formula. The initial sound pressure level refers to the sound pressure level of each output device at a reference distance. In this invention, the reference distance is set to 1m, which is the standard definition distance for the initial sound pressure level of the output device.

[0065] The formula for calculating the sound pressure level is: .

[0066] in Represents the initial sound pressure level. This represents the root mean square amplitude of the audio signal to be output. The standard reference root mean square amplitude represents the output device. This represents the speaker sensitivity, where values ​​such as 10 and 20 in the formula are set values ​​in the existing sound pressure level calculation formula.

[0067] in This can be obtained through output device calibration. Specifically, 16-bit audio is output through the output device. By adjusting the output gain of the device, the sound pressure level at a reference distance is measured using a sound pressure meter. When the sound pressure level is 0dB, the root mean square amplitude of the output audio within the current set time window is calculated and recorded as the standard reference root mean square amplitude of the output device.

[0068] Then, the difference between the initial sound pressure level and the minimum value of the speech recognition sound pressure range is obtained, and the effective coverage distance of the output device is calculated by combining them.

[0069] The specific formula for calculating the effective coverage distance is as follows: .

[0070] in , These represent the initial sound pressure level and the minimum speech recognition sound pressure level, respectively. The speech recognition sound pressure level refers to the range of sound pressure levels that participants can normally hear clearly, and in this invention, it is taken as 25 dB. , These represent the effective coverage distance of the output device and the reference distance of the initial sound pressure level, respectively, with the reference distance being 1m.

[0071] This formula is derived from the standard formula for sound source attenuation in a free field: It is derived by working backwards from the distance calculation.

[0072] Finally, the area encompassed by each output device as the center and the effective coverage distance as the radius is recorded as the effective coverage area.

[0073] In a specific embodiment of the present invention, if the output devices are located near the walls of the conference room, the area overlapping the interior of the conference room with the sphere centered on each output device and having an effective coverage distance as its radius is recorded as the effective coverage area.

[0074] This invention calculates the root mean square amplitude of the audio signal to be output and combines it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output from each output device; this provides data support for subsequent adjustment and judgment of the audio output device, and enables the improvement of the auditory effect through adaptive volume adjustment.

[0075] S5. Determine the overall audio coverage range based on the effective coverage range of each output device, and determine whether the overall audio coverage range needs to be adjusted based on the distribution characteristics of the participants.

[0076] The specific method for determining whether the overall audio coverage needs adjustment includes: recording the entire range consisting of the effective coverage ranges of all output devices as the overall audio coverage range, and extracting the location of each participant from the distribution characteristics of the participants.

[0077] Specifically, the union of the effective coverage ranges of all output devices is obtained based on their location in the conference room, and this union is used as the overall audio coverage range.

[0078] If a participant's location is outside the overall audio coverage area, then the overall audio coverage area needs to be adjusted.

[0079] Conversely, if all participants are within the overall audio coverage area, then the overall audio coverage area does not need to be adjusted.

[0080] S6. If the overall audio coverage needs to be adjusted, the devices to be adjusted are obtained based on the distribution characteristics of the participants, the pre-adjustment coverage distance of each device to be adjusted is calculated, and the output audio signal of each device to be adjusted is adjusted.

[0081] like Figure 3 As shown, the specific method for obtaining the device to be adjusted includes: W1, recording each participant who is not within the overall audio coverage area as an audio-uncovered person, and clustering each adjacent audio-uncovered person to form each coverage area.

[0082] The term "adjacent" refers to the people who are next to it in front, behind, left, and right. Specifically, each person whose audio is not covered is grouped into the same cluster as their adjacent uncovered people, and the smallest range that includes all uncovered people in the same cluster is recorded as the range to be covered.

[0083] W2. Obtain the device distance from each area to be covered to each output device, form a sequence of device distances for each area to be covered, and obtain the outlier ranges in the device distance sequence of each area to be covered.

[0084] The device distance refers to the farthest distance from the output device to the boundary of the area to be covered.

[0085] Furthermore, it should be noted that the method for obtaining the outlier range is as follows: calculate the average value of the device distance sequence for each area to be covered. and standard deviation ,Will The maximum value of the outlier range is used to determine the outlier range as... .

[0086] W3. Obtain the device distance sequence of each coverage area and the distance of all devices in the corresponding outlier range. Record the output devices corresponding to all the obtained device distances as the devices to be adjusted in each coverage area.

[0087] By selecting output devices that are close to the target coverage area from the outlier range as the adjustment devices, the target coverage area can be covered with the highest adjustment efficiency.

[0088] Considering that differences in the effective coverage distance of each device to be adjusted should be avoided during the adjustment process of each device in each coverage area, the pre-adjustment coverage distance of each device to be adjusted is kept uniform.

[0089] Based on this, the method for obtaining the pre-adjusted coverage distance includes: first, obtaining the device distances corresponding to each device to be adjusted in each coverage area, and recording the maximum value as the pre-adjusted coverage distance corresponding to each coverage area.

[0090] Setting the pre-adjustment coverage distance for each device to be adjusted within each coverage area according to the corresponding pre-adjustment coverage distance ensures that the overall audio coverage after adjustment covers the coverage area.

[0091] In a specific embodiment of the present invention, the specific steps for adjusting the output audio signal of each device to be adjusted include: calculating the deviation sound pressure level of each device to be adjusted based on the ratio of the pre-adjusted coverage distance to the effective coverage distance of each device corresponding to each coverage area.

[0092] In a specific embodiment of the present invention, the ratio of the pre-adjusted coverage distance to the effective coverage distance of each device to be adjusted is substituted into the deviation sound pressure level calculation formula, which is derived from the standard formula for sound source free field attenuation.

[0093] The formula for calculating the deviation sound pressure level is as follows: .

[0094] The above The deviation sound pressure level represents the difference between the target sound pressure level of the output audio after adjustment and the initial sound pressure level of the output audio before adjustment. It is obtained by substituting the target sound pressure level of the adjusted output audio and the initial sound pressure level of the output audio before adjustment into the standard formula for sound source free field attenuation, and then calculating the difference. This represents the ratio of the pre-adjusted coverage distance to the effective coverage distance for each device to be adjusted.

[0095] Secondly, the target root mean square amplitude is obtained by comprehensively calculating the deviation sound pressure level of each device to be adjusted within each coverage area and the root mean square amplitude of the audio signal to be output.

[0096] The specific relationship between the ground sound pressure level deviation and the root mean square amplitude is expressed as follows: .

[0097] in The root mean square amplitude represents the target amplitude.

[0098] This formula is derived from the sound pressure level calculation formula by reverse calculation, resulting in the expression for the standard reference root mean square amplitude of the output device: .

[0099] Similarly, the formula for the root mean square amplitude of the target can be derived by reverse calculation from the formula for sound pressure level. Substituting the standard reference root-mean-square amplitude formula of the output device into the expression for the target root-mean-square amplitude, the relationship between the output sound pressure level deviation and the root-mean-square amplitude is obtained, where... , These represent the target sound pressure level after adjustment and the initial sound pressure level before adjustment, respectively.

[0100] Then, the ratio of the target root mean square amplitude to the root mean square amplitude of the audio signal to be output is recorded as the amplitude scaling factor.

[0101] Finally, all amplitudes in each frame of the audio signal to be output from each device corresponding to each coverage area are scaled according to the corresponding amplitude scaling factor to obtain the adjusted audio signal to be output from each device, which is then used for output by the corresponding device.

[0102] This invention determines the overall audio coverage range based on the effective coverage range of each output device, and judges whether the overall audio coverage range needs adjustment based on the distribution characteristics of the participants. If adjustment is needed, the device to be adjusted is obtained based on the distribution characteristics of the participants, the pre-adjustment coverage distance of each device to be adjusted is calculated, and the output audio signal of the device to be adjusted is adjusted. This achieves automatic positioning and adjustment target, adaptive and intelligent adjustment, and takes into account both meeting efficiency and meeting experience.

[0103] like Figure 4 As shown, a digital conference audio processing system includes: an audio scene recognition module, a single-person speaking scene analysis module, a multi-person speaking scene analysis module, a coverage analysis module, a coverage adjustment judgment module, and a coverage adjustment module. The connections between these modules are as follows: the audio scene recognition module is connected to both the single-person speaking scene analysis module and the multi-person speaking scene analysis module; the coverage analysis module is connected to both the single-person speaking scene analysis module and the multi-person speaking scene analysis module, as well as the coverage adjustment judgment module; and the coverage adjustment module is connected to the coverage adjustment judgment module.

[0104] The audio scene recognition module collects conference audio signals through audio acquisition devices, identifies valid audio signals based on the conference audio signals, and determines the audio scene.

[0105] The single-person speaking scenario analysis module records the valid audio signal as the audio signal to be output when the audio scenario is a single person speaking.

[0106] The multi-person speaking scenario analysis module analyzes the relevance of the content of each valid audio signal to the meeting topic when the audio scenario involves multiple speakers. It also obtains the priority of each valid audio signal based on the speaking duration and selects the audio signal to be output.

[0107] The coverage analysis module calculates the root mean square amplitude of the audio signal to be output, and combines it with the device parameters of each output device to obtain the initial sound pressure level and effective coverage of the audio output by each output device.

[0108] The coverage adjustment judgment module determines the overall audio coverage based on the effective coverage of each output device, and judges whether the overall audio coverage needs to be adjusted based on the distribution characteristics of the participants.

[0109] The coverage adjustment module, if the overall audio coverage needs to be adjusted, obtains the devices to be adjusted based on the distribution characteristics of the participants, calculates the pre-adjustment coverage distance of each device, and adjusts the output audio signal of each device.

[0110] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0111] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0112] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0114] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method of digital conference audio processing, characterized by, The method comprises the following steps: S1, collecting conference audio signals through an audio collection device, identifying valid audio signals according to the conference audio signals, and determining an audio scene; S2, when the audio scene is single-person speech, the valid audio signals are recorded as to-be-output audio signals; S3, when the audio scene is multi-person speech, the content correlation degrees of the valid audio signals are analyzed, and the priority of each valid audio signal is obtained by combining the speech duration, so as to screen the to-be-output audio signals; S4, the root mean square amplitude of the to-be-output audio signals is calculated, and the initial sound pressure level and the effective coverage range of the audio output by each output device are obtained by combining the device parameters of each output device; S5, the overall audio coverage range is determined according to the effective coverage range of each output device, and whether the overall audio coverage range needs to be adjusted is determined according to the distribution characteristics of the participants; S6, if the overall audio coverage range needs to be adjusted, the to-be-adjusted devices are obtained according to the distribution characteristics of the participants, the pre-adjustment coverage distances of the to-be-adjusted devices are calculated, and the to-be-output audio signals of the to-be-adjusted devices are adjusted.

2. The method of claim 1, wherein, The specific method for determining the audio scene is as follows: Extract the conference audio signals of each audio collection device in the current set time window, and perform frame processing on each to obtain a frame signal, and perform noise reduction processing on each frame signal; According to the frame signals after noise reduction processing, the short-time average energy and the spectrum center frequency band of each frame signal are obtained, and the short-time average energy and the spectrum center frequency band are compared with the standard human voice energy threshold and the standard human voice frequency band range respectively; If the short-time average energy of a frame signal exceeds the standard human voice energy threshold and the spectrum center frequency band is within the standard human voice frequency band range, the frame signal is recorded as a valid frame; The number of valid frames of each audio collection device in the current set time window is counted, and the ratio of the number of valid frames to the total number of frames in the current set time window is recorded as the valid frame ratio, and if the valid frame ratio of the conference audio signals of an audio collection device is greater than a set ratio threshold, the conference audio signals of the audio collection device are determined as valid audio signals; The valid audio signals in the current set time window are counted, and if the valid audio signals are not unique, it is determined that the audio scene is multi-person speech, otherwise, it is determined that the audio scene is single-person speech.

3. The method of claim 1, wherein, The specific analysis method of the content correlation degree comprises: Performing speech recognition on each valid audio signal in the current set time window to obtain the text content of each valid audio signal; Recording the audio collection devices of each valid audio signal in the current set time window as speaking devices, tracing back the historical signals of each speaking device from the current set time window to the historical time, and determining whether there is a pause point in the historical signals of each speaking device; If there is a pause point in the historical signals of a speaking device, the audio signals between the nearest pause point and the current time point are subjected to speech recognition to obtain the historical text content of the speaking device, which is supplemented to the text content of the valid audio signals of the speaking device; Real-time acquisition of preset theme content of the conference from the background conference database, retrieval of theme-related words related to the preset theme content from a vocabulary library according to the preset theme content, and composition of the theme-related words into a theme-related word set; Count the number of the same association words in the text content of each valid audio signal and the theme association word set, and record the ratio of the number to the total number of words in the text content as the content association degree.

4. The method of claim 3, wherein, The specific method for screening the audio signal to be output comprises: Obtain the time difference between the latest pause point and the end of the current set time window of each valid audio signal, and record the time difference as the speaking duration of each valid audio signal; Perform normalization processing on the speaking duration of each valid audio signal, and perform weighted summation on the speaking duration and the corresponding content association degree to obtain the priority score corresponding to each valid audio signal; Perform priority sorting on each valid audio signal according to the priority score from high to low, and record the valid audio signal with the highest priority as the audio signal to be output.

5. The method of claim 2, wherein, The specific content of the step S4 comprises: Obtain the frame signal amplitude sequence of the audio signal to be output, and perform root mean square calculation to obtain the root mean square amplitude of the audio signal to be output; Extract the speaker sensitivity from the device parameters of each output device, and call the standard reference root mean square amplitude of each output device pre-calibrated from the background conference database, combine the root mean square amplitude of the audio signal to be output, and obtain the initial sound pressure level of the output audio of each output device through the sound pressure level calculation formula; Obtain the difference between the initial sound pressure level and the minimum value of the speech recognition sound pressure range, and comprehensively calculate the effective coverage distance of the output device; Record the range contained by taking each output device as the center and the effective coverage distance as the radius as the effective coverage range.

6. The method of claim 5, wherein, The specific method for judging whether the overall coverage range of the audio needs to be adjusted comprises: Record all the ranges composed of the effective coverage ranges of all the output devices as the overall coverage range of the audio, and extract the positions of all the participants from the participant distribution characteristics; If there is a participant position not in the overall coverage range of the audio, it is judged that the overall coverage range of the audio needs to be adjusted; On the contrary, if all the participant positions are in the overall coverage range of the audio, it is judged that the overall coverage range of the audio does not need to be adjusted.

7. The method of claim 6, wherein, The specific method for obtaining the device to be adjusted comprises: Record all the participants not in the overall coverage range of the audio as audio uncovered personnel, and cluster each adjacent audio uncovered personnel to form each coverage range to be covered; Obtain the device distances from each coverage range to be covered to each output device, and form the device distance sequence of each coverage range to be covered, and obtain the outlier range in the device distance sequence of each coverage range to be covered; Obtain all the device distances in the corresponding outlier range in the device distance sequence of each coverage range to be covered, and record the output devices corresponding to all the device distances as the devices to be adjusted of each coverage range to be covered.

8. The method of claim 7, wherein, The method for obtaining the pre-adjusted coverage distance comprises: Obtain the device distances corresponding to each device to be adjusted of each coverage range to be covered, and record the maximum value as the pre-adjusted coverage distance corresponding to each coverage range to be covered; Set the pre-adjusted coverage distance of each device to be adjusted of each coverage range to be covered according to the corresponding pre-adjusted coverage distance.

9. The method of claim 8, wherein, The specific steps for adjusting the audio signal to be output of each device to be adjusted comprise: The bias sound pressure level of each to-be-adjusted device in each to-be-covered range is calculated according to the ratio of the pre-adjustment coverage distance to the effective coverage distance of each to-be-adjusted device corresponding to each to-be-covered range; The target root mean square amplitude is calculated by comprehensively calculating the bias sound pressure level of each to-be-adjusted device in each to-be-covered range and the root mean square amplitude of the to-be-output audio signal; The ratio of the target root mean square amplitude to the root mean square amplitude of the to-be-output audio signal is recorded as an amplitude scaling factor; The amplitudes in each frame signal of the to-be-output audio signal of each to-be-adjusted device corresponding to each to-be-covered range are scaled according to the corresponding amplitude scaling factor to obtain the adjusted to-be-output audio signal of each to-be-adjusted device for output by the corresponding to-be-adjusted device.

10. A digital conference audio processing system, characterized by It comprises: An audio scene recognition module that collects conference audio signals through an audio collection device, recognizes effective audio signals according to the conference audio signals, and determines the audio scene; A single speaker scene analysis module that records the effective audio signal as the to-be-output audio signal when the audio scene is single speaker; A multi-speaker scene analysis module that analyzes the content correlation of each channel of the effective audio signal to the content of the conference theme, obtains the priority of each channel of the effective audio signal in combination with the speaking time, and screens the to-be-output audio signal therefrom; A coverage range analysis module that obtains the time-domain waveform of the to-be-output audio signal, calculates the root mean square amplitude of the to-be-output audio signal, and obtains the initial sound pressure level of the audio output by each output device and the effective coverage range in combination with the device parameters of each output device; A coverage range adjustment judgment module that determines the overall audio coverage range according to the effective coverage range of each output device, and judges whether the overall audio coverage range needs to be adjusted in combination with the participant distribution characteristics; A coverage range adjustment module that obtains the to-be-adjusted device according to the participant distribution characteristics if the overall audio coverage range needs to be adjusted, calculates the pre-adjustment coverage distance of each to-be-adjusted device, and adjusts the to-be-output audio signal of each to-be-adjusted device.