Acoustic loudspeaker correction enhancement method and system

By separating the speaker output from the on-site human voice, calculating the volume ratio and noise interference, and dynamically adjusting the speaker gain, the problem of inconsistent volume and noise interference in multi-person conference calls is solved, improving call clarity and experience.

CN121001018APending Publication Date: 2025-11-21GANZHOU DEHUIDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511170414.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing speakers have failed to effectively address the issues of inconsistent volume among different speakers and ambient noise interference in multi-person conference calls, resulting in decreased call clarity and user experience.

Method used

The system acquires audio signals from the meeting venue using a voice sensor, separates speech and non-speech signals using a sound activity detection algorithm, calculates the volume ratio and noise interference level, dynamically adjusts the speaker gain range, and forms a comprehensive correction coefficient to match human voices and resist noise interference.

Benefits of technology

It improves the clarity and efficiency of conference calls, ensuring a good listening experience in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121001018A_ABST
    Figure CN121001018A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of loudspeaker voice signal processing, in particular to a sound loudspeaker correction enhancement method and system, and the method comprises the steps: determining a first correction coefficient of a sound loudspeaker at a current moment; based on the audio similarity of the sound equipment voice signals, determining the noise interference degree of the sound equipment loudspeaker at the current moment; determining a second correction coefficient based on the volume of each non-voice signal and the sound equipment voice signal in combination with the noise interference degree, and determining a comprehensive correction coefficient of the sound equipment loudspeaker at the current moment in combination with the first correction coefficient; and correcting the gain range of the sound loudspeaker at the current moment based on the comprehensive correction coefficient. According to the method and the device, the definition and the experience of the conference call are improved by analyzing the condition that the volumes of different speakers are not uniform during the multi-person conference call and the noise in the conference environment interferes with the volume adjustment of the sound speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of loudspeaker speech signal processing technology, specifically to a method and system for correcting and enhancing audio loudspeakers. Background Technology

[0002] Portable Bluetooth speakers are widely used in various occasions such as small conference calls, speeches, and parties due to their portability and wireless connectivity. However, when using Bluetooth speakers for conference voice calls, the audio signal strength can fluctuate during transmission due to factors such as unstable signal sources and changes in the transmission medium. This can cause the sound from the speaker to vary in volume. Therefore, existing methods typically use Automatic Gain Control (AGC) technology to correct the gain of the received audio signal to enhance the consistency of the voice volume at the output audio signal, thereby solving the problem of fluctuating volume during voice calls.

[0003] However, while the existing methods can make the volume of the sound emitted by the speakers relatively uniform, they do not take into account the interference of different speakers' volumes and noise in the conference environment on the speaker volume adjustment during multi-person conference calls, which reduces the clarity and experience of the conference call. Summary of the Invention

[0004] To address the aforementioned technical problems, the purpose of this application is to provide a method and system for correcting and enhancing audio loudspeakers, the specific technical solution of which is as follows: In a first aspect, embodiments of this application provide a method for correcting and enhancing an audio loudspeaker, the method comprising the following steps: The system uses a voice sensor to acquire sound signals from the meeting room and audio signals from the speakers within a preset time period prior to the current moment. All independent signals are separated from the sound signal. A sound activity detection algorithm is used to classify all independent signals into speech signals and non-speech signals. The similarity between each speech signal and the audio signal in the frequency domain is analyzed to determine the audio similarity of each speech signal to obtain the audio speech signal. All speech signals other than the audio speech signal are recorded as conference speech signals. By measuring the volume of all conference speech signals, the volume characteristic value at the current moment is determined. Combined with the volume of the audio speech signal, the first correction coefficient of the speaker at the current moment is determined. Based on the audio similarity of the audio and speech signals, the noise interference level of the speaker at the current moment is determined; based on the volume of each non-speech signal and audio and speech signal, and in combination with the noise interference level, the second correction coefficient of the speaker at the current moment is determined; and in combination with the first correction coefficient, the comprehensive correction coefficient of the speaker at the current moment is determined. Based on the comprehensive correction coefficient, the gain range of the speaker at the current moment is corrected.

[0005] Preferably, the audio similarity of each speech signal is the similarity of the Mel spectrum between each speech signal and the audio signal.

[0006] Preferably, the audio voice signal is the voice signal with the highest audio similarity.

[0007] Preferably, the volume characteristic value at the current moment is the maximum volume among all conference voice signals.

[0008] Preferably, the first correction coefficient of the speaker at the current moment is the result of the volume characteristic value at the current moment being divided by the volume of the audio voice signal.

[0009] Preferably, the expression for the noise interference level of the speaker at the current moment is: In the formula, This indicates the noise level of the speaker at the current moment. This indicates the audio similarity of audio and speech signals.

[0010] Preferably, the expression for the second correction coefficient of the speaker at the current moment is: In the formula, This represents the second correction factor for the speaker at the current moment; This indicates the noise level of the speaker at the current moment. This represents the result of comparing the maximum volume of all non-speech signals to the volume of the audio speech signal.

[0011] Preferably, the comprehensive correction coefficient of the speaker at the current moment is the product of the first correction coefficient and the second correction coefficient of the speaker at the current moment.

[0012] Preferably, the step of correcting the gain range of the speaker at the current moment includes: Obtain the gain range of the speaker at the current moment, and multiply the upper and lower limits of the gain range by the comprehensive correction coefficient to obtain the corrected gain range.

[0013] Secondly, embodiments of this application also provide a speaker correction and enhancement system, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of any of the above-described speaker correction and enhancement methods.

[0014] This application has at least the following beneficial effects: This application, in this embodiment, separates the audio from the speaker output and the on-site human voice, and calculates the volume ratio of the two, i.e., the first correction coefficient. This effectively reduces the interference of background noise on the evaluation of the relative relationship between "maximum on-site human voice" and "actual speaker playback volume," enabling a more accurate judgment of whether the speaker volume is appropriate. It effectively avoids mutual interference between the speaker and human voice, improving the clarity and communication efficiency of conference calls. Furthermore, this application determines the noise interference level by calculating the audio similarity of the audio and speech signals, and dynamically generates a second correction coefficient by combining the volume ratio of non-speech signals to speech signals. This second correction coefficient is then combined with the first correction coefficient to form a comprehensive correction coefficient, which can be applied in real time. This application assesses the degree of background noise interference with speech and the relative relationship between the volume of the on-site human voice and the speaker volume, thereby intelligently adjusting the gain of the speaker to effectively improve speech clarity and ensure a good listening experience even in complex meeting environments. By analyzing the sound at the meeting venue, this application separates the speaker sound and the on-site human voice, analyzes the relative volume of the two and the degree of background noise interference with the speaker sound, and combines the first correction coefficient and the second correction coefficient to obtain a comprehensive correction coefficient, which is used to dynamically adjust the gain range of the speaker, thereby ensuring that the speaker volume can match the on-site human voice and effectively combat background noise interference, ultimately improving the clarity and experience of meeting calls. Attached Figure Description

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating the steps of a method for correcting and enhancing a loudspeaker, provided in one embodiment of this application; Figure 2 This is a schematic diagram of the comprehensive correction coefficient extraction process provided in one embodiment of this application. Detailed Implementation

[0017] To further illustrate the technical means and effects adopted by this application to achieve the intended inventive purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a sound speaker correction and enhancement method and system proposed in this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0019] The following description, in conjunction with the accompanying drawings, details a specific scheme for a method and system for correcting and enhancing an audio loudspeaker provided in this application.

[0020] Please see Figure 1 The diagram illustrates a flowchart of a method for correcting and enhancing a loudspeaker according to an embodiment of this application. The method includes the following steps: Step S1: Use a voice sensor to acquire the sound signals in the meeting room and the audio signals output by the speakers within a preset time period before the current moment.

[0021] When a user selects the multi-person conference call mode in the audio system control system, the audio signal for the audio speaker output is acquired by the voice sensor. The audio signal is a digital voice signal that has been processed by noise reduction, echo cancellation, and reverberation removal. The sound signal of the conference venue is collected by the voice sensor built into the audio system. The voice sensor uses a microphone array, and the sound signal is an array signal. The noise reduction, echo cancellation, and reverberation removal of the voice signal are all well-known technologies, and their specific processes will not be described in detail.

[0022] In this embodiment, the audio signal and sound signal within a preset time period before the current moment are obtained respectively, and used as the audio signal and sound signal at the current moment.

[0023] It should be noted that the preset duration is set manually. In this embodiment, the preset duration is 5 seconds. In actual applications, as other implementation methods, implementers can also set it according to specific circumstances. This embodiment does not impose any special restrictions.

[0024] The system uses a sound activity detection algorithm to detect whether human voices are present in the audio and sound signals at the current moment. If they are present, the system performs subsequent analysis and processing on the audio and sound signals at the current moment. If they are not present, the system does not perform analysis on the audio and sound signals at the current moment and continues to the next moment to make a judgment on the audio and sound signals at the next moment.

[0025] Among them, the sound activity detection algorithm is a well-known technology, and its specific principles will not be elaborated here.

[0026] Step S2: Separate all independent signals from the sound signal, and use a sound activity detection algorithm to divide all independent signals into speech signals and non-speech signals. Analyze the similarity between each speech signal and the audio signal in the frequency domain, determine the audio similarity of each speech signal, and obtain the audio speech signal. Record all speech signals other than the audio speech signal as the conference speech signal. By measuring the volume of all conference speech signals, determine the volume feature value at the current moment, and combine it with the volume of the audio speech signal to determine the first correction coefficient of the speaker at the current moment.

[0027] This embodiment analyzes the sound signals collected at the meeting venue and the audio signals output by the speakers to correct the gain range of the speakers and dynamically adjust the volume of the voice output by the speakers, thereby enhancing the call quality when using Bluetooth speakers for multi-person conference calls.

[0028] Generally, when the volume of the audio output from the speakers is too loud, drowning out the voices of those speaking in the meeting room, or vice versa, it becomes difficult for participants to hear the content of the conversation, thus affecting communication efficiency. Therefore, to avoid this, we separate all independent signals from the audio signal. A sound activity detection algorithm is used to classify all independent signals into speech and non-speech signals. The frequency similarity between each speech signal and the audio signal is analyzed to determine the audio similarity of each speech signal, thus obtaining the audio speech signal. All speech signals other than the audio speech signal are recorded as the meeting speech signal. By measuring the volume of all meeting speech signals, the volume characteristic value at the current moment is determined. Combined with the volume of the audio speech signal, the first correction coefficient of the speakers at the current moment is determined to accurately assess the relative relationship between the maximum volume of the speaker and the average volume of the audio playback. The specific process is as follows: (1) In this embodiment, all independent signals are separated from the sound signal, and a sound activity detection algorithm is used to divide all independent signals into speech signals and non-speech signals, specifically: In this embodiment, the sound signal at the current moment is used as the input to the source number estimation algorithm, and the number of sound sources is output.

[0029] It should be noted that there are many commonly used source number estimation algorithms. In this embodiment, a source estimation algorithm based on the AIC criterion is used to process the sound signal, and the FastICA blind source separation algorithm is used to process the sound signal to separate the speech signals generated by each person in the meeting and the speakers, so as to obtain all the independent signals of the sound signal at the current moment, which are used to represent the speech signals generated by all sound sources in the meeting.

[0030] Among them, the source estimation method based on the AIC criterion and the FastICA blind source separation algorithm are well-known technologies, and the specific principles and processes of using them to determine the source and quantity will not be elaborated here.

[0031] Furthermore, a sound activity detection algorithm is used to detect all independent signals of the sound signal at the current moment. All independent signals containing human speech are recorded as speech signals, which are used to characterize the speech signals generated by speakers and loudspeakers in the meeting room within a preset time period before the current moment. All remaining independent signal components without human speech are recorded as non-speech signals, which are used to characterize the background noise components that may exist in the meeting room within a preset time period before the current moment, such as keyboard typing, mouse clicking, air conditioning or fan operation.

[0032] (2) Further, in this embodiment, the audio similarity of each speech signal is determined by analyzing the similarity between each speech signal and the audio signal in the frequency domain, so as to obtain the audio speech signal. All speech signals other than the audio speech signal are recorded as the conference speech signal, specifically: In this embodiment, the Mel spectrum of the audio signal and all speech signals at the current moment is extracted, and the similarity of the Mel spectrum between each speech signal and the audio signal is used as the audio similarity of each speech signal to evaluate the degree of similarity of the audio content between each speech signal and the audio signal.

[0033] It should be noted that there are many methods to measure the similarity between spectra. In this embodiment, the cosine similarity of the Mel spectrum between each speech signal and the audio signal is used as the similarity of the Mel spectrum between each speech signal and the audio signal. In practical applications, as other implementation methods, implementers may also use other methods to measure the similarity of spectra, such as the reciprocal of the Euclidean distance, depending on the specific circumstances. This embodiment does not impose any special restrictions on the selection of methods to measure the similarity between spectra.

[0034] The extraction process of the Mel spectrum and the calculation method of the cosine similarity are well-known techniques, and the specific extraction process of the Mel spectrum and the specific calculation process of the cosine similarity will not be described in detail here.

[0035] Furthermore, the speech signal with the highest audio similarity is used as the audio speech signal to characterize the speech signal generated by the speaker within a preset time period before the current moment; Furthermore, all speech signals other than the speech signals are taken as conference speech signals, which are used to represent the speech signals produced by the speakers at the conference site within a preset time period before the current moment.

[0036] (3) Further, in this embodiment, by measuring the volume of all conference voice signals, the volume characteristic value at the current moment is determined, and combined with the volume of the audio voice signal, the first correction coefficient of the audio speaker at the current moment is determined, specifically as follows: As one implementation method, in this embodiment, the volume of each audio voice signal is calculated separately, and the maximum volume is used as the volume feature value at the current moment to evaluate the maximum volume of the speakers in the meeting room within a preset time period before the current moment. Furthermore, the average volume of the audio voice signal is used to evaluate the average volume of the audio speaker playing voice over a preset period of time prior to the current moment; Furthermore, in this embodiment, the result of dividing the volume characteristic value at the current moment by the volume of the audio voice signal is used as the first correction coefficient of the audio speaker at the current moment, which is used to correct the gain range of the audio speaker at the current moment.

[0037] The method for calculating volume is a well-known technique, and its specific calculation process will not be elaborated here.

[0038] Based on the first correction coefficient of the speaker at the current moment, it can be understood that if the volume characteristic value at the current moment is larger, it means that someone in the meeting room is speaking very loudly. The difference between the maximum volume and the average playback volume of the speaker is widening, indicating that the sound in the room may have drowned out the speaker's sound. The corresponding first correction coefficient is relatively large, meaning that the speaker volume may be insufficient. At the same time, the lower the volume of the audio voice signal, it means that the average volume of the audio voice played by the speaker has decreased. The intensity of the speaker sound is weakening relative to the maximum volume of the speaker in the room. That is, the larger the corresponding first correction coefficient, it means that the speaker volume may be insufficient. Conversely, if the volume characteristic value at the current moment is smaller, it means that the speaker's maximum volume in the current meeting room has decreased. The difference between the maximum volume in the room and the average playback volume of the speakers is narrowing, indicating that the sound in the room may not be enough to drown out the sound of the speakers. The corresponding first correction coefficient is relatively small, which means that the speaker volume may be too loud. At the same time, the larger the volume of the audio voice signal, it means that the average volume of the audio voice currently being played by the speakers has increased. The intensity of the speaker sound is increasing relative to the maximum volume of the speaker in the room. That is, the smaller the corresponding first correction coefficient, the more likely the speaker volume is too loud.

[0039] Thus, this embodiment dynamically assesses the volume balance by separating the sound from the speaker output and the voices in the room, and calculating the volume ratio of the two, i.e., the first correction coefficient. When the sound is too loud or the speaker volume is too low, the first correction coefficient increases, and vice versa. Based on this, the speaker gain is automatically adjusted, which effectively avoids interference between the speaker and the voices, and improves the clarity and communication efficiency of the conference call.

[0040] Step S3: Based on the audio similarity of the audio speech signal, determine the noise interference level of the speaker at the current moment; based on the volume of each non-speech signal and the audio speech signal, and in combination with the noise interference level, determine the second correction coefficient of the speaker at the current moment, and in combination with the first correction coefficient, determine the comprehensive correction coefficient of the speaker at the current moment.

[0041] When a speaker is playing audio, background noise in the meeting room often mixes with the sound, interfering with the auditory perception of the participants. When the background noise is too loud, even if the volume of the audio played by the speaker is greater than the maximum volume of the speakers in the room, it can still make it difficult for participants to clearly hear the content of the conversation. Therefore, to reduce the interference of excessive background noise in the meeting room on participants' ability to identify the content of the conversation output by the speaker, this embodiment determines the noise interference level of the speaker at the current moment based on the audio similarity of the audio signals; based on the volume of each non-speech signal and the audio signal, and in conjunction with the noise interference level, a second correction coefficient for the speaker at the current moment is determined; and combined with the first correction coefficient, a comprehensive correction coefficient for the speaker at the current moment is determined. The specific process is as follows: (1) In this embodiment, the noise interference level of the speaker at the current moment is determined based on the audio similarity of the audio voice signal, specifically as follows: As one implementation method, the noise interference level of the speaker at the current moment in this embodiment is... The expression is: In the formula, This indicates the audio similarity of audio and speech signals.

[0042] Based on the noise interference level of the speaker at the current moment, it can be understood that the noise interference level is used to evaluate the degree of interference of the audio signal generated by the speaker within a preset time period before the current moment with the background noise of the meeting venue. The greater the audio similarity, the smaller the corresponding noise interference level, indicating that the audio signal quality is better and less affected by noise interference, and therefore the noise interference level is smaller. Conversely, the smaller the audio similarity, the greater the corresponding noise interference level, indicating that the audio signal quality is worse and more severely affected by noise interference, and therefore the noise interference level is greater.

[0043] (2) Further, in this embodiment, based on the volume of each non-speech signal and audio speech signal, and in conjunction with the noise interference level, the second correction coefficient of the audio speaker at the current moment is determined as follows: In this embodiment, the second correction coefficient of the speaker at the current moment The expression is: In the formula, This indicates the noise level of the speaker at the current moment. This represents the result of comparing the maximum volume of all non-speech signals to the volume of the audio speech signal.

[0044] The value of 1 is to prevent the second correction coefficient from being invalidated if the noise interference level is 0.

[0045] Based on the second correction coefficient of the speaker at the current moment, it can be understood that the greater the noise interference of the speaker at the current moment, the more serious the interference of background noise on the speech is. Even if the volume of the speaker is increased, it is difficult to hear clearly. However, the speaker tends to try to cancel out this interference by significantly increasing the volume to ensure the audibility of the speech. Therefore, the greater the noise interference, the larger the corresponding second correction coefficient. At the same time, if the maximum value of all non-speech signal volumes is greater than the volume of the audio speech signal, the corresponding second correction coefficient is also greater. This reflects that the volume of background noise is very prominent relative to the volume of the speech played by the speaker. This means that the background noise itself is very loud and easily drowns out the speaker sound. Therefore, it is necessary to increase the speaker volume to counteract it and ensure that the speech can be heard. Conversely, if the noise interference from the speaker is lower at the current moment, it means that the background noise interferes less with the speech, and the speech is relatively clear. There's no need to significantly increase the volume to counteract the noise; therefore, the lower the noise interference, the smaller the corresponding second correction coefficient. Simultaneously, if the maximum value of all non-speech signal volumes is smaller than the volume of the speaker's speech signal, the corresponding second correction coefficient is also smaller. This reflects that the background noise volume is not prominent relative to the volume of the speech played by the speaker; the background noise itself is not very loud and is unlikely to drown out the speaker's sound. Therefore, there's no need to significantly increase the speaker volume to counteract it; maintaining the current volume or making slight adjustments is sufficient to ensure that the speech can be heard.

[0046] (3) Further, in this embodiment, based on the first correction coefficient and the second correction coefficient, the comprehensive correction coefficient of the speaker at the current moment is determined as follows: In this embodiment, the product of the first correction coefficient and the second correction coefficient of the speaker at the current moment is used as the comprehensive correction coefficient of the speaker at the current moment.

[0047] Preferably, the schematic diagram of the comprehensive correction coefficient extraction process provided in this embodiment is as follows: Figure 2 As shown.

[0048] Based on the overall correction coefficient of the speaker at the current moment, it can be understood that if the first correction coefficient of the speaker at the current moment is larger, it means that the voices of the people in the room are louder than the speaker. At the same time, combined with the second correction coefficient, if the second correction coefficient is larger, it means that the background noise interference is serious. At this time, the speaker needs to increase the volume to match the voices of the people in the room or to avoid the voices of the people in the room being drowned out. Therefore, the corresponding overall correction coefficient is relatively large. Conversely, if the first correction factor of the speaker is smaller at the current moment, it means that the voices in the room are relatively small compared to the speaker, and the speaker volume may already be sufficient or even too loud. At the same time, combined with the second correction factor, if the second correction factor is smaller, it means that the background noise interference is smaller and the proportion of background noise to the speaker volume is not high. In this case, the speaker does not need to significantly increase the volume to match the voices in the room or to counteract the background noise. It may need to maintain the current volume or slightly reduce it to avoid the volume being too loud. Therefore, the corresponding comprehensive correction factor is relatively small.

[0049] Thus, this embodiment determines the noise interference level by calculating the audio similarity of the audio and speech signals, and dynamically generates a second correction coefficient by combining the volume ratio of the non-speech signal to the speech signal. This second correction coefficient is then combined with the first correction coefficient to form a comprehensive correction coefficient. This allows for real-time assessment of the degree of background noise interference with speech and the relative relationship between the on-site human voice and the speaker volume. Consequently, the gain of the audio speaker is intelligently adjusted to effectively improve speech clarity and ensure a good listening experience even in complex meeting environments.

[0050] Step S4: Based on the comprehensive correction coefficient, correct the gain range of the speaker at the current moment.

[0051] Based on the comprehensive correction coefficient obtained in step S3, this embodiment corrects the gain range of the speaker at the current moment based on the comprehensive correction coefficient. Specifically, in this embodiment, the upper and lower limits of the gain range of the speaker at the current moment are obtained, and the upper and lower limits are multiplied by the comprehensive correction coefficient of the speaker at the current moment to obtain the corrected gain range. The corrected gain is used as the gain range at the next moment, and the audio microphone after the gain range is corrected is used to enhance the conference call.

[0052] The process of obtaining the gain range is a well-known technique, and its specific principles and processes will not be elaborated here.

[0053] Thus, this embodiment analyzes the sound at the meeting venue, separates the speaker sound from the on-site human voice, analyzes the relative volume of the two and the degree of interference of background noise on the speaker sound, and obtains a comprehensive correction coefficient by combining the first correction coefficient and the second correction coefficient. This comprehensive correction coefficient is used to dynamically adjust the gain range of the speaker, thereby ensuring that the speaker volume can match the on-site human voice and effectively combat background noise interference, ultimately improving the clarity and experience of the meeting call.

[0054] Based on the same inventive concept as the above method, this application embodiment also provides a speaker correction and enhancement system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described speaker correction and enhancement methods.

[0055] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0056] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0057] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for correcting and enhancing a loudspeaker, characterized in that, The method includes the following steps: The system uses a voice sensor to acquire sound signals from the meeting room and audio signals from the speakers within a preset time period prior to the current moment. All independent signals are separated from the sound signal. A sound activity detection algorithm is used to classify all independent signals into speech signals and non-speech signals. The similarity between each speech signal and the audio signal in the frequency domain is analyzed to determine the audio similarity of each speech signal to obtain the audio speech signal. All speech signals other than the audio speech signal are recorded as conference speech signals. By measuring the volume of all conference speech signals, the volume characteristic value at the current moment is determined. Combined with the volume of the audio speech signal, the first correction coefficient of the speaker at the current moment is determined. Based on the audio similarity of the audio and speech signals, the noise interference level of the speaker at the current moment is determined; based on the volume of each non-speech signal and audio and speech signal, and in combination with the noise interference level, the second correction coefficient of the speaker at the current moment is determined; and in combination with the first correction coefficient, the comprehensive correction coefficient of the speaker at the current moment is determined. Based on the comprehensive correction coefficient, the gain range of the speaker at the current moment is corrected.

2. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The audio similarity of each speech signal is the similarity of the Mel spectrum between each speech signal and the audio signal.

3. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The audio voice signal is the voice signal with the highest audio similarity.

4. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The volume characteristic value at the current moment is the maximum volume among all conference voice signals.

5. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The first correction coefficient of the speaker at the current moment is the result of the volume characteristic value at the current moment being divided by the volume of the audio voice signal.

6. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The expression for the noise interference level of the speaker at the current moment is: In the formula, This indicates the noise level of the speaker at the current moment. This indicates the audio similarity of audio and speech signals.

7. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The expression for the second correction coefficient of the speaker at the current moment is: In the formula, This represents the second correction factor for the speaker at the current moment; This indicates the noise level of the speaker at the current moment. This represents the result of comparing the maximum volume of all non-speech signals to the volume of the audio speech signal.

8. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The comprehensive correction coefficient of the speaker at the current moment is the product of the first correction coefficient and the second correction coefficient of the speaker at the current moment.

9. The method for correcting and enhancing a loudspeaker as described in claim 1, characterized in that, The correction of the gain range of the speaker at the current moment includes: Obtain the gain range of the speaker at the current moment, and multiply the upper and lower limits of the gain range by the comprehensive correction coefficient to obtain the corrected gain range.

10. A system for correcting and enhancing a loudspeaker, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio loudspeaker correction and enhancement method as described in any one of claims 1-9.