Signal optimization method, system and device of wireless cascade sound box and medium

Through the collaborative work of main and auxiliary equipment and signal optimization technology, wireless cascading speakers solve the problems of complex wiring and insufficient signal quality of traditional speakers, realize efficient and flexible audio signal processing and transmission, and improve conference audio quality and communication efficiency.

CN120302214APending Publication Date: 2025-07-11HUIZHOU ACOUSTIC BIT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403610.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional conference speakers rely on wired connections to lead to complex wiring, inflexible device deployment, limited microphone array acquisition range, limited signal enhancement effect, and difficult to guarantee signal quality.

Method used

Wireless cascading speaker technology is adopted to work collaboratively through the microphone array of the main device and the auxiliary device to collect, enhance, transmit and process audio signals, and combine adaptive filters and deep learning noise reduction models to achieve echo cancellation and noise suppression.

Benefits of technology

It improves the acquisition range and quality of audio signals, enhances voice clarity and intelligibility, improves the integrity and comfort of conference audio, reduces system complexity and cost, and improves conference communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302214A_ABST
    Figure CN120302214A_ABST
Patent Text Reader

Abstract

The invention discloses a signal optimization method, system and device of a wireless cascade sound box and a medium, and relates to the field of signal processing. In the method, when target equipment is started, the target equipment is allocated to main equipment or auxiliary equipment; audio signals are collected through microphone arrays of the auxiliary device and the main device to generate a first microphone enhancement signal and a second microphone enhancement signal, and the first microphone enhancement signal is sent to the main device in a wireless coding transmission mode; performing dynamic time alignment and spatial position weight overlapping on the first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; and receiving a downlink signal from the connected conference terminal through the master device, performing echo cancellation, noise reduction and gain processing on the downlink signal and the comprehensive microphone signal to obtain an uplink enhanced signal, and sending the uplink enhanced signal to the conference terminal. By implementing the technical scheme provided by the invention, the definition and the reduction degree of the voice are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of signal processing, and particularly to a signal optimization method, system, device and medium for a wireless cascaded speaker. Background Art

[0002] The core function of traditional conference speakers lies in voice collection and transmission. Most existing systems still rely on wired connections, and such a technical design has the following problems: The wired connection between conference speakers and extended microphones will lead to complex cabling in the conference room, affecting aesthetics and device mobility; The devices cannot be dynamically deployed according to the size of the conference room or the number of conference participants; The microphone array of a single device has a limited collection range, and the signal enhancement effect is limited by the physical structure and algorithm capabilities, and the signal quality is difficult to guarantee.

[0003] Therefore, a more efficient signal optimization method is needed to overcome the deficiencies in the prior art and achieve stable and reliable audio signal processing in a multi-device environment. Summary of the Invention

[0004] This application provides a signal optimization method, system, device and medium for a wireless cascaded speaker, effectively improving the clarity and restoration degree of voice.

[0005] In a first aspect of this application, a signal optimization method for a wireless cascaded speaker is provided, which is applied to a wireless cascaded speaker. The wireless cascaded speaker includes a master device and a slave device. The method includes: When a target device is started, the target device is assigned as a master device or a slave device according to the state at the last shutdown or a received control signal; Collect an audio signal through a first microphone array of the slave device to generate a first microphone enhanced signal, and send the first microphone enhanced signal to the master device by means of wireless coding transmission; Collect an audio signal through a second microphone array of the master device to generate a second microphone enhanced signal, and perform weighted superposition on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; Receive a downlink signal from a connected conference terminal through the master device, perform echo cancellation, noise reduction and gain processing on the downlink signal and the comprehensive microphone signal to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0006] Optionally, the collecting an audio signal through a first microphone array of the slave device to generate a first microphone enhanced signal includes: Perform a fast Fourier transform on the audio signal collected by a single first microphone of the first microphone array to obtain the spectrum of the signal, extract the target features related to noise from the spectrum, classify the noise components in the audio signal according to the target features, and the target features include amplitude spectrum, phase spectrum or spectral kurtosis; Process the noise components in the audio signal using the corresponding noise suppression method according to the classification result, and convert the processed spectrum back to the time domain through inverse fast Fourier transform to obtain the speech signal in the time domain; Perform post-processing on the speech signal to generate the first microphone enhanced signal, and the post-processing includes spectral flatness adjustment, fundamental frequency estimation and correction.

[0007] Optionally, collecting an audio signal through the first microphone array of the auxiliary device to generate a first microphone enhanced signal includes: Perform phase alignment on the multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determine the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and perform weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal; Perform enhancement processing on the second signal to generate the first microphone enhanced signal.

[0008] Optionally, performing phase alignment on the multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determining the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and performing weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal includes: Perform frame splitting on the audio signal to obtain multiple sub-frames. For each sub-frame, calculate the cross-correlation function between different first microphones to obtain multiple groups of time delay differences, and each group of time delay differences includes the time delay difference between two different first microphones; Perform phase correction on the multiple audio signals according to the multiple groups of time delay differences to obtain multiple first signals; Determine the direction of the sound source according to the multiple groups of time delay differences, and calculate the distance between the sound source and each first microphone according to the direction, the geometric structure of the first microphone array and the sound wave propagation speed; Allocate corresponding weights according to the distance. The closer the distance, the greater the weight of the first signal. Perform weighted superposition on the multiple first signals according to the weights to obtain a second signal.

[0009] Optionally, performing weighted superposition on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a combined microphone signal includes: Align the single first microphone enhanced signal and the second microphone enhanced signal to obtain an adjusted first microphone enhanced signal and an adjusted second microphone enhanced signal; Determine the spatial position weight corresponding to each target microphone array according to the spatial position relationship between the single first microphone array and the second microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal according to the spatial position weight.

[0010] Optionally, the weighted superposition of the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal includes: Align multiple first microphone enhanced signals and the second microphone enhanced signal to obtain multiple third signals; Determine the spatial position weight corresponding to each target microphone array according to the spatial position relationship between multiple first microphone arrays and the second microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on multiple third signals according to the spatial position weight.

[0011] Optionally, the obtaining of the uplink enhanced signal by performing echo cancellation, noise reduction, and gain processing on the downlink signal and the comprehensive microphone signal includes: Construct an adaptive filter based on the downlink signal, calculate the coefficients of the adaptive filter through the normalized least mean square algorithm, and generate an estimated echo signal coupled with the sound field played by the main device speaker; Perform time-domain cancellation on the comprehensive microphone signal and the estimated echo signal to obtain an error signal; Dynamically adjust the step factor in the normalized least mean square algorithm according to the ratio of the power spectral density of the error signal to the power spectral density of the comprehensive microphone signal to optimize the error signal; Perform noise reduction processing on the optimized error signal, and use a deep learning-based noise reduction model to suppress the environmental noise in the signal to obtain a noise-reduced signal; Adjust the amplitude of the noise-reduced signal according to a preset gain curve to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0012] In a second aspect of the present application, a signal optimization system for a wireless cascaded speaker is provided, including a distribution module, a transmission module, a superposition module, and an execution module, where: An allocation module, configured to allocate the target device as a master device or a slave device according to the state at the last shutdown or a received control signal when the target device starts up; A sending module, configured to collect an audio signal through a first microphone array of a slave device to generate a first microphone enhanced signal, and send the first microphone enhanced signal to the master device by means of wireless coding transmission; An overlay module, configured to collect an audio signal through a second microphone array of the master device to generate a second microphone enhanced signal, and perform weighted overlay on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; An execution module, configured to receive a downlink signal from a connected conference terminal through the master device, perform echo cancellation, noise reduction, and gain processing on the downlink signal and the comprehensive microphone signal to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0013] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.

[0014] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.

[0015] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By collecting an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal, and then wirelessly transmitting it to the master device, combined with the second microphone enhanced signal collected by the second microphone array of the master device itself, the collaborative work of multiple devices is realized. This expands the audio signal collection range from the coverage area of the microphone of a single device to the space range where multiple devices are located, enabling more comprehensive capture of the audio signal at the meeting site, ensuring that sounds at all positions can be effectively collected, and avoiding signal attenuation and voice quality degradation caused by the single microphone being too far from the sound source. It is particularly suitable for meeting scenarios in large conference rooms or open spaces, improving the integrity and intelligibility of conference audio; 2. By constructing an adaptive filter based on the downlink signal, using the NLMS algorithm to calculate the filter coefficients, generating an estimated echo signal coupled with the sound field played by the speaker, and performing time-domain cancellation with the integrated microphone signal, the common echo problem in meetings is effectively eliminated. A noise reduction model based on deep learning is used to process the signal after preliminary echo cancellation, which can effectively suppress environmental noise and further improve the purity of the audio signal. The amplitude of the signal after noise reduction is adjusted according to the preset gain curve to obtain the uplink enhanced signal. The gain processing can finely adjust the signal amplitude in different frequency ranges, enhance the loudness of the target speech signal, and make it clearer and more audible in the meeting. Through reasonable gain settings, the important frequency components of the speech signal can be highlighted, while avoiding the signal being too strong or too weak, ensuring that the volume of the meeting audio is moderate, and improving the comfort and professionalism of the meeting; 3. When the target device is started, the target device is assigned as the master device or the slave device according to the state at the last shutdown or the received control signal. This dynamic role assignment mechanism makes the system highly flexible. In different meeting scenarios and requirements, the role of the device can be flexibly adjusted according to the actual situation, making full use of the existing device resources, without the need to configure dedicated master devices or slave devices additionally, reducing the cost and complexity of the system. At the same time, it also facilitates the operation and management of users and improves the usability of the system. The first microphone enhanced signal of the slave device is sent to the master device through wireless coding transmission, getting rid of the limitation of wired connection, making the deployment of the device more free and flexible. At the meeting site, the slave device can be placed randomly according to needs without considering the wiring problem, and it can better adapt to different conference room layouts and meeting requirements, improving the adaptability and convenience of the system; 4. Through the acquisition, transmission, processing, and optimization of the audio signal, finally, the high-quality uplink enhanced signal is sent to the meeting terminal, effectively improving the quality and clarity of the meeting audio. The clear audio signal can ensure that the meeting participants can understand each other's speech content more accurately, reducing misunderstandings and repeated communications caused by unclear speech, and improving the communication efficiency and decision-making speed of the meeting. Brief Description of the Drawings

[0016] Figure 1 is a schematic flowchart of the signal optimization method for the wireless cascaded speaker disclosed in the embodiment of the present application; Figure 2 is a schematic architecture diagram of the wireless cascaded speaker disclosed in the embodiment of the present application; Figure 3 is a schematic module diagram of the signal optimization system for the wireless cascaded speaker disclosed in the embodiment of the present application; Figure 4 is a schematic structural diagram of an electronic device disclosed in the embodiment of the present application.

[0017] Explanation of the reference numerals: 301, allocation module; 302, sending module; 303, superposition module; 304, execution module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. DETAILED DESCRIPTION

[0018] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0019] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.

[0020] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0021] This embodiment discloses a signal optimization method for wireless cascade speakers, which is applied to wireless cascade speakers. Figure 1 is a flow chart of the signal optimization method of the wireless cascade speakers disclosed in the embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: S101, when a target device is started, the target device is allocated as a primary device or a secondary device according to the state when it was last shut down or a received control signal; S102 collects an audio signal through a first microphone array of the auxiliary device to generate a first microphone enhancement signal, and sends the first microphone enhancement signal to the main device through wireless coding transmission; S103, collecting an audio signal through a second microphone array of the master device to generate a second microphone enhancement signal, and performing weighted superposition on the received first microphone enhancement signal and the second microphone enhancement signal to obtain a comprehensive microphone signal; S104. Receive a downlink signal from the connected conference terminal through the master device, obtain an uplink enhanced signal by performing echo cancellation, noise reduction, and gain processing on the downlink signal and the integrated microphone signal, and send the uplink enhanced signal to the conference terminal.

[0022] Figure 2 It is a schematic diagram of the architecture of the wireless cascaded speaker disclosed in the embodiments of the present application. As Figure 2 shown, the master device is wirelessly cascaded with the slave device. The master device can be connected to the conference terminal through Bluetooth, USB, line, or wireless + dongle to achieve full-duplex communication for both uplink and downlink.

[0023] When the target device starts up, the target device is assigned as the master device or the slave device according to the state at the last shutdown or the received control signal. This role assignment mechanism enables the device to flexibly assume different roles according to the actual situation, thus achieving more efficient audio signal processing and transmission. The master device is responsible for connecting to the intelligent terminal or the conference terminal and comprehensively processing and optimizing the audio signal; the slave device is responsible for collecting the audio signal and sending it to the master device to assist the master device in completing the audio signal collection and processing tasks. This flexible role assignment and collaborative working mechanism can give full play to the advantages of each device and improve the overall performance and adaptability of the system. The first microphone array of the slave device collects the audio signal to generate the first microphone enhanced signal, and the first microphone enhanced signal is sent to the master device by means of wireless coding transmission. One slave device corresponds to one first microphone array and one first microphone enhanced signal. When there are multiple slave devices, multiple paths of first microphone enhanced signals will be generated. After the slave device starts up and establishes its role, it will turn off its speaker and turn on the microphone, and use its microphone array to collect the audio signal at the conference site. If there are multiple microphones on the slave device, array signal enhancement can be performed to generate an enhanced microphone signal; if there is only one microphone, noise suppression processing will be performed to output the microphone enhanced signal. These enhanced signals are sent to the master device through the low-latency wireless coding transmission chip, providing important basic data for the subsequent signal processing of the master device. This process enables the slave device to effectively collect the audio signal at the conference site and transmit it to the master device quickly and stably, expanding the audio signal collection range and improving the diversity and richness of the signal. The second microphone array of the master device collects the audio signal to generate the second microphone enhanced signal, and the received first microphone enhanced signal and the second microphone enhanced signal are dynamically time-aligned and spatially position-weighted and superimposed to obtain the comprehensive microphone signal. After the master device starts up, it will turn on its own speaker and microphone, use the second microphone array to collect the audio signal at the conference site, and perform corresponding signal enhancement processing to generate the second microphone enhanced signal. At the same time, after the master device receives the first microphone enhanced signal from the slave device, it will perform dynamic time alignment on these two signals to eliminate the time difference between the signals collected by different devices and ensure the synchronization of the signals. Then, according to the spatial position relationship between the master device and the slave device, the corresponding spatial position weight is determined, and the aligned signals are weighted and superimposed to obtain the comprehensive microphone signal. This process can effectively fuse the audio signals collected by the master device and the slave device, enhance the target voice signal, and at the same time suppress the noise signals from other directions, improving the quality and clarity of the audio signal and providing a better signal basis for the subsequent echo cancellation and noise reduction processing of the master device.The master device receives a downlink signal from the connected conference terminal, performs echo cancellation, noise reduction, and gain processing on the downlink signal and the integrated microphone signal to obtain an uplink enhanced signal, and sends the uplink enhanced signal to the conference terminal. The downlink signal received by the master device from the conference terminal contains the voice information of other participants in the conference. The master device will play this signal through the speaker, and at the same time, use an adaptive filter constructed based on the downlink signal to generate an estimated echo signal coupled with the sound field played by the speaker. The integrated microphone signal is subtracted from the estimated echo signal in the time domain to obtain a preliminary echo-cancelled error signal. Then, according to the ratio of the power spectral density of the error signal to the power spectral density of the integrated microphone signal, the step factor of the adaptive filter is dynamically adjusted to optimize the error signal. Next, the optimized error signal is subjected to noise reduction processing. A deep learning-based noise reduction model is used to suppress the ambient noise in the signal to obtain a noise-reduced signal. Finally, according to the preset gain curve, the noise-reduced signal is subjected to gain processing to adjust the amplitude of the signal to obtain an uplink enhanced signal. This process can effectively eliminate echoes, suppress noise, and enhance the loudness of the target voice signal, making the final audio signal sent to the conference terminal have higher quality, clarity, and intelligibility, and improving the communication effect of the conference and the auditory experience of the participants.

[0024] If there is a control signal, the current identity of each device can be directly adjusted according to the control signal to ensure that there is a master device currently. If there is no control signal, the slave device will periodically send a heartbeat signal to the master device. If the master device's response is not received continuously for multiple times, the slave device will confirm that the master device has been removed. An election mechanism will be initiated among the slave devices to determine a new master device. The election mechanism can be based on the following factors: Device priority: The preset device priority, for example, the priority of device A is higher than that of device B.

[0025] Signal strength: According to the signal strength among the slave devices, the device with the strongest signal is selected as the new master device.

[0026] Geographical location: According to the geographical locations of the slave devices, the most central or most suitable device is selected as the new master device.

[0027] Through the election mechanism, the slave devices will reach an agreement to determine a new master device. The selected slave device will switch to the master device and assume the responsibilities of the master device. The new master device will re-establish connections with other slave devices. The new master device will synchronize the status with other slave devices to ensure that the working status of all devices is consistent.

[0028] Optionally, the generating of the first microphone enhanced signal by collecting the audio signal through the first microphone array of the slave device includes: Perform a fast Fourier transform on the audio signal collected by a single first microphone of the first microphone array to obtain the spectrum of the signal, extract target features related to noise from the spectrum, classify the noise components in the audio signal according to the target features, where the target features include amplitude spectrum, phase spectrum or spectral kurtosis; Process the noise components in the audio signal by using a corresponding noise suppression method according to the classification result, and convert the processed spectrum back to the time domain through an inverse fast Fourier transform to obtain a speech signal in the time domain; Perform post-processing on the speech signal to generate a first microphone enhanced signal.

[0029] After the first microphone array of the auxiliary device collects the audio signal, a fast Fourier transform (FFT) is performed on the audio signal collected by a single first microphone to convert the time-domain signal into a frequency-domain signal, obtaining the spectrum of the signal. FFT is an efficient algorithm that can complete the spectrum analysis of the signal in a relatively short time, providing a basis for subsequent noise suppression and signal processing. Target features related to noise are extracted from the spectrum, and these features include the amplitude spectrum, phase spectrum, or spectral kurtosis, etc. The amplitude spectrum reflects the energy distribution of the signal at different frequencies. Noise usually has higher energy at certain frequencies, and the noise components can be identified by analyzing the amplitude spectrum. The phase spectrum reflects the phase information of the signal at different frequencies. The phase of the noise signal is usually relatively chaotic, different from the phase characteristics of the speech signal. Spectral kurtosis is an index to measure the severity of the change in the signal spectrum. The spectral kurtosis of the noise signal is usually higher, while that of the speech signal is relatively lower. By extracting these features, the noise components in the audio signal can be more accurately identified and classified. According to the extracted target features, the noise components in the audio signal are classified. Different types of noise have different characteristics. For example, Gaussian white noise shows a relatively flat distribution in the amplitude spectrum and is relatively random in the phase spectrum; impulse noise shows high energy peaks at certain specific frequencies in the amplitude spectrum and may have a certain periodicity in the phase spectrum. By classifying the noise, different noise suppression methods can be adopted specifically to improve the effect of noise suppression. According to the classification result, the corresponding noise suppression method is used to process the noise components in the audio signal. For Gaussian white noise, methods such as spectral subtraction and Wiener filtering can be used for suppression; for impulse noise, methods such as median filtering and adaptive filtering can be used for elimination. These methods process the spectrum to remove the noise components and retain the characteristics of the speech signal, thereby improving the purity of the audio signal. The processed spectrum is converted back to the time domain through the inverse fast Fourier transform (IFFT) to obtain the speech signal in the time domain. IFFT is the inverse operation of FFT, which can restore the frequency-domain signal to the time-domain signal, enabling the signal after noise suppression processing to be further processed and analyzed in the time domain. Post-processing includes spectral flatness adjustment, fundamental frequency estimation, and correction. The spectral flatness of the speech signal is adjusted to optimize the spectral characteristics of the signal. Spectral flatness refers to the degree of uniformity of the energy distribution of the signal spectrum at each frequency. The spectral flatness of the speech signal is usually lower, while that of the noise signal is higher. By adjusting the spectral flatness, the spectrum of the signal can be made more flat, reducing the influence of noise and improving the clarity and intelligibility of the speech signal. Fundamental frequency estimation and correction are performed to enhance the fundamental frequency component of the speech signal. The fundamental frequency is one of the most important features of the speech signal, which determines the pitch and timbre of the speech. By estimating and correcting the fundamental frequency, the accuracy of the fundamental frequency of the speech signal can be improved, making the speech more natural and clear.For example, methods such as adaptive filtering and harmonic enhancement can be used to enhance the fundamental frequency component while suppressing the interference of other frequency components.

[0030] Perform a fast Fourier transform on the audio signal collected by a single first microphone to obtain the spectrum of the signal, and extract target features related to noise from the spectrum, such as amplitude spectrum, phase spectrum, or spectral kurtosis, etc. Classify the noise components in the audio signal through these features, which can more accurately identify different types and sources of noise, such as environmental noise, device noise, etc. According to the classification results, adopt corresponding noise suppression methods to process the noise components, which can achieve targeted noise suppression, improve the effect and accuracy of noise suppression, effectively retain the useful components of the speech signal, while removing or weakening the noise components, and improve the quality of the audio signal. Convert the processed spectrum back to the time domain through the inverse fast Fourier transform to obtain the speech signal in the time domain. This process can accurately convert the spectrum information after noise suppression processing back to the time domain signal, retaining the time domain characteristics of the speech signal, and providing a high-quality speech signal for subsequent speech processing and transmission. Through post-processing of the speech signal, including operations such as spectral flatness adjustment, fundamental frequency estimation and correction, etc., the spectral characteristics of the speech signal can be further optimized to make it more flat, reduce spectral distortion, and at the same time accurately estimate and correct the fundamental frequency of the speech signal, improve the naturalness and intelligibility of the speech signal, and generate a high-quality first microphone enhanced signal. The first microphone enhanced signal generated through the above steps can more accurately reflect the speech content at the meeting site, reduce the interference of noise, and improve the clarity and quality of the audio signal.

[0031] Optionally, the generating the first microphone enhanced signal by collecting the audio signal through the first microphone array of the auxiliary device includes: Align the phases of the multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determine the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and perform weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal; Perform enhancement processing on the second signal to generate a first microphone enhanced signal.

[0032] Since the multiple microphones in the microphone array are at different distances from the sound source, there are time delays in the received audio signals, resulting in inconsistent phases. The purpose of phase alignment is to eliminate these time delays so that the same speech signal received by different microphones is consistent in phase, creating conditions for subsequent signal superposition. By performing frame splitting on the audio signals collected by multiple microphones, multiple sub-frames are obtained. Then, the cross-correlation function between different microphones is calculated to determine the time delay (time delay difference) between them. Based on these time delay differences, the phase of the audio signal of each microphone is corrected so that the same speech signal is aligned in the signals of different microphones. According to the spatial position and signal strength of the microphones, the weight of each microphone signal is determined. The weight allocation is to highlight the microphone signals that are closer to the sound source and have better signal quality while suppressing the microphone signals that are farther from the sound source and may be more affected by noise interference during the signal superposition process. First, obtain the spatial position information of each microphone in the microphone array. Then, based on this position information and the signal strength collected by each microphone, the weight of each microphone is calculated. Usually, the microphone closer to the sound source has a larger weight, and the microphone farther away has a smaller weight. The multiple microphone signals after phase alignment are superimposed according to the determined weights to generate an enhanced signal. Signal superposition can effectively enhance the amplitude of the target speech signal while suppressing the noise signal. Multiply the multiple microphone signals after phase alignment by their corresponding weights respectively, and then add these weighted signals to obtain the second signal. The enhancement process may include: adjusting the amplitude of the second signal through a preset gain factor, or enhancing the signals corresponding to some frequencies of the second signal, which may specifically include bass enhancement, anti-interference processing for specific frequencies, sound effect enhancement processing, equalization processing, etc. Adjust the amplitude of the second signal through a preset gain factor to make it reach an appropriate level. The gain adjustment can compensate for the attenuation of the signal during transmission and processing to ensure that the loudness of the signal is moderate. Multiply the second signal by a preset gain factor to obtain the signal after gain adjustment. The value of the gain factor can be set according to actual needs to meet different application scenarios. Use a filter bank to adjust the frequency of the signal and optimize the frequency characteristics of the signal. The frequency adjustment can further enhance the specific frequency components of the target speech signal while suppressing the frequency components of the noise signal, improving the clarity and intelligibility of the signal. Pass the signal after gain adjustment through a group of filters to process the signals in different frequency ranges. For example, a band-pass filter can be used to enhance the main frequency range of the speech signal (such as 300 Hz - 3400 Hz) while suppressing the low-frequency and high-frequency noise components. The signal after being processed by the filter bank is the first microphone enhanced signal.

[0033] Collect the audio signal through the first microphone array of the auxiliary device. By using multiple microphones for phase alignment, the target speech signal can be effectively enhanced while suppressing the noise signals from other directions. Phase alignment can ensure that the same speech signal received by different microphones is consistent in phase, so that the target signal can be effectively enhanced when weighted and superimposed. The effects of this signal enhancement and noise suppression can significantly improve the signal-to-noise ratio of the audio signal, making the speech signal clearer and more intelligible. Determine the corresponding weights according to the spatial positions and signal intensities of multiple microphones, and perform weighted superposition, which can improve the spatial resolution of the system. This way of weighted superposition can better utilize the spatial characteristics of the microphone array to distinguish and process the sound signals from different directions, so as to more accurately capture the target speech signal while suppressing the interference signals from other directions. Adjust the amplitude of the second signal by presetting the gain factor to ensure that the loudness of the signal is appropriate and avoid the signal being too strong or too weak. Adjust the frequency of the signal by using a filter bank to further optimize the spectral characteristics of the signal, remove unnecessary noise components, and enhance the frequency components of the target speech signal. This optimization of amplitude and frequency can make the generated first microphone enhanced signal clearer and more natural, improving the overall quality of the audio signal. The first microphone enhanced signal generated through the above steps can more accurately reflect the speech content at the meeting site, reduce the interference of noise, and improve the clarity and quality of the audio signal. This is of great significance for subsequent signal processing and transmission, which can ensure that the main device can receive a more accurate and clearer audio signal, so as to better perform operations such as generating the integrated microphone signal and echo cancellation, noise reduction, and gain processing, and finally obtain a higher-quality uplink enhanced signal, improving the audio quality and communication effect of the meeting.

[0034] Optionally, the step of performing phase alignment on the multiple audio signals collected by the multiple first microphones of the first microphone array to obtain multiple first signals, determining the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and performing weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal includes: Perform frame division processing on the audio signal to obtain multiple sub-frames. For each sub-frame, calculate the cross-correlation function between different first microphones to obtain multiple groups of time delay differences, and each group of time delay differences includes the time delay difference between two different first microphones; Perform phase correction on the multiple audio signals according to the multiple groups of time delay differences to obtain multiple first signals; Determine the direction of the sound source according to the multiple groups of time delay differences, and calculate the distance between the sound source and each first microphone according to the direction, the geometric structure of the first microphone array, and the sound wave propagation speed; Correspondingly assign weights according to the distance, where the closer the distance, the greater the weight of the first signal, and perform weighted superposition on the multiple first signals according to the weights to obtain a second signal.

[0035] Perform frame division processing on the audio signal to obtain multiple sub-frames. For each sub-frame, calculate the cross-correlation function between different first microphones, so as to obtain multiple groups of time delay differences, and each group of time delay differences includes the time delay differences between two different first microphones. The purpose of this step is to determine the time difference when the same voice signal is received by different microphones, providing a basis for subsequent phase correction. By calculating the cross-correlation function, the time delay between different microphone signals can be accurately found, thus achieving precise phase alignment. Perform phase correction on the multiple audio signals according to the multiple groups of time delay differences to obtain multiple first signals. Phase correction is to ensure that the same voice signal received by different microphones is consistent in phase, so that the target voice signal can be effectively enhanced during subsequent weighted superposition. Through phase correction, the phase deviation caused by the difference in the sound wave propagation time can be eliminated, enabling the signals of different microphones to be synchronously superimposed, improving the coherence and enhancement effect of the signals. Determine the direction of the sound source according to the multiple groups of time delay differences. By analyzing the time delay differences between different microphones, the position and direction of the sound source relative to the microphone array can be deduced. Combining the sound wave propagation speed and the geometric structure of the microphone array, calculate the distance between the sound source and each first microphone. The purpose of this step is to reasonably assign weights according to the position and distance of the sound source, so that the microphone signals closer to the sound source have a greater proportion in the weighted superposition, thereby better enhancing the target voice signal and suppressing noise signals from other directions at the same time. Correspondingly assign weights according to the distance between the sound source and each first microphone, where the closer the distance, the greater the weight of the first signal. The weight assignment principle is based on the reciprocal of the sound source distance or an exponential function, etc., so that the microphone signals closer to the sound source have a greater contribution in the weighted superposition. According to the assigned weights, perform weighted superposition on the multiple first signals to obtain a second signal. Weighted superposition can effectively enhance the target voice signal, suppress noise signals from other directions at the same time, and improve the signal-to-noise ratio and clarity of the audio signal. The second signal generated in this way can more accurately reflect the voice content at the meeting site, providing a high-quality audio signal for subsequent signal processing and transmission.

[0036] The audio signal is framed to obtain multiple sub-frames, and then the cross-correlation function between different microphones is calculated to obtain multiple groups of time-delay differences. This precise phase alignment method can ensure that the same speech signal received by different microphones is consistent in phase, so that the target speech signal can be effectively enhanced during the subsequent weighted superposition process, improving the clarity and intelligibility of the signal. Through phase alignment and weighted superposition, noise signals from other directions can be effectively suppressed. Since the phase of the noise signals received by different microphones is inconsistent, they will cancel each other out or weaken during the weighted superposition process, thereby improving the signal-to-noise ratio of the audio signal and making the target speech signal clearer. The direction of the sound source is determined based on multiple groups of time-delay differences, and the distance between the sound source and each microphone is calculated in combination with the geometric structure of the microphone array and the sound wave propagation speed. This sound source localization method can accurately determine the position of the sound source, providing important reference information for subsequent signal processing and helping to improve the spatial resolution and directivity of the system. Weights are assigned according to the distance between the sound source and each microphone, and the closer the distance, the greater the weight of the first signal. This weight assignment method can better highlight the target speech signal while suppressing noise signals from other directions and improving the quality of the audio signal. The second signal generated through the above steps is adjusted in amplitude by a preset gain factor and adjusted in frequency using a filter bank, which can further optimize the spectral characteristics of the audio signal, making it flatter, reducing spectral distortion, and at the same time enhancing the frequency components of the target speech signal to generate a high-quality first microphone enhanced signal.

[0037] Optionally, the weighted superposition of the received first microphone enhanced signal and the second microphone enhanced signal to obtain a combined microphone signal includes: Aligning a single first microphone enhanced signal and the second microphone enhanced signal to obtain an adjusted first microphone enhanced signal and an adjusted second microphone enhanced signal; Determining the spatial position weight corresponding to each target microphone array according to the spatial position relationship between the single first microphone array and the second microphone array, where the target microphone array is the first microphone array or the second microphone array; Performing weighted superposition on the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal according to the spatial position weight.

[0038] When there is only one auxiliary device, align the received first microphone enhanced signal and second microphone enhanced signal to obtain two aligned signals (i.e., the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal). Since the two microphones may collect the same voice signal at different times, it is necessary to perform temporal alignment, calculate the time difference (delay) between the two signals, and then adjust one or both signals according to the calculated delay so that the same voice component is consistent in time. According to the spatial position relationship between the first microphone array and the second microphone array (the position relationship between the microphone arrays can be determined based on the geometric center), determine the spatial position weight corresponding to each target microphone array (the first microphone array or the second microphone array). The determination of the weight is based on various factors, such as the distance between the target microphone array and the sound source, the relative position between the target microphone arrays, etc. According to the determined spatial position weight, perform weighted superposition on the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal. The purpose of weighted superposition is to fuse the signals collected by the two microphones. Weighted superposition can combine these two signals to form a composite signal, while highlighting the target voice signal and suppressing the noise signals from other directions. The composite microphone signal obtained through weighted superposition has a higher signal-to-noise ratio and better voice quality, can more accurately reflect the voice content at the meeting site, and provides a better signal basis for subsequent operations such as echo cancellation, noise reduction, and gain processing.

[0039] Optionally, the weighted superposition of the received first microphone enhanced signal and the second microphone enhanced signal to obtain a composite microphone signal includes: Align multiple first microphone enhanced signals and second microphone enhanced signals to obtain multiple third signals; According to the spatial position relationship between multiple first microphone arrays and the second microphone array, determine the spatial position weight corresponding to each target microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on the multiple third signals according to the spatial position weight.

[0040] Align the received multiple first microphone enhanced signals and second microphone enhanced signals to obtain multiple third signals. The adjusted multiple first microphone enhanced signals and the adjusted second microphone enhanced signals are both referred to as third signals. Since microphones of different devices may capture the same speech signal at different times, it is necessary to align these signals in time. Through the alignment operation, it can be ensured that the same speech components in different signals are consistent in time, thus providing an accurate signal basis for subsequent weighted superposition. Determine the spatial position weights corresponding to each target microphone array according to the spatial position relationship between the multiple first microphone arrays and the second microphone arrays. The target microphone array is any one of the first microphone arrays and the second microphone arrays. The determination of the spatial position weights is usually based on factors such as the distance between the target microphone array and the sound source, and the relative position between the target microphone arrays. The target microphone array closer to the sound source usually has a higher weight because the target speech components in the signals it captures are relatively stronger and the noise components are relatively weaker. The weights can be determined by calculating with a preset algorithm or model, or adjusted according to the actual measurement results. Determine the position relationship between the first microphone array and the second microphone array according to the single microphone or multiple microphones of the first microphone array and the single microphone or multiple microphones of the second microphone array. If there is only one microphone in the first microphone array or the second microphone array, its position can be simply represented as a point coordinate. If there are multiple microphones in the first microphone array or the second microphone array, then the geometric arrangement of these microphones needs to be considered. For example, a linear array may have multiple microphones arranged along a straight line, and a circular array may have multiple microphones arranged around a center point. The coordinates of the center of the array can be used as the position of the first microphone array or the second microphone array. According to the determined spatial position weights, perform weighted superposition on the multiple third signals. The purpose of weighted superposition is to fuse the signals captured by different microphones, while highlighting the target speech signal and suppressing the noise signals from other directions. The synthesized microphone signal obtained through weighted superposition has a higher signal-to-noise ratio and better speech quality, can more accurately reflect the speech content at the meeting site, and provides a better signal basis for subsequent operations such as echo cancellation, noise reduction, and gain processing.

[0041] Align the first microphone enhanced signal and the second microphone enhanced signal to obtain a plurality of third signals. This process ensures the temporal consistency of signals from different microphone arrays, thus providing an accurate basis for subsequent signal superposition. By alignment, the time delay between different signals can be eliminated, avoiding signal interference or distortion caused by time differences, and improving signal synchronization and consistency. Determine the spatial position weights corresponding to each microphone according to the spatial position relationship between the first microphone of the first microphone array and the second microphone of the second microphone array. This process takes into account the distribution and relative positions of microphones in space, making the weight assignment more reasonable and scientific. Microphone signals closer to the sound source have larger weights, which can better highlight the target speech signal, while microphone signals farther away have smaller weights, helping to suppress noise signals from other directions. Perform weighted superposition on the plurality of third signals according to the spatial position weights to obtain a combined microphone signal. This process effectively enhances the target speech signal while suppressing noise signals through weighted superposition. Weighted superposition can make full use of the signal information collected by different microphones, improve the signal-to-noise ratio and clarity, and make the combined microphone signal more accurately reflect the speech content at the meeting site. Through dynamic time alignment and spatial position weight superposition of the first microphone enhanced signal and the second microphone enhanced signal, the finally obtained combined microphone signal has higher quality and clarity. This process can effectively eliminate multipath interference, improve signal stability and reliability, provide a better signal basis for subsequent echo cancellation, noise reduction and gain processing, and thus improve the audio performance of the entire system. It is possible to flexibly adjust the configuration and weight assignment of the microphone array according to different meeting scenarios and requirements. Through dynamic time alignment and spatial position weight superposition, the system can adapt to different acoustic environments and meeting layouts, improving the adaptability and flexibility of the system. This flexibility enables the system to provide high-quality audio signals in various complex meeting environments, meeting the needs of different users.

[0042] Optionally, the obtaining of the uplink enhanced signal by performing echo cancellation, noise reduction, and gain processing on the downlink signal and the combined microphone signal includes: Construct an adaptive filter based on the downlink signal, calculate the coefficients of the adaptive filter through the normalized least mean square algorithm, and generate an estimated echo signal coupled to the sound field played by the main device speaker; Perform time-domain cancellation on the combined microphone signal and the estimated echo signal to obtain an error signal; Dynamically adjust the step factor in the normalized least mean square algorithm according to the ratio of the power spectral density of the error signal to the power spectral density of the combined microphone signal to optimize the error signal; Perform noise reduction processing on the optimized error signal. Use a deep learning-based noise reduction model to suppress the environmental noise in the signal and obtain a noise-reduced signal; Adjust the amplitude of the noise-reduced signal according to a preset gain curve to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0043] An adaptive filter is a digital filter that can automatically adjust its parameters according to the input signal. In this method, an adaptive filter based on the downlink signal is constructed, and the coefficients of the filter are calculated by the normalized least mean square (NLMS) algorithm. The NLMS algorithm is a commonly used adaptive filtering algorithm, and its core idea is to adjust the filter coefficients by minimizing the mean square error between the output signal and the desired signal. The weight update formula of the NLMS algorithm is: W(n + 1)=W(n)+μ·e(n)·x(n); where, W(n) represents the filter weight vector at the nth iteration, μ represents the step size factor (convergence factor), e(n) represents the instantaneous error, and x(n) represents the input signal vector (downlink signal).

[0044] The adaptive filter continuously updates the coefficients of the filter through an iterative process to achieve the best filtering effect. The error between the output signal of the filter and the desired signal is adjusted by the NLMS algorithm, so that the filter can track and eliminate echoes in real time. The combined microphone signal and the estimated echo signal are cancelled in the time domain to obtain the error signal e(n). The calculation formula of the error signal is: e(n)=d(n)-y(n); where, d(n) represents the combined microphone signal, and y(n) represents the output signal of the filter (estimated echo signal).

[0045] The error signal e(n) represents the part of the combined microphone signal that has not been eliminated by the estimated echo signal, that is, the combination of the target speech signal and the remaining echo signal. Through the error signal, the coefficients of the filter can be further optimized to improve the echo cancellation effect. Calculate the power spectral density of the error signal e(n) and the combined microphone signal d(n). The power spectral density reflects the energy distribution of the signal at different frequencies. According to the ratio of the power spectral density of the error signal and the combined microphone signal, dynamically adjust the step size factor μ in the NLMS algorithm. The adjustment formula of the step size factor is: ; where, μ0 represents the initial step size factor, α represents the adjustment coefficient, Pe(n) represents the power spectral density of the error signal, and Pd(n) represents the power spectral density of the combined microphone signal.

[0046] By dynamically adjusting the step size factor, the convergence speed and stability of the filter can be improved, thereby optimizing the error signal to more accurately reflect the target speech signal.

[0047] The signal after gain processing is the uplink enhanced signal x up (n): x up (n) = Gain(DNN(e(n)), Gain Curve); Among them, DNN(e(n)) represents the signal after being processed by the deep learning noise reduction model, and Gain(·, GainCurve) represents the gain processing according to the gain curve.

[0048] A deep learning-based noise reduction model is used to perform noise reduction processing on the optimized error signal. The deep learning noise reduction model is trained with a large amount of noisy speech data and corresponding clean speech data, and can effectively identify and suppress environmental noise. The noise reduction model generates a noise-reduced signal through feature extraction and signal processing. The model can adopt deep learning architectures such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to suppress the noise components in the signal and retain the target speech signal. According to the requirements of the meeting scenario and the user's preferences, a gain curve is preset. The gain curve defines the adjustment ratio of the signal amplitude in different frequency ranges. The noise-reduced signal is frequency segmented according to the preset gain curve, and the signal amplitude of each frequency segment is adjusted accordingly. Through gain processing, the loudness of the target speech signal can be enhanced, making it clearer and more audible in the meeting. The signal after gain processing is the uplink enhanced signal, which has high speech quality and appropriate loudness and can better meet the needs of the meeting terminal.

[0049] By constructing an adaptive filter based on the downlink signal and using the normalized least mean square (NLMS) algorithm to calculate the filter coefficients, an estimated echo signal coupled with the sound field played by the master device's speaker can be generated. Canceling the estimated echo signal from the combined microphone signal in the time domain can effectively eliminate the echo in the meeting and improve the clarity and intelligibility of the speech signal. Dynamically adjusting the step size factor in the NLMS algorithm according to the ratio of the power spectral density of the error signal to the power spectral density of the combined microphone signal can optimize the performance of the filter. This dynamic adjustment method can improve the convergence speed and stability of the filter, enabling it to better adapt to different signal environments and further enhancing the echo cancellation effect. Using a deep learning-based noise reduction model to process the optimized error signal can effectively suppress the environmental noise in the signal. The deep learning model is trained with a large amount of noisy speech data and corresponding clean speech data, and can accurately identify and remove the noise components while retaining the target speech signal, thereby improving the purity and quality of the speech signal. Adjusting the amplitude of the noise-reduced signal according to a preset gain curve can optimize the loudness of the speech signal. The gain curve can be set according to the requirements of the meeting scenario and the user's preference to ensure that the amplitude of the speech signal is appropriate within different frequency ranges, enhancing the clarity and intelligibility of the target speech signal and making it clearer and more audible in the meeting. By performing echo cancellation, noise reduction, and gain processing on the downlink signal and the combined microphone signal, the final uplink enhanced signal has higher speech quality and clarity. This process can effectively eliminate the echo and environmental noise, and at the same time enhance the loudness of the target speech signal through gain processing, improving the communication effect of the meeting and the auditory experience of the participants, meeting the requirements of the meeting terminal for high-quality audio signals.

[0050] Optionally, the adjusting the amplitude of the noise-reduced signal according to a preset gain curve to obtain an uplink enhanced signal includes: Setting a gain curve according to the requirements of the meeting scenario and the user's preference, where the gain curve defines the adjustment ratio of the signal amplitude within different frequency ranges; Segmenting the noise-reduced signal into frequency bands according to the gain curve and correspondingly adjusting the signal amplitude of each frequency band to obtain an uplink enhanced signal.

[0051] The gain curve is a pre-set function used to define the adjustment ratio of the signal amplitude in different frequency ranges. It can be set according to the requirements of the meeting scenario and the user's preferences to optimize the transmission and reception effects of the voice signal. In different meeting scenarios, there are different requirements for the clarity, loudness, and sound quality of the voice signal. For example, in a large meeting room, higher loudness and clearer voice signals may be required, while in a small meeting room or a quiet environment, more attention may be paid to the naturalness and comfort of the voice signal. Users can adjust the settings of the gain curve according to their personal preferences and usage habits. For example, some users may prefer clear and bright voice signals, while some users may prefer soft and natural voice signals. The noise-reduced signal is frequency-segmented according to the gain curve, that is, the signal is decomposed into sub-band signals in multiple different frequency ranges. Each sub-band signal corresponds to a frequency segment in the gain curve. The frequency segmentation can be achieved through a filter bank, which decomposes the signal into multiple sub-band signals, and the frequency range of each sub-band signal corresponds to the frequency segment in the gain curve. According to the adjustment ratio defined in the gain curve, the signal amplitude of each frequency segment is adjusted accordingly. Specifically, the amplitude of each sub-band signal is multiplied by the corresponding gain value in the gain curve to achieve the adjustment of the signal amplitude. Through the amplitude adjustment, the signal in the target frequency segment can be enhanced, and the signal in the unwanted frequency segment can be suppressed. For example, the signal in the mid-frequency band can be enhanced to improve the clarity and intelligibility of the voice signal, while the signals in the low-frequency band and high-frequency band can be suppressed to reduce noise and interference. The signal after frequency segmentation and amplitude adjustment is the uplink enhanced signal. The uplink enhanced signal has higher voice quality and appropriate loudness, and can better meet the needs of the meeting terminal. The uplink enhanced signal is sent to the meeting terminal, and the meeting terminal can receive and play the voice signal more clearly, improving the communication effect of the meeting and the auditory experience of the participants.

[0052] Setting the gain curve according to the requirements of the meeting scenario and the user's preferences can meet the audio optimization requirements in different scenarios. By presetting the gain curve, the system can automatically adjust the signal amplitude in different frequency ranges to adapt to the specific meeting scenario and the user's auditory preferences, thus providing a more personalized audio experience. Segmenting the frequency of the noise-reduced signal according to the gain curve and making corresponding adjustments to the signal amplitude of each frequency segment can achieve fine control of the audio signal. This method of segmented adjustment can optimize the signal characteristics in different frequency ranges, effectively improve the overall quality of the audio signal, and make it achieve the best auditory effect in different frequency ranges. Adjusting the signal amplitude of each frequency segment according to the gain curve can ensure that the signal loudness is moderate, avoiding the signal being too strong or too weak. By adjusting the gain curve, the loudness of the target voice signal can be effectively enhanced, making it clearer and more audible in the meeting. This is very important for voice communication in the meeting because clear voice signals can improve the communication efficiency of the meeting and the satisfaction of the participants. By enhancing the target voice signal, it can be ensured that every participant in the meeting can clearly hear the voice of the speaker, thus improving the overall effect of the meeting. By adjusting the amplitude of the noise-reduced signal, the finally obtained uplink enhanced signal has higher voice quality and clarity. This signal optimization method can effectively improve the overall quality of the meeting audio, making it more suitable for the needs of the meeting terminal. High-quality audio signals can improve the communication effect of the meeting, reduce misunderstandings and repeated communications caused by audio quality problems, and improve the efficiency and professionalism of the meeting.

[0053] This embodiment also discloses a signal optimization system for a wireless cascaded speaker. Figure 3 It is a schematic diagram of the modules of the signal optimization system for the wireless cascaded speaker disclosed in the embodiments of the present application. As Figure 3 shown, the system includes a distribution module: 301, a transmission module 302, a superimposition module 303, and an execution module 304, where: The distribution module: 301 is configured to, when the target device is started, allocate the target device as the master device or the slave device according to the state when it was last shut down or the received control signal; The transmission module 302 is configured to collect an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal, and transmit the first microphone enhanced signal to the master device by means of wireless coding transmission; The superimposition module 303 is configured to collect an audio signal through the second microphone array of the master device to generate a second microphone enhanced signal, and perform weighted superimposition on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; The execution module 304 is configured to receive a downlink signal from a connected conference terminal through a master device, perform echo cancellation, noise reduction, and gain processing on the downlink signal and the integrated microphone signal to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0054] Optionally, the sending module 302 is configured to: Perform a fast Fourier transform on the audio signal collected by a single first microphone of the first microphone array to obtain the spectrum of the signal, extract target features related to noise from the spectrum, classify the noise components in the audio signal according to the target features, where the target features include amplitude spectrum, phase spectrum, or spectral kurtosis; Process the noise components in the audio signal using a corresponding noise suppression method according to the classification result, and convert the processed spectrum back to the time domain through an inverse fast Fourier transform to obtain a speech signal in the time domain; Perform post-processing on the speech signal to generate a first microphone enhanced signal.

[0055] Optionally, the sending module 302 is configured to: Perform phase alignment on the multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determine corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and perform weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal; Perform enhancement processing on the second signal to generate a first microphone enhanced signal.

[0056] Optionally, the sending module 302 is configured to: Perform frame division processing on the audio signal to obtain multiple sub-frames. For each sub-frame, calculate the cross-correlation function between different first microphones to obtain multiple groups of time delay differences, and each group of time delay differences includes the time delay difference between two different first microphones; Perform phase correction on the multiple audio signals according to the multiple groups of time delay differences to obtain multiple first signals; Determine the direction of the sound source according to the multiple groups of time delay differences, and calculate the distance between the sound source and each first microphone according to the direction, the geometric structure of the first microphone array, and the sound wave propagation speed; Allocate corresponding weights according to the distance, the closer the distance, the greater the weight of the first signal, and perform weighted superposition on the multiple first signals according to the weights to obtain a second signal.

[0057] Optionally, the superimposing module 303 is configured to: Align the single first microphone enhanced signal and the second microphone enhanced signal to obtain an adjusted first microphone enhanced signal and an adjusted second microphone enhanced signal; Determine the spatial position weight corresponding to each target microphone array according to the spatial position relationship between the single first microphone array and the second microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal according to the spatial position weight.

[0058] Optionally, the superposition module 303 is configured to: Align multiple first microphone enhanced signals and the second microphone enhanced signal to obtain multiple third signals; Determine the spatial position weight corresponding to each target microphone array according to the spatial position relationship between multiple first microphone arrays and the second microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on multiple third signals according to the spatial position weight.

[0059] Optionally, the execution module 304 is configured to: Construct an adaptive filter based on the downlink signal, calculate the coefficients of the adaptive filter through the normalized least mean square algorithm, and generate an estimated echo signal coupled to the sound field played by the master device speaker; Perform time-domain cancellation on the combined microphone signal and the estimated echo signal to obtain an error signal; Dynamically adjust the step factor in the normalized least mean square algorithm according to the ratio of the power spectral density of the error signal to the power spectral density of the combined microphone signal to optimize the error signal; Perform noise reduction processing on the optimized error signal, use a deep learning-based noise reduction model to suppress environmental noise in the signal, and obtain a noise-reduced signal; Adjust the amplitude of the noise-reduced signal according to a preset gain curve to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

[0060] It should be noted that: when the device provided in the above embodiment realizes its functions, only the division of the above-mentioned functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0061] This embodiment also discloses an electronic device. Referring to Figure 4 , the electronic device may include: at least one processor 401, at least one communication bus 402, a user interface 403, a network interface 404, and at least one memory 405.

[0062] Among them, the communication bus 402 is used to realize the connection and communication between these components.

[0063] Among them, the user interface 403 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 403 may further include a standard wired interface and a wireless interface.

[0064] Among them, the network interface 404 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0065] Among them, the processor 401 may include one or more processing cores. The processor 401 connects various parts within the entire server through various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 405, and by calling data stored in the memory 405, it executes various functions of the server and processes data. Optionally, the processor 401 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 401 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content that needs to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 401 and may be implemented separately through a single chip.

[0066] Among them, the memory 405 may include a Random Access Memory (RAM), or may include a Read-Only Memory. Optionally, the memory 405 includes a non-transitory computer-readable storage medium. The memory 405 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 405 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments. Optionally, the memory 405 may also be at least one storage device located far from the aforementioned processor 401. As Figure 4 shown, the memory 405, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program of the signal optimization method for a wireless cascading speaker.

[0067] In Figure 4 the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 401 can be used to call the application program of the signal optimization method for a wireless cascading speaker stored in the memory 405. When executed by one or more processors 401, the electronic device is caused to execute the method of one or more of the above-mentioned embodiments.

[0068] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0069] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0070] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of devices or units can be in electrical or other forms.

[0071] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0072] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0073] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned memory 405 includes: various media such as USB flash drives, mobile hard disks, magnetic disks or optical discs that can store program codes.

[0074] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made according to the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the disclosure of the specification. The present application aims to cover any variations, uses or adaptive changes of the present disclosure, and these variations, uses or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A signal optimization method for a wireless cascaded speaker, characterized in that, Applied to a wireless cascaded speaker, the wireless cascaded speaker includes a master device and a slave device, and the method includes: When the target device is started, allocate the target device as the master device or the slave device according to the state at the last shutdown or the received control signal; Collect an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal, and send the first microphone enhanced signal to the master device by means of wireless coding transmission; Collect an audio signal through the second microphone array of the master device to generate a second microphone enhanced signal, and perform weighted superposition on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; Receive a downlink signal from the connected conference terminal through the master device, perform echo cancellation, noise reduction, and gain processing on the downlink signal and the comprehensive microphone signal to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

2. The signal optimization method of the wireless cascaded speaker according to claim 1, characterized in that The collecting an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal includes: Perform a fast Fourier transform on the audio signal collected by a single first microphone of the first microphone array to obtain the spectrum of the signal, extract the target features related to noise from the spectrum, classify the noise components in the audio signal according to the target features, and the target features include amplitude spectrum, phase spectrum or spectral kurtosis; Process the noise components in the audio signal by using the corresponding noise suppression method according to the classification result, and convert the processed spectrum back to the time domain through an inverse fast Fourier transform to obtain a speech signal in the time domain; Perform post-processing on the speech signal to generate a first microphone enhanced signal.

3. The signal optimization method of the wireless cascaded speaker according to claim 1, wherein, The collecting an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal includes: Perform phase alignment on multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determine the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and perform weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal; Perform enhancement processing on the second signal to generate a first microphone enhanced signal.

4. The signal optimization method of the wireless cascaded speaker according to claim 3, wherein The performing phase alignment on multiple audio signals collected by multiple first microphones of the first microphone array to obtain multiple first signals, determining the corresponding weights according to the spatial positions and signal intensities of the multiple first microphones, and performing weighted superposition on the multiple first signals according to the corresponding weights to obtain a second signal includes: Perform frame division processing on the audio signal to obtain multiple sub-frames. For each sub-frame, calculate the cross-correlation function between different first microphones to obtain multiple groups of time delay differences, and each group of time delay differences includes the time delay difference between two different first microphones; Perform phase correction on multiple audio signals according to multiple groups of time delay differences to obtain multiple first signals; Determine the direction of the sound source according to multiple groups of time delay differences, and calculate the distance between the sound source and each first microphone according to the direction, the geometric structure of the first microphone array, and the sound wave propagation speed; Allocate corresponding weights according to the distance. The closer the distance, the greater the weight of the first signal. Weight the multiple first signals according to the weights to obtain a second signal.

5. The signal optimization method of the wireless cascaded speaker according to claim 1, characterized in that, The weighted superposition of the received first microphone enhanced signal and the second microphone enhanced signal to obtain a combined microphone signal includes: Align a single first microphone enhanced signal and the second microphone enhanced signal to obtain an adjusted first microphone enhanced signal and an adjusted second microphone enhanced signal; According to the spatial position relationship between a single first microphone array and the second microphone array, determine the spatial position weight corresponding to each target microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on the adjusted first microphone enhanced signal and the adjusted second microphone enhanced signal according to the spatial position weight.

6. The signal optimization method of the wireless cascaded speaker according to claim 1, wherein The weighted superposition of the received first microphone enhanced signal and the second microphone enhanced signal to obtain a combined microphone signal includes: Align multiple first microphone enhanced signals and the second microphone enhanced signal to obtain multiple third signals; According to the spatial position relationship between multiple first microphone arrays and the second microphone array, determine the spatial position weight corresponding to each target microphone array, where the target microphone array is the first microphone array or the second microphone array; Perform weighted superposition on the multiple third signals according to the spatial position weight.

7. The signal optimization method of the wireless cascaded speaker according to claim 1, characterized in that The obtaining of the uplink enhanced signal by performing echo cancellation, noise reduction, and gain processing on the downlink signal and the combined microphone signal includes: Construct an adaptive filter based on the downlink signal, calculate the coefficients of the adaptive filter through the normalized least mean square algorithm, and generate an estimated echo signal coupled to the sound field played by the main device's speaker; Perform time-domain cancellation on the combined microphone signal and the estimated echo signal to obtain an error signal; Dynamically adjust the step factor in the normalized least mean square algorithm according to the ratio of the power spectral density of the error signal to the power spectral density of the combined microphone signal to optimize the error signal; Perform noise reduction processing on the optimized error signal, use a deep learning-based noise reduction model to suppress the environmental noise in the signal, and obtain a noise-reduced signal; Adjust the amplitude of the noise-reduced signal according to a preset gain curve to obtain an uplink enhanced signal, and send the uplink enhanced signal to the conference terminal.

8. A signal optimization system for a wireless cascaded speaker, characterized in that, It includes an allocation module, a sending module, a superposition module, and an execution module, where: The allocation module is configured to, when the target device is started, allocate the target device as the main device or the slave device according to the state at the last shutdown or the received control signal; The sending module is configured to collect an audio signal through the first microphone array of the slave device to generate a first microphone enhanced signal, and send the first microphone enhanced signal to the main device through wireless coding transmission; The superposition module is configured to collect an audio signal through a second microphone array of the master device to generate a second microphone enhanced signal, and perform weighted superposition on the received first microphone enhanced signal and the second microphone enhanced signal to obtain a comprehensive microphone signal; The execution module is configured to receive a downstream signal from a connected conference terminal through the master device, perform echo cancellation, noise reduction, and gain processing on the downstream signal and the comprehensive microphone signal to obtain an upstream enhanced signal, and send the upstream enhanced signal to the conference terminal.

9. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-7 is executed.