Local sound amplification method

By performing time-frequency domain feature extraction and human voice extraction model processing on the local sound reinforcement signal, acoustic factors are suppressed, thereby improving the clarity and audibility of the sound reinforcement signal.

WO2026077160A1PCT designated stage Publication Date: 2026-04-16ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/119626
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-08
Filing Date
2025-09-08
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

During local sound reinforcement, acoustic factors such as howling, ambient noise, room reverberation, and echo remnants severely affect the clarity and audibility of the sound.

Method used

By extracting features from the first time-frequency domain signal to be played and the subsequent second time-frequency domain signal, a time-frequency mask is generated using a human voice extraction model to suppress acoustic factors, and the signal is converted into a human voice time-frequency domain signal for local amplification.

Benefits of technology

It effectively suppresses feedback, ambient noise, room reverberation, and echo remnants, improving the clarity and audibility of local sound reinforcement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025119626_16042026_PF_FP_ABST
    Figure CN2025119626_16042026_PF_FP_ABST
Patent Text Reader

Abstract

A local sound amplification method and apparatus, a terminal device, a medium, and a product. The method may comprise: performing feature extraction on a first time-frequency domain signal to be played back, and obtaining a first time-frequency domain feature; performing feature extraction on a second time-frequency domain signal to be played back after the first time-frequency domain signal, and obtaining a second time-frequency domain feature; inputting the first time-frequency domain feature and the second time-frequency domain feature into a human voice extraction model, and obtaining a time-frequency mask outputted by the human voice extraction model, the time-frequency mask being set to represent the intensity of a human voice time-frequency domain signal at corresponding time and frequency points in the first time-frequency domain signal; on the basis of the time-frequency mask, suppressing acoustic factors in the first time-frequency domain signal, and obtaining a human voice time-frequency domain signal corresponding to the first time-frequency domain signal; and converting the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, and performing local sound amplification on the human voice time-domain signal.
Need to check novelty before this filing date? Find Prior Art

Description

Local sound reinforcement methods Cross-reference to related applications

[0001] This application claims priority to Chinese Patent Application No. 202411396077.5, filed with the Chinese Patent Office on October 8, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments of this application relate to, but are not limited to, the field of audio technology, and particularly to, but are not limited to, local sound reinforcement methods. Background Technology

[0003] Local sound reinforcement, a key technology for amplifying and propagating sound, is widely used in various scenarios such as education, conferences, and entertainment, including large classrooms, auditoriums, and karaoke rooms. Terminal devices capture sound source signals through microphones and then amplify and play the sound using speakers, enabling the sound to cover a wider area and enhancing the auditory experience. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims. According to a first aspect of any embodiment of this application, a local sound reinforcement method is provided, the method comprising: extracting features from a first time-frequency domain signal to be played to obtain first time-frequency domain features, wherein the first time-frequency domain features are feature data of the first time-frequency domain signal on multiple channels; extracting features from a second time-frequency domain signal to be played after the first time-frequency domain signal to obtain second time-frequency domain features, wherein the second time-frequency domain features are feature data of the second time-frequency domain signal on multiple channels; inputting the first time-frequency domain features and the second time-frequency domain features into a voice extraction model to obtain a time-frequency mask output by the voice extraction model, wherein the time-frequency mask is set to represent the intensity of the voice time-frequency domain signal at the corresponding time and frequency point in the first time-frequency domain signal; suppressing acoustic factors in the first time-frequency domain signal based on the time-frequency mask to obtain a voice time-frequency domain signal corresponding to the first time-frequency domain signal; wherein the acoustic factors include: howling, ambient noise, room reverberation, and echo remnant; converting the voice time-frequency domain signal corresponding to the first time-frequency domain signal into a voice time-domain signal, and performing local sound reinforcement on the voice time-domain signal.

[0005] According to a second aspect of any embodiment of this application, a local sound reinforcement device is provided, the device comprising: a first extraction module configured to extract features from a first time-frequency domain signal to be played, obtaining first time-frequency domain features, wherein the first time-frequency domain features are feature data of the first time-frequency domain signal in multiple channels; a second extraction module configured to extract features from a second time-frequency domain signal to be played after the first time-frequency domain signal, obtaining second time-frequency domain features, wherein the second time-frequency domain features are feature data of the second time-frequency domain signal in multiple channels; and a human voice extraction module configured to extract features from the first time-frequency domain features and the second time-frequency domain signal. The time-frequency domain features are input into the voice extraction model to obtain the time-frequency mask output by the voice extraction model. The time-frequency mask is set to represent the intensity of the human voice time-frequency domain signal at the corresponding time and frequency point in the first time-frequency domain signal. The factor suppression module is set to suppress acoustic factors in the first time-frequency domain signal based on the time-frequency mask to obtain the human voice time-frequency domain signal corresponding to the first video source signal. The acoustic factors include: howling, ambient noise, room reverberation, and echo remnant. The local sound reinforcement module is set to convert the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal and perform local sound reinforcement on the human voice time-domain signal.

[0006] According to a third aspect of any embodiment of this application, a terminal device is provided, comprising: one or more processors; and a memory configured to store executable instructions of the one or more processors; wherein the one or more processors implement the method described in any embodiment of this application by executing the executable instructions.

[0007] According to a fourth aspect of any embodiment of the present application, a computer-readable storage medium is provided having computer instructions stored thereon that, when executed by one or more processors, implement the method described in any of the embodiments of the present application described above.

[0008] According to a fifth aspect of any embodiment of this application, a computer program product is provided having a computer program / instructions stored thereon, which, when executed by one or more processors, implements the method described in any of the embodiments of this application described above.

[0009] The technical solution provided in this application can include the following beneficial effects: As can be seen from the above embodiments, by extracting features from the first time-frequency domain signal to be played, first time-frequency domain features are obtained; by extracting features from the second time-frequency domain signal to be played after the first time-frequency domain signal, second time-frequency domain features are obtained; the first time-frequency domain features and the second time-frequency domain features are input into the voice extraction model to obtain the time-frequency mask output by the voice extraction model; based on the time-frequency mask, acoustic factors in the first time-frequency domain signal are suppressed to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal; the human voice time-frequency domain signal corresponding to the first time-frequency domain signal is converted into a human voice time-frequency signal; and local amplification of the human voice time-frequency signal can simultaneously suppress howling, environmental noise, room reverberation, and echo remnants in the first time-frequency domain signal, making the processed human voice time-frequency signal closer to the user's original voice, thereby improving the clarity and audibility of local amplification.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Other aspects will become clear after reading and understanding the accompanying drawings and detailed description. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this application, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] Figure 1 is a flowchart illustrating a local sound reinforcement method according to an exemplary embodiment of this application.

[0013] Figure 2 is a schematic diagram illustrating a local sound reinforcement according to an exemplary embodiment of this application.

[0014] Figure 3 is a flowchart illustrating a training method for a human voice extraction model according to an exemplary embodiment of this application.

[0015] Figure 4 is a schematic diagram of the structure of a post-filtering module according to an exemplary embodiment of this application.

[0016] Figure 5 is a schematic diagram of the structure of a human voice extraction model according to an exemplary embodiment of this application.

[0017] Figure 6 is a flowchart illustrating another local sound reinforcement method according to an exemplary embodiment of this application.

[0018] Figure 7 is a schematic diagram of the structure of a terminal device according to an exemplary embodiment of this application.

[0019] Figure 8 is a block diagram of a local sound reinforcement device according to an exemplary embodiment of this application. Detailed Implementation

[0020] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0021] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0023] Currently, in local sound reinforcement processes, the digital signal processing system in the terminal equipment includes recording, amplification, and playback functions, which is equivalent to positive feedback of the signal. If the signal is directly amplified without processing, it will be continuously fed back and amplified, eventually resulting in howling.

[0024] In local sound reinforcement scenarios, in addition to the effects of feedback, there are usually unfavorable acoustic factors such as environmental noise, room reverberation, and echo remnants, which seriously affect the clarity and audibility of the sound.

[0025] This application proposes a local sound reinforcement method, and the following embodiments are provided to further illustrate this application.

[0026] Please refer to Figure 1, which is a flowchart illustrating a local sound reinforcement method according to an exemplary embodiment of this application. This local sound reinforcement method can be applied to terminal devices such as in-vehicle systems, smart speakers, mobile phones, computers, messaging devices, tablets, and personal digital assistants, and can be used in local sound reinforcement scenarios such as karaoke, classrooms, meetings, and performances.

[0027] As shown in Figure 1, the local sound reinforcement method may include steps 102 to 110.

[0028] Step 102: Extract features from the first time-frequency domain signal to be played to obtain the first time-frequency domain features.

[0029] In this step, the terminal device can use a microphone or other sound acquisition device to receive the first time-domain signal of the surrounding environment. Using methods such as Short Time Fourier Transform (STFT), Mel-Frequency Cepstral Coefficients (MFCC), and Constant-Q Transform (CQT), the first time-domain signal is converted to the time-frequency domain, resulting in the first time-frequency domain signal.

[0030] The terminal device extracts features from the first time-frequency domain signal by analyzing the spectrum and using deep learning models, thereby obtaining the first time-frequency domain features.

[0031] The first time-domain signal is a time-domain signal directly captured by the sound acquisition device. The time-domain signal is an audio waveform that changes over time and may include: a human voice image, ambient noise, and an echo signal. The echo signal is the sound played by the speaker again, captured by the microphone.

[0032] Voice mirroring is used to reflect the actual sound produced by a user in a specific room environment. Voice mirroring can include: the time domain signal of the voice and the room reverberation.

[0033] A human voice time-domain signal is the original sound signal directly emitted by the user, without significant alteration to the room environment. It can include the user's direct sound, i.e., the original human voice emitted directly by the user. It can also include early reverberation. Room reverberation is the later reverberation produced by the room environment on the original human voice.

[0034] The first time-frequency domain signal is used to describe the characteristics of the first time-domain signal when time and frequency information are considered simultaneously. In the time-frequency domain, the changes of the first time-domain signal at different time points and different frequencies can be observed, thereby providing a more comprehensive description of the signal characteristics.

[0035] For example, the frequency of the first time-domain signal changes over time. By using the first time-frequency domain signal, both aspects of information can be considered simultaneously, which helps to more accurately understand and analyze the time-frequency domain signal of human voice.

[0036] The first time-frequency domain signal may include a human voice time-frequency domain signal, which is a speech signal in the time-frequency domain and is used to describe the original sound signal directly emitted by the user in the time and frequency domains.

[0037] The first time-frequency domain feature is an audio feature that describes the characteristics of the first time-frequency domain signal. It is used to conduct in-depth analysis and processing of the human voice time-frequency domain signal in the first time-frequency domain signal. It can be the spectrum, power spectrum, filter bank (FBank) features, etc.

[0038] Step 104: Extract features from the second time-frequency domain signal played after the first time-frequency domain signal to obtain the second time-frequency domain features.

[0039] In this step, the terminal device uses methods such as short-time Fourier transform, Mel-frequency transform, and constant Q transform to transform the second time-domain signal to be played in the previous frame, thereby obtaining the second time-frequency domain signal in the time-frequency domain.

[0040] The terminal device extracts features from the second time-frequency domain signal by analyzing spectrograms, deep learning models, and other methods, thus obtaining the second time-frequency domain features.

[0041] The second time-domain signal is the time-domain signal of the loudspeaker or other sound reinforcement device in the terminal equipment that was to be played in the previous frame. It is used to suppress unfavorable acoustic factors in the first time-frequency domain signal. The second time-frequency domain signal is the time-frequency domain signal obtained by converting the second time-domain signal, and it is used to provide the distribution information of the second time-domain signal in the time-frequency domain.

[0042] The second time-frequency domain feature is an audio feature that describes the characteristics of the second time-frequency domain signal. It is used for in-depth analysis and processing of the second time-frequency domain signal and can be the spectrum, power spectrum, FBank features, etc.

[0043] In one embodiment, since the first time-domain signal includes an echo signal, the terminal device can employ an adaptive echo cancellation (AEC) algorithm to suppress linear echoes.

[0044] The terminal device uses an adaptive filter to adjust coefficients to simulate the echo path and generate a predicted echo path. Based on the second time-domain signal and the predicted echo path, the predicted echo path is convolved onto the second time-domain signal to calculate and eliminate linear echoes in the first time-domain signal. In this embodiment, convolving the predicted echo path onto the second time-domain signal means performing a convolution operation between the second time-domain signal and the predicted echo path.

[0045] The predicted echo path is the acoustic path from the speaker to the microphone.

[0046] For example, the terminal device can obtain the first time-domain signal after linear echo cancellation according to the following formula (1):

[0047] Where y represents the first time-domain signal after linear echo cancellation, and x represents the first time-domain signal captured by the microphone. This indicates the predicted echo path, r represents the second time-domain signal, and * indicates the convolution operation.

[0048] It is understood that the adaptive echo cancellation algorithm described in the embodiments of this application is only an example, and other methods can also be used to eliminate linear echoes in the first time domain signal. The embodiments of this application do not limit this.

[0049] As described above, by calculating and eliminating the linear echo in the first time domain signal based on the second time domain signal and the predicted echo path, the complexity of subsequent models in suppressing adverse acoustic factors can be reduced, thereby improving the overall efficiency and effectiveness of audio processing.

[0050] To further illustrate the local sound reinforcement process, Figure 2 shows a schematic diagram of a local sound reinforcement. As shown in Figure 2, the first time-domain signal x captured by the terminal device using a microphone includes: a human voice image s. img The ambient noise v and the echo signal e are obtained by convolving the second time-domain signal r with the unknown echo path b.

[0051] The terminal device convolves the second time-domain signal r to predict the echo path. Calculate the linear echo in the first time-domain signal x Eliminating linear echoes in the first time-domain signal x The first time-domain signal y after echo cancellation is obtained.

[0052] Step 106: Input the first time-frequency domain feature and the second time-frequency domain feature into the voice extraction model to obtain the time-frequency mask output by the voice extraction model. The time-frequency mask is used to represent the intensity of the voice time-frequency domain signal at the corresponding time and frequency points in the first time-frequency domain signal.

[0053] In this step, the terminal device inputs the first and second time-frequency domain features into the voice extraction model. The voice extraction model is a trained deep convolutional neural network (DCNN) model that can identify the human voice time-frequency domain signal in the first time-frequency domain signal and extract the time-frequency mask from the first time-frequency domain signal. The terminal device obtains the time-frequency mask output by the voice extraction model.

[0054] In this context, time-frequency masking is used to represent the intensity of the human voice time-frequency domain signal at the corresponding time and frequency point in the first time-frequency domain signal. Time-frequency masking is a two-dimensional matrix in which the value of each element represents the intensity or presence of the human voice time-frequency domain signal at the corresponding time and frequency point.

[0055] For example, the value of each element in the time-frequency mask can be normalized to between 0 and 1. Here, 1 represents a signal that belongs entirely to the time-frequency domain of human voice, and 0 represents a signal that does not belong entirely to the time-frequency domain of human voice.

[0056] Specifically, the terminal device can obtain the time-frequency mask according to the following formula (2):

[0057] Where m(k,τ) represents time-frequency masking, Let represent the first time-frequency domain feature and the second time-frequency domain feature, Λ represent the human voice extraction model operator, k represent the frequency band number, and τ represent the data frame number.

[0058] Step 108: Based on time-frequency masking, suppress acoustic factors in the first time-frequency domain signal to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal; acoustic factors include: howling, ambient noise, room reverberation and echo remnant.

[0059] In this step, the terminal device adjusts the first time-frequency domain signal based on time-frequency masking to suppress four acoustic factors in the first time-frequency domain signal: howling, ambient noise, room reverberation, and echo remnant, thereby obtaining the human voice time-frequency domain signal corresponding to the first time-frequency domain signal.

[0060] Acoustic factors are acoustic characteristics that negatively impact the clarity of local sound reinforcement. These factors can include: howling, ambient noise, room reverberation, and echo remnants. Echo remnants are the residual echo signals after eliminating linear echoes in the first time-domain signal, caused by factors such as discrepancies between predicted and actual echo paths and nonlinear echoes in the echo signal.

[0061] For example, the terminal device can obtain the human voice time-frequency domain signal according to the following formula (3): z(k,τ)=m(k,τ)y(k,τ) (3)

[0062] Where z(k,τ) represents the human voice time-frequency domain signal, m(k,τ) represents time-frequency masking, y(k,τ) represents the first time-frequency domain signal, k represents the frequency band number, and τ represents the data frame number.

[0063] Step 110: Convert the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, and amplify the human voice time-domain signal locally.

[0064] In this step, the terminal device uses methods such as Inverse Short-Time Fourier Transform (ISTFT), Overlap-Add Method, and Overlap-Save Method to convert the human voice time-frequency domain signal into a human voice time-domain signal. A loudspeaker is then used to amplify the human voice time-domain signal locally, achieving real-time local sound reinforcement.

[0065] Among them, the human voice time domain signal is the time domain signal that the loudspeaker is going to amplify and play in the current frame.

[0066] Please refer to Figure 2. For example, the post-filtering module 20 in the terminal device can obtain the human voice time-domain signal according to the following formula (4): z=H(y, r;θ) (4)

[0067] Where z represents the human voice time-domain signal, y represents the first time-domain signal after linear echo removal, r represents the second time-domain signal, H represents the post-filtering operator, and θ represents the human voice extraction model parameters.

[0068] Specifically, the post-filtering module 20 can convert and extract features from the second time-domain signal r and the echo-cancelled first time-domain signal y, respectively, to obtain first and second time-frequency domain features. The post-filtering module 20 inputs the first and second time-frequency domain features into the voice extraction model to obtain time-frequency masking.

[0069] The post-filtering module 20 applies time-frequency masking to the first time-frequency domain signal, suppressing four acoustic factors in the first time-frequency domain signal: howling, ambient noise, room reverberation, and echo remnants, to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal. The post-filtering module 20 then converts the human voice time-frequency domain signal into a human voice time-domain signal z.

[0070] In one embodiment, the terminal device can perform gain control and delay processing on the human voice time-domain signal to obtain a processed human voice time-domain signal. The processed human voice time-domain signal is superimposed with the background audio signal to obtain the target signal, and a loudspeaker is used to amplify the target signal locally.

[0071] The background audio signal can be a song accompaniment, background music, ambient sound effects, etc. The target signal is the audio signal to be amplified locally.

[0072] Please refer to Figure 2. Taking a local sound reinforcement scenario as an example (karaoke scenario), the terminal device can control the gain of the human voice time-domain signal z, where g represents the amplification gain. Due to the algorithm processing, recording, and playback operations, a certain signal delay occurs. The terminal device performs delay processing on the gain-controlled human voice time-domain signal to obtain the processed human voice time-domain signal, Z. -d This indicates that a delay was caused by d sampling points.

[0073] The terminal device will process the human voice time-domain signal and the background audio signal. The signal is then superimposed, adding new reverb and sound effects to obtain the target signal. A loudspeaker is then used to amplify the target signal locally.

[0074] As described above, by performing gain control and delay processing on the human voice time-domain signal, a processed human voice time-domain signal is obtained. The processed human voice time-domain signal is then superimposed with the background audio signal to obtain the target signal. By using a loudspeaker to amplify the target signal locally, the human voice time-domain signal and the background audio signal can be balanced in terms of volume and delay, thus achieving joint amplification of the human voice time-domain signal and the background audio signal.

[0075] The local sound reinforcement method in this embodiment extracts features from a first time-frequency domain signal to obtain first time-frequency domain features, extracts features from a second time-frequency domain signal to obtain second time-frequency domain features, inputs the first and second time-frequency domain features into a voice extraction model to obtain a time-frequency mask output by the voice extraction model, suppresses acoustic factors in the first time-frequency domain signal based on the time-frequency mask, obtains the corresponding human voice time-frequency domain signal, converts the human voice time-frequency domain signal into a human voice time-domain signal, and performs local sound reinforcement on the human voice time-domain signal. This can simultaneously suppress howling, environmental noise, room reverberation, and echo remnants in the first time-frequency domain signal, making the processed human voice time-domain signal closer to the user's original voice, thereby improving the clarity and audibility of the local sound reinforcement.

[0076] In the foregoing embodiments, it was described how to suppress howling, ambient noise, room reverberation, and echo remnants in the first time-frequency domain signal by using time-frequency masking extracted from the human voice extraction model based on first and second time-frequency domain features, thereby enabling local amplification of the clean human voice time-domain signal. In the following embodiments, the training process and structure of the human voice extraction model will be described in more detail, and these methods can be applied to any of the embodiments described above.

[0077] In one embodiment, the terminal device can obtain training data through data simulation, and use the training data and label data to train the initial model to obtain a trained human voice extraction model.

[0078] Specifically, the terminal device can superimpose the analog second time domain signal of the previous frame with the analog background audio signal to be amplified to obtain the superimposed analog second time domain signal.

[0079] For example, the terminal device can obtain the superimposed analog second time-domain signal according to the following formula (5):

[0080] Where r(τ-1) represents the analog second time-domain signal superimposed from the previous frame, and z(τ-1) represents the analog second time-domain signal from the previous frame. This represents the analog background audio signal, and τ represents the data frame number.

[0081] The terminal device can simulate the echo path of the superimposed analog second time-domain signal to obtain an analog echo signal. In sound reinforcement scenarios where no background audio signal needs to be added, the terminal device can also directly simulate the echo path of the previous frame's analog second time-domain signal to obtain an analog echo signal.

[0082] For example, the terminal device can obtain the analog echo signal according to the following formula (6): e(τ)=b*r(τ-1) (6)

[0083] Where e(τ) represents the analog echo signal, b represents the echo path, r(τ-1) represents the analog second time-domain signal superimposed from the previous frame, and τ represents the data frame number.

[0084] The terminal device can simulate a room function for a human voice source to obtain a simulated human voice image. This simulated human voice image can include: a simulated human voice time-domain signal and simulated room reverberation.

[0085] For example, the terminal device can obtain the simulated human voice image according to the following formula (7): s img (τ)=a early *s(τ)+a late *s(τ) (7)

[0086] Among them, s img (τ) represents the simulated human voice image, s(τ) represents the human voice source, and a early This represents the direct sound and early reverberation of the analog room transmission, a late a represents the late-stage simulated room reverberation of the simulated room transfer function. early *s(τ) represents the analog human voice time-domain signal, and τ represents the data frame number.

[0087] The terminal device can superimpose the analog echo signal, the analog human voice image, and the analog environmental noise to obtain the analog first time domain signal.

[0088] For example, the terminal device can obtain the analog first time-domain signal according to the following formula (8): x(τ)=e(τ-1)+s img (τ)+v(τ) (8)

[0089] Where x(τ) represents the simulated first time-domain signal, e(τ-1) represents the simulated echo signal, and s img (τ) represents the simulated human voice image, v(τ) represents the simulated environmental noise, and τ represents the data frame number.

[0090] The terminal equipment uses echo cancellation algorithms such as AEC to perform echo cancellation on the simulated first time domain signal to obtain an echo-cancelled signal.

[0091] Training data is generated based on the echo-cancelled signal and the simulated second time-domain signal from the previous frame (or the simulated second time-domain signal superimposed from the previous frame). Time-frequency domain transformation and feature extraction are performed on both the echo-cancelled signal and the simulated second time-domain signal from the previous frame (or the simulated second time-domain signal superimposed from the previous frame) to obtain simulated first time-frequency domain features and simulated second time-frequency domain features. These simulated first and second time-frequency domain features are used as training data. Since the echo-cancelled signal is obtained by echo cancellation of the simulated first time-domain signal, it is a time-domain signal; therefore, time-frequency domain transformation and feature extraction can be performed on the echo-cancelled signal to obtain the simulated first time-frequency domain features.

[0092] Based on the simulated human voice time-domain signal, label data corresponding to the training data is generated. The simulated human voice time-domain signal is then converted into the time-frequency domain to obtain the simulated human voice time-frequency domain signal, which is used as the corresponding label data.

[0093] The terminal device performs gain control and delay processing on the echo cancellation signal to obtain the simulated second time-domain signal of the current frame. By recursively simulating training data between different frames, and simultaneously simulating four adverse acoustic factors—feedback, ambient noise, room reverberation, and echo remnants—the model learns to process dynamically changing sound signals.

[0094] The simulated second time-domain signal is unaffected by four acoustic factors: howling, ambient noise, room reverberation, and echo remnants. The simulated echo signal is used to simulate the echoes generated in real-world environments due to sound from a speaker being reflected back to a microphone.

[0095] The human voice source is the raw, unprocessed, simulated real human voice. The simulated human voice image is the signal obtained by simulating the room transfer function (i.e., the frequency response characteristics of sound as it propagates within a room) of the human voice source, and is used to train the model to recognize and process simulated room reverberation.

[0096] The simulated human voice time-frequency domain signal is the portion of the simulated human voice image that comes directly from the human voice source, i.e., the pure human voice unaffected by room reverberation. Simulated room reverberation is used to simulate the sound effects produced by the reflection and superposition of the human voice source within a room.

[0097] Simulated environmental noise is the background noise that accompanies human voices in a real environment, and it is used to train models to learn how to extract human voices in noisy environments.

[0098] The simulated first time-domain signal is used to simulate the first time-domain signal actually received by the microphone. The echo-cancelled signal is the signal obtained after echo cancellation processing of the simulated first time-domain signal, and is used to train the model to identify and process echo remnants in the echo-cancelled signal.

[0099] It is understood that when generating training data and label data, other adverse acoustic factors can also be simulated, and the trained human voice extraction model can be used to predict time-frequency masking. Based on time-frequency masking, other acoustic factors in the first time-frequency domain signal can be suppressed. This application does not limit this.

[0100] As described above, by simulating the echo path of the simulated second time-domain signal of the previous frame (or the simulated second time-domain signal superimposed from the previous frame), a simulated echo signal is obtained. The simulated room transfer function of the human voice source is simulated to obtain a simulated human voice image. The simulated echo signal, the simulated human voice image, and the simulated ambient noise are superimposed to obtain a simulated first time-domain signal. The simulated first time-domain signal is then subjected to echo cancellation to obtain an echo cancellation signal. Based on the echo cancellation signal and the simulated second time-domain signal of the previous frame (or the simulated second time-domain signal superimposed from the previous frame), training data is generated. The label data corresponding to the training data is generated based on the simulated human voice time-domain signal. The echo cancellation signal is subjected to gain control and delay processing to obtain the simulated second time-domain signal of the current frame. Since the howling process is recursive and progressive, the training data adopts a recursive data simulation process to better simulate howling. The training data also contains simulated ambient noise, simulated room reverberation, and echo remnants to ensure that the training data is closer to the real sound reinforcement data, so that the trained human voice extraction model has a better suppression effect.

[0101] In one embodiment, Figure 3 shows a flowchart of a training method for a human voice extraction model. As shown in Figure 3, the training method may include steps 302 to 316.

[0102] Step 302: Initialize the initial model.

[0103] In this step, the terminal device selects or defines an initial neural network model, which may include the number of layers in the neural network model, the number of neurons in each layer, the activation function, etc. This initial model will serve as the starting point for training.

[0104] The terminal device allocates a portion of the available training data as a validation set. The validation set is used to evaluate the model's performance during training, but not for training the model itself, to prevent overfitting.

[0105] Step 304: Based on the training progress of the initial model, select training data of different difficulty levels.

[0106] In this step, the terminal device dynamically selects training data of varying difficulty based on the initial model's training progress. For example, simpler training data is used in the early stages of training to help the initial model quickly learn basic features, while the difficulty of the training data is gradually increased as training progresses.

[0107] The training progress can be determined based on the number of completed epochs. The difficulty of the training data is reflected in the signal-to-noise ratio of the training data. For example, the lower the signal-to-noise ratio of the training data, the greater the difficulty of the training data.

[0108] Step 306: Generate a small batch of training data.

[0109] In this step, the terminal device randomly selects a small batch of data from the training data for this training iteration. The "small batch" is a hyperparameter that may need to be adjusted based on the specific circumstances.

[0110] Step 308: Input the small batch of training data into the initial model to obtain the prediction results output by the initial model.

[0111] In this step, the terminal device inputs the small batch of training data generated in step 306 into the model and performs forward propagation to obtain the initial model's prediction results for each input data.

[0112] Step 310: Calculate the loss value based on the prediction results and label data, and update the parameters of the initial model.

[0113] In this step, the terminal device uses loss functions such as mean squared error and cross-entropy to calculate the loss value between the predicted result and the actual label data. Based on the loss value, the model parameters are updated using the backpropagation algorithm to reduce the loss value.

[0114] Repeat steps 306 to 310 until all the training data in the current small batch has been processed.

[0115] Step 312: Calculate the epoch loss of the initial model.

[0116] In this step, after the terminal device completes one epoch (i.e., after traversing all training data once), it calculates and records the loss for the entire epoch (e.g., the average loss value of the epoch) to evaluate the training effect of that epoch.

[0117] Step 314: Save the parameters and state of the initial model.

[0118] In this step, after each round, the terminal device saves the parameters and state of the current initial model as a checkpoint so that training can be resumed or further analysis can be performed later.

[0119] Step 316: Adjust the training progress of the initial model, select training data of the next difficulty level for model training, until the initial model meets the preset conditions, and obtain the human voice extraction model.

[0120] In this step, the terminal device adjusts the training progress of the initial model, selects training data of the next difficulty level, and continues to execute steps 304 to 316 until the initial model meets preset conditions such as the loss value being lower than a certain threshold and the accuracy of the validation set no longer improving, thus obtaining the final human voice extraction model.

[0121] In this embodiment, the voice extraction model is obtained by progressively training the model using training data of varying difficulty based on the initial model's training progress. As described above, by selecting training data of different difficulties based on the initial model's training progress, inputting the training data into the initial model, obtaining the prediction results output by the initial model, calculating the loss value based on the prediction results and label data, updating the parameters of the initial model, adjusting the initial model's training progress, selecting training data of the next difficulty for model training, until the initial model meets the preset conditions, thus obtaining the voice extraction model. Progressively advancing the training process using training data of varying difficulty helps the model gradually master basic knowledge and steadily improve its ability to handle complex tasks. While ensuring training effectiveness, this avoids the possibility of slow model learning due to directly using high-difficulty data, reducing unnecessary resource consumption.

[0122] In one embodiment, please refer to Figure 4, which shows a schematic diagram of a post-filtering module. The post-filtering module 20 may include: an STFT module 200, a feature extraction module 201, a human voice extraction model 202, a suppression module 203, and an ISTFT module 204.

[0123] The STFT module 200 can perform a short-time Fourier transform on the first time-domain signal received by the microphone to obtain a first time-frequency domain signal. The STFT module 200 can also perform a short-time Fourier transform on the second time-domain signal to be played by the speaker to obtain a second time-frequency domain signal.

[0124] The STFT module 200 inputs the first time-frequency domain signal to the suppression module 203. The STFT module 200 inputs the first time-frequency domain signal and the second time-frequency domain signal to the feature extraction module 201. The feature extraction module 201 extracts the first time-frequency domain feature corresponding to the first time-frequency domain signal and the second time-frequency domain feature corresponding to the second time-frequency domain signal.

[0125] The feature extraction module 201 inputs the first time-frequency domain features and the second time-frequency domain features into the voice extraction model 202, and the voice extraction model 202 inputs the extracted time-frequency mask into the suppression module 203.

[0126] The suppression module 203 applies time-frequency masking to the first time-frequency domain signal, suppressing acoustic factors in the first time-frequency domain signal to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal. The suppression module 203 inputs the human voice time-frequency domain signal to the ISTFT module 204.

[0127] The ISTFT module 204 performs a short-time inverse Fourier transform on the human voice time-frequency domain signal to obtain the human voice time-domain signal.

[0128] As described above, by using a microphone to receive a first time-domain signal, performing a short-time Fourier transform on the first time-domain signal to obtain a first time-frequency domain signal, performing a short-time Fourier transform on the second time-domain signal to be played by the speaker to obtain a second time-frequency domain signal, and performing an inverse short-time Fourier transform on the human voice time-frequency domain signal to obtain the human voice time-domain signal, the characteristics of the first and second time-frequency domain signals in both time and frequency dimensions are captured. Furthermore, the short-time Fourier transform and inverse short-time Fourier transform can improve the computational efficiency of converting from the time domain to the time-frequency domain and then back to the time domain, ensuring the real-time performance of local sound reinforcement.

[0129] In one embodiment, the terminal device can stitch together the first time-frequency domain features of multiple channels, such as stereo channels, through methods such as serialization and fusion to obtain the stitched first stitched feature. Then, it can stitch together the second time-frequency domain features of multiple channels to obtain the stitched second stitched feature.

[0130] Among them, the first time-frequency domain feature is the feature data of the first time-frequency domain signal in multiple channels, and the second time-frequency domain feature is the feature data of the second time-frequency domain signal in multiple channels.

[0131] Please refer to Figure 5, which shows a schematic diagram of a human voice extraction model. The human voice extraction model may include: a nonlinear transformation unit 50, a multi-layer convolutional unit 51, and a masking processing unit 52.

[0132] The nonlinear transformation unit 50 may include a fully connected layer (Affine) and a Rectified Linear Unit (ReLU) activation function. The masking processing unit 52 may include a fully connected layer (Affine) and a tanh activation function.

[0133] The multi-layer convolutional unit 51 may include a multi-layer feedforward sequential memory network (FSMN) unit, and each FSMN unit may include: a linear function 510, an FSMN module 511, and a ReLU (Affine) 512.

[0134] Linear function 510 is configured to receive input data and generate output through linear transformation. ReLU (Affine) 512 is configured to perform a nonlinear transformation on the output of FSMN module 511.

[0135] The FSMN module 511 consists of convolutional structures, which is equivalent to the frequency-band filtering operation in frequency domain signal processing algorithms, in concatenating features. Perform one-dimensional convolution on each feature dimension. Concatenate the features. It can refer to the first splicing feature and the second splicing feature.

[0136] Terminal devices will splice features The input is fed into the nonlinear transformation unit 50, which performs splicing on the features. A nonlinear transformation is performed to obtain nonlinear transformed data. This nonlinear transformed data is used to enhance the model's nonlinear expressive power, helping the model learn more complex audio feature representations.

[0137] The nonlinear transformation unit 50 inputs the nonlinear transformation data into the multi-layer convolutional unit 51. The multi-layer convolutional unit 51 performs convolution processing on the nonlinear transformation data, and builds a deeper network structure by stacking multiple convolutional layers to extract local features and reduce the spatial dimension of the features, thereby obtaining filtered data.

[0138] Multi-layer convolutional unit 51 inputs filtered data to masking unit 52, which calculates and outputs a time-frequency mask based on the filtered data. In the fully connected layer, each input node in masking unit 52 is connected to the output node via weights and a bias term is added.

[0139] Through forward propagation, the values ​​of each output node are calculated, corresponding to the real and imaginary parts of the time-frequency mask. An activation function is then applied to the output of the fully connected layer to obtain the final time-frequency mask composed of the real and imaginary parts.

[0140] The time-frequency masking term is in complex form, consisting of a real part and an imaginary part. The real and imaginary parts in the time-frequency masking term jointly describe the characteristics of the human voice signal in the time-frequency domain.

[0141] It is understood that the structure of the voice extraction model shown in Figure 5 is only an example, and other neural network structures can also be used, as long as the voice extraction model can implement the local sound reinforcement method described in the embodiments of this application. The embodiments of this application do not limit this.

[0142] As described above, by splicing the first time-frequency domain features of the first time-frequency domain signal in multiple channels, a spliced ​​first spliced ​​feature is obtained. Similarly, by splicing the second time-frequency domain features of the second time-frequency domain signal in multiple channels, a spliced ​​second spliced ​​feature is obtained. Multi-channel information fusion helps to more accurately capture the directionality, spatial distribution, and other characteristics of the sound source, thereby improving the effect of human voice separation and enhancement. The nonlinear transformation unit 50 performs nonlinear transformation on the first and second spliced ​​features to obtain nonlinear transformation data. The multi-layer convolution unit 51 performs convolution processing on the nonlinear transformation data to obtain filtered data, which can more accurately locate the position of the human voice time-frequency domain signal in the time-frequency domain and effectively suppress interference factors such as environmental noise and echo remnants. Based on the filtered data, the masking processing unit 52 calculates and outputs a time-frequency mask composed of real and imaginary parts. The time-frequency mask can more accurately control the preservation and suppression of the human voice time-frequency domain signal in the time-frequency domain, thereby improving the suppression effect of the four acoustic factors.

[0143] To further illustrate the local sound reinforcement process, Figure 6 shows a flowchart of another local sound reinforcement method. This local sound reinforcement method may include steps 602 to 620.

[0144] Step 602: Recursively simulate training data and label data.

[0145] In this step, the terminal device simulates the echo path of the simulated second time-domain signal (or the simulated second time-domain signal superimposed from the previous frame) of the previous frame to obtain a simulated echo signal. The simulated echo signal, the simulated human voice mirror image, and the simulated ambient noise are superimposed to obtain a simulated first time-domain signal. Echo cancellation processing is performed on the simulated first time-domain signal to obtain an echo-cancelled signal. Training data is generated based on the echo-cancelled signal and the simulated second time-domain signal (or the simulated second time-domain signal superimposed from the previous frame). Tag data is generated based on the simulated human voice time-domain signal in the simulated human voice mirror image.

[0146] Gain control and delay processing are applied to the echo cancellation signal to simulate its changes in the next frame, resulting in the simulated second time-domain signal of the current frame. Then, the simulated second time-domain signal of the current frame is used to obtain the simulated echo signal of the next frame, simulating howling, environmental noise, room reverberation, and echo remnants in a real environment, thus realizing recursive simulation training data.

[0147] Step 604: Use the training data and label data to train the initial model to obtain a trained human voice extraction model.

[0148] In this step, the terminal device selects training data of varying difficulty from the simulated training data based on the initial model's training progress. The training data is then input into the initial model to obtain its prediction results. The loss value is calculated based on the prediction results and the label data, and the parameters of the initial model are updated.

[0149] The terminal device adjusts the training progress of the initial model, selects training data of the next difficulty level for model training, until the preset conditions are met, and obtains a trained human voice extraction model.

[0150] Step 606: Eliminate the linear echo in the first time domain signal to obtain the first time domain signal after eliminating the linear echo.

[0151] In this step, the terminal device convolves the second time-domain signal with the predicted echo path, calculates the linear echo in the first time-domain signal, and eliminates the linear echo in the first time-domain signal.

[0152] Step 608: Perform short-time Fourier transform on the first time-domain signal and the second time-domain signal after eliminating linear echo.

[0153] In this step, the terminal device performs a short-time Fourier transform on the first time-domain signal and the second time-domain signal after eliminating linear echo, to obtain the first time-frequency domain signal and the second time-frequency domain signal.

[0154] Step 610: The terminal device extracts features from the first time-frequency domain signal to obtain the first time-frequency domain features.

[0155] Step 612: The terminal device extracts features from the second time-frequency domain signal to obtain the second time-frequency domain features.

[0156] Step 614: Input the first time-frequency domain features and the second time-frequency domain features into the voice extraction model to obtain the time-frequency mask output by the voice extraction model.

[0157] In this step, the terminal device splices the first time-frequency domain features to obtain the spliced ​​first spliced ​​feature, and splices the second time-frequency domain features to obtain the spliced ​​second spliced ​​feature.

[0158] In the human voice extraction model, the nonlinear transformation unit performs a nonlinear transformation on the first and second spliced ​​features to obtain nonlinear transformed data. A multi-layer convolutional unit performs convolution processing on the nonlinear transformed data to obtain filtered data. Based on the filtered data, the masking unit calculates and outputs a time-frequency mask.

[0159] Step 616: Based on time-frequency masking, suppress acoustic factors in the first time-frequency domain signal to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal. Convert the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, and perform gain control and delay processing on the human voice time-domain signal.

[0160] In this step, the terminal device applies time-frequency masking to the first time-frequency domain signal to suppress acoustic factors in the first time-frequency domain signal, thereby obtaining the human voice time-frequency domain signal.

[0161] Step 618: Superimpose the human voice time-domain signal after gain control and delay processing with the background audio signal to obtain the target signal.

[0162] In this step, the terminal device performs gain control and delay processing on the human voice time-domain signal, and then superimposes the human voice time-domain signal after gain control and delay processing with the background audio signal to obtain the target signal.

[0163] Step 620: Use a loudspeaker to amplify the target signal locally.

[0164] In this step, the terminal device uses a loudspeaker to amplify the target signal locally, achieving the local amplification function that simultaneously suppresses four acoustic factors: howling, ambient noise, room reverberation, and echo remnant.

[0165] Figure 7 is a schematic diagram of a terminal device according to an exemplary embodiment of this application. This terminal device may be, for example, a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, personal digital assistant, server, smart home appliance, in-vehicle system, etc. Referring to Figure 7, at the hardware level, the terminal device includes a processor 702, an internal bus 704, a network interface 706, memory 708, and non-volatile memory 710, and may also include other hardware required for services. The processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it, forming a local sound amplification device at the logical level. Of course, besides software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but may also be hardware or logic devices.

[0166] Figure 8 is a block diagram of a local sound reinforcement device according to an exemplary embodiment of this application. Referring to Figure 8, the device may include: a first extraction module 802, a second extraction module 804, a voice extraction module 806, a factor suppression module 808, and a local sound reinforcement module 810, wherein: the first extraction module 802 is configured to extract features from a first time-frequency domain signal to be played, obtaining first time-frequency domain features, wherein the first time-frequency domain features are feature data of the first time-frequency domain signal on multiple channels; the second extraction module 804 is configured to extract features from a second time-frequency domain signal to be played after the first time-frequency domain signal, obtaining second time-frequency domain features, wherein the second time-frequency domain features are feature data of the second time-frequency domain signal on multiple channels; the voice extraction module 806... The system is configured to input the first time-frequency domain features and the second time-frequency domain features into the voice extraction model to obtain the time-frequency mask output by the voice extraction model. The time-frequency mask is used to represent the intensity of the voice time-frequency domain signal at the corresponding time and frequency points in the first time-frequency domain signal. The factor suppression module 808 is configured to suppress acoustic factors in the first time-frequency domain signal based on the time-frequency mask to obtain the voice time-frequency domain signal corresponding to the first time-frequency domain signal. The acoustic factors include: howling, ambient noise, room reverberation, and echo remnant. The local sound reinforcement module 810 is configured to convert the voice time-frequency domain signal corresponding to the first time-frequency domain signal into a voice time-domain signal and perform local sound reinforcement on the voice time-domain signal.

[0167] In one example, the voice extraction module 806 is further configured to: train an initial model using training data and label data to obtain a trained voice extraction model; when generating training data and label data, the voice extraction module 806 is further configured to: simulate an echo path for the simulated second time-domain signal of the previous frame to obtain a simulated echo signal; simulate a room transfer function for the voice source to obtain a simulated voice image, the simulated voice image including: a simulated voice time-domain signal and simulated room reverberation; superimpose the simulated echo signal, the simulated voice image, and simulated ambient noise to obtain a simulated first time-domain signal; perform echo cancellation on the simulated first time-domain signal to obtain an echo cancellation signal; generate the training data based on the echo cancellation signal and the simulated second time-domain signal of the previous frame; generate the label data corresponding to the training data based on the simulated voice time-domain signal; and perform gain control and delay processing on the echo cancellation signal to obtain the simulated second time-domain signal of the current frame.

[0168] In one example, the voice extraction module 806 trains an initial model using training data and label data, configured to: select training data of different difficulty based on the training progress of the initial model; input the training data into the initial model to obtain the prediction result output by the initial model; calculate the loss value based on the prediction result and the label data, and update the parameters of the initial model; adjust the training progress of the initial model, select training data of the next difficulty for model training, until the initial model meets the preset conditions, and obtain the voice extraction model.

[0169] In one example, the local sound reinforcement module 810, when performing local sound reinforcement on the human voice time-domain signal, is configured to: perform gain control and delay processing on the human voice time-domain signal to obtain a processed human voice time-domain signal; superimpose the processed human voice time-domain signal with the background audio signal to obtain a target signal; and use a loudspeaker to perform local sound reinforcement on the target signal.

[0170] In one example, the first extraction module 802 is further configured to: receive a first time-domain signal using a microphone; perform a short-time Fourier transform on the first time-domain signal to obtain a first time-frequency domain signal; perform a short-time Fourier transform on the second time-domain signal to be played by the speaker to obtain a second time-frequency domain signal; and the local sound reinforcement module 810, when converting the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, is configured to: perform an inverse short-time Fourier transform on the human voice time-frequency domain signal corresponding to the first time-frequency domain signal to obtain the human voice time-domain signal in the time domain.

[0171] In one example, the first extraction module 802 is further configured to: calculate the linear echo in the first time domain signal based on the second time domain signal and the predicted echo path; and eliminate the linear echo in the first time domain signal.

[0172] In one example, the voice extraction model includes: a nonlinear transformation unit, a multi-layer convolution unit, and a masking unit; the voice extraction module 806, when inputting the first time-frequency domain features and the second time-frequency domain features into the voice extraction model to obtain the time-frequency mask output by the voice extraction model, is configured to: concatenate the first time-frequency domain features to obtain a first concatenated feature; concatenate the second time-frequency domain features to obtain a second concatenated feature; the nonlinear transformation unit performs a nonlinear transformation on the first concatenated feature and the second concatenated feature to obtain nonlinear transformation data; the multi-layer convolution unit performs convolution processing on the nonlinear transformation data to obtain filtered data; the masking unit calculates and outputs the time-frequency mask based on the filtered data, wherein the time-frequency mask consists of a real part and an imaginary part.

[0173] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0174] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0175] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory including instructions, is also provided, which can be executed by one or more processors of a local sound reinforcement device to implement the method as described in any of the above embodiments.

[0176] The non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc., and this application does not limit it.

[0177] In an exemplary embodiment, a computer program product including a computer program / instructions is also provided, which can be executed by one or more processors of a local sound reinforcement device to implement the method described in any of the above embodiments.

[0178] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0179] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments herein. This application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0180] The above descriptions are some embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A local sound reinforcement method, comprising: Feature extraction is performed on the first time-frequency domain signal to be played to obtain the first time-frequency domain features, wherein the first time-frequency domain features are the feature data of the first time-frequency domain signal in multiple channels; Feature extraction is performed on the second time-frequency domain signal played after the first time-frequency domain signal to obtain the second time-frequency domain features, wherein the second time-frequency domain features are the feature data of the second time-frequency domain signal on multiple channels; The first time-frequency domain feature and the second time-frequency domain feature are input into the human voice extraction model to obtain the time-frequency mask output by the human voice extraction model. The time-frequency mask is set to represent the intensity of the human voice time-frequency domain signal at the corresponding time and frequency point in the first time-frequency domain signal. Based on the time-frequency masking, acoustic factors in the first time-frequency domain signal are suppressed to obtain the human voice time-frequency domain signal corresponding to the first time-frequency domain signal; the acoustic factors include: howling, ambient noise, room reverberation and echo remnant; The human voice time-frequency domain signal corresponding to the first time-frequency domain signal is converted into a human voice time-domain signal, and the human voice time-domain signal is amplified locally.

2. The method according to claim 1, wherein, The human voice extraction model is obtained by training an initial model using training data and labeled data. The methods for generating the training data and the labeled data include: The simulated echo path of the simulated second time-domain signal of the previous frame is simulated to obtain the simulated echo signal; A simulated room transfer function is applied to a human voice source to obtain a simulated human voice image, wherein the simulated human voice image includes a simulated human voice time-domain signal and a simulated room reverberation; The simulated echo signal, the simulated human voice image, and the simulated environmental noise are superimposed to obtain the simulated first time domain signal; Echo cancellation is performed on the simulated first time-domain signal to obtain an echo-cancelled signal; The training data is generated based on the echo cancellation signal and the analog second time-domain signal of the previous frame; Based on the simulated human voice time-domain signal, the label data corresponding to the training data is generated; Gain control and delay processing are applied to the echo cancellation signal to obtain the analog second time-domain signal of the current frame.

3. The method according to claim 2, wherein, The process of training the initial model using training data and label data includes: Based on the training progress of the initial model, training data of different difficulties are selected; The training data is input into the initial model to obtain the prediction result output by the initial model; The loss value is calculated based on the prediction results and the label data, and the parameters of the initial model are updated. Adjust the training progress of the initial model, select training data of the next difficulty level for model training, until the initial model meets the preset conditions, and obtain the human voice extraction model.

4. The method according to any one of claims 1 to 3, wherein, The local amplification of the human voice time-domain signal includes: Gain control and delay processing are applied to the human voice time-domain signal to obtain the processed human voice time-domain signal; The processed human voice time-domain signal is superimposed with the background audio signal to obtain the target signal; The target signal is amplified locally using a loudspeaker.

5. The method according to any one of claims 1 to 4, further comprising: Receive the first time-domain signal using a microphone; Perform a short-time Fourier transform on the first time-domain signal to obtain the first time-frequency domain signal; A short-time Fourier transform is performed on the second time-domain signal to be played by the speaker to obtain the second time-frequency domain signal; The step of converting the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal includes: The human voice time-frequency domain signal corresponding to the first time-frequency domain signal is subjected to a short-time inverse Fourier transform to obtain the human voice time-domain signal.

6. The method according to claim 5, further comprising: Based on the second time-domain signal and the predicted echo path, calculate the linear echo in the first time-domain signal; Eliminate linear echoes in the first time-domain signal.

7. The method according to any one of claims 1 to 6, wherein, The human voice extraction model includes a nonlinear transformation unit, a multi-layer convolutional unit, and a masking unit; the step of inputting the first time-frequency domain features and the second time-frequency domain features into the human voice extraction model to obtain the time-frequency mask output by the human voice extraction model includes: The first time-frequency domain features are concatenated to obtain the first concatenated features; The second time-frequency domain features are concatenated to obtain the second concatenated features; The nonlinear transformation unit performs a nonlinear transformation on the first splicing feature and the second splicing feature to obtain nonlinear transformation data; The multi-layer convolutional unit performs convolution processing on the nonlinear transformed data to obtain filtered data; The masking processing unit calculates and outputs the time-frequency mask based on the filtered data, wherein the time-frequency mask consists of a real part and an imaginary part.

8. A local sound reinforcement device, comprising: The first extraction module is configured to extract features from the first time-frequency domain signal to be played, and obtain the first time-frequency domain features, wherein the first time-frequency domain features are the feature data of the first time-frequency domain signal in multiple channels; The second extraction module is configured to extract features from the second time-frequency domain signal played after the first time-frequency domain signal to obtain second time-frequency domain features, wherein the second time-frequency domain features are feature data of the second time-frequency domain signal in multiple channels; The human voice extraction module is configured to input the first time-frequency domain features and the second time-frequency domain features into the human voice extraction model to obtain the time-frequency mask output by the human voice extraction model. The time-frequency mask is configured to represent the intensity of the human voice time-frequency domain signal at the corresponding time and frequency points in the first time-frequency domain signal. The factor suppression module is configured to suppress acoustic factors in the first time-frequency domain signal based on the time-frequency masking, thereby obtaining the human voice time-frequency domain signal corresponding to the first time-frequency domain signal; the acoustic factors include: howling, ambient noise, room reverberation, and echo remnant; The local amplification module is configured to convert the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, and to amplify the human voice time-domain signal locally.

9. A terminal device, comprising: One or more processors; The memory is configured to store one or more processor-executable instructions; The one or more processors implement the method as described in any one of claims 1-7 by executing the executable instructions.

10. A computer-readable storage medium having stored thereon computer instructions, wherein, When executed by one or more processors, this instruction implements the method as described in any one of claims 1-7.

11. A computer program product having a computer program / instructions stored thereon, wherein, When the computer program / instructions are executed by one or more processors, they implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Echo cancellation method under noise environment and echo cancellation system thereof

    CN105825865A

  • Echo cancellation method, echo cancellation device and computer readable storage medium

    CN113689878A

  • Data enhancement method and system for acoustic scene classification

    CN117037824A

  • Processing method and device for communication enhancement voice signal in same acoustic space, and medium

    CN117116282A

  • Audio signal processing method and device, equipment and storage medium

    CN117392994A