A signal processing method, apparatus, device and computer-readable storage medium
By extracting the logarithmic energy spectrum characteristics of the audio signal and generating correction coefficients using the noise optimization model, the problem of echo noise signal interference is solved and the quality of voice interaction is improved.
Patent Information
- Application Number
- CN202010597937.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-06-28
AI Technical Summary
In multi-party remote meeting scenarios, the echo noise signal interference caused by the conference room environment reduces the quality of voice interaction.
By collecting the audio signal to be processed, extracting its logarithmic energy spectrum characteristics, and calling a pre-trained noise optimization model, a noise correction coefficient is generated, and the audio signal is corrected to reduce or eliminate the impact of the echo noise signal.
Effectively reduce or eliminate the influence of echo noise signals, improve the quality of voice interaction, and make the audio signals in voice conferences clearer and more discernible.
Smart Images

Figure CN111710344B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a signal processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] With the continuous development of communication technologies, people's requirements for signal quality are constantly increasing. Especially in scenarios such as holding network conferences using computer networks or mobile communication networks, it is desired that the call signals of the conference are clear and distinguishable, and at the same time, some unnecessary signals input along with the voices of the participants can be minimized.
[0003] In one scenario, the unnecessary signal mainly refers to a noise signal, which may be some unwanted echo audio signals. In the scenario of a multi-party remote conference, there will be a situation where participants at multiple ends speak simultaneously. At this time, the local voice communication device not only needs to play the voices of the participants in other regions, but also collect the local voices of the local participants. Due to the influence of factors such as the conference room environment, there will be a part of special noise signals in the local voices collected by the voice communication device, such as the echo of the voice played by the voice communication device reflected in the conference room.
[0004] These echo signals will have an adverse impact on the conference voice signals such as interaction. For example, these echoes may bring noises such as "hiss" in the voice conference, reducing the quality of voice interaction. Summary of the Invention
[0005] Embodiments of the present invention provide a signal processing method, apparatus, device, and computer-readable storage medium, which can improve the quality of voice interaction.
[0006] On the one hand, embodiments of the present application provide a signal processing method, which includes:
[0007] Collect an audio signal to be processed, and extract the spectral features of the audio signal to be processed, where the spectral features include N-dimensional logarithmic energy spectral features;
[0008] Call a noise optimization model to process the logarithmic energy spectral features, and obtain M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral features, where N and M are positive integers;
[0009] Calculate the N-dimensional logarithmic energy spectral features and the M-dimensional noise correction coefficients to obtain a processed audio signal;
[0010] Among them, the noise optimization model is trained according to audio training data including noisy audio signals. The M-dimensional noise correction coefficients output by the noise optimization model include: p-dimensional coefficients for correcting the features of the noisy audio signals in the input logarithmic energy spectrum features, where p is less than M.
[0011] On the other hand, the present application provides a signal processing device, which includes:
[0012] An acquisition unit, configured to collect an audio signal to be processed and extract the spectrum features of the audio signal to be processed, where the spectrum features include N-dimensional logarithmic energy spectrum features;
[0013] A processing unit, configured to call a noise optimization model to process the logarithmic energy spectrum features, obtain M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectrum features, where N and M are positive integers; calculate the N-dimensional logarithmic energy spectrum features and the M-dimensional noise correction coefficients to obtain a processed audio signal;
[0014] Among them, the noise optimization model is trained according to audio training data including noisy audio signals. The M-dimensional noise correction coefficients output by the noise optimization model include: p-dimensional coefficients for correcting the features of the noisy audio signals in the input logarithmic energy spectrum features, where p is less than M.
[0015] Correspondingly, an embodiment of the present application further provides a signal processing device, including a processor, a memory, and a communication interface. The processor, the memory, and the communication interface are connected to each other. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the above signal processing method.
[0016] Correspondingly, the present application provides a computer-readable storage medium, which stores one or more instructions, and the one or more instructions are suitable for being loaded and executed by a processor to execute the above signal processing method.
[0017] Correspondingly, the present application provides a computer program product or a computer program, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above signal processing method.
[0018] In the embodiments of the present application, for the to-be-processed audio signals generated in situations such as audio-video conferences and audio-video calls, starting from the logarithmic energy spectrum features, the noise correction coefficients for the to-be-processed audio signals generated by a pre-trained and optimized noise optimization model can be used to effectively optimize and correct the collected to-be-processed audio signals, reduce or even eliminate the adverse effects of noise audio signals such as echoes on the collected audio signals in the to-be-processed audio signals, thereby improving the quality of voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1a It is a scene architecture diagram of signal processing provided by an embodiment of the present invention;
[0021] Figure 1b It is a signal processing flowchart provided by an embodiment of the present application;
[0022] Figure 2 It is a flowchart of a signal processing method provided by an embodiment of the present application;
[0023] Figure 3 It is a flowchart of extracting frequency-domain spectrum features from time-domain audio signals provided by an embodiment of the present application;
[0024] Figure 4a It is a flowchart of a model training method provided by an embodiment of the present application;
[0025] Figure 4b It is a flowchart of another model training method provided by an embodiment of the present application;
[0026] Figure 5 It is a brief schematic diagram of the training of a noise optimization model provided by an embodiment of the present application;
[0027] Figure 6 It is a flowchart of another signal processing method provided by an embodiment of the present application;
[0028] Figure 7 It is a conference session interface diagram provided by an embodiment of the present application;
[0029] Figure 8 It is a structural schematic diagram of a signal processing device provided by an embodiment of the present application;
[0030] Figure 9 This is a schematic structural diagram of an intelligent device provided by an embodiment of the present application. Specific implementation manners
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0032] The embodiments of the present application relate to Artificial Intelligence (AI) and Machine Learning (ML). By combining AI and ML, the features in the audio signal can be mined and analyzed, enabling the device to more accurately identify and process the audio signal, and determining the spectral features of noise signals such as echoes, so as to reduce or even eliminate the adverse effects of this part of the noise signal on the original audio signal. Among them, AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, and is a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning and decision-making.
[0033] AI technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, processing technologies for large application programs, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The embodiments of the present application mainly relate to the language processing technology among them.
[0034] ML is a multi-disciplinary cross-discipline, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills, and reorganizes the existing knowledge structure to continuously improve its own performance. ML is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0035] Statistical estimation echo cancellation algorithms based on traditional machine learning can be used to analyze and process the audio signals to be processed. Such algorithms may include, for example, echo cancellation algorithms based on Adaptive Filter. For these traditional statistical learning cancellation algorithms, specific algorithms can be used to estimate the coefficients of the filter and automatically adjust the weighting coefficients according to the statistical characteristics of the input and output signals to achieve echo cancellation. For the estimation of the filter coefficients, the Least Mean Square (LMS) is usually the optimization goal.
[0036] Echo cancellation algorithms based on neural network methods can also be used to analyze and process the audio signals to be processed. Such echo cancellation algorithms collect the signals at the far-end (the other party joining the meeting, Far-end) and the near-end (the party joining the meeting, Near-End), extract their spectral features respectively, splice them as the input of the neural network, and use the spectral features at the near-end as the output of the neural network. Some mainstream network models, such as Convolution Neural Network (CNN) and Recurrent Neural Network, can be applied to the cancellation of noise signals such as echo.
[0037] To address the above problems, the present application proposes a signal processing method. For the collected audio signals to be processed, first, the spectral features of the audio signals to be processed are extracted, and then the pre-trained noise optimization model is called to process the logarithmic energy spectral features of the audio signals to obtain the noise correction coefficients corresponding to the logarithmic energy spectral features, and the logarithmic energy spectral features of the audio signals are corrected by the noise correction coefficients, so as to reduce or even eliminate the noise in the audio signals to be processed and improve the quality of voice interaction.
[0038] Please refer to Figure 1a , Figure 1a which is a scenario architecture diagram of signal processing provided by an embodiment of the present invention. As Figure 1a shown, the scenario architecture diagram includes the party joining the meeting, the participating parties, and the terminal device 101. Among them, the party joining the meeting and the participating parties participate in a remote meeting through their respective terminal devices. For example, the party joining the meeting uses the terminal device 101 to participate in the remote meeting. During the remote meeting, the terminal device 101 will collect the sound waves of the party joining the meeting and send them to the participating parties; and, the terminal device 101 will also play the voice sent by the participating parties. During Figure 1aIn the schematic diagram, the opposite-end sound wave is the sound wave emitted when the terminal device 101 plays the voice of the participating party; when the opposite-end sound wave encounters a reflector (such as a wall), it will be reflected to form an echo sound wave. Therefore, after the echo sound wave is generated, if the user of the participating party is also speaking and generating the sound wave of the participating party user, the audio signal collected by the terminal device 101 may include: the audio signal corresponding to the human voice sound wave of the participating party and the audio signal corresponding to the echo sound wave. Of course, the scenario of the participating party signal processing and the scenario of the participating party signal processing may be the same.
[0039] The number of the terminal devices 101 may be one or more. The form of the terminal device 101 is only for illustration. The terminal device 101 may include but is not limited to: smart phones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, Mobile Internet Devices (MID for short), voice collectors (players), and other devices with voice playback and collection functions. The number of the participating parties may be one or more, which is not limited in the embodiments of the present application.
[0040] Figure 1b This is a signal processing flowchart provided by the embodiments of the present application. As Figure 1b shown, the signal processing process mainly includes: the terminal device 101 collects the human voice sound wave and the noise sound wave of the participating party. These noise sound waves may be, for example, the echo sound waves mentioned above, and obtains the corresponding audio input signal according to the collected sound waves, that is, obtains the audio signal to be processed; then extracts the spectral features of the audio signal to be processed. The spectral features may be N-dimensional logarithmic energy spectral features; calls the noise optimization model to process the N-dimensional logarithmic energy spectral features, and obtains the M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral features. In one embodiment, the noise optimization model may be a model constructed based on a convolutional neural network of Long Short-Term Memory (LSTM). The M-dimensional noise correction coefficients are estimated spectral correction coefficients; performs operations on each dimension of the logarithmic energy spectral features and the noise correction coefficients of the corresponding dimensions, and the processed audio signal can be obtained. The processed audio signal weakens or eliminates signals such as echoes and can be transmitted to one or more participating parties through the conference system.
[0041] In one embodiment, when building and training the model, the N-dimensional logarithmic energy spectrum feature includes the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal, that is, N = n + p. The logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal can be arranged successively. For example, the first n dimensions are defined as the logarithmic energy spectrum of the human voice audio signal, and the last p dimensions are defined as the logarithmic energy spectrum of the echo audio signal, or they can be arranged in a cross-mixed manner. Correspondingly, the M-dimensional noise correction coefficient includes the noise correction coefficient of the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficient of the logarithmic energy spectrum of the p-dimensional echo audio signal, that is, M = N = n + p. The arrangement of the noise correction coefficient of the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficient of the logarithmic energy spectrum of the p-dimensional echo audio signal corresponds to the arrangement of the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal in the N-dimensional logarithmic energy spectrum feature. For example, assume that the j-th dimensional noise correction coefficient represents the correction coefficient corresponding to the human voice audio signal, and the i-th dimensional noise correction coefficient represents the correction coefficient corresponding to the noise signal; then the j-th dimensional logarithmic energy spectrum feature also correspondingly represents a feature of the human voice audio signal, and the i-th dimensional logarithmic energy spectrum feature correspondingly represents a feature of the noise signal. And, in the calculation, if the noise correction coefficient of the j-th dimension in the obtained M-dimensional noise correction coefficient is 1, and the noise correction coefficient of the i-th dimension is 0.01, then the operation of the logarithmic energy spectrum feature and the noise correction coefficient of the corresponding dimension means: multiplying the value of the i-th dimensional logarithmic energy spectrum feature in the N-dimensional logarithmic energy spectrum feature by the noise correction coefficient 0.01 of the i-th dimension in the M-dimensional noise correction coefficient to obtain a new value of the logarithmic energy spectrum feature. It can be understood that since the logarithmic energy spectrum of the echo is greatly reduced after being multiplied by the corresponding noise correction coefficient 0.01, while the logarithmic energy spectrum of the human voice remains unchanged after being multiplied by the corresponding noise correction coefficient 1, the influence of the echo audio signal on the human voice audio signal will be greatly reduced after the operation.
[0042] Please refer to Figure 2 , Figure 2 which is a flowchart of a signal processing method provided by an embodiment of the present application. This method can be executed by an intelligent device, which can specifically be Figure 1a the terminal device 101 shown in
[0043] S201: Collect the audio signal to be processed and extract the spectral features of the audio signal to be processed. The spectral features include N-dimensional logarithmic energy spectral features. In the embodiments of the present invention, the N-dimensional logarithmic energy spectral features can uniquely represent the collected audio signal to be processed, and the duration of the audio signal to be processed can be predetermined. For example, if the duration of the audio signal to be processed is 10 ms, then every 10 ms of the audio signal will correspond to N-dimensional logarithmic energy spectral features; another example, if the duration of the audio signal to be processed is 100 ms of the audio signal, then every 100 ms of the audio signal will correspond to N-dimensional logarithmic energy spectral features; of course, the duration of the audio signal to be processed can also be other values. For the audio signal corresponding to the continuous human voice input by the user, audio signals to be processed with a duration of 10 ms (or 100 ms) etc. will be obtained in chronological order. The audio signal to be processed is obtained by dividing the collected time-domain audio signal. The audio signal to be processed mentioned in the embodiments of the present invention includes at least one of the following audio signals: the human voice audio signal corresponding to the human voice emitted by the user participating in the meeting, the audio signal of the noise in the current environment of the user participating in the meeting. In the embodiments of the present invention, the noise signal mainly refers to the echo signal. This application considers the echo signal as a noise signal, so the echo audio signal is used as the noise signal for description. The spectral features of the audio signal to be processed are frequency-domain spectral features, and the frequency-domain spectral features include N-dimensional logarithmic energy spectrum (Log Power Spectrum, LPS).
[0044] Figure 3 This is a flowchart for extracting frequency-domain spectral features from a time-domain audio signal provided by an embodiment of this application. As Figure 3 shown, first, perform time-frame processing on the sound corresponding to the time-domain audio signal and add a sliding window operation. For example, assume that the duration of the time-domain audio signal 1 is 10 seconds and the duration of the audio signal to be processed is 100 ms, then the time-domain audio signal 1 is divided into 100 frames of audio signals to be processed in chronological order; then perform a Fast Fourier Transform (FFT) on each frame of the divided sound segment to obtain the spectral energy distribution of each frequency bin (i.e., the frequency-domain discrete spectrum); then perform a squaring operation on the frequency-domain discrete spectrum (such as inputting the frequency-domain discrete spectrum into a spectrum squaring operator); finally, perform a logarithmic operation on the result of the squaring operation to obtain the logarithmic energy spectral features corresponding to the time-domain audio signal.
[0045] S202: Call the noise optimization model to process the logarithmic energy spectrum features to obtain the M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectrum features, where N and M are positive integers. The noise correction coefficients are used to reduce or eliminate the noise audio signals in the audio signal to be processed. In one embodiment, the values of M and N can be between 64 dimensions and 512 dimensions. For example, 500 dimensions can be taken. For each piece of audio signal to be processed (a segment of speech between 10 ms and 100 ms) collected, it corresponds to logarithmic energy spectrum features of 500 dimensions and other equal dimensions.
[0046] In one implementation manner, M = N, that is, after the noise optimization model processes the N-dimensional logarithmic energy spectrum features, the noise correction coefficients corresponding to each dimension of the logarithmic energy spectrum features are obtained. Based on the noise optimization model, the greater the energy of the echo contained in each dimension of the logarithmic energy spectrum features, the smaller the value of the desired noise correction coefficient, that is, the value of the desired noise correction coefficient is inversely proportional to the energy of the noise, so as to minimize or even remove the echo as much as possible. The value ranges of the noise correction coefficients are [0, 1]. The coefficient values corresponding to the human voice part in the noise correction coefficients are 1 or approaching 1, and the coefficient values corresponding to the echo part are 0 or approaching 0. For example, assume that among the 500-dimensional logarithmic energy spectrum, 400 dimensions are used to represent the features of the human voice audio signal, and 100 dimensions are used to represent the features of the echo audio signal. Then the noise correction coefficient values of the 400-dimensional logarithmic energy spectrum used to represent the features of the human voice audio signal are 1 or approaching 1, and the noise correction coefficient values of the 100-dimensional logarithmic energy spectrum used to represent the features of the echo audio signal are 0 or approaching 0. It can be understood that if the noise correction coefficient values of the 400-dimensional logarithmic energy spectrum used to represent the features of the human voice audio signal in the 500-dimensional logarithmic energy spectrum are all less than the energy threshold or are 0, it means that the audio signal to be processed does not include the human voice audio signal of the participant; similarly, if the noise correction coefficient values of the 100-dimensional logarithmic energy spectrum used to represent the features of the echo audio signal in the 500-dimensional logarithmic energy spectrum are all less than the energy threshold or are 0, it means that the audio signal to be processed does not include the echo audio signal.
[0047] S203: Calculate the N-dimensional logarithmic energy spectrum features and the M-dimensional noise correction coefficients to obtain the processed audio signal. In one implementation manner, multiply each dimension of the logarithmic energy spectrum features by the corresponding noise correction coefficient (for example, multiply the i-th dimension of the logarithmic energy spectrum features by the i-th dimension of the noise correction coefficient) to obtain the processed audio signal. By reducing the energy of the noise, the effect of residual noise elimination or weakening is achieved.
[0048] In the embodiments of the present application, for the to-be-processed audio signals generated in situations such as voice conferences and voice calls, starting from the logarithmic energy spectrum features, the noise correction coefficient for the to-be-processed audio signals generated by a pre-trained and optimized noise optimization model can be used to effectively optimize and correct the to-be-processed audio signals, reducing or even eliminating the adverse effects of the features of the noise audio signals on the to-be-processed audio signals in the to-be-processed audio signals, thereby improving the quality of voice interaction.
[0049] Please refer to Figure 4a , Figure 4a which is a flowchart of a model training method provided by an embodiment of the present application. This method can be executed by an intelligent device, which can specifically be Figure 1a the terminal device 101 shown in
[0050] S401: Collect echo audio signals in a target environment where an audio signal is being played. Here, the target environment refers to some selected environments that can generate echoes, and there can be multiple target environments. For example, offices, conference rooms, etc. Echo audio signals can be collected in these environments to further obtain audio training data for training the model.
[0051] In one embodiment, in order to obtain more suitable audio training data subsequently, during the process of recording the echo audio signals, different target environments can be selected. In different target environments, the echo effects are different. That is, in some target environments, the echo sound is louder, while in some target environments, the echo sound is smaller. Then, the collected echo audio signals include echo audio signals with different sound intensities. In this way, after mixing these echo audio signals with clean human voice audio signals subsequently, a mixed audio signal including echoes with different sound intensities can be obtained, and thus the model can be trained more comprehensively under multiple echo sound intensities.
[0052] Moreover, in some embodiments, in these target environments that can generate different sound intensities, the device for playing the audio signal can also be adjusted in terms of sound, and the audio can be played at different volume levels to obtain richer echo audio signals. In addition, the echo audio signals can also be collected when different audio signal playing devices are playing in the same or different target environments. Then, subsequently, audio training data can be generated for different clients (such as intelligent terminals like mobile phones, and dedicated conference devices like octopuses), and a noise optimization model for different clients can be trained.
[0053] S402: Obtain the human voice audio signal. In one implementation, the human voice audio signal is generated by directly collecting the user's voice. In another implementation, the human voice audio signal is obtained from a pre-stored voice audio signal library.
[0054] S403: Superimpose the obtained human voice audio signal and the echo audio signal in the time domain to obtain a mixed audio signal, and generate audio training data based on the mixed audio signal. In one implementation, multiple human voice audio signals are respectively superimposed with the collected echo audio signal to obtain audio training data. Among them, each segment of the mixed audio signal includes an echo audio signal and a human voice audio signal. In another implementation, the mixed audio signal may also include an echo audio signal and a human voice audio signal in some time periods, while in some time periods, it does not include a human voice audio signal and only includes an echo audio signal. That is to say, the audio training data may include X segments of mixed audio signals, each segment of the mixed audio signal includes an echo audio signal, the i-th segment of the mixed audio signal includes a human voice audio signal and an echo audio signal, and the j-th segment of the mixed audio signal only includes an echo audio signal, where i, j, X are positive integers, i≠j, and i, j are less than or equal to X. For example, the audio training data includes 100 segments of mixed audio signals, 50 segments of mixed audio signals only include echo audio signals, and 50 segments of mixed audio signals include echo audio signals and human voice audio signals.
[0055] S404: Use the audio training data to trigger the training of the initial model to obtain a noise optimization model. The noise optimization model is obtained by training the initial model with the audio training data and the loss function. The audio training data is calculated and converted to obtain the corresponding logarithmic energy spectrum features, and after normalizing the logarithmic energy spectrum features, they are used as the input of the initial model. For the result output by the initial model, loss calculation is performed based on the loss function, and the parameters of the initial model are adjusted according to the result of the loss calculation, so as to finally obtain a noise optimization model.
[0056] In one implementation, the loss function adopted by the noise optimization model can be:
[0057] minE[(Y mix:clean+echo (w)H model_coef (w)-X clean (w)) 2
[0058] Among them, w represents the specific dimensional value, Y mix:clean+echo (w) is the input of the noise optimization model, that is, the N-dimensional logarithmic energy spectrum features corresponding to a segment of the mixed audio signal in the audio training data; H model_coef (w) is the coefficient estimated by the noise optimization model, that is, the aforementioned noise correction coefficient; Ymix:clean+echo (w)H model_coef (w) is the first clean spectral feature. X clean (w) is the logarithmic energy spectral feature corresponding to the human voice audio signal, that is, the second clean spectral feature. Based on the value obtained from this loss function, the parameters (such as convolution parameters) in the noise correction model are optimized, so that after the parameters of the noise optimization model are optimized, for the target audio training data (any one of all the audio training data), the noise optimization model can generate the corresponding noise correction coefficient, and the difference between the value obtained by multiplying the logarithmic energy spectral feature corresponding to the target audio training data by this noise correction coefficient and the logarithmic energy spectral feature corresponding to the human voice audio signal in this audio training data is minimized.
[0059] Further, G model_coef (w) is expressed as:
[0060]
[0061] The noise optimization model can be considered to be constructed based on the above expression form, where S clean (w) is the logarithmic spectral energy of the human voice audio signal corresponding to the second clean logarithmic spectral feature; S echo (w) is the logarithmic spectral energy of the echo audio signal; S clean (w)+S echo (w) is the logarithmic spectral energy corresponding to the mixed audio signal. It can be considered that the logarithmic spectral energy corresponding to the mixed audio signal is the sum of the logarithmic spectral energy of the human voice audio signal and the logarithmic spectral energy of the echo audio signal.
[0062] Further, S clean (w) = log{|F[x clean (t)]| 2}, S echo (w) = log{|F[x echo (t)]| 2}, that is, the logarithmic spectral energy of the human voice audio signal and the logarithmic spectral energy of the echo audio signal are obtained respectively by performing a fast Fourier transform and taking the logarithm.
[0063] S405: Test the noise optimization model obtained by optimized training. The performance of the noise optimization model is determined through the test, where this test step is an optional step.
[0064] The above S401 to S404 describe the training process of the noise optimization model of the present application. For details, please refer to Figure 5 , Figure 5 which is a brief schematic diagram of the training of a noise optimization model provided by an embodiment of the present application. As Figure 5As shown, first, the echo is collected and the corresponding echo audio signal (Echo Signal) is generated; then, different human voice audio signals (such as the pre-collected human voice or the voice of the current user being collected) are respectively superimposed on the echo audio signal in the time domain to obtain multiple different mixed audio signals, and these mixed audio signals are used as the audio training data of the initial model; then, the initial model is trained and optimized using the audio training data to obtain a noise optimization model, and the loss function used is as described above.
[0065] Continue to refer to Figure 5 , after the training and optimization of the noise optimization model are completed, enter the test link. Collect audio test data, perform calculation and conversion on the audio test data to obtain the corresponding logarithmic energy spectrum features, input the logarithmic energy spectrum features of the audio test data into the noise optimization model, obtain the noise correction coefficient output by the noise optimization model, multiply the noise correction coefficient output by the noise optimization model by the logarithmic energy spectrum features of the audio test data to obtain the test result. If the test result does not contain the "hissing" noise during playback, and / or the value of the loss function in the noise optimization model is less than the noise cancellation threshold, it is determined that the noise optimization model passes the test. Deploy the noise optimization model to terminal clients such as conference applications for the following Figure 6 corresponding embodiments. Correspondingly, if the test result contains the "hissing" noise during playback, and / or the value of the loss function in the noise optimization model is greater than the noise cancellation threshold, it is determined that the noise optimization model fails the test, and the noise optimization model is continuously trained by the method in S401 - S404 above until the test passes.
[0066] Please refer to Figure 4b , Figure 4b is a flowchart of another model training method provided by an embodiment of the present application. This method can be executed by an intelligent device, and the intelligent device can specifically be Figure 1a the terminal device 101 shown in, or a server for model training and optimization. The method of the embodiment of the present invention includes the following steps.
[0067] S411: Perform audio recording operations in multiple target environments where audio signals are being played to obtain multiple segments of noisy audio information, and each segment of noisy audio information includes a noisy audio signal and recording device information. Among them, the target environment refers to some selected environments that can generate echoes, and there can be multiple target environments. For example, offices, conference rooms, etc. In these environments, echo audio signals can be collected, and further audio training data can be obtained to train the model.
[0068] In one implementation, different audio recording devices (such as smart terminals, multi-party conference phones, dedicated conference devices like octopuses, etc.) can also be used to collect audio signals played at different volume levels (for example, setting the playback volume levels of the voice playback device to levels 1 - 10 respectively) in different target environments, so as to collect the generated echo audio signals and obtain multiple segments of echo audio information. Each segment of echo audio information includes the echo audio signal and the recording device information. For example, the first segment of echo audio information is obtained from the echo during the playback of voice by the voice playback device at a playback volume level of 2 in Conference Room 1 using a microphone; the second segment of echo audio information is obtained from the echo during the playback of voice by the voice playback device at a playback volume level of 2 in Conference Room 1 using a multi-party conference phone; the third segment of echo audio information is obtained from the echo during the playback of voice by the voice playback device at a playback volume level of 2 in Conference Room 2 using a microphone; the fourth segment of echo audio information is obtained from the echo during the playback of voice by the voice playback device at a playback volume level of 5 in Conference Room 1 using a smart terminal such as a mobile phone.
[0069] S412: Generate audio training data corresponding to each recording device information based on multiple segments of noise audio information; wherein, the audio training data includes Y segments of noise audio signals, where Y is a positive integer. In S412, audio training data corresponding to different recording devices can be generated based on multiple segments of noise audio information. It can be understood that, for example, audio training data 1 corresponding to the smart terminal is generated based on the echo audio signal recorded by the smart terminal, and audio training data 2 corresponding to the multi-party conference phone (such as dedicated conference devices like octopuses) is generated based on the echo audio signal recorded by the multi-party conference phone.
[0070] S413: Trigger the training of the initial model using the audio training data to obtain a noise optimization model. By collecting echo audio signals with different recording devices, audio training data for different recording devices can be obtained, and thus a more targeted noise optimization model can be obtained; for example, when the recording device in a certain target environment is a smart terminal such as a mobile phone, a noise optimization model for the smart terminal can be trained based on the audio training data generated from the echo audio signals with different sound intensities collected by the smart terminal; when the recording device in a certain target environment is a dedicated conference device like an octopus, a noise optimization model corresponding to the dedicated conference device can be trained based on the audio training data generated from the echo audio signals with different sound intensities collected by the dedicated conference device.
[0071] It can be understood that the recording device during model training corresponds to the voice input device corresponding to the client during model use, and a mapping table can be established. This mapping table records the relationship between the device type identifier (the type identifier corresponding to the recording device or the voice input device) and the identifier of the noise optimization model.
[0072] S414: Test the noise optimization model obtained through optimized training. Determine the performance of the noise optimization model through the test. Herein, this test step is an optional step. Performance tests can be respectively conducted on corresponding noise optimization models based on different types of devices. It can be understood that for the specific descriptions of S413 and S414 above, reference can also be made to Figure 4a the descriptions of the relevant content of the corresponding S404 and S405, which will not be elaborated herein.
[0073] Since different noise optimization models correspond to different types of devices, during the subsequent process of using the model, the type of the device for collecting sound can be referred to select the corresponding noise optimization model to perform optimization processing such as weakening or even removing noises such as echoes.
[0074] In the embodiment of the present application, echo is eliminated as a specific type of noise (i.e., one type of noise), and during the training process, it is not necessary to collect the signals of the participating parties, effectively reducing the model dimension (size), thereby improving the calculation efficiency.
[0075] Figure 6 , Figure 6 is a flowchart of another signal processing method provided by the embodiment of the present application. This method can be executed by an intelligent device, and the intelligent device can specifically be Figure 1a the terminal device 101 shown in, and a meeting application is installed on the terminal device, and the noise optimization model that has passed the test is deployed in the meeting application. The method of the embodiment of the present invention includes the following steps.
[0076] S601: Collect the audio signal to be processed, and extract the spectral features of the audio signal to be processed. The spectral features include N-dimensional logarithmic energy spectral features. When the end user opens the meeting application to participate in an Internet meeting, whether it is a video conference or a pure voice conference, the microphone of the terminal device can be called to collect the audio signal to be processed. At this time, the audio signal to be processed is the meeting audio signal. The meeting audio signal may include human voice audio signals, may also include echo audio signals, and may also include both human voice audio signals and echo audio signals at the same time. Specifically, the meeting audio signal is collected when detecting entry into the Figure 7 meeting session interface shown in (i.e., during the multi-terminal conference communication process).
[0077] In one embodiment, before the spectral characteristics of the conference audio signal, the terminal device can determine the spatial type of the current location, so as to determine whether to call the noise optimization model deployed in the conference application to optimize echoes and the like according to the spatial type. Among them, if the spatial type of the current location of the terminal device belongs to the first type (in this application, the first type refers to an environmental space type where the location is open and the echo is generally small or even non-existent), the collected conference audio signal is encoded and the encoded conference audio signal is sent to the participants, that is, there is no need to perform echo cancellation processing on the conference audio signal in the first type of space. If the spatial type of the current location of the terminal device belongs to the second type (in this application, the second type mainly refers to indoor environments, such as when it is found through positioning that the user is in a certain building, it is considered that the spatial type of the conference user belongs to the second type), the step of extracting the spectral characteristics of the conference audio signal is performed, so as to perform subsequent related steps of echo cancellation. It can be seen that by judging the spatial type of the current location, it is possible to avoid performing echo cancellation processing on the audio signal to be processed with an echo less than the threshold, thereby reducing the waste of memory resources and improving the efficiency of signal processing.
[0078] S602: Normalize the extracted spectral characteristics. In one embodiment, by normalizing the extracted spectral characteristics, the extracted spectral characteristics are mapped to the processing numerical range of [0, 1].
[0079] S603: Call the noise optimization model to process the logarithmic energy spectral characteristics to obtain the M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral characteristics, where N and M are positive integers. Specifically, the logarithmic energy spectral characteristics are the logarithmic energy spectral characteristics corresponding to a section of speech between 10 ms and 100 ms, which are 64-dimensional or 500-dimensional, etc., and the output is the coefficient with the same dimension (64 - 512) as the energy spectral characteristics.
[0080] In one embodiment, when at least two noise optimization models are recorded in the stored noise optimization model set, the noise optimization model associated with the location attribute can be selected from the noise optimization model set according to the location attribute of the location in the current conference environment, and the associated noise optimization model is called to process the logarithmic energy spectral characteristics of the current conference audio signal as the input. Among them, the location attribute refers to the location identifier. For example, location A is located in Building YY in Area XX, and the corresponding identifier is 567678; location B is located in Building XY in Area ZZ, and the corresponding identifier is 877454.
[0081] Furthermore, a relationship mapping table between the location attribute and the noise optimization model is established. Table 1 is an exemplary relationship mapping table provided by an embodiment of this application:
[0082] Table 1
[0083] Address Location identifier Noise optimization model identifier YY Building, XX District 567678 MX-68465 XY Building, ZZ District 877454 MX-68968 … … …
[0084] In Table 1 above, the address, location identifier, and noise optimization model identifier all have indexing functions, and each address, location identifier, and noise optimization model identifier is associated with each other. When the terminal device locates the position in the current meeting environment through the positioning function and the position belongs to a certain address in Table 1 above, the associated noise optimization model is called to process the logarithmic energy spectrum feature of the meeting audio signal to be processed, and the noise correction coefficient of the logarithmic energy spectrum feature of the meeting audio signal is obtained. Optionally, the terminal device adds the logarithmic energy spectrum feature of the current meeting audio signal to an audio training dataset, and re-optimizes the noise optimization model corresponding to the position in the current meeting environment through the updated audio training data. The specific optimization method can refer to Figure 4a Steps S404 and S405 in, which will not be elaborated here.
[0085] Correspondingly, if the terminal device does not detect through the positioning function that the position in the current meeting environment belongs to a certain address in Table 1 above, it collects the environmental audio signal in the current meeting environment; for example, when the user participating in the meeting is not detected to speak, it collects the environmental audio signal (i.e., the echo of the current meeting environment). And the environmental audio signal is superimposed with the human voice audio signal to obtain new audio training data (i.e., the audio training data for the current meeting environment), and the noise optimization model is optimized and trained through the new audio training data to obtain an optimized noise optimization model (i.e., the noise optimization model for the current meeting environment). The specific implementation process of the optimization training can refer to the process of training the noise optimization model in Steps S401 - S405, which will not be elaborated here. The position attribute of the current meeting environment is associated and stored with the optimized noise optimization model in the above relationship mapping table.
[0086] As described in the foregoing embodiments, for different types of recording devices, that is, the voice input devices corresponding to the corresponding clients, there can also be different types of noise optimization models. Therefore, in one embodiment, when the user participates in a meeting through the meeting application, it can also be determined the type of the voice input device that currently collects the user's voice. If the type is a smart terminal such as a mobile phone, the noise optimization model called in S603 is the optimized noise optimization model corresponding to the smart terminal. If the type is a meeting - specific device type such as an octopus, the noise optimization model called in S603 is the optimized noise optimization model corresponding to the meeting - specific device type.
[0087] In one embodiment, a smart terminal such as a mobile phone may externally connect other voice input devices through wireless connection methods such as Bluetooth. Then, it is possible to determine the device currently connected to the smart terminal wirelessly. If it is a voice input device (such as a microphone or a dedicated conference device, etc.), the noise optimization model is also selected based on the type of the externally connected voice input device. If the type of the voice input device corresponding to the current conference application is not detected, the corresponding processing can be performed using the default or randomly selected noise optimization model. The mapping relationship between the device type identifier and the noise optimization model can be established in the form of a mapping table to find and call the corresponding noise optimization model.
[0088] S604: Calculate the N-dimensional logarithmic energy spectrum feature and the M-dimensional noise correction coefficient to obtain the processed audio signal.
[0089] In one implementation manner, after obtaining the processed audio signal, the embodiment of the present invention may further perform the following steps.
[0090] S605: Perform antilogarithmic transformation on the processed audio signal to obtain the conference audio signal.
[0091] S606: Encode the conference audio signal and send the encoded conference audio signal to the terminal devices logged in by each participating account on the conference session interface.
[0092] In the embodiments of the present application, the noise optimization model is secondarily optimized according to different multi-terminal communication environments, so that the optimized noise optimization model is more targeted, thereby further improving the echo cancellation effect of the optimized noise optimization model on the current conference environment. In addition, by associatively storing the environmental attributes of different environments with the corresponding noise optimization models, when the user conducts a remote conference communication next time, the corresponding noise optimization model can be quickly called, further enhancing the user experience.
[0093] The above has elaborated in detail the method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the following provides the device of the embodiments of the present application.
[0094] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a signal processing device provided by the embodiments of the present application. This device can be mounted on the smart device in the above method embodiments. The smart device can specifically be Figure 1a the terminal device 101 shown in Figure 8 The shown signal processing device can be used to execute the above Figure 2 , Figure 4a , Figure 4b andFigure 6 Some or all of the functions in the described method embodiments. Among them, the detailed descriptions of each unit are as follows:
[0095] An acquisition unit 801, configured to collect an audio signal to be processed and extract spectral features of the audio signal to be processed, where the spectral features include N-dimensional logarithmic energy spectral features;
[0096] A processing unit 802, configured to call a noise optimization model to process the logarithmic energy spectral features to obtain M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral features, where N and M are positive integers; calculate the N-dimensional logarithmic energy spectral features and the M-dimensional noise correction coefficients to obtain a processed audio signal;
[0097] Among them, the noise optimization model is trained according to audio training data including noisy audio signals, and the M-dimensional noise correction coefficients output by the noise optimization model include: p-dimensional coefficients for correcting features of the input logarithmic energy spectral features related to the noisy audio signals, where p is less than M.
[0098] In one embodiment, the processing unit 802 is further configured to:
[0099] Collect a noisy audio signal in a target environment where an audio signal is being played;
[0100] Obtain a human voice audio signal;
[0101] Superimpose the obtained human voice audio signal and the noisy audio signal in the time domain to obtain a mixed audio signal, and generate audio training data according to the mixed audio signal;
[0102] Among them, the audio training data includes X segments of mixed audio signals, and the i-th segment of mixed audio signal includes a human voice audio signal and a noisy audio signal, where i and X are positive integers, and i is less than or equal to X.
[0103] In one embodiment, the processing unit 802 is further configured to:
[0104] Perform an audio recording operation in multiple target environments where an audio signal is being played to obtain multiple segments of noisy audio information, and each segment of noisy audio information includes a noisy audio signal and recording device information;
[0105] Generate audio training data corresponding to the recording device information according to the multiple segments of noisy audio information;
[0106] Among them, the audio training data includes Y segments of noisy audio signals, where Y is a positive integer.
[0107] In one embodiment, the noise optimization model is obtained by optimizing an initial model with a loss function constructed based on the mean square error between the first clean logarithmic spectrum feature and the second clean logarithmic spectrum feature; the first clean logarithmic spectrum feature is obtained by multiplying the mixed audio signal in the audio training data by a training noise correction coefficient output after processing the mixed audio signal in the audio training data through the initial model, and the second clean spectrum feature is obtained from the human voice audio signal.
[0108] In one embodiment, the training noise correction coefficient output by the constructed initial model is used to reflect the ratio of the logarithmic spectrum energy of the human voice audio signal corresponding to the second clean logarithmic spectrum feature to the logarithmic spectrum energy of the mixed audio signal; wherein, the logarithmic spectrum energy of the mixed audio signal is the sum of the logarithmic spectrum energy of the noise audio in the mixed audio signal and the logarithmic spectrum energy of the human voice audio signal in the mixed audio signal.
[0109] In one embodiment, the audio signal to be processed is collected when it is detected that the meeting session interface is entered, and the processed audio signal refers to the signal obtained by multiplying the N-dimensional logarithmic energy spectrum feature and the M-dimensional noise correction coefficient. The processing unit 802 is further configured to:
[0110] Perform an inverse logarithmic transform on the processed audio signal to obtain a conference audio signal;
[0111] Encode the conference audio signal and send the encoded conference audio signal to each participating account corresponding to the meeting session interface.
[0112] In one embodiment, the processing unit 802 is further configured to:
[0113] Collect an environmental audio signal when a voice signal from a participating account is detected;
[0114] Use the environmental audio signal as new audio training data to optimize and train the noise optimization model to obtain an optimized noise optimization model;
[0115] Record the optimized noise optimization model for subsequent processing of the logarithmic energy spectrum feature of the audio signal to be processed collected according to the optimized noise optimization model.
[0116] In one embodiment, at least two noise optimization models are recorded in the stored noise optimization model set; the processing unit 802 is specifically configured to: call a noise optimization model to process the logarithmic energy spectrum feature;
[0117] Select a noise optimization model from the set of noise optimization models according to the location attribute of the location in the current conference environment;
[0118] Call the selected noise optimization model to process the logarithmic energy spectrum feature.
[0119] In one embodiment, before extracting the spectrum feature of the audio signal to be processed, the processing unit 802 is further configured to:
[0120] Determine the space type of the current location;
[0121] If the space type is the first type, perform encoding processing on the collected audio signal to be processed to obtain an encoded audio signal;
[0122] If the space type is the second type, trigger the step of extracting the spectrum feature of the audio signal to be processed.
[0123] According to an embodiment of the present application, Figure 2 , Figure 4a , Figure 4b and Figure 6 Some of the steps involved in the signal processing method shown can be executed by each unit in the signal processing device shown in Figure 8 . For example, Figure 2 The step S201 shown in can be executed by the acquisition unit 801 shown in Figure 8 , and the steps S202 and S203 can be executed by the processing unit 802 shown in Figure 8 . Figure 4a The steps S401 and S402 shown in can be executed by the acquisition unit 801 shown in Figure 8 , and the steps S403 to S405 can be executed by the processing unit 802 shown in Figure 8 . Figure 4b The step S411 shown in can be executed by the acquisition unit 801 shown in Figure 8 , and the steps S412 to S414 can be executed by the processing unit 802 shown in Figure 8 . Figure 6 The step S601 shown in can be executed by the acquisition unit 801 shown in Figure 8 , and the steps S602 to S606 can be executed by the processing unit 802 shown in Figure 8 . Figure 8Each unit in the signal processing device shown can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the signal processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.
[0124] According to another embodiment of this application, it can be achieved by running, on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 2 , Figure 4a , Figure 4b and Figure 6 to construct the signal processing device shown in Figure 8 and to implement the signal processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0125] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the signal processing device provided in the embodiments of this application are similar to those of the signal processing device in the method embodiments of this application. For the principle and beneficial effects of the method implementation, refer to them. For the sake of brevity, they will not be elaborated here.
[0126] Please refer to Figure 9 , Figure 9A schematic structural diagram of an intelligent device provided by an embodiment of the present application. The intelligent device at least includes a processor 901, a communication interface 902, and a memory 903. Among them, the processor 901, the communication interface 902, and the memory 903 can be connected through a bus or other means. Among them, the processor 901 (or Central Processing Unit, CPU) is the computing core and control core of the terminal. It can parse various instructions in the terminal and process various data of the terminal. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the terminal and control the terminal to perform power-on and power-off operations. Another example is that the CPU can transmit various interaction data between the internal structures of the terminal, and so on. The communication interface 902 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.). Under the control of the processor 901, it can be used to send and receive data. The communication interface 902 can also be used for the transmission and interaction of internal data of the terminal. The memory 903 (Memory) is a memory device in the terminal, used to store programs and data. It can be understood that the memory 903 here can include both the built-in memory of the terminal and, of course, the extended memory supported by the terminal. The memory 903 provides a storage space, and this storage space stores the operating system of the terminal, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.
[0127] In the embodiment of the present application, the processor 901 performs the following operations by running the executable program code in the memory 903:
[0128] Collect an audio signal to be processed through the communication interface 902, and extract the spectral features of the audio signal to be processed. The spectral features include N-dimensional logarithmic energy spectral features;
[0129] Call a noise optimization model to process the logarithmic energy spectral features to obtain M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral features. N and M are positive integers;
[0130] Calculate the N-dimensional logarithmic energy spectral features and the M-dimensional noise correction coefficients to obtain a processed audio signal;
[0131] Among them, the noise optimization model is trained according to audio training data including noisy audio signals. The M-dimensional noise correction coefficients output by the noise optimization model include: p-dimensional coefficients used to correct the features of the input logarithmic energy spectral features regarding the noisy audio signals, where p is less than M.
[0132] As an optional embodiment, the processor 901 also performs the following operations:
[0133] Collect the noise audio signal in the target environment where the audio signal is being played;
[0134] Obtain the human voice audio signal;
[0135] Superimpose the obtained human voice audio signal and the noise audio signal in the time domain to obtain a mixed audio signal, and generate audio training data according to the mixed audio signal;
[0136] Wherein, the audio training data includes X segments of mixed audio signals, and the i-th segment of mixed audio signal includes a human voice audio signal and a noise audio signal, where i and X are positive integers, and i is less than or equal to X.
[0137] As an optional embodiment, the processor 901 also performs the following operations:
[0138] Perform an audio recording operation in multiple target environments where the audio signal is being played to obtain multiple segments of noise audio information, and each segment of noise audio information includes a noise audio signal and recording device information;
[0139] Generate the audio training data corresponding to the recording device information according to the multiple segments of noise audio information;
[0140] Wherein, the audio training data includes Y segments of noise audio signals, where Y is a positive integer.
[0141] As an optional embodiment, the noise optimization model is obtained by optimizing the initial model with a loss function constructed based on the mean square error between the first clean logarithmic spectral feature and the second clean logarithmic spectral feature; the first clean logarithmic spectral feature is obtained by multiplying the mixed audio signal in the audio training data by the training noise correction coefficient output after processing the mixed audio signal in the audio training data through the initial model, and the second clean spectral feature is obtained according to the human voice audio signal.
[0142] As an optional embodiment, the training noise correction coefficient output by the constructed initial model is used to reflect the ratio of the logarithmic spectral energy of the human voice audio signal corresponding to the second clean logarithmic spectral feature to the logarithmic spectral energy of the mixed audio signal; wherein, the logarithmic spectral energy of the mixed audio signal is: the sum of the logarithmic spectral energy of the noise audio in the mixed audio signal and the logarithmic spectral energy of the human voice audio signal in the mixed audio signal.
[0143] As an optional embodiment, the audio signal to be processed is collected when it is detected that the meeting session interface is entered, and the processed audio signal refers to the signal obtained by multiplying the N-dimensional logarithmic energy spectral feature and the M-dimensional noise correction coefficient. The processor 901 also performs the following operations:
[0144] Perform an antilogarithmic transformation on the processed audio signal to obtain a conference audio signal;
[0145] Encode the conference audio signal and send the encoded conference audio signal to each participating account corresponding to the conference session interface.
[0146] As an optional embodiment, the processor 901 further performs the following operations:
[0147] When detecting a voice signal from a participating account, collect an environmental audio signal;
[0148] Use the environmental audio signal as new audio training data to optimize and train the noise optimization model, and obtain an optimized noise optimization model;
[0149] Record the optimized noise optimization model for subsequent processing of the logarithmic energy spectrum features corresponding to the collected audio signal to be processed according to the optimized noise optimization model.
[0150] As an optional embodiment, at least two noise optimization models are recorded in the stored noise optimization model set; a specific embodiment in which the processor 901 calls a noise optimization model to process the logarithmic energy spectrum features is as follows:
[0151] Select a noise optimization model from the noise optimization model set according to the location attribute of the location where the current conference environment is located;
[0152] Call the selected noise optimization model to process the logarithmic energy spectrum features.
[0153] As an optional embodiment, before extracting the spectrum features of the audio signal to be processed, the processor 901 further performs the following operations:
[0154] Judge the space type of the current location;
[0155] If the space type is the first type, perform an encoding process on the collected audio signal to be processed to obtain an encoded audio signal;
[0156] If the space type is the second type, trigger the step of extracting the spectrum features of the audio signal to be processed.
[0157] Based on the same inventive concept, the principle of solving problems and the beneficial effects of the intelligent device provided in the embodiments of the present application are similar to the principle of solving problems and the beneficial effects of the signal processing method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity, they will not be described here again.
[0158] An embodiment of the present application further provides a computer-readable storage medium, in which one or more instructions are stored, and the one or more instructions are adapted to be loaded and executed by a processor to perform the signal processing method described in the foregoing method embodiment.
[0159] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the various methods mentioned in the foregoing embodiments.
[0160] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0161] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0162] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0163] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The readable storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0164] The foregoing disclosure is only a preferred embodiment of the present application, and of course cannot be used to limit the scope of rights of the present application. Those of ordinary skill in the art can understand the entire or part of the processes of the foregoing embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. A signal processing method, characterized in that, The method includes: Collecting an audio signal to be processed and determining the spatial type of the environment at the current location; If the spatial type of the environment at the current location is the first type, encoding the collected audio signal to be processed and sending the encoded conference audio signal to the participants; If the spatial type of the environment at the current location is the second type, extracting the spectral features of the audio signal to be processed, where the spectral features include N-dimensional logarithmic energy spectral features; Invoking a noise optimization model to process the N-dimensional logarithmic energy spectral features to obtain M-dimensional noise correction coefficients corresponding to the N-dimensional logarithmic energy spectral features, where N and M are positive integers; Calculating the N-dimensional logarithmic energy spectral features and the M-dimensional noise correction coefficients to obtain a processed audio signal; Wherein, the noise optimization model is trained according to audio training data including noisy audio signals, and the audio training data is obtained according to the collected echo audio signals and human voice audio signals. The echo audio signals refer to any one or more of the echo audio signals collected in different environments, the echo audio signals collected when playing audio at different volume levels, and the echo audio signals collected when different playback devices play audio in the same or different target environments. The M-dimensional noise correction coefficients output by the noise optimization model include: p-dimensional coefficients for correcting the features of the noisy audio signal in the input N-dimensional logarithmic energy spectral features, where p is less than M; Wherein, during the process of training the noise optimization model with the audio training data, the logarithmic energy spectral features for training include the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal. The noise correction coefficients during training correspondingly include the noise correction coefficients for the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficients for the logarithmic energy spectrum of the p-dimensional echo audio signal. The logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal are arranged in sequence or cross-mixed, and the arrangement mode of the noise correction coefficients for the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficients for the logarithmic energy spectrum of the p-dimensional echo audio signal corresponds to the arrangement mode of the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal in the N-dimensional logarithmic energy spectral features.
2. The method according to claim 1, characterized in that, The method further includes: Collecting a noisy audio signal in a target environment where an audio signal is being played; Obtaining a human voice audio signal; Superimposing the obtained human voice audio signal and the noisy audio signal in the time domain to obtain a mixed audio signal, and generating audio training data according to the mixed audio signal; Wherein, the audio training data includes X segments of mixed audio signals, and the i-th segment of the mixed audio signal includes a human voice audio signal and a noisy audio signal, where i and X are positive integers, and i is less than or equal to X.
3. The method according to claim 1, wherein The method further includes: Performing an audio recording operation in multiple target environments where an audio signal is being played to obtain multiple segments of noisy audio information, and each segment of noisy audio information includes a noisy audio signal and recording device information; Generating audio training data corresponding to each recording device information according to the multi-segment noisy audio information; Among them, the audio training data includes Y segments of noisy audio signals, where Y is a positive integer.
4. The method according to claim 2, characterized in that, The noise optimization model is obtained by optimizing the initial model with a loss function constructed based on the mean square error between the first clean logarithmic spectrum feature and the second clean logarithmic spectrum feature; The first clean logarithmic spectrum feature is obtained by multiplying the mixed audio signal in the audio training data by the training noise correction coefficient output after processing the mixed audio signal in the audio training data through the initial model, and the second clean logarithmic spectrum feature is obtained according to the human voice audio signal.
5. The method according to claim 4, characterized in that, The training noise correction coefficient output by the constructed initial model is used to reflect the ratio of the logarithmic spectrum energy of the human voice audio signal corresponding to the second clean logarithmic spectrum feature to the logarithmic spectrum energy of the mixed audio signal; Among them, the logarithmic spectrum energy corresponding to the mixed audio signal is: the sum of the logarithmic spectrum energy of the noisy audio in the mixed audio signal and the logarithmic spectrum energy of the human voice audio signal in the mixed audio signal.
6. The method according to claim 1, wherein The audio signal to be processed is collected when it is detected that the meeting session interface is entered, and the processed audio signal refers to the signal obtained by multiplying the N-dimensional logarithmic energy spectrum feature and the M-dimensional noise correction coefficient. The method further includes: Performing an inverse logarithmic transformation on the processed audio signal to obtain a conference audio signal; Encoding the conference audio signal and sending the encoded conference audio signal to each participating account corresponding to the conference session interface.
7. The method according to claim 6, wherein The method further includes: Collecting an environmental audio signal when a sound signal from a participating account is detected; Using the environmental audio signal as new audio training data to optimize and train the noise optimization model to obtain an optimized noise optimization model; Recording the optimized noise optimization model for subsequent processing of the logarithmic energy spectrum feature corresponding to the audio signal to be collected according to the optimized noise optimization model.
8. The method according to claim 7, wherein At least two noise optimization models are recorded in the stored noise optimization model set; The step of using the noise optimization model to process the logarithmic energy spectrum feature includes: Selecting a noise optimization model from the noise optimization model set according to the location attribute of the current position in the conference environment; Invoking the selected noise optimization model to process the logarithmic energy spectrum feature.
9. A signal processing device, characterized in that, Including: An acquisition unit for collecting an audio signal to be processed and extracting the spectrum feature of the audio signal to be processed, where the spectrum feature includes an N-dimensional logarithmic energy spectrum feature; A processing unit for using a noise optimization model to process the logarithmic energy spectrum feature to obtain an M-dimensional noise correction coefficient corresponding to the N-dimensional logarithmic energy spectrum feature, where N and M are positive integers; calculating the N-dimensional logarithmic energy spectrum feature and the M-dimensional noise correction coefficient to obtain a processed audio signal; The processing unit is further configured to determine the spatial type of the environment at the current location; if the spatial type of the environment at the current location is the first type, encode the collected audio signal to be processed, and send the encoded conference audio signal to the participants; if the spatial type of the environment at the current location is the second type, trigger the obtaining unit to extract the spectral features of the audio signal to be processed. Wherein, the noise optimization model is trained according to audio training data including the characteristics of the noise audio signal. The audio training data is obtained from the collected echo audio signal and human voice audio signal. The echo audio signal refers to any one or more of the echo audio signals collected in different environments, the echo audio signals collected when playing audio at different volume levels, and the echo audio signals collected when different playback devices play audio in the same or different target environments. The M-dimensional noise correction coefficients output by the noise optimization model include: a p-dimensional coefficient for correcting the characteristics of the noise audio signal in the input logarithmic energy spectral features, where p is less than M. Wherein, during the process of training the noise optimization model with the audio training data, the logarithmic energy spectral features for training include the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal. The noise correction coefficients during training correspondingly include the noise correction coefficients for the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficients for the logarithmic energy spectrum of the p-dimensional echo audio signal. The logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal are arranged in sequence one after another or arranged in a cross-mixed manner. The arrangement manner of the noise correction coefficients for the logarithmic energy spectrum of the n-dimensional human voice audio signal and the noise correction coefficients for the logarithmic energy spectrum of the p-dimensional echo audio signal corresponds to the arrangement manner of the logarithmic energy spectrum of the n-dimensional human voice audio signal and the logarithmic energy spectrum of the p-dimensional echo audio signal in the N-dimensional logarithmic energy spectral features.
10. A signal processing device, characterized in that, Including: A memory that stores computer-readable instructions. A processor connected to the memory, and the processor is configured to execute the computer-readable instructions to implement the signal processing method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to implement the signal processing method according to any one of claims 1-8.
12. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the signal processing method according to any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Audio optimization method and device
CN109087659A
Voice signal noise reduction processing method, microphone and electronic equipment
CN111223493A