Artificial intelligence-based audio processing method, apparatus, electronic device, computer-readable storage medium, and computer program product

MY214863AActive Publication Date: 2026-08-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MYPI2023003148
Authority / Receiving Office
MY · MY
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-11-17
Publication Date
2026-08-18
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

The existing technology will reduce the quality of audio signals when suppressing noise in audio processing, and the noise reduction algorithm based on universal noise is not robust enough in different scenarios and cannot effectively adapt to the quality fluctuations of the network environment.

Method used

An audio processing method based on artificial intelligence is used to obtain the audio clips of the audio scene, perform audio scene classification processing, determine the target audio processing mode, and apply adaptive noise reduction and code rate switching processing modes according to the degree of noise interference.

Benefits of technology

It improves the accuracy and quality of audio processing, retains useful signals to the greatest extent, adapts to noise environments in different scenarios, and improves user experience and call quality.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An artificial intelligence-based audio processing method includes: obtaining (101) an audio clip of an audio scene, the audio clip including noise; performing (102) audio scene classification processing based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip; and determining (103) a target audio processing mode corresponding to the audio scene type, and applying the target audio processing mode to the audio clip of the audio scene according to a degree of interference caused by the noise in the audio clip. Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence-based audio processing method, device, electronic device, computer-readable storage medium, and computer program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The embodiments of this application are based on the Chinese patent application with application number 202011410814.4 and application date December 3, 2020, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into the embodiments of this application as a reference. Technical Field

[0003] The present application relates to cloud technology and artificial intelligence technology, and in particular to an audio processing method, device, electronic device, computer-readable storage medium, and computer program product based on artificial intelligence. Background Art

[0004] Artificial Intelligence (AI) is a comprehensive field in computer science. By studying the design principles and implementation methods of various intelligent machines, AI enables them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, including natural language processing and machine learning / deep learning. With technological advancements, AI will be applied in more areas and play an increasingly important role. For example, in cloud-based online conferencing scenarios, AI technology can be introduced to improve audio quality.

[0005] However, the audio processing method in the related art is relatively simple, which can suppress the noise in the audio, but inevitably reduces the quality of the useful signal (such as the voice signal) in the audio.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide an artificial intelligence-based audio processing method, device, electronic device, computer-readable storage medium, and computer program product, which can improve audio quality based on targeted audio processing of audio scenes.

[0008] The technical solution of the embodiment of the present application is implemented as follows:

[0009] The present invention provides an artificial intelligence-based audio processing method, including:

[0010] Acquire an audio segment of an audio scene, wherein the audio segment includes noise;

[0011] performing audio scene classification processing based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment;

[0012] A target audio processing mode corresponding to the audio scene type is determined, and the target audio processing mode is applied to the audio segment according to the interference degree caused by the noise in the audio segment.

[0013] The present invention provides an audio processing device, including:

[0014] an acquisition module configured to acquire an audio segment of an audio scene, wherein the audio segment includes noise;

[0015] a classification module configured to perform audio scene classification processing based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment;

[0016] The processing module is configured to determine a target audio processing mode corresponding to the audio scene type, and apply the target audio processing mode to the audio segment according to the interference degree caused by the noise in the audio segment.

[0017] An embodiment of the present application provides an electronic device for audio processing, the electronic device comprising:

[0018] a memory for storing executable instructions;

[0019] The processor is used to implement the artificial intelligence-based audio processing method provided in the embodiment of the present application when executing the executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute and implement the artificial intelligence-based audio processing method provided in the embodiment of the present application.

[0021] An embodiment of the present application provides a computer program product, including a computer program or instructions, which enable a computer to execute the artificial intelligence-based audio processing method provided by the embodiment of the present application.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] By associating noise with audio scene types, the audio scene type corresponding to the audio is identified, and targeted audio processing is performed based on the audio scene type. This allows the audio processing mode introduced into the audio scene to adapt to the noise included in the audio scene, thereby preserving the useful information in the audio to the greatest extent and improving the accuracy of audio processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic diagram of an application scenario of an audio processing system provided in an embodiment of the present application;

[0025] FIG2 is a schematic structural diagram of an electronic device for audio processing provided in an embodiment of the present application;

[0026] Figures 3 to 5 are flowcharts of an artificial intelligence-based audio processing method according to an embodiment of the present application;

[0027] FIG6 is a schematic diagram of the structure of a neural network model provided in an embodiment of the present application;

[0028] FIG7 is a schematic diagram of the overall process of audio scene recognition provided by an embodiment of the present application;

[0029] FIG8 is a schematic diagram of a process for extracting spectral features from a time-domain sound signal according to an embodiment of the present application;

[0030] FIG9 is a schematic diagram of the structure of a neural network model provided in an embodiment of the present application;

[0031] FIG10 is a schematic diagram of the structure of the ResNet unit provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0033] 1) Convolutional Neural Networks (CNN): A type of feedforward neural network (FNN) with a deep structure that incorporates convolutional computations, CNNs are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification of input images based on their hierarchical structure.

[0034] 2) Residual Network (ResNet): A convolutional neural network that is easy to optimize and can improve accuracy by increasing its depth. Its internal residual blocks use skip connections, which alleviates the vanishing gradient problem caused by increasing the depth of deep neural networks.

[0035] 3) Audio Processing Mode: A mode for performing audio processing. Applying an audio processing mode to an audio clip can optimize the audio to obtain clear and smooth audio. The audio processing modes in the embodiments of the present application include a noise reduction processing mode and a bit rate switching processing mode.

[0036] Among related technologies, Adaptive Bitrate Streaming (ABR) can adaptively adjust the video bitrate. The bitrate adjustment algorithm is mostly used for video playback, automatically adjusting the video bitrate (i.e., clarity) according to network conditions or client playback buffer conditions. The noise reduction algorithm based on universal noise uses the spectral characteristics of noisy speech as the neural network input and clean speech as the reference output of the neural network to train the noise reduction model. The minimum mean square error (LMS) is used as the optimization target. After the noise reduction function is turned on, the same noise reduction method is used for various scene environments.

[0037] During the implementation of this application, the applicant discovered that due to the frequent fluctuations in network quality, the bitrate switching based on network speed also fluctuates frequently. Frequent switching of clarity significantly affects the user experience. In real-world environments, the specific noise in specific scenarios places higher demands and challenges on the robustness of noise reduction algorithms based on universal noise.

[0038] In order to solve the above problems, the embodiments of the present application provide an audio processing method, device, electronic device and computer-readable storage medium based on artificial intelligence, which can improve audio quality based on targeted audio processing of audio scenes.

[0039] The artificial intelligence-based audio processing method provided in the embodiments of the present application can be implemented by the terminal / server alone; it can also be implemented collaboratively by the terminal and the server. For example, the terminal alone undertakes the artificial intelligence-based audio processing method described below, or terminal A sends an optimization request for audio (including an audio clip) to the server, and the server executes the artificial intelligence-based audio processing method based on the received optimization request for audio. In response to the optimization request for audio, the server applies the target audio processing mode (including a noise reduction processing mode and a bit rate switching processing mode) to the audio clip of the audio scene, and sends the processed audio clip to terminal B, so that a clear voice call can be made between terminal A and terminal B.

[0040] The electronic device for audio processing provided in the embodiments of the present application can be various types of terminal devices or servers, wherein the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services; the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0041] Taking servers as an example, it can be a server cluster deployed in the cloud, opening artificial intelligence cloud services (AiaaS, AI as a Service) to users. The AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to an AI theme mall. All users can access and use one or more artificial intelligence services provided by the AIaaS platform through the application programming interface.

[0042] For example, one of the artificial intelligence cloud services may be an audio processing service, that is, a server in the cloud encapsulates an audio processing program provided by an embodiment of the present application. A user calls the audio processing service in the cloud service through a terminal (running a client, such as a recording client or an instant messaging client), so that the server deployed in the cloud calls the encapsulated audio processing program, based on a target audio processing mode that matches the audio scene type, and applies the target audio processing mode to the audio clips of the audio scene.

[0043] As an application example, for a recording client, a user may be a contracted anchor on an audio platform who needs to regularly publish audio of an audiobook. However, the anchor's recording scenes may vary, such as recording at home, in a library, or even outdoors, each of which has different noise levels. By recording the current audio scene, audio scene recognition is performed on the recorded audio clip to determine the audio scene type. Based on the audio scene type, targeted noise reduction processing is performed on the audio clip to store the denoised audio clip, thus realizing the denoised recording function.

[0044] As another application example, in an instant messaging client, users can send voice messages to a specific friend or a specific group. However, the user's current scene may vary, such as an office or a shopping mall, and different scenes may have different noise levels. By performing voice scene recognition on the current scene to determine the voice scene type, and then performing targeted noise reduction processing on the voice based on the voice scene type, the denoised audio clip can be sent, thus realizing the denoised voice sending function.

[0045] As another application example, for a conference client, users participating in a conference may be in different environments. For example, user A may be in an office, while user B may be on a high-speed train. Different scenarios present different noise levels and communication signals, such as the sound of trains and a poor communication signal. Voice scene recognition is performed on each participant's voice call to determine the type of voice scene. Based on the voice scene type, targeted bitrate switching is performed for each participant's voice, achieving adaptive bitrate switching and improving conference call quality.

[0046] Refer to Figure 1, which is a schematic diagram of an application scenario of an audio processing system 10 provided in an embodiment of the present application. The terminal 200 is connected to the server 100 through a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0047] Terminal 200 (running a client, such as a recording client, an instant messaging client, a call client, etc.) can be used to obtain an optimization request for audio. For example, when a user inputs or records an audio clip of an audio scene through terminal 200, terminal 200 automatically obtains the audio clip of the audio scene and automatically generates an optimization request for audio.

[0048] In some embodiments, an audio processing plug-in may be implanted in the client running in the terminal to implement an audio processing method based on artificial intelligence locally on the client. For example, after the terminal 200 obtains an optimization request for audio (including an audio clip of an audio scene), it calls the audio processing plug-in to implement an audio processing method based on artificial intelligence, identifies the audio scene type corresponding to the noise in the audio clip, and based on the target audio processing mode that matches the audio scene type, applies the target audio processing mode to the audio clip of the audio scene. For example, for a recording application, the user records in the current audio scene, performs audio scene recognition on the recorded audio clip to determine the audio scene type, and performs targeted noise reduction processing on the audio clip based on the audio scene type to store the denoised audio clip to implement a denoising recording function.

[0049] In some embodiments, after the terminal 200 obtains the optimization request for the audio, it calls the audio processing interface of the server 100 (which can be provided in the form of a cloud service, i.e., an audio processing service). The server 100 identifies the audio scene type corresponding to the noise in the audio clip, and based on the target audio processing mode that matches the audio scene type, applies the target audio processing mode to the audio clip of the audio scene, and sends the audio clip that passes the target audio processing mode (the optimized audio clip) to the terminal 200 or other terminals. For example, for a recording application, the user records in the current audio scene, the terminal 200 obtains the corresponding audio clip, and automatically generates an optimization request for the audio, and sends the optimization request for the audio to the server 100. The server 100 performs audio scene recognition on the recorded audio clip based on the optimization request for the audio to determine the audio scene type, and performs targeted noise reduction processing on the audio clip based on the audio scene type to store the denoised audio clip and realize the denoised recording function. For an instant messaging application, the user sends voice in the current voice scene, and the terminal 200 obtains the audio scene type. The terminal 200 obtains the audio clip corresponding to the user A, and automatically generates an optimization request for the audio, and sends the optimization request for the audio to the server 100. The server 100 performs voice scene recognition on the audio clip based on the optimization request for the audio to determine the voice scene type, and performs targeted noise reduction processing on the audio clip based on the voice scene type to send the denoised audio clip, thereby realizing the denoised voice sending function, and performs targeted bit rate switching processing on the audio clip based on the voice scene type to realize adaptive bit rate switching and improve the voice call quality; for the call application, user A and user B have a voice call, and user A has a voice call in the current voice scene. The terminal 200 obtains the audio clip corresponding to user A, and automatically generates an optimization request for the audio, and sends the optimization request for the audio to the server 100. The server 100 performs voice scene recognition on the audio clip of user A based on the optimization request for the audio to determine the voice scene type, and performs targeted noise reduction processing on the audio clip based on the voice scene type, and sends the denoised audio clip of user A to user B, thereby realizing the denoised voice call function.

[0050] The structure of an electronic device for audio processing provided by an embodiment of the present application is described below. Referring to FIG. 2 , FIG. 2 is a schematic diagram of the structure of an electronic device 500 for audio processing provided by an embodiment of the present application. Taking the electronic device 500 as a server as an example, the electronic device 500 for audio processing shown in FIG. 2 includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It will be understood that the bus system 540 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the various buses are all labeled as the bus system 540 in FIG.

[0051] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0052] The memory 550 includes a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory. The memory 550 may optionally include one or more storage devices physically remote from the processor 510.

[0053] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0054] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0055] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;

[0056] In some embodiments, the audio processing device provided in the embodiments of the present application can be implemented in software, for example, it can be an audio processing plug-in in the terminal described above, or it can be an audio processing service in the server described above. Of course, without limitation, the audio processing device provided in the embodiments of the present application can be provided as various software embodiments, including various forms such as application programs, software, software modules, scripts, or code.

[0057] Figure 2 shows an audio processing device 555 stored in a memory 550, which can be software in the form of a program and a plug-in, such as an audio processing plug-in, and includes a series of modules, including an acquisition module 5551, a classification module 5552, a processing module 5553 and a training module 5554; wherein, the acquisition module 5551, the classification module 5552, and the processing module 5553 are used to implement the audio processing function provided in the embodiment of the present application, and the training module 5554 is used to train a neural network model, wherein the audio scene classification processing is implemented through the neural network model.

[0058] As previously mentioned, the artificial intelligence-based audio processing method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG3 , FIG3 is a flow chart of the artificial intelligence-based audio processing method provided in the embodiments of the present application, which will be described in conjunction with the steps shown in FIG3 .

[0059] In the following steps, the audio scene refers to the environment in which audio is generated, such as the external environment of a home environment, an office environment, or various types of transportation such as a high-speed train.

[0060] In step 101, an audio segment of an audio scene is obtained, where the audio segment includes noise.

[0061] As an example of obtaining an audio clip, the user inputs audio through the terminal (running a client) in the current audio scene, the terminal 200 obtains the corresponding audio clip, and automatically generates an optimization request for the audio, and sends the optimization request for the audio to the server. The server parses the optimization request for the audio to obtain the audio clip of the audio scene, so as to subsequently perform audio scene recognition based on the audio clip.

[0062] In step 102, an audio scene classification process is performed based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment.

[0063] For example, after obtaining an audio clip of an audio scene, an audio scene classification process can be performed based on the audio clip through a neural network model to obtain an audio scene type corresponding to the noise in the audio clip. The audio clip can be input into the neural network model, and the time domain features or frequency domain features of the audio clip can also be input into the neural network model. The neural network model performs audio scene classification process based on the time domain features or frequency domain features of the audio clip to obtain an audio scene type corresponding to the noise in the audio clip. Taking the frequency domain features of the audio clip as an example, after obtaining the audio clip, the time domain signal of the audio clip is first framed to obtain a multi-frame audio signal, and then the multi-frame audio signal is windowed, and the audio signal after the windowing process is Fourier transformed to obtain the frequency domain signal of the audio clip, and the Mel frequency band of the frequency domain signal is logarithmically processed to obtain the frequency domain features of the audio clip, that is, the audio clip used for audio scene classification.

[0064] In order to enable the neural network model to process multi-channel input, the frequency domain features of the audio clip obtained by logarithmic processing can be differentiated to obtain the first-order derivative of the audio clip, and then the first-order derivative can be differentiated to obtain the second-order derivative of the audio clip. Finally, the frequency domain features, first-order derivatives and second-order derivatives of the audio clip are combined into a three-channel input signal, and the three-channel input signal is used as the audio clip for audio scene classification.

[0065] In some embodiments, the audio scene classification processing is implemented through a neural network model, which learns the association between the noise included in the audio clip and the audio scene type; the audio scene classification processing is performed based on the audio clip to obtain the audio scene type corresponding to the noise in the audio clip, including: calling the neural network model based on the audio clip to perform audio scene classification processing to obtain the audio scene type that has an association with the noise included in the audio clip.

[0066] For example, as shown in Figure 6, the neural network model includes a mapping network, a residual network and a pooling network; the audio clip is subjected to feature extraction processing through the mapping network to obtain a first feature vector of the noise in the audio clip; the first feature vector is mapped through the residual network to obtain a mapping vector of the audio clip; the mapping vector of the audio clip is subjected to feature extraction processing through the mapping network to obtain a second feature vector of the noise in the audio clip; the second feature vector is pooled through the pooling network to obtain a pooling vector of the audio clip; the pooling vector of the audio clip is subjected to nonlinear mapping processing to obtain an audio scene type that is associated with the noise included in the audio clip.

[0067] Continuing with the above example, the mapping network includes multiple cascaded mapping layers; the audio clip is subjected to feature extraction processing through the mapping network to obtain the first feature vector of the noise in the audio clip, including: performing feature mapping processing on the audio clip through the first mapping layer in the multiple cascaded mapping layers; outputting the mapping result of the first mapping layer to the subsequent cascaded mapping layers, continuing to perform feature mapping and output mapping results through the subsequent cascaded mapping layers until the output is to the last mapping layer, and using the mapping result output by the last mapping layer as the first feature vector of the noise in the audio clip.

[0068] Among them, the mapping network can effectively extract scene noise features in audio clips, and the mapping layer can be a convolutional neural network, but the embodiments of the present application are not limited to convolutional neural networks, and can also be other neural networks.

[0069] In some embodiments, the residual network includes a first mapping network and a second mapping network; mapping processing is performed on the first feature vector through the residual network to obtain a mapping vector of the audio segment, including: mapping processing is performed on the first feature vector through the first mapping network to obtain a first mapping vector of the audio segment; nonlinear mapping processing is performed on the first mapping vector to obtain a non-mapping vector of the audio segment; mapping processing is performed on the non-mapping vector of the audio segment through the first mapping network to obtain a second mapping vector of the audio segment; and the sum of the first feature vector of the audio segment and the second mapping vector of the audio segment is used as the mapping vector of the audio segment.

[0070] Among them, the residual network can effectively prevent the gradient vanishing problem in the error transmission of neural network training to speed up the training of neural network models.

[0071] In some embodiments, it is necessary to train a neural network model so that the trained neural network model can perform audio scene classification. The training method is as follows: based on a noise-free audio signal and background noise corresponding to multiple different audio scenes, audio samples corresponding to multiple different audio scenes are constructed; based on the audio samples corresponding to multiple different audio scenes, the neural network model is trained to obtain a neural network model for audio scene classification.

[0072] In order to enhance the diversity of sample data, the method for constructing audio samples is as follows: performing the following processing for any audio scene among multiple different audio scenes: based on the fusion ratio of the background noise and the noise-free audio signal of the audio scene, the background noise of the audio scene and the noise-free audio signal are fused to obtain a first fused audio signal of the audio scene; the background noise of the audio scene corresponding to the first random coefficient is fused into the first fused audio signal to obtain a second fused audio signal of the audio scene; and the noise-free audio signal corresponding to the second random coefficient is fused into the second fused audio signal to obtain an audio sample of the audio scene.

[0073] For example, after retaining the human voice (noise-free audio signal) and fusing it with the background noise at a fusion ratio of 1:1, a partial random ratio is generated to superimpose the data, for example, the noise superposition coefficient (first random coefficient) is a random number between 0.3 and 0.5, and the human voice superposition coefficient (second random coefficient) is a random number between 0.5 and 0.7.

[0074] In some embodiments, a neural network model is trained based on audio samples corresponding to multiple audio scenes to obtain a neural network model for audio scene classification, including: performing audio scene classification processing on audio samples corresponding to multiple different audio scenes through the neural network model to obtain predicted audio scene types of the audio samples; constructing a loss function of the neural network model based on the predicted audio scene types of the audio samples, the audio scene labels of the audio samples, and the weights of the audio samples; updating the parameters of the neural network model until the loss function converges, and using the updated parameters of the neural network model when the loss function converges as the parameters of the neural network model for audio scene classification.

[0075] For example, after determining the value of the loss function of the neural network model based on the predicted audio scene type of the audio sample, the audio scene label of the audio sample, and the weight of the audio sample, it can be determined whether the value of the loss function of the neural network model exceeds a preset threshold. When the value of the loss function of the neural network model exceeds the preset threshold, the error signal of the neural network model is determined based on the loss function of the neural network model, the error information is back-propagated in the neural network model, and the model parameters of each layer are updated during the propagation process.

[0076] Here, back propagation is explained. The training sample data is input into the input layer of the neural network model, passes through the hidden layer, and finally reaches the output layer and outputs the result. This is the forward propagation process of the neural network model. Since there is an error between the output result of the neural network model and the actual result, the error between the output result and the actual value is calculated, and the error is backpropagated from the output layer to the hidden layer until it propagates to the input layer. During the back propagation process, the value of the model parameter is adjusted according to the error; the above process is continuously iterated until convergence.

[0077] In step 103 , a target audio processing mode corresponding to the audio scene type is determined, and the target audio processing mode is applied to the audio segment according to the interference degree caused by the noise in the audio segment.

[0078] For example, after obtaining the audio scene type, a target audio processing mode matching the audio scene type is first determined, and then the target audio processing mode is applied to the audio clips of the audio scene to perform targeted audio optimization and improve the accuracy of audio processing.

[0079] Refer to Figure 4, which is an optional flow chart of the artificial intelligence-based audio processing method provided in an embodiment of the present application. Figure 4 shows that step 103 in Figure 3 can be implemented through steps 1031A to 1032A shown in Figure 4: the target audio processing mode includes a noise reduction processing mode; in step 1031A, based on the audio scene type corresponding to the audio scene, the correspondence between different candidate audio scene types and candidate noise reduction processing modes is queried to obtain a noise reduction processing mode matching the audio scene type; in step 1032A, the noise type matching the audio scene type is matched with the noise in the audio clip; the noise that successfully matches the noise type is suppressed to obtain a suppressed audio clip; wherein, the ratio of the speech signal strength to the noise signal strength in the suppressed audio clip is lower than the signal-to-noise ratio threshold.

[0080] For example, according to the actual application scenario, a mapping table including the correspondence between different candidate audio scene types and candidate noise reduction processing modes is pre-constructed, and the mapping table is stored in a storage space. By reading the correspondence between different candidate audio scene types and candidate noise reduction processing modes included in the mapping table in the storage space, the noise reduction processing mode that matches the audio scene type can be quickly queried based on the audio scene type corresponding to the audio scene, so as to apply the noise reduction processing mode to the audio clip of the audio scene to remove the noise of the audio clip of the audio scene, thereby achieving targeted denoising and improving the audio quality of the audio clip (i.e., the clarity of the audio).

[0081] In some embodiments, the target audio processing mode includes a noise reduction processing mode; determining the target audio processing mode corresponding to the audio scene type includes: determining the noise type that matches the audio scene type based on the audio scene type corresponding to the audio scene; based on the noise type that matches the audio scene type, querying the correspondence between different candidate noise types and candidate noise reduction processing modes to obtain the noise reduction processing mode corresponding to the audio scene type; wherein, the noise types that match different audio scene types are not exactly the same.

[0082] For example, first determine the noise type that matches the audio scene type through the audio scene type corresponding to the audio scene, and then obtain the noise reduction processing mode that matches the audio scene type through the noise type that matches the audio scene type, that is, decoupling the audio scene type and the candidate noise reduction processing mode, so that the correspondence between the audio scene type and the candidate noise reduction processing mode can be flexibly adjusted subsequently.

[0083] For example, since the client developer's strategy for assigning noise reduction processing modes to different noises may change, or the needs of different users for noise reduction processing modes for different noises may change, therefore, if the mapping relationship between the audio scene type of the audio clip and the noise reduction processing mode is implemented through a neural network model, a large number of models need to be trained specifically, and once the noise reduction processing mode assigned to different noises changes, the neural network model needs to be retrained, which will consume a lot of computing resources.

[0084] However, if the mapping relationship between the audio scene type and noise of the audio clip is realized only through the neural network model, then training a neural network model can meet various requirements for noise reduction processing modes in actual applications. It is only necessary to implement the policy settings of noise type and noise reduction processing mode in the client. Even if the noise reduction processing mode changes for different noise allocations, it is only necessary to adjust the policy settings of noise type and noise reduction processing mode in the client, thereby avoiding consuming a large amount of computing resources to train the neural network model.

[0085] In some embodiments, a target audio processing mode is applied to the audio segment based on the interference level caused by the noise in the audio segment, including: determining the interference level caused by the noise in the audio segment; when the interference level is greater than an interference level threshold, applying a noise reduction processing mode corresponding to the audio scene type to the audio segment of the audio scene.

[0086] For example, when the noise in the audio clip has little impact on the audio clip, noise reduction processing may not be performed. Only when the noise in the audio clip affects the audio clip, noise reduction processing is performed on the audio clip. For example, when the user is recording, although some noise in the audio scene will be included in the recording, this noise does not affect the effect of the recording, so noise reduction processing may not be performed on the recording; when this noise affects the effect of the recording (for example, the content of the recording cannot be heard clearly), noise reduction processing may be performed on the recording.

[0087] Refer to Figure 5, which is an optional flow chart of the artificial intelligence-based audio processing method provided in an embodiment of the present application. Figure 5 shows that step 103 in Figure 3 can be implemented through steps 1031B-1032B shown in Figure 4: the target audio processing mode includes a rate switching processing mode; in step 1031B, based on the audio scene type corresponding to the audio scene, the correspondence between different candidate audio scene types and candidate rate switching processing modes is queried to obtain the rate switching processing mode corresponding to the audio scene type; in step 1032B, the rate switching processing mode corresponding to the audio scene type is applied to the audio clip.

[0088] For example, according to the actual application scenario, a mapping table including the correspondence between different audio scene types and candidate bit rate switching processing modes is pre-constructed, and the mapping table is stored in a storage space. By reading the correspondence between different audio scene types and candidate bit rate switching processing modes included in the mapping table in the storage space, the bit rate switching processing mode that matches the audio scene type can be quickly queried based on the audio scene type corresponding to the audio scene, so as to apply the bit rate switching processing mode to the audio clip of the audio scene to switch the bit rate of the audio clip, thereby realizing targeted bit rate switching and improving the fluency of the audio clip.

[0089] In some embodiments, the target audio processing mode includes a rate switching processing mode; determining the target audio processing mode corresponding to the audio scene type includes: comparing the audio scene type with a preset audio scene type; when the comparison determines that the audio scene type is the preset audio scene type, using the rate switching processing mode associated with the preset audio scene type as the rate switching processing mode corresponding to the audio scene type.

[0090] For example, not all audio scenes require bitrate switching. For example, the communication signal in an office environment is relatively stable and does not require bitrate switching, while the signal in a high-speed rail environment is weak and unstable, so bitrate switching is required. Therefore, before determining the bitrate switching processing mode, it is necessary to compare the audio scene type with the preset audio scene type that requires bitrate switching. When the comparison determines that the audio scene type belongs to the preset audio scene type that requires bitrate switching, the bitrate switching processing mode associated with the preset audio scene type is determined to be the bitrate switching processing mode that matches the audio scene type, so as to avoid resource waste caused by bitrate switching in all scenes.

[0091] In some embodiments, a target audio processing mode is applied to an audio segment of an audio scene, including: when the communication signal strength of the audio scene is less than a communication signal strength threshold, reducing the audio bit rate of the audio segment according to a first set ratio or a first set value; when the communication signal strength of the audio scene is greater than or equal to the communication signal strength threshold, increasing the audio bit rate of the audio segment according to a second set ratio or a second set value.

[0092] Taking a voice call scenario as an example, multiple people are conducting voice calls in different environments and send audio clips to a server through a client. The server receives the audio clips sent by each client, performs audio scene classification processing based on the audio clips to obtain an audio scene type corresponding to the noise in the audio clip, and determines the bitrate switching processing mode that matches the audio scene type. The server then determines the communication signal strength of the audio scene. When the communication signal strength of the audio scene is less than the communication signal strength threshold, it indicates that the signal of the current audio scene is weak and the bitrate needs to be reduced. Therefore, the audio bitrate of the audio clip is reduced according to the first set ratio or first set value in the bitrate switching processing mode that matches the audio scene type, so as to enable subsequent smooth audio interaction and avoid voice call interruption. When the communication signal strength of the audio scene is greater than or equal to the communication signal strength threshold, it indicates that the communication signal of the current audio scene is strong and the call will not be interrupted even if the bitrate is not lowered. Therefore, there is no need to reduce the bitrate. Therefore, the audio bitrate of the audio clip is increased according to the second set ratio or second set value in the bitrate switching processing mode that matches the audio scene type, thereby improving the smoothness of the audio interaction. It should be noted that the first set ratio and the second set ratio can be the same or different; the first set value ratio and the second set value ratio can be the same or different. The first set ratio, the second set ratio, the first set value, and the second set value can be set according to actual needs.

[0093] An example method for obtaining the communication signal strength of an audio scene is as follows: averaging the communication signal strengths obtained from multiple samplings in the audio scene and using the average result as the communication signal strength of the audio scene. For example, the communication signal strength of a user's voice call from the beginning to the present can be averaged, and the average result can be used as the communication signal strength of the user in the audio scene.

[0094] In some embodiments, a target audio processing mode is applied to an audio segment of an audio scene, including: determining jitter information of the communication signal strength in the audio scene based on the communication signal strength obtained by multiple samplings in the audio scene; when the jitter information indicates that the communication signal is in an unstable state, reducing the audio bit rate of the audio segment according to a third set ratio or a third set value.

[0095] For example, the communication signal strength is sampled multiple times in the audio scene, and the jitter change of the communication signal strength in the audio scene (i.e., jitter information) is obtained by normal distribution. When the variance in the normal distribution that characterizes the jitter change is greater than the variance threshold, it means that the data in the normal distribution (i.e., the communication signal) is relatively dispersed, and the communication signal strength jitters violently, which means that the communication signal is unstable. In order to avoid subsequent switching of the audio bit rate back and forth while ensuring audio fluency, the audio bit rate of the audio clip can be reduced according to the preset ratio or preset value in the bit rate switching processing mode that matches the audio scene type. Among them, the third set ratio and the third set value can be set according to actual needs.

[0096] By judging the jitter changes in the communication signal strength in the audio scene, we can further determine whether the bit rate needs to be switched. This can avoid frequent switching of the audio bit rate while ensuring audio smoothness, thereby improving the user experience.

[0097] In some embodiments, applying a target audio processing mode to an audio segment of an audio scene includes reducing the audio bit rate of the audio segment according to a fourth set ratio or a fourth set value when the type of the communication network used to transmit the audio segment belongs to a set type.

[0098] For example, after determining the bitrate switching processing mode that matches the audio scene type, it is also possible to first determine whether the type of communication network used to transmit the audio clip belongs to a set type (e.g., WiFi network, cellular network, etc.). For example, if it is determined that the type of communication network used to transmit the audio clip belongs to a WiFi network, it means that the current audio clip is in an unstable environment. To ensure audio fluency, the audio bitrate of the audio clip can be reduced according to a preset ratio or preset value in the bitrate switching processing mode that matches the audio scene type. The third set ratio and the third set value can be set according to actual needs.

[0099] Below, an exemplary application of the embodiment of the present application in a practical application scenario will be described.

[0100] The embodiments of the present application can be applied to various voice application scenarios. For example, for a recording application, a user records in the current voice scene through a recording client running in a terminal, and the recording client performs voice scene recognition on the recorded audio clip to determine the voice scene type, and performs targeted noise reduction processing on the audio clip based on the voice scene type to store the denoised audio clip and realize the denoised recording function; for an instant messaging application, a user sends a voice in the current voice scene through an instant messaging client running in a terminal, and the instant messaging client obtains the corresponding audio clip, performs voice scene recognition on the audio clip to determine the voice scene type, and performs targeted noise reduction processing on the audio clip based on the voice scene type. The denoised audio clip is sent by the instant communication client to realize the denoised voice sending function; for the call application, user A has a voice call with user B, and user A has a voice call in the current voice scene through the call client running in the terminal. The call client obtains the audio clip of user A, and automatically generates an optimization request for the audio based on the audio clip of user A, and sends the optimization request for the audio to the server. The server performs voice scene recognition on the audio clip of user A based on the received optimization request for audio to determine the voice scene type, and performs targeted noise reduction processing on the audio clip based on the voice scene type, and sends the denoised audio clip of user A to user B to realize the denoised voice call function.

[0101] The embodiment of the present application provides an audio processing method based on artificial intelligence, which extracts the Mel-frequency logarithmic energy features for audio clips, inputs the normalized features into a neural network, and obtains the scene prediction corresponding to the audio clip. Since call scenes are relatively stable, scene-based bit rate control has better stability. According to the noise characteristics of different scenes, adaptive learning, transfer learning, etc. can be used to obtain personalized noise reduction solutions for specific scenes. Switching to a dedicated noise reduction mode for a specific scene based on the results of scene recognition will achieve better noise reduction performance and improve call quality and user experience.

[0102] For example, in real-time communication conferences, with the continuous improvement of mobile conference terminals, users may join conferences in a variety of different environments (audio and voice scenarios), such as offices, homes, and mobile transportation such as subways and high-speed trains. These different scenarios bring unique challenges to real-time audio signal processing. For example, in scenarios like high-speed trains, the signal is weak and unstable, and audio communication often freezes, seriously affecting communication quality. The unique background noise characteristics of different scenarios (such as children playing, TV background noise, and kitchen noise in a home environment) place higher demands on the robustness of noise reduction algorithms.

[0103] In order to meet the needs of users for holding meetings in various scenarios and improve the user's conference audio experience in complex environments, providing scene-personalized solutions based on environmental characteristics is an important trend in the optimization of audio processing algorithms. Accurately identifying the scene where the audio occurs is an important basis and foundation for implementing scene-personalized solutions. The embodiment of the present application proposes an audio scene classification solution, which enables scene-personalized audio processing algorithms based on the classification results of the audio scene. For example, automatic bit rate switching is performed for high-speed rail scenes with weak and unstable signals to reduce the audio bit rate to avoid freezes; noise reduction solutions for specific scenes are applied according to the identified scenes to improve the user's meeting experience.

[0104] Among them, the targeted noise reduction solution mainly targets scene-specific noise, such as keyboard sounds and paper friction sounds in office environments; kitchen noises, children playing, and TV background sounds in home environments; and station announcements in mobile vehicles. Based on the general noise reduction model, noise reduction models for each scene are trained through adaptive methods (models that focus on eliminating scene-specific noise). After the scene is identified, the noise reduction model corresponding to the scene is enabled for noise reduction. Targeted bitrate switching is mainly aimed at specific scenarios. For example, in mobile transportation environments with relatively weak signals such as high-speed rail, the bitrate of conference communications is reduced (for example, from 16k to 8k), reducing the transmission burden, reducing lag, and improving the conference experience.

[0105] The audio scene classification scheme proposed in the embodiment of the present application first extracts the corresponding frequency domain spectrum features, namely the Mel Frequency Log Filterbank Energy, based on the collected time domain audio signal, and then normalizes these spectrum features. After normalization, these normalized spectrum features are input into a neural network model, such as a deep residual network (ResNet, Residue Network) based on a convolutional neural network (CNN, Convolutional Neural Network), and the normalized spectrum features are modeled by the neural network model. During the actual test, the logarithmic energy spectrum of the input audio signal is first normalized and input into the established neural network model. The neural network model outputs the scene classification result for each audio clip (Audio clip) input. Based on the scene results identified by the neural network model, the conference system can automatically switch the adaptive audio bit rate and enable targeted noise reduction solutions suitable for the scene, etc., to improve the overall voice call quality and user experience.

[0106] As shown in Figure 7, Figure 7 is a schematic diagram of the overall process of audio scene recognition provided by an embodiment of the present application. Audio scene recognition includes two stages: training and testing, which include 5 modules, namely: 1) scene noise corpus collection; 2) training data construction; 3) training data feature extraction; 4) neural network model training; 5) scene prediction.

[0107] 1) Scene noise corpus collection

[0108] Collect background noise in different scenarios, such as keyboard sounds and paper friction sounds in office environments; kitchen noises, children playing, and TV background sounds in home environments; and noises with scene characteristics such as station announcements in mobile vehicles.

[0109] 2) Training data construction

[0110] The collected background noise from different scenarios is superimposed with different clean audio (noise-free speech) in the time domain to generate a mixed signal of scene noise and clean audio, which serves as the input for neural network model training. To prevent the superimposed speech amplitude from exceeding the system threshold, enhance data diversity, and better simulate real-world audio, while maintaining the original 1:1 ratio of human voice to noise, some randomized superimposed data is generated. For example, the human voice superimposition coefficient is a random number between 0.5 and 0.7, and the noise superimposition coefficient is a random number between 0.3 and 0.5.

[0111] 3) Training data feature extraction

[0112] The audio signals in the training data are framed, windowed, and Fourier transformed to obtain Mel logarithmic energy spectrum features.

[0113] As shown in Figure 8, Figure 8 is a flow chart of extracting spectral features from time-domain sound signals provided by an embodiment of the present application. First, the time-domain signal is framed based on the short-time transient hypothesis to convert the continuous signal into a discrete vector; thereafter, each frame of the audio signal is windowed and smoothed to eliminate edge discontinuities; each frame is then subjected to a Fourier transform (FT) to obtain a frequency-domain signal; the frequency-domain signal is then subjected to a Mel-band operation to obtain the energy within each frequency band. Based on the nonlinear response of the human ear to audio, the Mel-band is used here instead of the linear band to better simulate the human ear response, and finally a logarithmic operation is performed to obtain the Mel-log spectrum feature.

[0114] 4) Neural network model training

[0115] The input of the neural network model is the three-channel Mel logarithmic energy spectrum features of the scene noise superimposed on the clean audio, and the output of the neural network model is the classification result of the scene recognition. During the training process, the cross entropy error (Cross Entropy Loss) is used as the loss function, and the minimization of the loss function is the training goal: min θ L(θ)=-∑ i t i log o i Among them, t i represents the correct scene annotation of the input audio, o i The scene category predicted by the neural network model.

[0116] As shown in Figure 9, Figure 9 is a structural diagram of the neural network model provided in an embodiment of the present application. The neural network model consists of 2 ResNet units (residual networks), multiple CNN networks and an average pooling layer (Pooling Layer). The Mel logarithmic energy feature and its first-order derivative and second-order derivative constitute a three-channel input signal, and finally output the scene classification result.

[0117] Among them, the neural network model uses a ResNet unit. As shown in Figure 10, Figure 10 is a structural diagram of the ResNet unit provided in an embodiment of the present application. Each ResNet unit contains two layers of CNN, where x and y are the input and output of the residual unit, respectively. f1 and f2 represent the function mapping of the two CNN layers, respectively. W1 and W2 represent the weight parameters (Weights) corresponding to the two CNN layers, respectively. The CNN layer can effectively capture the scene noise characteristics in the spectral information, and the residual network can effectively prevent the gradient disappearance problem in the error transmission of neural network training.

[0118] 5) Scenario Prediction

[0119] After training the neural network model, the optimal model parameters are selected and saved as the trained model. During testing, the noisy speech is normalized, and the spectral features are extracted and input into the trained model, which then outputs the predicted audio scene. Subsequently, based on the audio scene classification results, scenario-specific audio processing algorithms are enabled. For example, automatic bit rate switching is performed for high-speed rail scenarios with weak and unstable signals, reducing the audio bit rate to avoid lag. Based on the identified scenario, scenario-specific noise reduction solutions are applied to improve the user experience when joining a meeting.

[0120] In summary, the embodiments of this application construct a lightweight audio scene recognition model with low storage space requirements and fast prediction speed. As a front-end algorithm, it can serve as the basis and foundation for subsequent optimization of complex algorithms. It can also provide personalized audio solutions based on the audio scene recognition results, such as adjusting the audio bitrate and enabling scene-specific noise reduction schemes.

[0121] So far, the audio processing method based on artificial intelligence provided by the embodiment of the present application has been described in combination with the exemplary application and implementation of the server provided by the embodiment of the present application. The embodiment of the present application also provides an audio processing device. In actual applications, the various functional modules in the audio processing device can be implemented in collaboration with the hardware resources of an electronic device (such as a terminal device, a server or a server cluster), such as computing resources such as a processor, communication resources (such as for supporting various communication modes such as optical cables and cellular), and a memory. Figure 2 shows an audio processing device 555 stored in a memory 550, which can be software in the form of programs and plug-ins, for example, software modules designed in programming languages ​​such as C / C++ and Java, application software designed in programming languages ​​such as C / C++ and Java, or special software modules in large software systems, application program interfaces, plug-ins, cloud services, etc. The following examples illustrate different implementation methods.

[0122] Example 1: The audio processing device is a mobile application and module

[0123] The audio processing device 555 in the embodiment of the present application can be provided as a software module designed using a programming language such as software C / C++, Java, etc., and embedded in various mobile applications based on systems such as Android or iOS (stored in the storage medium of the mobile terminal as executable instructions and executed by the processor of the mobile terminal), thereby directly using the mobile terminal's own computing resources to complete related information recommendation tasks, and regularly or irregularly transmit the processing results to a remote server through various network communication methods, or save them locally on the mobile terminal.

[0124] Example 2: The audio processing device is a server application and platform

[0125] The audio processing device 555 in the embodiment of the present application can be provided as an application software designed using programming languages ​​such as C / C++ and Java, or a dedicated software module in a large software system, running on the server side (stored in the storage medium on the server side in the form of executable instructions and run by the processor on the server side), and the server uses its own computing resources to complete related information recommendation tasks.

[0126] The embodiments of the present application can also be provided as a distributed, parallel computing platform composed of multiple servers, equipped with a customized, easy-to-interact network (Web) interface or other user interfaces (UI, User Interface), to form an information recommendation platform (for recommendation lists) for use by individuals, groups or units.

[0127] Example 3: The audio processing device is a server-side API (Application Program Interface) and plug-in

[0128] The audio processing device 555 in the embodiment of the present application can be provided as a server-side API or plug-in for users to call to execute the artificial intelligence-based audio processing method of the embodiment of the present application and be embedded in various applications.

[0129] Example 4: Audio processing device is a mobile device client API and plug-in

[0130] The audio processing device 555 in the embodiment of the present application can be provided as an API or plug-in on the mobile device for users to call to execute the artificial intelligence-based audio processing method in the embodiment of the present application.

[0131] Example 5: Audio processing device is a cloud-based open service

[0132] The audio processing device 555 in the embodiment of the present application can provide an information recommendation cloud service developed for users, so that individuals, groups or units can obtain recommendation lists.

[0133] The audio processing device 555 includes a series of modules, including an acquisition module 5551, a classification module 5552, a processing module 5553, and a training module 5554. The following further describes how the various modules in the audio processing device 555 provided in the embodiment of the present application cooperate to implement the audio processing solution.

[0134] The acquisition module 5551 is configured to acquire an audio clip of an audio scene, wherein the audio clip includes noise; the classification module 5552 is configured to perform audio scene classification processing based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip; the processing module 5553 is configured to determine a target audio processing mode corresponding to the audio scene type, and apply the target audio processing mode to the audio clip according to the degree of interference caused by the noise in the audio clip.

[0135] In some embodiments, the target audio processing mode includes a noise reduction processing mode; the processing module 5553 is also configured to query the correspondence between different candidate audio scene types and candidate noise reduction processing modes based on the audio scene type corresponding to the audio scene, and obtain the noise reduction processing mode corresponding to the audio scene type.

[0136] In some embodiments, the target audio processing mode includes a noise reduction processing mode; the processing module 5553 is also configured to determine the noise type that matches the audio scene type based on the audio scene type corresponding to the audio scene; based on the noise type that matches the audio scene type, query the correspondence between different candidate noise types and candidate noise reduction processing modes to obtain the noise reduction processing mode corresponding to the audio scene type; wherein, the noise types matching different audio scene types are not exactly the same.

[0137] In some embodiments, before applying the target audio processing mode to the audio segment of the audio scene, the processing module 5553 is further configured to determine the interference level caused by the noise in the audio segment; when the interference level is greater than an interference level threshold, applying a noise reduction processing mode corresponding to the audio scene type to the audio segment.

[0138] In some embodiments, the processing module 5553 is further configured to match the noise type that matches the audio scene type with the noise in the audio segment; suppress the noise that successfully matches the noise type to obtain the suppressed audio segment; wherein the ratio of the speech signal strength to the noise signal strength in the suppressed audio segment is lower than the signal-to-noise ratio threshold.

[0139] In some embodiments, the target audio processing mode includes a rate switching processing mode; the processing module 5553 is further configured to query the correspondence between different candidate audio scene types and candidate rate switching processing modes based on the audio scene type corresponding to the audio scene, and obtain the rate switching processing mode corresponding to the audio scene type.

[0140] In some embodiments, the target audio processing mode includes a rate switching processing mode; the processing module 5553 is further configured to compare the audio scene type corresponding to the audio scene with a preset audio scene type; when the comparison determines that the audio scene type is the preset audio scene type, the rate switching processing mode associated with the preset audio scene type is used as the rate switching processing mode corresponding to the audio scene type.

[0141] In some embodiments, the processing module 5553 is further configured to obtain the communication signal strength of the audio scene; when the communication signal strength of the audio scene is less than a communication signal strength threshold, the audio bit rate of the audio segment is reduced according to a first set ratio or a first set value; when the communication signal strength of the audio scene is greater than or equal to the communication signal strength threshold, the audio bit rate of the audio segment is increased according to a second set ratio or a second set value.

[0142] In some embodiments, the processing module 5553 is further configured to determine jitter information of the communication signal strength in the audio scene based on the communication signal strength obtained by multiple sampling in the audio scene; when the jitter information indicates that the communication signal is in an unstable state, the audio bit rate of the audio clip is reduced according to a third set ratio or a third set value.

[0143] In some embodiments, the processing module 5553 is further configured to reduce the audio bit rate of the audio segment according to a fourth set ratio or set value when the type of the communication network used to transmit the audio segment belongs to a set type.

[0144] In some embodiments, the audio scene classification processing is implemented through a neural network model, and the neural network model learns the association between the noise included in the audio clip and the audio scene type; the classification module 5552 is also configured to call the neural network model based on the audio clip to perform audio scene classification processing, and obtain the audio scene type that is associated with the noise included in the audio clip.

[0145] In some embodiments, the neural network model includes a mapping network, a residual network and a pooling network; the classification module 5552 is also configured to perform feature extraction processing on the audio clip through the mapping network to obtain a first feature vector of the noise in the audio clip; perform mapping processing on the first feature vector through the residual network to obtain a mapping vector of the audio clip; perform feature extraction processing on the mapping vector of the audio clip through the mapping network to obtain a second feature vector of the noise in the audio clip; perform pooling processing on the second feature vector through the pooling network to obtain a pooling vector of the audio clip; perform nonlinear mapping processing on the pooling vector of the audio clip to obtain an audio scene type that is associated with the noise included in the audio clip.

[0146] In some embodiments, the mapping network includes multiple cascaded mapping layers; the classification module 5552 is also configured to perform feature mapping processing on the audio segment through the first mapping layer among the multiple cascaded mapping layers; output the mapping result of the first mapping layer to the subsequent cascaded mapping layers, continue feature mapping and mapping result output through the subsequent cascaded mapping layers until it is output to the last mapping layer, and use the mapping result output by the last mapping layer as the first feature vector of the noise in the audio segment.

[0147] In some embodiments, the residual network includes a first mapping network and a second mapping network; the classification module 5552 is also configured to perform mapping processing on the first feature vector through the first mapping network to obtain a first mapping vector of the audio segment; perform nonlinear mapping processing on the first mapping vector to obtain a non-mapping vector of the audio segment; perform mapping processing on the non-mapping vector of the audio segment through the first mapping network to obtain a second mapping vector of the audio segment; and use the sum of the first feature vector of the audio segment and the second mapping vector of the audio segment as the mapping vector of the audio segment.

[0148] In some embodiments, the device also includes: a training module 5554, configured to construct audio samples corresponding to multiple different audio scenes based on a noise-free audio signal and background noise corresponding to multiple different audio scenes; and train a neural network model based on the audio samples corresponding to the multiple different audio scenes to obtain a neural network model for audio scene classification.

[0149] In some embodiments, the training module 5554 is also configured to perform the following processing for any audio scene among the multiple different audio scenes: based on the fusion ratio of the background noise of the audio scene and the noise-free audio signal, the background noise of the audio scene and the noise-free audio signal are fused to obtain a first fused audio signal of the audio scene; the background noise of the audio scene corresponding to the first random coefficient is fused into the first fused audio signal to obtain a second fused audio signal of the audio scene; and the noise-free audio signal corresponding to the second random coefficient is fused into the second fused audio signal to obtain an audio sample of the audio scene.

[0150] In some embodiments, the training module 5554 is further configured to perform audio scene classification processing on the audio samples corresponding to the multiple different audio scenes through the neural network model to obtain the predicted audio scene type of the audio sample; construct a loss function of the neural network model based on the predicted audio scene type of the audio sample, the audio scene label of the audio sample and the weight of the audio sample; update the parameters of the neural network model until the loss function converges, and use the updated parameters of the neural network model when the loss function converges as the parameters of the neural network model for audio scene classification.

[0151] In some embodiments, before performing audio scene classification processing based on the audio clip, the acquisition module 5551 is further configured to perform frame processing on the time domain signal of the audio clip to obtain multiple frames of audio signals; perform windowing processing on the multiple frames of audio signals, and perform Fourier transform on the windowed audio signals to obtain the frequency domain signal of the audio clip; and perform logarithmic processing on the Mel frequency band of the frequency domain signal to obtain the audio clip used for performing the audio scene classification.

[0152] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the artificial intelligence-based audio processing method described in the present invention.

[0153] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the artificial intelligence-based audio processing method provided by the embodiment of the present application, for example, the artificial intelligence-based audio processing method shown in Figures 3-5.

[0154] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0155] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0156] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0157] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0158] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An audio processing method based on artificial intelligence, the method comprising: Acquire an audio segment of an audio scene, wherein the audio segment includes noise; performing audio scene classification processing based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment; A target audio processing mode corresponding to the audio scene type is determined, and the target audio processing mode is applied to the audio segment according to the interference degree caused by the noise in the audio segment.

2. The method according to claim 1, wherein The target audio processing mode includes a noise reduction processing mode; The determining of the target audio processing mode corresponding to the audio scene type includes: Based on the audio scene type corresponding to the audio scene, a correspondence between different candidate audio scene types and candidate noise reduction processing modes is queried to obtain the noise reduction processing mode corresponding to the audio scene type.

3. The method according to claim 1, wherein The target audio processing mode includes a noise reduction processing mode; The determining of the target audio processing mode corresponding to the audio scene type includes: determining, based on an audio scene type corresponding to the audio scene, a noise type matching the audio scene type; Based on the noise type matching the audio scene type, querying the correspondence between different candidate noise types and candidate noise reduction processing modes to obtain the noise reduction processing mode corresponding to the audio scene type; The noise types matching different audio scene types are not completely the same.

4. The method according to claim 2 or 3, wherein: The applying the target audio processing mode to the audio segment according to the interference degree caused by the noise in the audio segment includes: Determining the level of interference caused by noise in the audio segment; When the interference level is greater than an interference level threshold, a noise reduction processing mode corresponding to the audio scene type is applied to the audio segment.

5. The method according to claim 2 or 3, wherein: Applying the target audio processing mode to the audio segment includes: Matching the noise type that matches the audio scene type with the noise in the audio clip; Suppressing the noise that successfully matches the noise type to obtain the suppressed audio segment; The ratio of the voice signal strength to the noise signal strength in the suppressed audio segment is lower than a signal-to-noise ratio threshold.

6. The method according to claim 1, wherein The target audio processing mode includes a rate switching processing mode; The determining of the target audio processing mode corresponding to the audio scene type includes: Based on the audio scene type corresponding to the audio scene, a correspondence between different candidate audio scene types and candidate bitrate switching processing modes is queried to obtain the bitrate switching processing mode corresponding to the audio scene type.

7. The method according to claim 1, wherein The target audio processing mode includes a rate switching processing mode; The determining of the target audio processing mode corresponding to the audio scene type includes: Comparing the audio scene type corresponding to the audio scene with a preset audio scene type; When the comparison determines that the audio scene type is the preset audio scene type, the rate switching processing mode associated with the preset audio scene type is used as the rate switching processing mode corresponding to the audio scene type.

8. The method according to claim 6 or 7, wherein: Applying the target audio processing mode to the audio segment includes: Obtaining the communication signal strength of the audio scene; When the communication signal strength of the audio scene is less than a communication signal strength threshold, reducing the audio bit rate of the audio segment according to a first set ratio or a first set value; When the communication signal strength of the audio scene is greater than or equal to the communication signal strength threshold, the audio bit rate of the audio segment is increased according to a second set ratio or a second set value.

9. The method according to claim 6 or 7, wherein: Applying the target audio processing mode to the audio segment includes: Determining jitter information of the communication signal strength in the audio scene based on communication signal strengths obtained by multiple samplings in the audio scene; When the jitter information indicates that the communication signal is in an unstable state, the audio bit rate of the audio segment is reduced according to a third set ratio or a third set value.

10. The method according to claim 6 or 7, wherein: Applying the target audio processing mode to the audio segment includes: When the type of the communication network used to transmit the audio segment belongs to the set type, the audio bit rate of the audio segment is reduced according to a fourth set ratio or set value.

11. The method according to claim 1, wherein The audio scene classification process is implemented by a neural network model, and the neural network model learns the correlation between the noise included in the audio clip and the audio scene type; The performing audio scene classification processing based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment includes: The neural network model is called based on the audio segment to perform audio scene classification processing to obtain an audio scene type that is associated with the noise included in the audio segment.

12. The method according to claim 11, wherein The neural network model includes a mapping network, a residual network and a pooling network; The calling the neural network model based on the audio clip to perform audio scene classification processing includes: performing feature extraction processing on the audio segment through the mapping network to obtain a first feature vector of noise in the audio segment; Performing mapping processing on the first feature vector through the residual network to obtain a mapping vector of the audio clip; performing feature extraction processing on the mapping vector of the audio segment through the mapping network to obtain a second feature vector of the noise in the audio segment; performing pooling processing on the second feature vector through the pooling network to obtain a pooling vector of the audio clip; Nonlinear mapping processing is performed on the pooled vector of the audio segment to obtain an audio scene type associated with the noise included in the audio segment.

13. The method according to claim 12, wherein: The mapping network includes a plurality of cascaded mapping layers; The performing feature extraction processing on the audio segment by the mapping network to obtain a first feature vector of noise in the audio segment includes: performing feature mapping processing on the audio segment through a first mapping layer of the plurality of cascaded mapping layers; Output the mapping result of the first mapping layer to the subsequent cascade mapping layer, continue to perform feature mapping and mapping result output through the subsequent cascade mapping layer until it is output to the last mapping layer, and The mapping result output by the last mapping layer is used as the first feature vector of the noise in the audio segment.

14. The method according to claim 12, wherein: The residual network includes a first mapping network and a second mapping network; The mapping process of the first feature vector by the residual network to obtain the mapping vector of the audio segment includes: Performing mapping processing on the first feature vector through the first mapping network to obtain a first mapping vector of the audio clip; performing nonlinear mapping processing on the first mapping vector to obtain a non-mapping vector of the audio segment; performing mapping processing on the non-mapped vector of the audio clip by using the first mapping network to obtain a second mapped vector of the audio clip; A sum of the first feature vector of the audio segment and the second mapping vector of the audio segment is used as the mapping vector of the audio segment.

15. The method according to claim 1, wherein The method further comprises: Constructing audio samples corresponding to the multiple different audio scenes respectively based on the noise-free audio signal and the background noises corresponding to the multiple different audio scenes; The neural network model is trained based on the audio samples corresponding to the multiple different audio scenes to obtain a neural network model for audio scene classification.

16. The method according to claim 15, wherein The step of constructing audio samples corresponding to the plurality of different audio scenes based on the noise-free audio signal and the background noises corresponding to the plurality of different audio scenes comprises: The following processing is performed for any audio scene among the multiple different audio scenes: fusing the background noise of the audio scene and the noise-free audio signal based on a fusion ratio of the background noise of the audio scene to the noise-free audio signal to obtain a first fused audio signal of the audio scene; fusing background noise of the audio scene corresponding to the first random coefficient into the first fused audio signal to obtain a second fused audio signal of the audio scene; The noise-free audio signal corresponding to the second random coefficient is fused into the second fused audio signal to obtain an audio sample of the audio scene.

17. The method according to claim 15, wherein: The step of training a neural network model based on audio samples corresponding to the plurality of different audio scenes to obtain a neural network model for audio scene classification includes: performing audio scene classification processing on the audio samples corresponding to the multiple different audio scenes respectively using the neural network model to obtain predicted audio scene types of the audio samples; constructing a loss function of the neural network model based on the predicted audio scene type of the audio sample, the audio scene label of the audio sample, and the weight of the audio sample; The parameters of the neural network model are updated until the loss function converges, and the updated parameters of the neural network model when the loss function converges are used as the parameters of the neural network model for audio scene classification.

18. The method according to claim 1, wherein Before performing audio scene classification processing based on the audio segment, the method further includes: Performing frame processing on the time domain signal of the audio clip to obtain multiple frames of audio signals; Performing windowing processing on the multiple frames of audio signals, and performing Fourier transform on the windowed audio signals to obtain frequency domain signals of the audio segments; Logarithmic processing is performed on the Mel frequency band of the frequency domain signal to obtain the audio segment for performing the audio scene classification.

19. An audio processing device, comprising: an acquisition module configured to acquire an audio segment of an audio scene, wherein the audio segment includes noise; a classification module configured to perform audio scene classification processing based on the audio segment to obtain an audio scene type corresponding to the noise in the audio segment; The processing module is configured to determine a target audio processing mode corresponding to the audio scene type, and apply the target audio processing mode to the audio segment according to the interference degree caused by the noise in the audio segment.

20. An electronic device, comprising: a memory for storing executable instructions; The processor is configured to implement the artificial intelligence-based audio processing method according to any one of claims 1 to 18 when executing the executable instructions stored in the memory.

21. A computer-readable storage medium storing executable instructions for implementing the artificial intelligence-based audio processing method according to any one of claims 1 to 18 when executed by a processor.

22. A computer program product, comprising a computer program or instructions, wherein the computer program or instructions enable a computer to execute the audio processing method based on artificial intelligence according to any one of claims 1 to 18.