Voice processing method, apparatus, electronic device, and computer program based on artificial intelligence

The AI-based voice processing method addresses the limitations of existing technologies by classifying voice scenes and applying targeted processing modes, resulting in improved voice quality and reduced noise interference.

JP7700236B2Active Publication Date: 2025-06-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023530056
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-03
Filing Date
2021-11-17
Publication Date
2025-06-30
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

Existing voice processing methods in cloud technology and artificial intelligence are limited in their ability to improve voice quality while minimizing noise suppression, which often results in reduced quality of useful audio signals.

Method used

A voice processing method based on artificial intelligence that involves acquiring a voice clip with noise, classifying the voice scene type based on the noise, and applying a targeted voice processing mode to improve voice quality, while minimizing noise interference.

Benefits of technology

This approach effectively preserves useful information in voice signals by matching the voice processing mode to the specific noise characteristics of the voice scene, thereby enhancing voice quality and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700236000002
    Figure 0007700236000002
  • Figure 0007700236000003
    Figure 0007700236000003
  • Figure 0007700236000004
    Figure 0007700236000004
Patent Text Reader

Abstract

The present application provides an artificial intelligence-based audio processing method, device, electronic device, computer-readable storage medium, and computer program product, which relate to cloud technology and artificial intelligence technology. The method includes the steps of obtaining an audio clip of an audio scene, where the audio clip contains noise; performing an audio scene classification process based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip; determining a target audio processing mode corresponding to the audio scene type, and applying the target audio processing mode to the audio clip of the audio scene based on the degree of interference caused by the noise in the audio clip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to cloud technology and artificial intelligence technology, and particularly to a voice processing method, apparatus, electronic device, computer-readable storage medium, and computer program product based on artificial intelligence.

[0002] The embodiments of the present application are filed based on a Chinese patent application with an application number of No. 202011410814.4 and an application date of December 3, 2020, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby incorporated by reference into the embodiments of the present application.

Background Art

[0003] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent devices, it enables devices to have functions of perception, inference, and decision-making. Artificial intelligence technology is a comprehensive discipline with a wide range of related fields, such as natural language processing technology and machine learning / deep learning in multiple directions. With the development of technology, artificial intelligence technology is applied in more fields and plays an increasingly important role. For example, in a network meeting scenario based on cloud technology, introducing artificial intelligence technology can improve voice quality.

[0004] However, in related technologies, the processing method for voice is relatively single. Although it forms a suppression effect on the noise in the voice, it inevitably reduces the quality of useful signals (such as audio signals) in the voice.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The embodiments of the present application provide a voice processing method, apparatus, electronic device, computer-readable storage medium, and computer program product based on artificial intelligence, which can obtain a target based on a voice scene and perform voice processing to improve voice quality.

Means for Solving the Problem

[0006] The technical solution of the embodiments of the present application is realized as follows:

[0007] The embodiments of the present application provide a voice processing method based on artificial intelligence, obtaining a voice clip of a voice scene, where noise is included in the voice clip, and performing voice scene classification processing based on the voice clip to obtain a voice scene type corresponding to the noise in the voice clip, and determining a target voice processing mode corresponding to the voice scene type, and applying the target voice processing mode to the voice clip based on the interference degree caused by the noise in the voice clip.

[0008] The embodiments of the present application provide a voice processing device, an acquisition module configured to acquire a voice clip of a voice scene, where noise is included in the voice clip, a classification module configured to perform voice scene classification processing based on the voice clip to obtain a voice scene type corresponding to the noise in the voice clip, and a processing module configured to determine a target voice processing mode corresponding to the voice scene type, and apply the target voice processing mode to the voice clip based on the interference degree caused by the noise in the voice clip.

[0009] The embodiments of the present application provide an electronic device for voice processing, and the electronic device includes a memory used to store executable instructions, a processor used to implement the voice processing method based on artificial intelligence provided by the embodiments of the present application when executing the executable instructions stored in the memory.

[0010] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which are used to implement the artificial intelligence-based voice processing method provided by the embodiment of the present application when causing a processor to execute.

[0011] An embodiment of the present application provides a computer program product including a computer program or instructions, and the computer program or instructions cause a computer to execute the artificial intelligence-based voice processing method provided by the embodiment of the present application.

[0012] The embodiment of the present application has the following beneficial effects: By associating noise with a voice scene type, the voice scene type corresponding to the voice is identified, and thereby targeted voice processing is performed based on the voice scene type, enabling the voice processing mode introduced in the voice scene to match the noise included in the voice scene, thereby maximizing the preservation of useful information in the voice and improving the accuracy of voice processing.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying out the Invention

[0014] Before further elaborating on the embodiments of the present application, relevant nouns and terms in the embodiments of the present application will be explained. The relevant nouns and terms in the embodiments of the present application shall be applied to the following interpretations.

[0015] 1) Convolutional Neural Networks (CNN): A feedforward neural network (FNN) that includes convolutional calculations and has a deep structure, and is one of the representative algorithms of deep learning. The convolutional neural network has the ability of representation learning and can perform shift-invariant classification on the input image according to its hierarchical structure.

[0016] 2) Residual Network (ResNet): A convolutional neural network that is easy to optimize and improves the accuracy rate through a fairly deep increase. The internal residual block uses a jump connection to mitigate the problem of gradient disappearance caused by increasing the depth in a deep neural network.

[0017] 3) Audio processing mode: This is a mode used for audio processing. When the audio processing mode is applied to an audio clip, the audio can be optimized, thereby obtaining clear and smooth audio. The audio processing mode in the embodiments of the present application includes a noise reduction processing mode and a bitrate switching processing mode.

[0018] In related technologies, the video bitrate adaptation technology (ABR, Adaptive Bitrate Streaming) adaptively adjusts the video bitrate. The bitrate adjustment algorithm is mostly applied to video playback, and automatically adjusts the video bitrate (i.e., resolution) according to the network situation or the playback buffer situation of the client. Also, based on the noise reduction algorithm of universal noise, the frequency spectrum characteristics of noisy audio are used as the input of the neural network, and clean speech is used as the reference output of the neural network to train the noise reduction model. Using the least mean square (LMS) as the optimization target, after activating the noise reduction function, the same type of noise reduction method is used for various scene environments.

[0019] The applicant discovered that in the process of implementing the present application, due to the relatively frequent variation and change of the quality of the network environment, the bitrate switching based on the network speed also frequently varies accordingly. However, frequent switching of the resolution greatly affects the user experience. In the actual environment, the specific noise of a specific scene presents higher requirements and challenges to the robustness of the noise reduction algorithm based on universal noise.

[0020] To solve the above problems, the embodiments of the present application provide an audio processing method, apparatus, electronic device, and computer-readable storage medium based on artificial intelligence, which can obtain and process audio based on the audio scene to improve the audio quality.

[0021] The voice processing method based on artificial intelligence provided by the embodiments of the present application may be independently implemented by a terminal / server, or may be jointly implemented by a terminal and a server. For example, the terminal may independently perform the voice processing method based on artificial intelligence described below, or terminal A may send an optimization request for the voice (including the voice clip) to the server, and the server executes the voice processing method based on artificial intelligence according to the received optimization request for the voice, and in response to the optimization request for the voice, applies a target voice processing mode (including a noise reduction processing mode and a bit rate switching processing mode) to the voice clip of the voice scene, and sends the processed voice clip to terminal B, whereby a clear audio call can be made between terminal A and terminal B.

[0022] The electronic device for voice processing provided by the embodiments of the present application may be various types of terminal devices or a server. Here, the server may be an independent physical server, a server cluster composed of a plurality of physical servers, or a distributed system, and may further be a cloud server providing cloud computing services. The terminal may be a smart mobile phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through a wired or wireless communication method, and the present application does not limit this here.

[0023] Taking the server as an example, for example, it may be a server cluster arranged at the cloud end, which publishes artificial intelligence cloud services (AiaaS, AI as a Service) to users, and the AIaaS platform decomposes several types of general AI services and provides services that are independent or packaged at the cloud end. Such a service mode is similar to a mall themed on one AI, and any user can access and use one or more types of artificial intelligence services provided by the AIaaS platform through the application programming interface method.

[0024] For example, one type of artificial intelligence cloud service may be a voice processing service, that is, the voice processing program provided by the embodiments of the present application is encapsulated in the server at the cloud end. The user can call the voice processing service in the cloud service through a terminal (where a client is operating, such as a recording client or an instant messaging client), so as to call the voice processing program encapsulated in the server arranged at the cloud end, and apply the target voice processing mode to the voice clip of the voice scene based on the target voice processing mode that matches the voice scene type.

[0025] As an application example, for a recording client, the user may be a subscribed live streamer of a certain audio platform who needs to regularly release the voice of an audiobook. The scene of live streamer recording may change. For example, it may be recorded at home, in a library, or even outdoors, and there are different noises in these scenes. By recording the current voice scene and performing voice scene identification on the recorded voice clip, the voice scene type is determined, and based on the voice scene type, targeted noise reduction processing is performed on the voice clip to obtain a noise-removed voice clip and store it, thereby realizing the noise removal recording function.

[0026] As another application example, for an instant messaging client, the user may send audio to a certain designated friend or a certain designated group. However, the scene where the user is currently located may change. For example, there are different noises in different scenes such as an office or a shopping mall. By performing audio scene identification on the audio of the current scene, the audio scene type is determined, and based on the audio scene type, targeted noise reduction processing is performed on the audio to obtain a noise-removed voice clip and send it, thereby realizing the noise removal audio sending function.

[0027] As another application example, for a meeting client, users participating in a meeting can make audio calls in different environments. For example, user A participating in the meeting is in the office, and user B participating in the meeting is on a high-speed train. Different scenes have different noises and different communication signals. For example, there is traffic noise on the high-speed train and the communication signal is relatively poor. By performing audio scene identification on the audio calls of each user participating in the meeting, the audio scene type of each user participating in the meeting is determined, and based on the audio scene type, targeted bitrate switching processing is performed on the audio of each user participating in the meeting to achieve adaptive bitrate switching and improve the call quality of the meeting call.

[0028] Referring to FIG. 1, FIG. 1 is a schematic diagram of an application scene of an audio processing system 10 provided according to an embodiment of the present application. The terminal 200 is connected to the server 100 through the network 300. The network 300 may be a wide area network, a local area network, or a combination of both.

[0029] The terminal 200 (with a client operating, such as a recording client, an instant messaging client, a call client, etc.) can be used to obtain optimization requirements for audio. For example, when a user inputs or registers an audio clip of an audio scene through the terminal 200, the terminal 200 automatically obtains the audio clip of the audio scene and automatically generates optimization requirements for the audio.

[0030] In some embodiments, in a client operating on a terminal, a voice processing plugin assembly can be embedded, thereby realizing an artificial intelligence-based voice processing method locally on the client. For example, after the terminal 200 obtains an optimization requirement for voice (including a voice clip of a voice scene), it calls the voice processing plugin assembly to realize an artificial intelligence-based voice processing method, identify the noise in the voice clip and the corresponding voice scene type, and apply the target voice processing mode to the voice clip of the voice scene based on the target voice processing mode matching the voice scene type. For example, for a recording application, the user records in the current voice scene, performs voice scene identification on the recorded voice clip to determine the voice scene type, and performs noise reduction processing on the voice clip based on the voice scene type to obtain a denoised voice clip, and stores the denoised voice clip to realize a noise removal recording function.

[0031] In some embodiments, after obtaining the optimization requirement for the voice, the terminal 200 calls the voice processing interface of the server 100 (which can be provided in the form of a cloud service, that is, a voice processing service). The server 100 identifies the noise in the voice clip and the corresponding voice scene type, and based on the target voice processing mode that matches the voice scene type, applies the target voice processing mode to the voice clip of the voice scene, and transmits the voice clip (the optimized voice clip) of the target voice processing mode to the terminal 200 or other terminals. For example, for a recording application, the user records in the current voice scene, the terminal 200 obtains the corresponding voice clip, automatically generates an optimization requirement for the voice, and transmits the optimization requirement for the voice to the server 100. The server 100 performs voice scene identification on the recorded voice clip based on the optimization requirement for the voice, determines the voice scene type, and performs noise reduction processing on the voice clip based on the voice scene type to obtain the noise-removed voice clip, and realizes the noise removal recording function. For an instant messaging application, the user performs audio transmission in the current audio scene, the terminal 200 obtains the corresponding voice clip, automatically generates an optimization requirement for the voice, and transmits the optimization requirement for the voice to the server 100. The server 100 performs audio scene identification on the voice clip based on the optimization requirement for the voice, determines the audio scene type, and performs noise reduction processing on the voice clip based on the audio scene type to obtain the noise-removed voice clip, and transmits it to realize the noise removal audio transmission function. And based on the audio scene type, performs bit rate switching processing on the voice clip to obtain the optimized bit rate switching, and improves the audio call quality. For a call application, user A makes an audio call with user B, and user A makes an audio call in the current audio scene.The terminal 200 acquires the voice clip corresponding to user A, automatically generates an optimization requirement for the voice, and transmits the optimization requirement for the voice to the server 100. The server 100 performs audio scene identification on the voice clip of user A based on the optimization requirement for the voice, determines the audio scene type, performs noise reduction processing on the voice clip based on the audio scene type to obtain a target voice clip, and transmits the noise-removed voice clip of user A to user B, thereby realizing the noise removal audio call function.

[0032] Hereinafter, the structure of the electronic device for voice processing provided by the embodiments of the present application will be described. Referring to FIG. 2, FIG. 2 is a structural schematic diagram of an electronic device 500 for voice processing provided by the embodiments of the present application. Taking the case where the electronic device 500 is a server as an example, the electronic device 500 for voice processing shown in FIG. 2 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each assembly in the electronic device 500 is integrally coupled through a bus system 540. As can be understood, the bus system 540 is used to realize connection communication between these assemblies. The bus system 540 further includes a power bus, a control bus, and a status signal bus in addition to the data bus. However, for the sake of clear description, all various buses in FIG. 2 are marked as the bus system 540.

[0033] The processor 510 may be a kind of integrated circuit chip and has signal processing capabilities. For example, it may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, individual gates, or transistor logic devices, individual hardware assemblies, etc. Here, the general-purpose processor may be a microprocessor or any ordinary processor, etc.

[0034] Memory 550 may include volatile memory, non-volatile memory, or both volatile and non-volatile memory. Here, the non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory. Memory 550 optionally includes one or more storage devices that are physically located away from the processor 510.

[0035] In some embodiments, memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or subsets or supersets thereof, which will be exemplarily described below.

[0036] The operating system 551 includes system programs such as a frame layer, a core library layer, and a driver layer that are used to process various basic system services and execute hardware-related tasks, realize various basic services, and process tasks based on hardware.

[0037] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include Bluetooth (registered trademark), wireless fidelity (WiFi), and universal serial bus (USB).

[0038] In some embodiments, the voice processing apparatus provided by the embodiments of the present application may be implemented in a software manner. For example, it may be the voice processing plugin assembly in the terminal described above, or the voice processing service in the server described above. Of course, it is not limited thereto. The voice processing apparatus provided by the embodiments of the present application can be provided as various software embodiments, including various forms including application programs, software, software modules, scripts, or code.

[0039] FIG. 2 shows a voice processing apparatus 555 stored in the memory 550, which may be software in the form of a program, a plugin assembly, etc. For example, it is a voice processing plugin assembly, and includes a series of modules including an acquisition module 5551, a classification module 5552, a processing module 5553, and a training module 5554. Here, the acquisition module 5551, the classification module 5552, and the processing module 5553 are used to implement the voice processing function provided by the embodiments of the present application, and the training module 5554 is used to train the neural network model. Here, the voice scene classification process is realized through the neural network model.

[0040] As described above, the artificial intelligence-based voice processing method provided by the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG. 3, FIG. 3 is a flowchart schematic diagram of the artificial intelligence-based voice processing method provided by the embodiments of the present application. The steps shown in FIG. 3 will be described in combination.

[0041] In the following steps, the voice scene refers to the environment where the voice is generated, such as a home environment, an office environment, an external environment when riding various vehicles such as a high-speed railway, etc.

[0042] In step 101, a voice clip of the voice scene is acquired, where noise is included in the voice clip.

[0043] As an example of obtaining an audio clip, a user inputs audio through a terminal (where the client is operating) in the current audio scene. The terminal 200 obtains the corresponding audio clip, automatically generates an optimization request for the audio, and sends the optimization request for the audio to the server. The server analyzes the optimization request for the audio, obtains the audio clip of the audio scene, and thereby subsequently performs audio scene identification based on the audio clip.

[0044] In step 102, perform audio scene classification processing based on the audio clip to obtain the noise in the audio clip and the corresponding audio scene type.

[0045] For example, after obtaining the audio clip of the audio scene, perform audio scene classification processing based on the audio clip through a neural network model to obtain the noise in the audio clip and the corresponding audio scene type. Here, the audio clip can be input into the neural network model, and furthermore, the time-domain features or frequency-domain features of the audio clip can be input into the neural network model. The neural network model performs audio scene classification processing based on the time-domain features or frequency-domain features of the audio clip to obtain the noise in the audio clip and the corresponding audio scene type. Taking the frequency-domain features of the audio clip as an example, after obtaining the audio clip, first perform frame splitting processing on the time-domain signal of the audio clip to obtain a multi-frame audio signal, then perform windowing processing on the multi-frame audio signal, and perform Fourier transform on the audio signal after windowing processing to obtain the frequency-domain signal of the audio clip, and perform logarithmic processing on the mel frequency band of the frequency-domain signal to obtain the frequency-domain features of the audio clip. That is, it is the audio clip used for performing audio scene classification.

[0046] In order to enable the neural network model to process multi-channel inputs, differential processing can be performed on the frequency-domain features of the audio clip obtained by logarithmic processing to obtain the first derivative of the audio clip. Next, differential processing is performed on the first derivative to obtain the second derivative of the audio clip. Finally, the frequency-domain features, the first derivative, and the second derivative of the audio clip are synthesized into a three-channel input signal, and the three-channel input signal is used as the audio clip for performing audio scene classification.

[0047] In some embodiments, the audio scene classification process is implemented through a neural network model. For the neural network model to learn the correlation between the noise contained in the audio clip and the audio scene type, and execute the audio scene classification process based on the audio clip to obtain the noise in the audio clip and the corresponding audio scene type, it includes calling the neural network model based on the audio clip to execute the audio scene classification process to obtain the audio scene type having a correlation with the noise contained in the audio clip.

[0048] For example, as shown in FIG. 6, the neural network model includes a mapping network, a residual network, and a pooling network. Feature extraction processing is performed on the audio clip through the mapping network to obtain the first feature vector of the noise in the audio clip, and mapping processing is performed on the first feature vector through the residual network to obtain the mapping vector of the audio clip. Feature extraction processing is performed on the mapping vector of the audio clip through the mapping network to obtain the second feature vector of the noise in the audio clip, and pooling processing is performed on the second feature vector through the pooling network to obtain the pooling vector of the audio clip. Nonlinear mapping processing is performed on the pooling vector of the audio clip to obtain the audio scene type having a correlation with the noise contained in the audio clip.

[0049] Based on the above example, the mapping network includes a plurality of cascaded mapping layers. Obtaining the first feature vector of the noise in the audio clip by performing feature extraction processing on the audio clip through the mapping network involves performing feature mapping processing on the audio clip through the first mapping layer in the plurality of cascaded mapping layers, outputting the mapping result of the first mapping layer to the subsequent cascaded mapping layers, continuously performing feature mapping and output of the mapping result through the subsequent cascaded mapping layers until it is output to the last mapping layer, and using the mapping result output from the last mapping layer as the first feature vector of the noise in the audio clip.

[0050] Here, the mapping network can effectively extract the features of the noise in the scene of the audio clip. The mapping layer may be a convolutional neural network, but the embodiments of the present application are not limited to convolutional neural networks and may be other neural networks.

[0051] In some embodiments, the residual network includes a first mapping network and a second mapping network. Obtaining the mapping vector of the audio clip by performing mapping processing on the first feature vector through the residual network involves performing mapping processing on the first feature vector through the first mapping network to obtain the first mapping vector of the audio clip, performing non-linear mapping processing on the first mapping vector to obtain the non-mapping vector of the audio clip, performing mapping processing on the non-mapping vector of the audio clip through the first mapping network to obtain the second mapping vector of the audio clip, and using the sum result of the first feature vector of the audio clip and the second mapping vector of the audio clip as the mapping vector of the audio clip.

[0052] Here, the residual network can effectively prevent the problem of gradient disappearance in neural network training error transmission, thereby accelerating the training of the neural network model.

[0053] In some embodiments, due to the need to train a neural network model, the trained neural network model can perform voice scene classification. The training method includes constructing voice samples corresponding to a plurality of different voice scenes based on a noise-free voice signal and background noises respectively corresponding to the plurality of different voice scenes, and training a neural network model based on the voice samples respectively corresponding to the plurality of different voice scenes to obtain a neural network model used for classifying voice scenes.

[0054] To enhance the diversity of sample data, the voice sample construction method includes performing a process of fusing the background noise of a voice scene and a noise-free voice signal based on the fusion ratio of the background noise of the voice scene and the noise-free voice signal for any voice scene in a plurality of different voice scenes to obtain a first fused voice signal of the voice scene, a process of fusing the background noise of the voice scene corresponding to a first random coefficient in the first fused voice signal to obtain a second fused voice signal of the voice scene, and a process of fusing the noise-free voice signal corresponding to a second random coefficient in the second fused voice signal to obtain a voice sample of the voice scene.

[0055] For example, after fusing the saved human voice (noise-free voice signal) and background noise at a fusion ratio of 1:1, a partial random ratio is generated to superimpose the data. For example, the noise superimposition coefficient (first random coefficient) is a random number in the range of 0.3 to 0.5, and the human voice superimposition coefficient (second random coefficient) is a random number in the range of 0.5 to 0.7.

[0056] In some embodiments, obtaining a neural network model used for classifying audio scenes by training a neural network model based on audio samples respectively corresponding to a plurality of audio scenes includes performing audio scene classification processing on audio samples respectively corresponding to a plurality of different audio scenes through the neural network model to obtain a predicted audio scene type of the audio samples, constructing a loss function of the neural network model based on the predicted audio scene type of the audio samples, the audio scene mark of the audio samples, and the weighting of the audio samples, continuously updating the parameters of the neural network model until the loss function converges, and using the updated parameters of the neural network model when the loss function converges as the parameters of the neural network model used for classifying audio scenes.

[0057] For example, after determining the value of the loss function of the neural network model based on the predicted audio scene type of the audio samples, the audio scene mark of the audio samples, and the weighting of the audio samples, it is possible to determine whether the value of the loss function of the neural network model exceeds a preset threshold. When the value of the loss function of the neural network model exceeds the preset threshold, an error signal of the neural network model is determined based on the loss function of the neural network model, the error information is propagated in the reverse direction in the neural network model, and the model parameters of each layer are updated during the propagation process.

[0058] Here, to explain backpropagation, training sample data is input into the input layer of the neural network model, passes through the hidden layer, and finally reaches the output layer to output the result. This is the forward propagation process of the neural network model. If there is an error between the output result of the neural network model and the actual result, the error between the output result and the actual value is calculated, and until it propagates to the input layer, this error continues to be propagated in the reverse direction from the output layer to the hidden layer. In the backpropagation process, the above process is continuously repeated until the values of the model parameters are adjusted according to the error and converge.

[0059] In step 103, determine the target voice processing mode corresponding to the voice scene type, and apply the target voice processing mode to the voice clip based on the interference degree caused by the noise in the voice clip.

[0060] For example, after obtaining the voice scene type, first determine the target voice processing mode that matches the voice scene type, and then apply the target voice processing mode to the voice clip of the voice scene to perform targeted voice optimization and improve the accuracy of voice processing.

[0061] Referring to FIG. 4, FIG. 4 is an optional flow schematic diagram of an artificial intelligence-based voice processing method provided by an embodiment of the present application. FIG. 4 shows that step 103 in FIG. 3 can be realized through steps 1031A to 1032A shown in FIG. 4. The target voice processing mode includes a noise reduction processing mode. In step 1031A, based on the voice scene type corresponding to the voice scene, query the correspondence between different candidate voice scene types and candidate noise reduction processing modes to obtain the noise reduction processing mode corresponding to the voice scene type. In step 1032A, perform matching processing on the noise type that matches the voice scene type with the noise in the voice clip, and perform suppression processing on the noise that successfully matches the noise type to obtain the voice clip after suppression. Here, the ratio of the audio signal intensity to the noise signal intensity in the voice clip after suppression is lower than the threshold of the signal-to-noise ratio.

[0062] For example, based on the actual application scenario, a mapping table including the correspondence between different candidate voice scene types and candidate noise reduction processing modes is constructed in advance, and the mapping table is stored in the storage space. By reading the correspondence between different candidate voice scene types and candidate noise reduction processing modes included in the mapping table in the storage space, the noise reduction processing mode corresponding to the voice scene type can be quickly queried based on the voice scene type corresponding to the voice scene. Thereby, by applying the noise reduction processing mode to the voice clip of the voice scene, the noise of the voice clip of the voice scene is removed, so as to realize targeted noise removal and improve the voice quality (i.e., the resolution of the voice) of the voice clip.

[0063] In some embodiments, the target voice processing mode includes a noise reduction processing mode. Determining the target voice processing mode corresponding to the voice scene type includes determining a noise type that matches the voice scene type based on the voice scene type corresponding to the voice scene, and querying the correspondence between different candidate noise types and candidate noise reduction processing modes based on the noise type that matches the voice scene type to obtain the noise reduction processing mode corresponding to the voice scene type. Here, the noise types that match different voice scene types are not exactly the same.

[0064] For example, first determine the noise type that matches the voice scene type through the voice scene type corresponding to the voice scene, and then obtain the noise reduction processing mode corresponding to the voice scene type through the noise type that matches the voice scene type. That is, non-interference between the voice scene type and the candidate noise reduction processing mode is realized, so that subsequently, the correspondence between the voice scene type and the candidate noise reduction processing mode can be flexibly adjusted.

[0065] For example, the strategy for the client developer to assign noise reduction processing modes to different noises may change, or the needs for noise reduction processing modes for different noises of different users may change. Therefore, if the mapping relationship between the voice scene type and the noise reduction processing mode of the voice clip is realized through a neural network model, a large number of models need to be accurately obtained and trained, and if the noise reduction processing modes assigned to different noises change, the neural network model needs to be retrained, consuming a large amount of computing resources.

[0066] However, if the mapping relationship between the audio scene type and noise of the audio clip is realized only through the neural network model, by training one neural network model, various needs for the noise reduction processing mode in actual applications can be satisfied. If it is necessary to realize the strategic setting of the noise type and the noise reduction processing mode on the client side, even if the noise reduction processing mode assigned to different noises changes, it is only necessary to adjust the strategic setting of the noise type and the noise reduction processing mode on the client side, thereby avoiding consuming a large amount of computing resources to train the neural network model.

[0067] In some embodiments, applying a target audio processing mode to an audio clip based on the interference degree caused by the noise in the audio clip includes determining the interference degree caused by the noise in the audio clip, and when the interference degree is greater than the interference degree threshold, applying the noise reduction processing mode corresponding to the audio scene type to the audio clip of the audio scene.

[0068] For example, when the noise in the audio clip hardly affects the audio clip, noise reduction processing may not be performed, and noise reduction processing may be performed on the audio clip only when the noise in the audio clip affects the audio clip. For example, when the user records, some noises in the audio scene are recorded during recording, but if these noises do not affect the recording effect, noise reduction processing may not be performed on the recording, and when these noises affect the recording effect (for example, the recording content cannot be heard), noise reduction processing can be performed on the recording.

[0069] Referring to FIG. 5, FIG. 5 is an optional flowchart of an artificial intelligence-based voice processing method provided by an embodiment of the present application. FIG. 5 shows that step 103 in FIG. 3 can be realized through steps 1031B-1032B shown in FIG. 4. The target voice processing mode includes a bitrate switching processing mode. In step 1031B, based on the voice scene and the corresponding voice scene type, the correspondence between different candidate voice scene types and candidate bitrate switching processing modes is queried to obtain the bitrate switching processing mode corresponding to the voice scene type. In step 1032B, the bitrate switching processing mode corresponding to the voice scene type is applied to the voice clip.

[0070] For example, based on the actual application scenario, a mapping table including the correspondence between different voice scene types and candidate bitrate switching processing modes is constructed in advance, and the mapping table is stored in the storage space. By reading the correspondence between different voice scene types and candidate bitrate switching processing modes included in the mapping table in the storage space, the bitrate switching processing mode matching the voice scene type can be quickly queried based on the voice scene type corresponding to the voice scene. Thereby, by applying the bitrate switching processing mode to the voice clip of the voice scene to switch the bitrate of the voice clip, the desired bitrate switching is realized, and the smoothness of the voice clip is improved.

[0071] In some embodiments, the target voice processing mode includes a bitrate switching processing mode. Determining the target voice processing mode corresponding to the voice scene type includes comparing the voice scene type with a preset voice scene type, and when it is determined by comparison that the voice scene type is the preset voice scene type, setting the bitrate switching processing mode related to the preset voice scene type as the bitrate switching processing mode corresponding to the voice scene type.

[0072] For example, it is not necessary for every voice scene to perform bitrate switching. For example, the communication signals in an office environment are relatively stable and do not require bitrate switching, while the signals in a high-speed railway environment are relatively weak and unstable and require bitrate switching. Therefore, before determining the bitrate switching processing mode, it is necessary to compare the voice scene type with a preset voice scene type that requires bitrate switching. When it is determined through comparison that the voice scene type belongs to the preset voice scene type that requires bitrate switching, the bitrate switching processing mode related to the preset voice scene type is determined as the bitrate switching processing mode that matches the voice scene type. This avoids resource waste caused by all scenes performing bitrate switching.

[0073] In some embodiments, applying the target voice processing mode to the voice clip of the voice scene includes reducing the voice bitrate of the voice clip according to the first setting ratio or the first setting value when the communication signal strength of the voice scene is smaller than the communication signal strength threshold, and increasing the voice bitrate of the voice clip according to the second setting ratio or the second setting value when the communication signal strength of the voice scene is greater than or equal to the communication signal strength threshold.

[0074] Taking an audio call scenario as an example, multiple people make an audio call in different environments and send voice clips to the server through the client. The server receives the voice clips sent from each client, executes voice scene classification processing based on the voice clips to obtain the noise in the voice clips and the corresponding voice scene type, and determines the bitrate switching processing mode that matches the voice scene type. After that, it determines the communication signal strength of the voice scene. When the communication signal strength of the voice scene is smaller than the communication signal strength threshold, it means that the signal of the current voice scene is weak and it is necessary to reduce the bitrate. Therefore, by reducing the audio bitrate of the voice clip according to the first setting ratio or the first setting value in the bitrate switching processing mode that matches the voice scene type, subsequent smooth voice interaction can be performed to avoid interruption of the audio call. When the communication signal strength of the voice scene is greater than or equal to the communication signal strength threshold, it means that the communication signal of the current voice scene is strong and there is no need to reduce the bitrate to interrupt the call. Therefore, according to the second setting ratio or the second setting value in the bitrate switching processing mode that matches the voice scene type, the audio bitrate of the voice clip is increased to improve the fluency of the voice interaction. It should be noted that the first setting ratio and the second setting ratio may be the same or different, and the first setting value ratio and the second setting value may be the same or different. The first setting ratio, the second setting ratio, the first setting value, and the second setting value may be set according to actual needs.

[0075] Here, for an example of the method for obtaining the communication signal strength of the voice scene, the communication signal strengths obtained by multiple samplings in the voice scene are averaged, and the average result is used as the communication signal strength of the voice scene. For example, the multiple sampling results from when a certain user starts an audio call until now are averaged, and the average result is used as the communication signal strength in the voice scene where the user is located.

[0076] In some embodiments, applying a target audio processing mode to an audio clip of an audio scene includes determining jitter information of a communication signal strength in the audio scene based on the communication signal strength obtained by sampling multiple times in the audio scene, and when the jitter information indicates that the communication signal is in an unstable state, reducing the audio bit rate of the audio clip according to a third setting ratio or a third setting value.

[0077] For example, obtain the communication signal strength by sampling multiple times in the audio scene, and acquire the jitter change situation (i.e., jitter information) of the communication signal strength in the audio scene through a normal distribution method. When the variance in the normal distribution representing the jitter change situation is greater than the variance threshold, it means that the data (i.e., the communication signal) in the normal distribution is relatively dispersed, indicating that the communication signal strength jitters severely and the communication signal is unstable. Based on ensuring the fluency of the audio, in order to avoid repeatedly switching the audio bit rate subsequently, the audio bit rate of the audio clip can be reduced according to a preset ratio or a default value in the bit rate switching processing mode that matches the audio scene type. Here, the third setting ratio and the third setting value can be set according to actual needs.

[0078] By determining the jitter change situation of the communication signal strength in the audio scene, it is further determined whether it is necessary to switch the bit rate, thereby avoiding frequently switching the audio bit rate based on ensuring the fluency of the audio and improving the user experience.

[0079] In some embodiments, applying a target audio processing mode to an audio clip of an audio scene includes reducing the audio bit rate of the audio clip according to a fourth setting ratio or a fourth setting value when the type of the communication network used to transmit the audio clip belongs to the set type.

[0080] For example, after determining the bitrate switching processing mode that matches the audio scene type, it is further determined whether the type of the communication network previously used to transmit the audio clip belongs to the set type (for example, WiFi network, cellular network, etc.). For example, if it is determined that the type of the communication network used to transmit the audio clip belongs to the WiFi network, it means that the current audio clip is in an unstable environment. In order to ensure the smoothness of the audio, based on the preset ratio or the default value in the bitrate switching processing mode that matches the audio scene type, the audio bitrate of the audio clip can be reduced. Here, the third setting ratio and the third setting value can be set according to actual needs.

[0081] Next, an exemplary application in one actual application scenario of the embodiments of the present application will be described.

[0082] The embodiments of the present application can be applied to various audio application scenarios. For example, for a recording application, a user can record in the current audio scene through a recording client operating on a terminal. The recording client performs audio scene identification on the recorded audio clip, determines the audio scene type, and performs targeted noise reduction processing on the audio clip based on the audio scene type to obtain a noise-reduced audio clip, which is then stored to realize the noise reduction recording function. For an instant messaging application, a user can perform audio transmission in the current audio scene through an instant messaging client operating on the terminal. The instant messaging client acquires the corresponding audio clip, performs audio scene identification on the audio clip to determine the audio scene type, and performs targeted noise reduction processing on the audio clip based on the audio scene type, and then transmits the noise-reduced audio clip through the instant messaging client to realize the noise reduction audio transmission function. For a call application, user A makes an audio call with user B. User A makes an audio call in the current audio scene through a call client operating on the terminal. The call client acquires user A's audio clip, automatically generates an optimization requirement for the audio based on user A's audio clip, and transmits the optimization requirement for the audio to the server. The server performs audio scene identification on user A's audio clip based on the received optimization requirement for the audio, determines the audio scene type, and performs targeted noise reduction processing on the audio clip based on the audio scene type, and then transmits the noise-reduced audio clip of user A to user B to realize the noise reduction audio call function.

[0083] The embodiments of the present application provide a voice processing method based on artificial intelligence, which extracts Mel-frequency logarithmic energy features for a voice clip, inputs the features after normalization into a neural network, and obtains a scene prediction corresponding to the voice clip. Since the call scene is relatively stable, the bitrate control based on the scene has better stability. The noise characteristics for different scenes can obtain personalized noise reduction measures for a specific scene by using methods such as adaptive learning and transfer learning, and can switch the dedicated noise reduction mode of a specific scene accordingly based on the result of scene identification, so as to obtain higher noise reduction performance and improve the call quality and user experience.

[0084] For example, in real-time communication meetings, with the continuous improvement of meeting mobile terminals, users may participate in meetings in various different environments (voice scenes, audio scenes), such as office environments, home environments, mobile vehicle environments such as subways and high-speed railways. Different scenes pose challenges with scene characteristics for real-time voice signal processing. For example, the scene signals such as high-speed railways are relatively weak and unstable, and always cause delays during voice communication, seriously affecting the communication quality. The specific background noises of different scenes (for example, children playing in a home environment, the background sound of a TV, the noise in a kitchen, etc.) put forward higher requirements for the robustness of the noise reduction algorithm.

[0085] To meet the need for users to hold meetings in each scene and improve the user's meeting voice experience in a complex environment, providing a solution that enables scene personalization based on environmental characteristics is an important trend in optimizing voice processing algorithms. Accurately identifying the voice generation scene is an important basis and foundation for realizing scene personalization measures. The embodiments of this application propose an Audio Scene Classification strategy and activate a voice processing algorithm that enables scene personalization for the classification results of voice scenes. For example, perform automatic bit rate switching for the high-speed railway scene where the signal is relatively weak and unstable, reduce the voice bit rate to avoid delays, apply noise reduction measures for specific scenes according to the identified scene, and improve the experience of users participating in meetings.

[0086] Here, the obtained noise reduction measures mainly target the noise of scene characteristics, such as the sound of keyboards in an office environment, the friction sound of paper, the noise in the kitchen of a home environment, playing children, the background sound of a TV, and the sound indicating the station name of a moving vehicle, etc. Train a noise reduction model for each scene (a model that focuses on deleting scene characteristic noise) through methods such as adaptation based on a general noise reduction model, and after identifying the scene, activate the noise reduction model corresponding to the scene to perform noise reduction. The obtained bit rate switching mainly reduces the bit rate of meeting communication (for example, reducing it from 16k to 8k) for a specific scene, such as a moving vehicle environment where the signal is relatively weak like a high-speed railway, reduces the transmission load and delays, and improves the experience of participating in the meeting.

[0087] The voice scene classification method proposed by the embodiments of the present application first extracts the corresponding frequency-domain frequency spectrum features, i.e., Mel Frequency Log Filterbank Energy, based on the collected time-domain voice signals, and then performs normalization processing on these frequency spectrum features. After the normalization processing, these normalized frequency spectrum features are input into a neural network model, such as a deep residual network (ResNet) based on a convolutional neural network (CNN), to model the normalized frequency spectrum features through the neural network model. During actual testing, first, the logarithmic energy spectrum of the input voice signal is normalized and input into the already created neural network model, and the neural network model outputs the scene classification results for each input audio clip. According to the scene results identified by the neural network model, the meeting system can automatically switch the adapted voice bitrate and activate noise reduction measures and the like suitable for the scene, thereby improving the overall audio call quality and user experience.

[0088] As shown in FIG. 7, FIG. 7 is an overall flow schematic diagram of voice scene identification provided by the embodiments of the present application. Voice scene identification includes two stages: training and testing, which include five modules, namely, 1) scene noise corpus collection, 2) training data construction, 3) training data feature extraction, 4) neural network model training, and 5) scene prediction.

[0089] 1) Scene noise corpus collection Collect the noise of scene characteristics, such as background noise under different scenes, for example, the sound of the keyboard, the friction sound of paper in an office environment, kitchen noise in a home environment, playing children, the background sound of the TV, the sound of the station name of a moving vehicle, etc.

[0090] 2) Construction of training data The background noise under different collected scenes is superimposed on different clean voices (audio without noise) in the time domain to generate a mixed signal in which the scene noise is superimposed on the clean voice, and this is used as the input corpus for neural network model training. When superimposing, it is necessary to prevent the audio amplitude after superimposition from exceeding the system threshold, and to enhance the diversity of the data. In order to better simulate the voice in the actual environment, data in which the human voice and noise are superimposed at a ratio of 1:1 in the original ratio is ensured, and some superimposed data with random ratios is generated. For example, the random number of the human voice superimposition coefficient is within 0.5 to 0.7, and the random number of the noise superimposition coefficient is within 0.3 to 0.5.

[0091] 3) Extraction of training data features Operations such as frame splitting, windowing, and Fourier transform are performed on the voice signals in the training data to obtain the mel logarithmic energy frequency spectrum features.

[0092] As shown in Figure 8, Figure 8 is a flow schematic diagram for extracting frequency spectrum features from the acoustic signal in the time domain provided by the embodiment of the present application. First, a frame splitting operation is performed on the time domain signal according to the short-term transient hypothesis to convert the continuous signal into a discrete vector. Then, windowing smoothing is performed on the voice signal of each frame to eliminate the discontinuity at the edge. Next, a Fourier transform (FT, Fourier Transform) is performed on each frame to obtain the frequency domain signal, and the mel frequency band is applied to the frequency domain signal to obtain the energy within each band. Based on the non-linear response of the human ear to voice, the mel frequency band is used here to replace the linear frequency band to better simulate the human ear response, and finally a logarithmic operation is performed to obtain the mel logarithmic frequency spectrum features.

[0093] 4) Training of neural network model The input of the neural network model is the 3-channel mel-logarithmic energy frequency spectrum feature with scene noise superimposed on the clean speech, and the output of the neural network model is the classification result of scene identification. In the training process, the cross entropy loss is adopted as the loss function, and the minimized loss function is used as the training target, that is, [Equation 1]. Here, t i represents the accurate scene mark of the input speech, and O i is the scene category predicted by the neural network model.

[0094]

Equation

[0095] As shown in Figure 9, Figure 9 is a structural schematic diagram of the neural network model provided by the embodiment of the present application. The neural network model consists of two ResNet units (residual neural network), a plurality of CNN networks, and an average pooling layer. The mel-logarithmic energy feature, its first derivative, and second derivative consist of 3-channel input signals, and finally output the scene classification result.

[0096] Here, the neural network model uses ResNet units. As shown in Figure 10, Figure 10 is a structural schematic diagram of the ResNet unit provided by the embodiment of the present application. Each ResNet unit includes two layers of CNN. Here, x and y are the input and output of the residual unit respectively, f1 and f2 represent the function mappings of the two CNN layers respectively, and W1 and W2 represent the weight parameters corresponding to the two CNN layers respectively. The CNN layer can effectively capture the scene noise features in the frequency spectrum information, and the residual neural network can effectively prevent the problem of gradient disappearance in the neural network training error transmission.

[0097] 5) Scene prediction After the training of the neural network model is completed, select the optimal model parameters and save them as the trained model. During testing, after normalizing the noisy audio, extract the frequency spectrum features and input them into the trained model, and the trained model outputs the predicted audio scene. Subsequently, activate an audio processing algorithm that enables scene personalization for the classification result of the audio scene. For example, perform automatic bit rate switching for the high-speed railway scene where the signal is relatively weak and unstable, reduce the audio bit rate to avoid latency, apply noise reduction measures for specific scenes according to the identified scene, and improve the user experience of participating in a meeting.

[0098] Therefore, the embodiment of the present application constructs a lightweight audio scene recognition model, which has low requirements for memory space and high prediction speed. As a front-end algorithm, it can be the basis and foundation for subsequent complex algorithm optimization. Adjust and control the audio bit rate according to the audio scene recognition result, and activate an audio solution that enables scene personalization such as scene-specific noise reduction measures.

[0099] So far, an exemplary application and implementation of the server provided by the embodiments of the present application have been combined to describe the voice processing method based on artificial intelligence provided by the embodiments of the present application. The embodiments of the present application further provide a voice processing device. In actual applications, each functional module in the voice processing device can be jointly realized by the hardware resources of an electronic device (such as a terminal device, a server, or a server cluster), such as computing resources like a processor, communication resources (such as those used to support and realize various communication methods such as cables and cellular), and memory. FIG. 2 shows a voice processing device 555 stored in a memory 550, which may be software in the form of a program and a plugin assembly, etc. For example, a software module designed by programming languages such as software C / C++, Java, etc., application software designed by programming languages such as C / C++, Java, etc., or a dedicated software module, application program interface, plugin assembly, cloud service, etc. in a large software system. Hereinafter, examples will be given and described for different implementation methods.

[0100] Example 1. The voice processing device is a mobile terminal application program and a module

[0101] The voice processing device 555 in the embodiments of the present application can be provided as a software module designed using programming languages such as software C / C++, Java, etc., and incorporated into various mobile terminal applications based on systems such as Android or iOS (stored in the storage medium of the mobile terminal as executable instructions and executed by the processor of the mobile terminal), thereby directly using the computing resources of the mobile terminal itself to complete related information recommendation tasks, and periodically or irregularly transferring the processing results to a remote server through various network communication methods, or storing them locally on the mobile terminal.

[0102] Example 2. The voice processing device is a server application program and a platform

[0103] The voice processing device 555 in the embodiment of the present application can be provided as application software designed using programming languages such as C / C++ and Java, or as a dedicated software module in a large software system. It operates on the server side (stored in the server-side storage medium in an executable instruction format and operated by the server-side processor), and the server uses its own computing resources to complete related information recommendation tasks.

[0104] The embodiment of the present application can further provide a customized and user-friendly network (Web) interface or other user interfaces (UI) on a distributed and parallel computing platform composed of multiple servers to form an information recommendation platform for personal, group, or unit use (used in the recommendation list), etc.

[0105] Example 3. The voice processing device is a server-side application program interface (API) and a plugin assembly

[0106] The voice processing device 555 in the embodiment of the present application can be provided as a server-side API or a plugin assembly. By being called by the user, the voice processing method based on artificial intelligence in the embodiment of the present application is executed and incorporated into various application programs.

[0107] Example 4. The voice processing device is a mobile device client API and a plugin assembly

[0108] The voice processing device 555 in the embodiments of the present application can be provided as an API at the mobile device end or as a plug-in assembly, so that when called by a user, the voice processing method based on artificial intelligence in the embodiments of the present application can be executed.

[0109] Example 5. The voice processing device is a cloud-end public service

[0110] The voice processing device 555 in the embodiments of the present application can be provided as an information recommendation cloud service developed for users, so that individuals, groups, or units can obtain a recommendation list.

[0111] Here, the voice processing device 555 includes a series of modules, including an acquisition module 5551, a classification module 5552, a processing module 5553, and a training module 5554. Hereinafter, it will continue to be described how each module in the voice processing device 555 provided by the embodiments of the present application cooperates to realize a voice processing strategy.

[0112] The acquisition module 5551 is configured to acquire a voice clip of a voice scene. Here, noise is included in the voice clip. The classification module 5552 is configured to perform voice scene classification processing based on the voice clip to obtain a voice scene type corresponding to the noise in the voice clip. The processing module 5553 is configured to determine a target voice processing mode corresponding to the voice scene type, and apply the target voice processing mode to the voice clip based on the interference degree caused by the noise in the voice clip.

[0113] In some embodiments, the target voice processing mode includes a noise reduction processing mode. The processing module 5553 is further configured to query the correspondence between different candidate voice scene types and candidate noise reduction processing modes based on the voice scene type corresponding to the voice scene, so as to obtain a noise reduction processing mode corresponding to the voice scene type.

[0114] In some embodiments, the target voice processing mode includes a noise reduction processing mode, and the processing module 5553 is further configured to determine a noise type that matches the voice scene type based on the voice scene type corresponding to the voice scene, and query the correspondence between different candidate noise types and candidate noise reduction processing modes based on the noise type that matches the voice scene type, so as to obtain a noise reduction processing mode corresponding to the voice scene type. Here, the noise types that match different voice scene types are not exactly the same.

[0115] In some embodiments, before the step of applying the target voice processing mode to the voice clip of the voice scene, the processing module 5553 is further configured to determine an interference degree caused by noise in the voice clip, and when the interference degree is greater than an interference degree threshold, apply a noise reduction processing mode corresponding to the voice scene type to the voice clip.

[0116] In some embodiments, the processing module 5553 is further configured to perform a matching process on the noise in the voice clip with respect to the noise type that matches the voice scene type, perform a suppression process on the noise that has successfully matched the noise type, and obtain a voice clip after suppression. Here, the ratio of the audio signal strength to the noise signal strength in the voice clip after suppression is lower than a signal-to-noise ratio threshold.

[0117] In some embodiments, the target voice processing mode includes a bitrate switching processing mode, and the processing module 5553 is further configured to query the correspondence between different candidate voice scene types and candidate bitrate switching processing modes based on the voice scene type corresponding to the voice scene, so as to obtain a bitrate switching processing mode corresponding to the voice scene type.

[0118] In some embodiments, the target voice processing mode includes a bitrate switching processing mode, and the processing module 5553 is further configured to compare the voice scene type corresponding to the voice scene with a preset voice scene type, and when it is determined by comparison that the voice scene type is the preset voice scene type, set the bitrate switching processing mode related to the preset voice scene type as the bitrate switching processing mode corresponding to the voice scene type.

[0119] In some embodiments, the processing module 5553 is further configured to obtain the communication signal strength of the voice scene, and when the communication signal strength of the voice scene is less than the communication signal strength threshold, reduce the voice bitrate of the voice clip according to a first setting ratio or a first setting value, and when the communication signal strength of the voice scene is greater than or equal to the communication signal strength threshold, increase the voice bitrate of the voice clip according to a second setting ratio or a second setting value.

[0120] In some embodiments, the processing module 5553 is further configured to determine jitter information of the communication signal strength in the voice scene based on the communication signal strength obtained by multiple samplings in the voice scene, and when the jitter information indicates that the communication signal is in an unstable state, reduce the voice bitrate of the voice clip according to a third setting ratio or a third setting value.

[0121] In some embodiments, the processing module 5553 is further configured to reduce the voice bitrate of the voice clip according to a fourth setting ratio or a fourth setting value when the type of the communication network used to transmit the voice clip belongs to a set type.

[0122] In some embodiments, the above voice scene classification process is realized through a neural network model. The neural network model learns the correlation between the noise included in the voice clip and the voice scene type. The classification module 5552 is further configured to call the neural network model based on the voice clip to execute a voice scene classification process, so as to obtain a voice scene type having a correlation with the noise included in the voice clip.

[0123] In some embodiments, the neural network model includes a mapping network, a residual network, and a pooling network. The classification module 5552 is further configured to perform feature extraction processing on the voice clip through the mapping network to obtain a first feature vector of the noise in the voice clip, perform mapping processing on the first feature vector through the residual network to obtain a mapping vector of the voice clip, perform feature extraction processing on the mapping vector of the voice clip through the mapping network to obtain a second feature vector of the noise in the voice clip, perform pooling processing on the second feature vector through the pooling network to obtain a pooling vector of the voice clip, and perform non-linear mapping processing on the pooling vector of the voice clip, so as to obtain a voice scene type having a correlation with the noise included in the voice clip.

[0124] In some embodiments, the mapping network includes a plurality of cascaded mapping layers, and the classification module 5552 further performs feature mapping processing on the audio clip through the first mapping layer in the plurality of cascaded mapping layers, outputs the mapping result of the first mapping layer to the subsequent cascaded mapping layers, and continuously performs feature mapping and output of the mapping result through the subsequent cascaded mapping layers until it is output to the last mapping layer, and is configured such that the mapping result output from the last mapping layer is used as the first feature vector of the noise in the audio clip.

[0125] In some embodiments, the residual network includes a first mapping network and a second mapping network, and the classification module 5552 further performs mapping processing on the first feature vector through the first mapping network to obtain the first mapping vector of the audio clip, performs non-linear mapping processing on the first mapping vector to obtain the non-mapping vector of the audio clip, performs mapping processing on the non-mapping vector of the audio clip through the first mapping network to obtain the second mapping vector of the audio clip, and is configured such that the total result of the first feature vector of the audio clip and the second mapping vector of the audio clip is used as the mapping vector of the audio clip.

[0126] In some embodiments, the apparatus further includes a training module 5554, which constructs audio samples corresponding to the plurality of different audio scenes based on a noise-free audio signal and background noises respectively corresponding to the plurality of different audio scenes, and trains a neural network model based on the audio samples respectively corresponding to the plurality of different audio scenes to obtain a neural network model used for classifying audio scenes.

[0127] In some embodiments, the training module 5554 further performs, for any one of the plurality of different audio scenes, a process of fusing the background noise of the audio scene and the noise-free audio signal based on the fusion ratio between the background noise of the audio scene and the noise-free audio signal to obtain a first fused audio signal of the audio scene, a process of fusing the background noise of the audio scene corresponding to a first random coefficient in the first fused audio signal to obtain a second fused audio signal of the audio scene, and a process of fusing the noise-free audio signal corresponding to a second random coefficient in the second fused audio signal to obtain an audio sample of the audio scene.

[0128] In some embodiments, the training module 5554 further performs audio scene classification processing on the audio samples respectively corresponding to the plurality of different audio scenes through the neural network model to obtain the predicted audio scene type of the audio samples, constructs a loss function of the neural network model based on the predicted audio scene type of the audio samples, the audio scene mark of the audio samples, and the weighting of the audio samples, continuously updates the parameters of the neural network model until the loss function converges, and when the loss function converges, sets the updated parameters of the neural network model as the parameters of the neural network model used for classifying audio scenes.

[0129] In some embodiments, before the step of performing the audio scene classification process based on the audio clip, the acquisition module 5551 is further configured to perform frame splitting processing on the time domain signal of the audio clip to obtain a multi-frame audio signal, perform windowing processing on the multi-frame audio signal, and perform Fourier transform on the audio signal after windowing processing to obtain the frequency domain signal of the audio clip, and perform logarithmic processing on the mel frequency band of the frequency domain signal to obtain the audio clip used for performing the audio scene classification.

[0130] The embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and by executing the computer instructions, causes the computer device to execute the audio processing method based on artificial intelligence described in the embodiments of the present application.

[0131] The embodiments of the present application provide a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by a processor, cause the processor to execute the audio processing method based on artificial intelligence provided by the embodiments of the present application, for example, the audio processing method based on artificial intelligence shown in FIGS. 3-5.

[0132] In some embodiments, the computer-readable storage medium may be a memory such as FRAM (registered trademark), ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM, or various devices including one or any combination of the above memories.

[0133] In some embodiments, the executable instructions may take the form of a program, software, software module, script, or code, and can be created in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and it can be arranged in any form, including being arranged in an independent program, or being arranged in a module, assembly, subroutine, or other assembly suitable for use in a computing environment.

[0134] By way of example, the executable instructions may or may not correspond to a file in a file system, and may be stored as part of another program or file storing data, for example, stored in one or more scripts in a hypertext markup language (HTML) document, stored in a single file dedicated to the program under consideration, or stored in multiple co-files (e.g., files storing one or more modules, subroutines, or code portions).

[0135] By way of example, the executable instructions can be arranged and executed on one computing device, or can be arranged and executed on multiple computing devices located in one place, or can also be arranged and executed on multiple computing devices distributed in multiple locations and coupled to each other through a communication network.

[0136] What is described above are only examples of the present application and are not used to limit the protection scope of the present application. Any revisions, equivalent substitutions, improvements, etc. made within the spirit and scope of the present application are all included within the protection scope of the present application.

Description of Reference Numerals

[0137] 10 Voice Processing System 100 Server 200 Terminal 300 Network 500 Electronic Device 510 Processor 520 Network Interface 530 User Interface 540 Bus System 550 Memory 551 Operating System 552 Network Communication Module 555 Voice Processing Device 5551 Acquisition Module 5552 Classification Module 5553 Processing Module 5554 Training Module

Claims

1. An artificial intelligence-based voice processing method executed by a computer, the method comprising: obtaining a voice clip of a voice scene, wherein the voice clip includes noise; performing voice scene classification processing based on the voice clip to obtain a voice scene type corresponding to the noise in the voice clip; determining a target voice processing mode corresponding to the voice scene type; applying the target voice processing mode to the voice clip, wherein the target voice processing mode includes a noise reduction processing mode; the step of determining the target voice processing mode corresponding to the voice scene type includes: determining a noise type that matches the voice scene type based on the voice scene type corresponding to the voice scene; querying a correspondence relationship between different candidate noise types and candidate noise reduction processing modes based on the noise type that matches the voice scene type to obtain a noise reduction processing mode corresponding to the voice scene type; An artificial intelligence-based voice processing method, wherein the noise types that match different voice scene types are not exactly the same.

2. The step of applying the target voice processing mode to the voice clip includes: determining an interference degree caused by the noise in the voice clip; when the interference degree is greater than an interference degree threshold, applying the noise reduction processing mode corresponding to the voice scene type to the voice clip. The method according to claim 1.

3. An artificial intelligence-based voice processing method executed by a computer, the method comprising: obtaining a voice clip of a voice scene, wherein the voice clip includes noise; performing voice scene classification processing based on the voice clip to obtain a voice scene type corresponding to the noise in the voice clip; determining a target voice processing mode corresponding to the voice scene type; applying the target voice processing mode to the voice clip, wherein the step of applying the target voice processing mode to the voice clip is: performing a matching process on the noise in the audio clip with respect to the noise type matching the audio scene type; performing a suppression process on the noise successfully matched with the noise type to obtain the suppressed audio clip, including: An artificial intelligence-based audio processing method, wherein the ratio of the audio signal strength to the noise signal strength in the suppressed audio clip is lower than the signal-to-noise ratio threshold.

4. The audio scene classification process is realized through a neural network model, and the neural network model learns the correlation between the noise included in the audio clip and the audio scene type. The step of performing an audio scene classification process based on the audio clip to obtain the noise in the audio clip and the corresponding audio scene type is: The method according to claim 1 or 3, including the step of calling the neural network model based on the audio clip to perform an audio scene classification process to obtain the audio scene type having a correlation with the noise included in the audio clip.

5. The neural network model includes a mapping network, a residual network, and a pooling network. The step of calling the neural network model based on the audio clip to perform an audio scene classification process to obtain the audio scene type having a correlation with the noise included in the audio clip is: performing a feature extraction process on the audio clip through the mapping network to obtain a first feature vector of the noise in the audio clip; performing a mapping process on the first feature vector through the residual network to obtain a mapping vector of the audio clip; performing a feature extraction process on the mapping vector of the audio clip through the mapping network to obtain a second feature vector of the noise in the audio clip; performing a pooling process on the second feature vector through the pooling network to obtain a pooling vector of the audio clip. Performing non-linear mapping processing on the pooling vector of the voice clip to obtain a voice scene type having an associated relationship with the noise included in the voice clip, the method according to claim 4, comprising:

6. The mapping network includes a plurality of cascaded mapping layers, The step of performing feature extraction processing on the voice clip through the mapping network to obtain a first feature vector of the noise in the voice clip is: Performing feature mapping processing on the voice clip through the first mapping layer in the plurality of cascaded mapping layers; Outputting the mapping result of the first mapping layer to the subsequent cascaded mapping layers, and continuously performing feature mapping and output of the mapping result through the subsequent cascaded mapping layers until the output of the last mapping layer; and Taking the mapping result output from the last mapping layer as the first feature vector of the noise in the voice clip, the method according to claim 5, comprising:

7. The residual network includes a first mapping network and a second mapping network, The step of performing mapping processing on the first feature vector through the residual network to obtain a mapping vector of the voice clip is: Performing mapping processing on the first feature vector through the first mapping network to obtain a first mapping vector of the voice clip; Performing non-linear mapping processing on the first mapping vector to obtain a non-mapping vector of the voice clip; Performing mapping processing on the non-mapping vector of the voice clip through the first mapping network to obtain a second mapping vector of the voice clip; Taking the total result of the first feature vector of the voice clip and the second mapping vector of the voice clip as the mapping vector of the voice clip, the method according to claim 5, comprising:

8. A voice processing method based on artificial intelligence, executed by a computer, the method comprising: A step of obtaining an audio clip of an audio scene, wherein the audio clip contains noise, and Executing an audio scene classification process based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip, and Determining a target audio processing mode corresponding to the audio scene type, and Applying the target audio processing mode to the audio clip, and The method is Constructing audio samples corresponding to the plurality of different audio scenes based on a noise-free audio signal and background noises respectively corresponding to the plurality of different audio scenes, and Training a neural network model based on the audio samples respectively corresponding to the plurality of different audio scenes to obtain a neural network model used for classifying audio scenes, and The step of constructing audio samples respectively corresponding to the plurality of different audio scenes based on a noise-free audio signal and background noises respectively corresponding to the plurality of different audio scenes is For any audio scene among the plurality of different audio scenes Based on a fusion ratio between the background noise of the audio scene and the noise-free audio signal, fusing the background noise of the audio scene and the noise-free audio signal to obtain a first fused audio signal of the audio scene, and Fusing the background noise of the audio scene corresponding to a first random coefficient in the first fused audio signal to obtain a second fused audio signal of the audio scene, and Fusing the noise-free audio signal corresponding to a second random coefficient in the second fused audio signal to obtain an audio sample of the audio scene, and

9. The step of training a neural network model based on the audio samples respectively corresponding to the plurality of different audio scenes to obtain a neural network model used for classifying audio scenes is Performing an audio scene classification process on the audio samples respectively corresponding to the plurality of different audio scenes through the neural network model to obtain a predicted audio scene type of the audio sample, and A step of constructing a loss function of the neural network model based on a predicted audio scene type of the audio sample, an audio scene mark of the audio sample, and weighting of the audio sample; continuously updating parameters of the neural network model until the loss function converges, and setting the updated parameters of the neural network model when the loss function converges as parameters of a neural network model used for classifying an audio scene, the method according to claim 8.

10. Before the step of performing audio scene classification processing based on the audio clip to obtain a noise and a corresponding audio scene type within the audio clip, the method further includes: performing frame division processing on a time domain signal of the audio clip to obtain a multi-frame audio signal; performing windowing processing on the multi-frame audio signal, and performing Fourier transform on the audio signal after windowing processing to obtain a frequency domain signal of the audio clip; performing logarithmic processing on a mel frequency band of the frequency domain signal to obtain the audio clip used for performing the audio scene classification, the method according to claim 1 or 8.

11. An audio processing apparatus, the apparatus includes: an acquisition module configured to acquire an audio clip of an audio scene, where noise is included in the audio clip; a classification module configured to perform audio scene classification processing based on the audio clip to obtain a noise and a corresponding audio scene type within the audio clip; a processing module configured to determine a target audio processing mode corresponding to the audio scene type, and apply the target audio processing mode to the audio clip, where the target audio processing mode includes a noise reduction processing mode; the processing module is configured to: determine a noise type matching the audio scene type based on the audio scene type corresponding to the audio scene; query a correspondence relationship between different candidate noise types and candidate noise reduction processing modes based on the noise type matching the audio scene type to obtain a noise reduction processing mode corresponding to the audio scene type. The noise types matching the different voice scene types are not exactly the same, a voice processing device.

12. An electronic device, wherein the electronic device a memory used to store executable instructions, and a processor used to implement the voice processing method based on artificial intelligence according to any one of claims 1 to 10 when executing the executable instructions stored in the memory. An electronic device comprising:

13. A computer program for causing a computer to execute the voice processing method based on artificial intelligence according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Audio denoising method based on scene, device, electronic equipment and storage medium

    CN109817236A

  • Noise suppressor and method for suppressing noise

    JP2002073066A

  • Noise detecting device and noise detecting method

    JP2008145988A

  • Voice recognition device, voice recognition program and voice recognition method

    JP2020034683A

  • Optimization of network microphone devices using noise classification

    US10602268B1