Audio Processing Method, Apparatus, Computer Device, Storage Medium, and Program Product

By performing feature encoding and classification learning of the band signals of audio frames, combined with frame type filtering processing, the problem of insufficient detection accuracy in music scenes in the prior art is solved, and higher detection accuracy and better audio processing performance are achieved.

CN115130569BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210719734.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-05-30
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

The prior art lacks accuracy in music scene detection, making it difficult to effectively identify the business model of audio.

Method used

By acquiring multiple band signals of audio frames, feature encoding, classification learning and feature decoding are performed, the frame type of audio frame is determined, and filtering is performed based on the frame type to improve the accuracy of music scene detection.

Benefits of technology

It improves the accuracy of music scene detection, enhances the ability to recognize audio business models, and improves the performance of the audio processing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130569B_ABST
    Figure CN115130569B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method, apparatus, computer device, storage medium, and program product. The method relates to the artificial neural network technology in the machine learning technology in the field of artificial intelligence. The method includes: performing feature encoding on each of the multiple band signals of a target audio frame in the audio to be processed to obtain a feature vector corresponding to each of the multiple band signals; performing classification learning on the multiple feature vectors obtained by the feature encoding to obtain a classification indication vector; performing feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determining the frame type of the target audio frame according to the classification result; based on the frame type of the target audio frame, performing filtering processing on the audio to be processed to obtain a service mode of the audio to be processed, and performing service processing on the audio to be processed according to the determined service mode; by using the embodiments of the present application, the accuracy of music scene detection can be improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to an audio processing method, an audio processing device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the continuous development of computer technologies, audio processing technologies have gradually become a research focus in the field of computer technologies. As an important research direction in audio processing technologies, music scene detection is of great significance for audio-related usage scenarios; the so-called music scene detection refers to a technology for detecting audio to determine whether the service mode of the audio is a music mode or a non-music mode. Currently, some solutions for music scene detection based on traditional digital signal technologies can basically meet the requirements of audio-related usage scenarios, but the accuracy of music scene detection still needs to be improved. Summary of the Invention

[0003] Embodiments of this application provide an audio processing method, device, computer device, storage medium, and program product, which can improve the accuracy of music scene detection to a certain extent.

[0004] On the one hand, embodiments of this application provide an audio processing method, which includes:

[0005] Obtain multiple band signals of a target audio frame in the audio to be processed;

[0006] Perform feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals;

[0007] Perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector;

[0008] Perform feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determine the frame type of the target audio frame according to the classification result;

[0009] Based on the frame type of the target audio frame, perform filtering processing on the audio to be processed to obtain the service mode of the audio to be processed, and perform service processing on the audio to be processed according to the determined service mode; the filtering processing includes statistical filtering processing and decision filtering processing.

[0010] Correspondingly, embodiments of this application provide an audio processing device, which includes:

[0011] An obtaining unit, configured to obtain multiple band signals of a target audio frame in the audio to be processed;

[0012] A processing unit for performing feature encoding on each of a plurality of band signals to obtain a feature vector corresponding to each of the plurality of band signals;

[0013] The processing unit is further configured to perform classification learning on the plurality of feature vectors obtained by feature encoding to obtain a classification indication vector;

[0014] The processing unit is further configured to perform feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determine the frame type of the target audio frame according to the classification result;

[0015] The processing unit is further configured to perform filtering processing on the audio to be processed based on the frame type of the target audio frame to obtain a service mode of the audio to be processed, and perform service processing on the audio to be processed according to the determined service mode; the filtering processing includes statistical filtering processing and decision filtering processing.

[0016] In one implementation, when the obtaining unit is configured to obtain a plurality of band signals of a target audio frame in the audio to be processed, it is specifically configured to perform the following steps:

[0017] Obtain the original time-domain signal of the target audio frame in the audio to be processed, and obtain the target time-domain signal from the original time-domain signal according to the sampling information, where the sampling rate of the target time-domain signal meets the signal analysis conditions corresponding to the sampling information;

[0018] Perform frequency-domain conversion processing on the target time-domain signal to obtain the frequency-domain signal of the target audio frame;

[0019] Perform frequency band division processing on the frequency-domain signal to obtain a plurality of band signals.

[0020] In one implementation, when the obtaining unit is configured to obtain the target time-domain signal from the original time-domain signal according to the sampling information, it is specifically configured to perform the following steps:

[0021] Obtain the sampling rate when performing sampling processing on the original time-domain signal;

[0022] If the sampling rate of the original time-domain signal meets the signal analysis conditions, determine the original time-domain signal as the target time-domain signal;

[0023] If the sampling rate of the original time-domain signal does not meet the signal analysis conditions, perform frequency separation processing on the original time-domain signal based on the sampling information to obtain the target time-domain signal.

[0024] In one implementation, the number of band signals is M, where M is an integer greater than 1; the feature encoding process is performed by the feature encoding layer of the classification model. The feature encoding layer includes M feature encoding units, and the M feature encoding units correspond one-to-one with the M band signals. The M band signals are sequentially input into the corresponding feature encoding units in the order of the magnitudes of the band information corresponding to the M band signals respectively; when the processing unit is used to perform feature encoding on each of the multiple band signals to obtain the feature vector corresponding to each of the multiple band signals, it is specifically used to perform the following steps:

[0025] Call the m-th feature encoding unit among the M feature encoding units to perform feature encoding on the m-th band signal corresponding to the m-th feature encoding unit, so as to obtain the feature vector corresponding to the m-th band signal, where m is a positive integer less than or equal to M.

[0026] In one implementation, the feature learning process is performed by the feature learning layer of the classification model. The feature learning layer includes N feature learning units; the number of feature vectors is M, and the M feature vectors are divided into N vector groups, each vector group includes at least two feature vectors, both N and M are integers greater than 1, and N is less than or equal to M; when the processing unit is used to perform classification learning on the multiple feature vectors obtained by feature encoding to obtain the classification indication vector, it is specifically used to perform the following steps:

[0027] Call the first feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the first vector group among the N vector groups, so as to obtain the feature learning result of the first feature learning unit;

[0028] The feature learning layer calls the second feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the second vector group among the N vector groups and the feature learning result of the first feature learning unit, so as to obtain the feature learning result of the second feature learning unit;

[0029] Among them, the number of classification indication vectors is P, and the P classification indication vectors are determined according to the feature learning results of the last P feature learning units among the N feature learning units, where P is a positive integer less than or equal to N.

[0030] In one implementation, the number of classification indication vectors is P, where P is a positive integer; the feature decoding process is performed by the feature decoding layer of the classification model. The feature decoding layer includes Y feature decoding units, and different feature decoding units among the Y feature decoding units correspond to different frame types, where Y is an integer greater than 1; when the processing unit is used to perform feature decoding on the classification indication vector to obtain the classification result of the target audio frame, it is specifically used to perform the following steps:

[0031] Invoke the y-th feature decoding unit among the Y feature decoding units to perform dimensionality reduction processing on the P classification indication vectors, and obtain the dimensionality reduction result of the y-th feature decoding unit; and perform regression processing on the dimensionality reduction result of the y-th feature decoding unit to obtain the probability that the target audio frame belongs to the frame type corresponding to the y-th feature decoding unit, where y is a positive integer less than or equal to Y;

[0032] Among them, the classification result includes: the probability that the target audio frame belongs to the frame type corresponding to each of the Y feature decoding units.

[0033] In one implementation, the audio frames in the audio to be processed are divided into multiple audio blocks, and each audio block includes at least two audio frames; the audio sliding window slides and detects in the audio to be processed with one audio block as the sliding step, and the audio sliding window allows accommodating at least two audio blocks; when the processing unit is used to perform filtering processing on the audio to be processed based on the frame type of the target audio frame to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0034] Based on the frame type of the target audio frame, perform statistical filtering processing on the parent audio block of the target audio frame to determine the state of the parent audio block, where the parent audio block refers to the audio block containing the target audio frame;

[0035] If the audio sliding window slides to a position containing the parent audio block, perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed.

[0036] In one implementation, when the processing unit is used to perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0037] Statistically count the states of the respective audio blocks currently included in the audio sliding window, and the parent audio block is within the audio sliding window;

[0038] If the state statistical result indicates that the number of target audio blocks in the respective audio blocks currently included in the audio sliding window meets the audio block number condition, determine that the service mode of the audio to be processed is the music mode, where the target audio block refers to the audio block in the respective audio blocks currently included in the audio sliding window that is in the active state;

[0039] If the state statistical result indicates that the number of target audio blocks in the respective audio blocks currently included in the audio sliding window does not meet the audio block number condition, determine that the service mode of the audio to be processed is the non-music mode.

[0040] In one implementation, the processing unit is configured to perform statistical filtering on the parent audio block of the target audio frame based on the frame type of the target audio frame, and when determining the state of the parent audio block, specifically perform the following steps:

[0041] Perform statistics on the frame types of each audio frame in the parent audio block;

[0042] If the frame type statistics result indicates that the number of audio frames belonging to the target frame type among each audio frame included in the parent audio block meets the audio frame quantity condition, determine that the state of the parent audio block is the active state.

[0043] In one implementation, when the service mode is the music mode, the processing unit, when performing service processing on the to-be-processed audio according to the determined service mode, specifically performs the following steps:

[0044] Perform noise reduction processing on the to-be-processed audio according to the low noise reduction processing rule corresponding to the music mode; or, output a music mode selection interface, and in response to a music mode selection operation input in the music mode selection interface, perform noise reduction processing on the to-be-processed audio according to the low noise reduction processing rule corresponding to the music mode;

[0045] When the service mode is a non-music mode, the processing unit, when performing service processing on the to-be-processed audio according to the determined service mode, specifically performs the following steps: Perform noise reduction processing on the to-be-processed audio according to the high noise reduction processing rule corresponding to the non-music mode.

[0046] In one implementation, the classification result includes the probability that the target audio frame belongs to each of the Y frame types, where Y is an integer greater than 1; different frame types among the Y frame types correspond to different probability ranges; the processing unit, when determining the frame type of the target audio frame according to the classification result, specifically performs the following steps:

[0047] Determine the audio scene type of the to-be-processed audio;

[0048] According to the audio scene type, determine the weight of each of the Y frame types;

[0049] Perform weighted summation processing on the probabilities of the corresponding frame types among the Y frame types according to the weight of each of the Y frame types to obtain a weighted average probability;

[0050] Determine the frame type of the target audio frame according to the probability range to which the weighted average probability belongs.

[0051] Correspondingly, an embodiment of the present application provides a computer device, which includes a processor and a computer-readable storage medium, where:

[0052] A processor, adapted to implement a computer program;

[0053] A computer-readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the above audio processing method.

[0054] Accordingly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when read and executed by a processor of a computer device, causes the computer device to perform the above audio processing method.

[0055] Accordingly, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the above audio processing method.

[0056] In an embodiment of the present application, after obtaining multiple band signals of a target audio frame in the audio to be processed, by performing feature encoding on the multiple band signals, performing feature learning on the feature vectors obtained by the feature encoding, and performing feature decoding on the classification indication vectors obtained by the classification learning, the accuracy of the classification result obtained by the feature decoding can be improved, and the accuracy of the frame type of the target audio frame determined according to the classification result can be improved, thereby improving the accuracy of music scene detection. And based on the frame type of the target audio frame, performing statistical filtering processing and decision filtering processing on the audio to be processed can further improve the accuracy of music scene detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0058] Figure 1 is a schematic diagram of frame division of an audio provided by an embodiment of the present application;

[0059] Figure 2a is a schematic diagram of the architecture of an audio processing system provided by an embodiment of the present application;

[0060] Figure 2b is a schematic diagram of the architecture of another audio processing system provided by an embodiment of the present application;

[0061] Figure 3It is a schematic flowchart of an audio processing method provided by an embodiment of the present application;

[0062] Figure 4 It is a schematic structural diagram of a classification model provided by an embodiment of the present application;

[0063] Figure 5 It is a schematic flowchart of another audio processing method provided by an embodiment of the present application;

[0064] Figure 6 It is a schematic diagram of a frequency band division result provided by an embodiment of the present application;

[0065] Figure 7 It is a schematic flowchart of yet another audio processing method provided by an embodiment of the present application;

[0066] Figure 8a It is a schematic diagram of an audio block in an active state provided by an embodiment of the present application;

[0067] Figure 8b It is a schematic diagram of an audio block in a non - active state provided by an embodiment of the present application;

[0068] Figure 9a It is a schematic diagram of a sliding detection process provided by an embodiment of the present application;

[0069] Figure 9b It is a schematic diagram of another sliding detection process provided by an embodiment of the present application;

[0070] Figure 10a It is a schematic diagram of a music mode selection interface provided by an embodiment of the present application;

[0071] Figure 10b It is a schematic diagram of a non - music mode selection interface provided by an embodiment of the present application;

[0072] Figure 11 It is a schematic flowchart of an audio processing method provided by an embodiment of the present application;

[0073] Figure 12 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;

[0074] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0075] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0076] In order to better understand the technical solutions provided by the embodiments of the present application, the following introduces the key terms related to the embodiments of the present application:

[0077] (1) The embodiments of the present application relate to audio. Audio refers to all sounds that humans can hear, and the audio may include contents such as music and speech. Classifying audio according to the usage scenarios of audio, audio can be divided into call audio, conference audio, film and television audio, song audio, etc.; classifying audio according to the generation methods of audio, audio can be divided into real-time audio and non-real-time audio. Among them, real-time audio refers to audio that is generated and transmitted in real time, such as call audio and conference audio, and non-real-time audio refers to audio that is pre-produced, such as film and television audio and song audio. The audio involved in the embodiments of the present application is real-time audio, which is hereby stated. Dividing the time-domain signal of audio according to the time range can obtain multiple audio frames, that is to say, the audio may include multiple audio frames. For example, Figure 1 for the audio shown, the audio content from 0 to 20 ms (milliseconds) in the audio is divided into one audio frame, and the audio content from 20 to 40 ms is divided into one audio frame, and so on until the last audio frame of the audio is obtained.

[0078] (2) The embodiments of the present application relate to music scene detection. Music scene detection belongs to a type of ASC (Acoustic Scene Classify), which refers to the technology of detecting audio to determine whether the service mode of the audio is a music mode or a non-music mode. The service mode of the audio being a music mode means that the audio contains a large amount of music content, and the service mode of the audio being a non-music mode means that the audio contains a small amount of music content or the audio does not contain audio content.

[0079] (3) The embodiments of the present application relate to DSP (Digital Signal Process) technology. For audio, digital signal processing technology refers to the technology of using a computer or a dedicated processing device to collect, transform, filter, estimate, enhance, compress, identify, etc. the audio signal in digital form to obtain a signal form that meets people's needs.

[0080] (4) The embodiments of the present application relate to the ANNs (Artificial Neural Networks) technology in the ML (Machine Learning) technology in the field of AI (Artificial Intelligence). Among them, AI is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making. ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. ANNs, also simply referred to as neural networks or connection models, are algorithmic mathematical models that imitate the behavioral characteristics of animal neural networks and perform distributed parallel information processing. ANNs rely on the complexity of the system and adjust the relationships between a large number of internal nodes to achieve the purpose of processing information.

[0081] Based on the above descriptions of key terms such as audio, music scene detection, digital signal processing, and artificial neural networks, the embodiments of the present application combine digital signal processing technology with artificial neural networks and propose an audio processing method that can improve the efficiency and accuracy of music scene detection.

[0082] In the digital signal processing stage of the audio processing method, for the original time-domain signal of any audio frame in the audio, for the original time-domain signal of high frequency (the "high frequency" mentioned here refers to high sampling frequency) whose sampling frequency (which can also be simply referred to as sampling rate) does not meet the conditions, the target time-domain signal of low frequency (the "low frequency" mentioned here refers to low sampling frequency) is separated from the original time-domain signal of high frequency of the audio, and then the target time-domain signal is subjected to frequency-domain conversion processing to obtain the frequency-domain signal of the audio frame, and the frequency-domain signal is divided into multiple frequency-band signals. For the original time-domain signal of high frequency, the digital signal processing stage of the audio processing method does not need to perform frequency-domain conversion processing on all the original time-domain signals, which can reduce the amount of data for frequency-domain conversion processing, thereby reducing the amount of calculation in the music scene detection process, and thus improving the efficiency of music scene detection.

[0083] In the neural network learning stage of the audio processing method, a special classification and recognition neural network is designed to classify and recognize the multiple frequency-band signals of the audio frame, improving the classification accuracy of the target audio frame, and thus improving the accuracy of music scene detection. Moreover, the classification and recognition neural network is composed of a fully connected layer and a recurrent neural network, with a simple network structure and low complexity, which can further reduce the amount of calculation in the music scene detection process and further improve the efficiency of music scene detection.

[0084] After obtaining the classification result of the audio frame, the frame type of the audio frame can be determined according to the classification result of the audio frame. Based on the frame types of each audio frame in the audio, in the embodiment of the present application, the service mode of the audio is determined to be a music mode or a non-music mode through music mode decision, which can further improve the accuracy of music scene detection.

[0085] The audio processing method provided by the embodiments of the present application can be applied to audio-visual communication scenarios (such as audio-visual call scenarios (such as VOIP (Voice over Internet Protocol, voice transmission based on IP (Internet Protocol, Internetworking Protocol)) call scenarios) or audio-visual conference scenarios). In traditional audio-visual communication scenarios, the music content in the audio is generally suppressed as noise to enhance the voice content in the audio. Such a solution seriously affects the audio-visual communication experience for the case where music is played during audio-visual communication, that is, the music content in the audio is an effective content in the audio-visual communication scenario. In the embodiments of the present application, when it is determined that the service mode of the audio is the music mode in the audio-visual communication scenario, low noise reduction processing can be performed on the audio. Low noise reduction processing means weakening the degree of noise reduction so that the music played in the audio-visual communication scenario is retained, or means removing other noise content in the audio except for the voice content and the music content so that the music played in the audio-visual communication scenario is retained; when it is determined that the service mode of the audio is the non-music mode in the audio-visual communication scenario, the same as in the traditional audio-visual communication scenario, high noise reduction processing can be performed on the audio, that is, suppressing the music content in the audio as noise; it is not difficult to see that the audio processing method provided by the embodiments of the present application can improve the audio-visual communication experience for the audio-visual communication scenario where music is played during a call or a meeting.

[0086] The following will introduce the audio processing system suitable for implementing the audio processing method provided by the embodiments of the present application in conjunction with Figure 2a and Figure 2b

[0087] In one implementation, as Figure 2a shown, the audio processing system may include a first terminal 201 and a second terminal 202. The first terminal 201 or the second terminal 202 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto; a direct connection may be established between the first terminal 201 and the second terminal 202 through a wired communication method, or an indirect connection may be established through a wireless communication method, which is not limited in the present application; moreover, the number of the first terminals 201 may be one or more, and the number of the second terminals 202 may be one or more, which is not limited in the embodiments of the present application. Among them:

[0088] The first terminal can be an audio generation end, and the second terminal can be an audio receiving end. Alternatively, the first terminal can be an audio receiving end, and the second terminal can be an audio generation end. Here, an example where the first terminal is an audio generation end and the second terminal is an audio receiving end is used for introduction: The first terminal (i.e., the audio generation end) can collect the audio generated by the first terminal in an audio-video communication scenario, then can determine the service mode of the audio by using the audio processing method provided in the embodiments of the present application, perform service processing on the audio according to the determined service mode. After that, the first terminal can perform encoding processing on the audio after service processing and send the encoded audio stream to the second terminal (i.e., the audio receiving end); The second terminal can decode and play the audio stream to implement audio-video communication from the first terminal to the second terminal.

[0089] In this implementation manner, the audio processing method provided in the embodiments of the present application can run in an embedded device (such as an embedded chip), and the embedded device is installed at the audio generation end; Alternatively, the audio processing method provided in the embodiments of the present application can be integrated into an audio-video communication client, an audio-video communication application program, or an audio-video communication software, etc., and the audio-video communication client, the audio-video communication application program, or the audio-video communication software, etc. runs in the audio generation end; Alternatively, the audio processing method provided in the embodiments of the present application can directly run in the audio generation end.

[0090] In another implementation manner, as Figure 2b shown, the audio processing system can include a first terminal 201, a second terminal 202, and a server 203. The server 203 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms; The first terminal 201 or the second terminal 202 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto; A direct connection can be established between the first terminal 201, the second terminal 202, and the server 203 through a wired communication method, or an indirect connection can be established through a wireless communication method. The present application does not make any restrictions here; Moreover, the number of the first terminals 201 can be one or more, and the number of the second terminals 202 can be one or more. The embodiments of the present application do not limit this. Among them:

[0091] The first terminal can be an audio generation end, and the second terminal can be an audio receiving end. Alternatively, the first terminal can be an audio receiving end, and the second terminal is an audio generation end. The server can be a data interaction server between the first terminal and the second terminal. Here, taking the first terminal as the audio generation end and the second terminal as the audio receiving end as an example, the first terminal (i.e., the audio generation end) can collect the audio generated by the first terminal in an audio-video communication scenario, and then send the audio to the server. The server can determine the service mode of the audio by using the audio processing method provided in the embodiments of the present application, perform service processing on the audio according to the determined service mode. After that, the server can perform encoding processing on the audio after service processing and send the encoded audio stream to the second terminal (i.e., the audio receiving end). The second terminal can decode and play the audio stream to implement audio-video communication from the first terminal to the second terminal.

[0092] In this implementation manner, the audio processing method provided in the embodiments of the present application can run in an embedded device (such as an embedded chip), and the embedded device is installed in the server; or, the audio processing method provided in the embodiments of the present application can directly run in the server.

[0093] In the embodiments of the present application, by introducing the audio processing method into the audio processing system, the audio processing system can efficiently and accurately perform music scene detection, improving the audio-video communication experience of the audio processing system. It can be understood that the audio processing system described in the embodiments of the present application is for more clearly explaining the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art know that with the evolution of the system architecture and the emergence of new service scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0094] The following introduces the audio processing method provided in the embodiments of the present application in more detail with reference to the accompanying drawings.

[0095] The embodiments of the present application provide an audio processing method, which mainly introduces content such as the feature encoding process, the classification learning process, the feature decoding process, and determining the frame type of the target audio frame. This audio processing method can be executed by a computer device, and the computer device can be the audio generation end in the above audio processing system (corresponding to the above Figure 2a ), or the server (corresponding to the above Figure 2b ). Please refer to Figure 3 , this audio processing method can include the following steps S301 to step S307:

[0096] S301, obtain multiple band signals of the target audio frame in the audio to be processed.

[0097] The audio to be processed can be real-time audio collected in an audio-video communication scenario (which can include an audio-video call scenario or an audio-video conference scenario) between an audio generation end and an audio reception end. The audio to be processed can include multiple audio frames, and the target audio frame is any one of the audio frames in the audio to be processed. The multiple band signals of the target audio frame can be obtained by performing band division processing on the frequency domain signal of the target audio frame, and the frequency domain signal of the target audio frame can be obtained by performing frequency domain conversion processing on the time domain signal of the target audio frame.

[0098] Before introducing steps S302 - S304, it should be noted here that the embodiments of this application can call a classification model (i.e., the classification and recognition neural network mentioned above) to classify and recognize multiple band signals to obtain a classification result. Among them, the classification model can include a feature encoding layer, a feature learning layer, and a feature decoding layer. The classification and recognition can include a feature encoding process, a classification learning process, and a feature decoding process. The feature encoding process can be executed by the feature encoding layer of the classification model, the classification learning process can be executed by the feature learning layer of the classification model, and the feature decoding process can be executed by the feature decoding layer. The following combines Figure 4 steps S302 - S304 to introduce the structure of the classification model and the classification and recognition process of multiple band signals:

[0099] S302, perform feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals.

[0100] The feature encoding layer (Encode-DNN) is a DNN (Deep Neural Network). The feature encoding layer can include M feature encoding units. The number of band signals is M. The M feature encoding units correspond one-to-one with the M band signals. The M band signals are input into the corresponding feature encoding units in the order of the magnitudes of the band information corresponding to each of the M band signals. The process of the feature encoding layer performing feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals can include: calling the m-th feature encoding unit among the M feature encoding units to perform feature encoding on the m-th band signal corresponding to the m-th feature encoding unit to obtain a feature vector corresponding to the m-th band signal, where M is an integer greater than 1 and m is a positive integer less than or equal to M.

[0101] As Figure 4In the structure of the classification model shown, a circle in the feature encoding layer represents a feature encoding unit. After any band signal is input into its corresponding feature encoding unit for feature encoding, the feature vector corresponding to the band signal can be obtained. Generally speaking, the purpose of the feature encoding layer for feature encoding is to extract the key information in the band signal and convert the key information in the extracted band signal into a feature vector that the feature learning layer can understand.

[0102] S303. Perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector.

[0103] After calling the feature encoding layer to perform feature encoding on each of the multiple band signals to obtain the feature vector corresponding to each of the multiple band signals, the feature learning layer can be called to perform feature learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector. Specifically, the feature learning layer is an RNN (Recurrent Neural Network), and specifically can be any one of the following RNNs: GRU (Gate Recurrent Unit) or LSTM (Long-Short Term Memory). The embodiments of the present application do not limit this; the feature learning layer can include N feature learning units, and the M feature vectors obtained by feature encoding by the feature encoding layer can be divided into N vector groups, each vector group can include at least two feature vectors, both N and M are integers greater than 1, and N is less than or equal to M. The process of calling the feature learning layer to perform feature learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector can include: the feature learning layer can call the first feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the first vector group among the N vector groups to obtain the feature learning result of the first feature learning unit; the feature learning layer can call the second feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the second vector group among the N vector groups and the feature learning result of the first feature learning unit to obtain the feature learning result of the second feature learning unit; the feature learning layer can continue to perform feature learning by calling the subsequent feature learning units to obtain the feature learning results of the subsequent feature learning units until the feature learning result of the last feature learning unit among the N feature learning units is obtained. Among them, the number of classification indication vectors is P, and the P classification indication vectors can be determined according to the feature learning results of the last P feature learning units among the N feature learning units. The feature learning result of one feature learning unit is used to determine one classification indication vector, and P is a positive integer less than or equal to N.

[0104] As Figure 4In the structure of the classification model shown, one block in the feature learning layer represents a feature learning unit; the first and second feature vectors among the M feature vectors are divided into the first vector group, the first, second, and third feature vectors among the M feature vectors are divided into the second vector group, and so on, and the M feature vectors are divided into N vector groups; the first vector group is input into the first feature learning unit, the second vector group is input into the second feature learning unit, and so on, and the Nth vector group is input into the Nth feature learning unit; the number of classification indication vectors is 3, and the 3 classification indication vectors are determined according to the feature learning results of the last 3 feature learning units among the N feature learning units. Generally speaking, the purpose of the feature learning layer for feature learning is to perform feature analysis and generalization on the feature vectors.

[0105] S304. Perform feature decoding on the classified indication vectors to obtain the classification result of the target audio frame.

[0106] After calling the feature learning layer to perform feature learning on the multiple feature vectors obtained by feature encoding to obtain the classification indication vectors, the feature decoding layer can be called to perform feature decoding on the classification indication vectors to obtain the classification result. Specifically, the feature decoding layer (Decode-DNN) is also a DNN (Deep Neural Network), and the feature decoding layer can include Y feature decoding units. Each feature decoding unit among the Y feature decoding units corresponds to a different frame type. As can be seen from the content of the foregoing step S303, the number of classification indication vectors obtained by calling the feature learning layer to perform feature learning on the multiple feature vectors obtained by feature encoding is P. The process of calling the feature decoding layer to perform feature decoding on the classification indication vectors to obtain the classification result can include: the feature decoding layer can call the yth feature decoding unit among the Y feature decoding units to perform dimensionality reduction processing on the P classification indication vectors to obtain the dimensionality reduction result of the yth feature decoding unit, and further, regression processing can be performed on the dimensionality reduction result of the yth feature decoding unit to obtain the probability that the target audio frame belongs to the frame type corresponding to the yth feature decoding unit. Y is an integer greater than 1, and y is a positive integer less than or equal to Y; wherein, the classification result can include: the probability that the target audio frame belongs to the frame type corresponding to each feature decoding unit among the Y feature decoding units.

[0107] As Figure 4In the structure of the classification model shown, a circle in the feature decoding layer represents a feature decoding unit, and the feature decoding layer includes 3 feature decoding units; the frame type corresponding to the first feature decoding unit is a music frame. After the 3 feature decoding units input into the first feature decoding unit for dimensionality reduction regression, the probability that the frame type of the target audio frame is a music frame is output; the frame type corresponding to the second feature decoding unit is a speech frame. After the 3 feature decoding units input into the second feature decoding unit for dimensionality reduction regression, the probability that the frame type of the target audio frame is a speech frame is output; the frame type corresponding to the third feature decoding unit is a noise frame. After the 3 feature decoding units input into the third feature decoding unit for dimensionality reduction regression, the probability that the frame type of the target audio frame is a noise frame is output.

[0108] Generally speaking, the feature decoding layer plays the role of Mask Predict, and the output probability can be called Mask (mask value). Through the dimensionality reduction process of the feature decoding layer, the classification indication vector can be interpreted as an understandable numerical pattern (i.e., the dimensionality reduction result). Through the regression process of the feature decoding layer, the dimensionality reduction result can be regressed to a mask value (i.e., probability) within the range of [0, 1]. The greater the probability that the target audio frame belongs to a certain frame type (i.e., the probability is closer to 1), the higher the possibility that the target audio frame belongs to that frame type; the smaller the probability that the target audio frame belongs to a certain frame type (i.e., the probability is closer to 0), the lower the possibility that the target audio frame belongs to that frame type; the regression process can be implemented through an activation function (such as the Sigmoid activation function).

[0109] It should be noted that the classification model can be obtained by training an initial model with multiple sample audios. The sample audios can include multiple sample audio frames, and each sample audio frame in the sample audio carries a marked frame type. By classifying and identifying the sample audio frames according to the initial model, the classification result of the sample audio frames can be obtained. According to the classification result of the sample audio frames, the predicted frame type of the sample audio frames can be determined. The purpose of training the initial model is to make the predicted frame type of the sample audio frames the same as the marked frame type of the sample audio frames.

[0110] S305, determine the frame type of the target audio frame according to the classification result.

[0111] As described above, the classification result may include the probabilities that the target audio frame belongs to each of the Y frame types. After obtaining the classification result, based on the probabilities that the target audio frame belongs to each of the Y frame types, a frame type can be determined as the frame type to which the target audio frame belongs. Among them, the frame type of the target audio frame may include, but is not limited to, any one of the following: music frame, speech frame, and noise frame. The embodiments of the present application do not limit this. When the frame type of the target audio frame is a music frame, it can indicate that the audio content included in the target audio frame is music content; when the frame type of the target audio frame is a speech frame, it can indicate that the audio content included in the target audio frame is speech content; when the frame type of the target audio frame is a noise frame, it can indicate that the audio content included in the target audio frame is noise content.

[0112] In one implementation, the frame type of the target audio frame may be the frame type corresponding to the maximum probability in the classification result. That is to say, the maximum probability in the classification result can be determined, and then the frame type corresponding to the maximum probability can be determined as the frame type of the target audio frame. For example, the classification result includes: the probability that the target audio frame belongs to a music frame is 0.7, the probability that the target audio frame belongs to a speech frame is 0.6, and the probability that the target audio frame belongs to a noise frame is 0.5. The maximum probability in the classification result is 0.7, that is, the frame type of the target audio frame is a music frame.

[0113] In another implementation, the frame type of the target audio frame may be determined according to the classification result and the audio scene type of the target audio. Specifically, the audio scene type of the audio to be processed can be determined. After determining the audio scene type of the audio to be processed, the weight of each of the Y frame types can be determined according to the audio scene type, and then, according to the weight of each of the Y frame types, the probabilities of the corresponding frame types among the Y frame types can be weighted and summed to obtain a weighted average probability; different frame types among the Y frame types correspond to different probability ranges, and the frame type of the target audio frame can be determined according to the probability range to which the weighted average probability belongs, that is, the frame type corresponding to the probability range to which the weighted average probability belongs is determined as the frame type of the target audio frame. For example, when the audio scene type of the audio to be processed is an audio-video communication scene, the weight of the speech frame can be set higher than the weights of the music frame and the noise frame. For example, the weight of the speech frame is set to 1, and the weights of the music frame and the speech frame are both set to 0.5. In this way, the audio scene type to which the target audio frame belongs is considered in the process of determining the frame type of the target audio frame, and weights are adapted for different frame types according to the determined audio scene type, which can improve the accuracy of the determined frame type of the target audio frame, thereby improving the accuracy of music scene detection.

[0114] In another implementation, the frame type of the target audio frame can be determined based on the classification result and the audio scene type of the target audio. Specifically, the audio scene type of the audio to be processed can be determined. After determining the audio scene type of the audio to be processed, the weight of each of the Y frame types can be determined according to the audio scene type. Then, according to the weight of each of the Y frame types, the probability of the corresponding frame type among the Y frame types can be weighted to obtain the weighted probability of each of the Y frame types. Further, the maximum weighted probability among the weighted probabilities of the Y frame types can be determined, and the frame type corresponding to the maximum weighted probability can be determined as the frame type of the target audio frame. For example, the classification result includes: the probability that the target audio frame belongs to a music frame is 0.7, the probability that the target audio frame belongs to a speech frame is 0.6, and the probability that the target audio frame belongs to a noise frame is 0.5; when the audio scene type of the audio to be processed is an audio-video communication scene, the weight of the speech frame can be set to be higher than the weights of the music frame and the noise frame. For example, the weight of the speech frame is set to 1, and the weights of the music frame and the speech frame are both set to 0.5; in this case, the weighted probability that the target audio frame belongs to a music frame is 0.35, the weighted probability that the target audio frame belongs to a speech frame is 0.6, the weighted probability that the target audio frame belongs to a noise frame is 0.25, and the maximum weighted probability is 0.6, that is, the frame type of the target audio frame is a speech frame. In this way, the audio scene type to which the target audio frame belongs is considered in the process of determining the frame type of the target audio frame, and weights are adapted for different frame types according to the determined audio scene type, so that the accuracy of the frame type of the determined target audio frame can be improved, thereby improving the accuracy of music scene detection.

[0115] S306. Filter the audio to be processed based on the frame type of the target audio frame to obtain the service mode of the audio to be processed.

[0116] After determining the frame type of the target audio frame, the audio to be processed can be filtered to obtain the service mode of the audio to be processed. Among them, the filtering process can include statistical filtering and decision filtering; in the statistical filtering process, the audio frames in the audio to be processed can be divided into multiple audio blocks, and each audio block can include at least two audio frames. The statistical filtering process can be used to determine the state of the audio block according to the frame types of the respective audio frames included in the audio block. The state of the audio block can include an active state or a non-active state; in the decision filtering process, an audio sliding window is set, and the audio sliding window can include at least two audio blocks. The decision filtering process can perform a sliding detection on the audio to be processed based on the audio sliding window. Each sliding detection is used to determine the service mode of the audio to be processed according to the states of the respective audio blocks currently included in the audio sliding window. The service mode can include a music mode or a non-music mode. As described above, the service mode of the audio to be processed being the music mode means that the audio to be processed contains a large amount of music content, and the service mode of the audio to be processed being the non-music mode means that the audio to be processed contains a small amount of music content or does not contain audio content.

[0117] S307, perform service processing on the audio to be processed according to the determined service mode.

[0118] After determining the service mode of the audio to be processed, service processing can be performed on the audio to be processed according to the determined service mode. The service processing methods for the audio to be processed in the music mode and the audio to be processed in the non-music mode are different.

[0119] In the embodiments of the present application, a classification model with a model structure of "feature encoding layer - feature learning layer - feature decoding layer" is used to classify and identify the multi-band signals of the target audio frame. The feature encoding layer and the feature decoding layer are deep learning networks (i.e., fully connected layers), and the feature learning layer is a recurrent neural network (i.e., a recursive neural network), which can improve the accuracy of classification and identification, thereby improving the accuracy of music scene detection to a certain extent. Moreover, the model structure is simple and does not adopt an end-to-end mode, reducing the model complexity, and can improve the classification and identification efficiency of the band signals, thereby improving the efficiency of music scene detection to a certain extent. In addition, in the process of determining the frame type of the target audio frame in the embodiments of the present application, the audio scene type to which the target audio frame belongs can be considered, and weights can be adapted for different frame types according to the determined audio scene type, so as to improve the accuracy of the determined frame type of the target audio frame, thereby improving the accuracy of music scene detection to a certain extent.

[0120] In the above Figure 3Based on the illustrated embodiments, the embodiments of the present application provide an audio processing method, which mainly introduces the time-domain signal separation process based on the sampling rate, the frequency-domain conversion process, and the frequency band division process, etc. This audio processing method can be executed by a computer device, and the computer device can be the audio generation end in the above audio processing system (corresponding to the above Figure 2a ) or the server (corresponding to the above Figure 2b ). Please refer to Figure 5 , this audio processing method can include the following steps S501 to S510:

[0121] S501, obtain the original time-domain signal of the target audio frame in the audio to be processed.

[0122] S502, obtain the target time-domain signal from the original time-domain signal according to the sampling information.

[0123] After obtaining the original time-domain signal of the target audio frame in the audio to be processed, the target time-domain signal can be obtained from the original time-domain signal according to the sampling information, and the sampling rate of the target time-domain signal meets the signal analysis conditions corresponding to the sampling information. Specifically, obtaining the target time-domain signal from the original time-domain signal according to the sampling information may include: obtaining the sampling rate when the original time-domain signal is sampled. If the sampling rate of the original time-domain signal meets the signal analysis conditions, the original time-domain signal can be determined as the target time-domain signal. If the sampling rate of the original time-domain signal does not meet the signal analysis conditions, frequency separation processing can be performed on the original time-domain signal based on the sampling information to obtain the target time-domain signal.

[0124] More specifically, the sampling information may refer to the target sampling rate. In this way, the sampling rate of the original time-domain signal meeting the signal analysis conditions means that the sampling rate of the original time-domain signal is less than or equal to the target sampling rate. If the sampling rate of the original time-domain signal is less than or equal to the target sampling rate, the original time-domain signal can be directly determined as the target time-domain signal; the sampling rate of the original time-domain signal not meeting the signal analysis conditions means that the sampling rate of the original time-domain signal is greater than the target sampling rate. If the sampling rate of the original time-domain signal is greater than the target sampling rate, frequency-domain separation processing can be performed on the original time-domain signal based on the target sampling rate to obtain the low-frequency time-domain signal and the high-frequency time-domain signal, and then the low-frequency time-domain signal can be determined as the target time-domain signal, where the sampling rate of the low-frequency time-domain signal may be less than or equal to the target sampling rate. That is to say, the sampling rate of the target time-domain signal meeting the signal analysis conditions corresponding to the sampling information means that the sampling rate of the target time-domain signal is less than or equal to the target sampling rate, and the sampling rate of the high-frequency time-domain signal is greater than the target sampling rate.

[0125] For example, the sampling rate of the original time-domain signal of the target audio frame is 32 kHz (kilohertz), the target sampling rate is 16 kHz, and the sampling rate of the original time-domain signal of the target audio frame is greater than the target sampling rate. Then, frequency separation processing can be performed on the original time-domain signal to obtain a low-frequency time-domain signal of 16 kHz (0 - 16 kHz) and a high-frequency time-domain signal of 16 kHz (16 - 32 kHz). Then, the low-frequency time-domain signal of 16 kHz can be determined as the target time-domain signal. For another example, the sampling rate of the original time-domain signal of the target audio frame is 16 kHz, the target sampling rate is 16 kHz, and the sampling rate of the original time-domain signal of the target audio frame is equal to the target sampling rate. Then, the original time-domain signal of 16 kHz can be directly determined as the target time-domain signal.

[0126] In this way, for the original time-domain signal with a sampling rate greater than the target sampling rate, through frequency separation processing, the low-frequency time-domain signal in the original time-domain signal can be obtained as the target time-domain signal for processing, without the need to process the entire original time-domain signal, which can save the computational consumption in the music scene detection process and thus improve the efficiency of music scene detection. For the original time-domain signal with a sampling rate less than or equal to the target sampling rate, the original time-domain signal itself is the low-frequency time-domain signal and can be directly processed without frequency domain separation processing, which can also save the computational consumption in the music scene detection process and thus improve the efficiency of music scene detection.

[0127] S503. Perform frequency domain conversion processing on the target time-domain signal to obtain the frequency domain signal of the target audio frame.

[0128] After obtaining the target time-domain signal that meets the signal analysis conditions corresponding to the sampling information from the original time-domain signal of the target audio frame according to the sampling information, frequency domain conversion processing can be performed on the target time-domain signal to obtain the frequency domain signal of the target audio frame. Specifically, frequency domain conversion processing can be performed on the target time-domain signal according to the frequency domain conversion algorithm to obtain the frequency domain signal of the target audio frame. Among them, the frequency domain conversion algorithm can include but is not limited to any one of the following: FT (Fourier Transform), FFT (Fast Fourier Transform), and STFT (Short-Time Fourier Transform or Short-Term Fourier Transform). The embodiments of the present application do not limit this.

[0129] S504. Perform frequency band division processing on the frequency domain signal to obtain multiple frequency band signals.

[0130] The band division process can be implemented using a band division algorithm. Among them, the band division algorithm can be, but is not limited to, any one of the following: BFCC (Bark-Frequency Cepstral Coefficients) algorithm or MFCC (Mel-Frequency Cepstral Coefficients) algorithm. The embodiments of the present application do not limit this. Among them, the process of the band division process can include the following: The frequency domain signal can include multiple frequency points (which can also be called frequency points), each frequency point corresponding to a different frequency. Multiple band ranges can be obtained, and the frequency points whose frequencies belong to the same frequency range form the band signal corresponding to that frequency range.

[0131] As Figure 6 shown in the band division result, Figure 6 One small rectangular block in it represents a band signal, and the frequency domain signal is divided into 18 band signals. In this way, the frequency points whose frequencies belong to the same frequency range (that is, the frequency information mentioned above) are unified as a band signal for processing, rather than processing each frequency point separately, which can reduce the calculation amount in the music scene detection process, thereby improving the efficiency of music scene detection.

[0132] In particular, the width of the band range can be related to the richness of the frequency domain characteristics of the frequency band where the band range is located. As Figure 6 shown in the band division result, the frequency domain characteristics contained in the sub-frequency domain signal corresponding to the frequency band of 0 - 2700 HZ are relatively rich. The number of band signals obtained by dividing the sub-frequency domain signal corresponding to this frequency band is relatively large, and the width of the frequency range corresponding to the band signal is relatively narrow. The frequency domain characteristics contained in the sub-frequency domain signal corresponding to the frequency band of 8500 - 16000 HZ are not very rich. The number of band signals obtained by dividing the sub-frequency domain signal corresponding to this frequency band is relatively small, and the width of the frequency range corresponding to the band signal is relatively wide. In this way, the classification and recognition granularity of the band signal with relatively rich frequency domain characteristics can be fine, and the classification and recognition granularity of the band signal with less rich frequency domain characteristics can be coarse, which can improve the efficiency of the music scene detection process.

[0133] S505, perform feature encoding on each of the multiple band signals to obtain the feature vector corresponding to each of the multiple band signals.

[0134] S506, perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector.

[0135] S507, perform feature decoding on the classified indication vector to obtain the classification result of the target audio frame.

[0136] S508. Determine the frame type of the target audio frame according to the classification result.

[0137] S509. Based on the frame type of the target audio frame, perform filtering processing on the audio to be processed to obtain the service mode of the audio to be processed.

[0138] S510. Perform service processing on the audio to be processed according to the determined service mode.

[0139] In the embodiments of the present application, sampling information is set. For the original time-domain signal of the target audio frame that does not meet the signal analysis conditions corresponding to the sampling information (i.e., the high-frequency original time-domain signal), the target time-domain signal (i.e., the low-frequency target time-domain signal) can be separated from the original time-domain signal according to the sampling information, and only the low-frequency time-domain signal needs to be processed, rather than the entire high-frequency time-domain signal. This can reduce the computational complexity in the music scene detection process, thereby improving the efficiency of music scene detection; for the original time-domain signal of the target audio frame that meets the signal analysis conditions corresponding to the sampling information (i.e., the low-frequency original time-domain signal), the original time-domain signal can be directly determined as the target time-domain signal without performing frequency-domain separation processing on the original time-domain signal. This can also reduce the computational complexity in the music scene detection process, thereby improving the efficiency of music scene detection.

[0140] Based on the above Figure 5 shown embodiments, the embodiments of the present application provide an audio processing method. This audio processing method mainly introduces the audio mode decision process (i.e., the process of determining the service mode of the audio to be processed), as well as the service processing process of the audio to be processed under different service modes, etc. This audio processing method can be executed by a computer device, and the computer device can be the audio generation end in the above audio processing system (corresponding to the above Figure 2a ), or a server (corresponding to the above Figure 2b ). Please refer to Figure 7 , this audio processing method may include the following steps S701 to step S711:

[0141] S701. Obtain the original time-domain signal of the target audio frame in the audio to be processed.

[0142] S702. Obtain the target time-domain signal from the original time-domain signal according to the sampling information.

[0143] S703. Perform frequency-domain conversion processing on the target time-domain signal to obtain the frequency-domain signal of the target audio frame.

[0144] S704. Perform frequency band division processing on the frequency-domain signal to obtain multiple frequency band signals.

[0145] S705 encodes the features of each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals.

[0146] S706 performs classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector.

[0147] S707 performs feature decoding on the classified indication vector to obtain the classification result of the target audio frame.

[0148] S708 determines the frame type of the target audio frame according to the classification result.

[0149] S709 performs statistical filtering on the parent audio block of the target audio frame based on the frame type of the target audio frame to determine the state of the parent audio block.

[0150] As described above, the audio to be processed may include multiple audio frames, and the multiple audio frames are arranged in sequence according to the time order of generation of the audio frames. The audio frames in the audio to be processed may be divided into multiple audio blocks, and each audio block may include at least two audio frames. Statistical filtering (which may also be referred to as short filtering) is performed for each audio block to determine the state of the audio block according to the types of the respective audio frames included in the audio block. The state of the audio block may be an active state or a non-active state.

[0151] Specifically, taking the parent audio block of the target audio frame (which refers to the audio block containing the target audio frame) as an example, the statistical filtering may include the following: The frame types of the respective audio frames in the parent audio block may be statistically counted. If the frame type statistical result indicates that the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block satisfies the audio frame number condition, the state of the parent audio block may be determined to be the active state. If the frame type statistical result indicates that the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block does not satisfy the audio frame number condition, the state of the parent audio block may be determined to be the non-active state. Among them, the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block satisfying the audio frame number condition means that the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block is greater than the audio frame number threshold. The number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block not satisfying the audio frame number condition means that the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block is less than or equal to the audio frame number threshold.

[0152] For example, the target frame type is a music frame. As Figure 8a shown, the number of audio frames of the music frame type in the audio block is 5, exceeding the audio frame number threshold of 4, then it can be determined that Figure 8aThe status of the audio block shown is the active status. As Figure 8b shown, the number of audio frames with the frame type of music frames in the audio block is 3, which does not exceed the audio frame number threshold of 4. Then it can be determined that Figure 8b the status of the audio block shown is the inactive status.

[0153] S710, if the audio sliding window slides to a position including the parent audio block, decision filtering processing is performed on the statuses of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed.

[0154] The decision filtering processing can also be referred to as long filtering processing. In the decision filtering processing, an audio sliding window (duration) is set, and the audio sliding window can allow at least two audio blocks to be accommodated; the audio sliding window can include two states, one state is the window loading state, and the other is the sliding detection state; since the audio to be processed is real-time audio, the window loading state means: the state where the currently generated audio frames are not enough to fill the entire audio sliding window; the sliding detection state means: starting from the first audio block of the audio to be processed, the audio sliding window slides in the audio to be processed with one audio block as the sliding step length for sliding detection. Among them, the sliding detection can be understood as any one of the following two situations:

[0155] (1) The first understanding of the sliding detection: It can be determined whether the audio blocks currently included in the audio sliding window meet the music mode condition according to the status of the audio blocks currently included in the audio sliding window; if the music mode condition is met, the sliding detection can be stopped, and it can be determined that the service mode of the audio to be processed is the music mode; if the music mode condition is not met, it can be determined that the service mode of the audio to be processed is the non-music mode, and the sliding detection continues until the sliding detection stop condition is triggered. The sliding detection stop condition can include: the audio blocks currently included in the audio sliding window meet the music mode condition, or the sliding detection reaches the last audio block of the audio to be processed.

[0156] Taking the case where the parent audio block is within the audio sliding window as an example, for the first understanding of the sliding detection: The decision filtering process may include the following: It is possible to count the states of the respective audio blocks currently included in the audio sliding window. If the state count result indicates that the number of target audio blocks among the respective audio blocks currently included in the audio sliding window meets the audio block quantity condition (i.e., the audio blocks currently included in the audio sliding window meet the music mode condition), then the sliding detection can be stopped, and the service mode of the audio to be processed can be determined as the music mode. If the state count result indicates that the number of target audio blocks among the respective audio blocks currently included in the audio sliding window does not meet the audio block quantity condition (i.e., the audio blocks currently included in the audio sliding window do not meet the music mode condition), then the service mode of the audio to be processed can be determined as the non-music mode, and the sliding detection can be continued until the sliding detection stop condition is triggered.

[0157] Among them, the target audio block refers to the audio block in the activated state among the respective audio blocks currently included in the audio sliding window; the number of target audio blocks among the respective audio blocks currently included in the audio sliding window meeting the audio block quantity condition means that: the number of target audio blocks among the respective audio blocks currently included in the audio sliding window is greater than the audio block quantity threshold. The number of target audio blocks among the respective audio blocks currently included in the audio sliding window not meeting the audio block quantity condition means that: the number of target audio blocks among the respective audio blocks currently included in the audio sliding window is less than or equal to the audio block quantity threshold. That is to say, if the number of audio blocks in the activated state among the respective audio blocks currently included in the audio sliding window is greater than the audio block quantity threshold, the service mode of the audio to be processed can be determined as the music mode. If the number of audio blocks in the activated state among the respective audio blocks currently included in the audio sliding window is less than or equal to the audio block quantity threshold, the service mode of the audio to be processed can be determined as the non-music mode.

[0158] As Figure 9a shown in the decision filtering process example, the audio sliding window is allowed to accommodate 15 audio blocks. When the audio sliding window slides to the first position, the number of audio blocks in the activated state among the audio blocks currently included in the audio sliding window ( Figure 9a the audio blocks in the activated state in are represented as rectangular blocks filled with lines, and the audio blocks in the non-activated state are represented as rectangular blocks filled with pure white) is 5, which is equal to the audio block quantity threshold of 5. It can be determined that the service mode of the audio to be processed is the non-music mode, and the sliding detection can be continued; when the audio sliding window slides to the second position, the number of audio blocks in the activated state among the audio blocks currently included in the audio sliding window is 6, which is greater than the audio block quantity threshold of 5. It can be determined that the service mode of the audio to be processed is the music mode, and the sliding detection can be stopped.

[0159] (2) The second understanding of sliding detection: It is possible to determine whether the audio blocks currently included in the audio sliding window meet the music mode condition based on the status of the audio blocks currently included in the audio sliding window; if the music mode condition is met, sliding detection can continue, and it can be determined that the service mode of the audio to be processed is the music mode until the sliding detection stop condition is triggered; if the music mode condition is not met, it can be determined that the service mode of the audio to be processed is the non-music mode, and sliding detection continues until the sliding detection stop condition is triggered; among them, the sliding detection stop condition may include: sliding detection reaches the last audio block of the audio to be processed.

[0160] It should be noted that the second understanding of sliding detection is exactly the same as the first understanding of sliding detection in determining whether the audio blocks currently included in the audio sliding window meet the music mode condition based on the status of the audio blocks currently included in the audio sliding window. The difference between the second understanding of sliding detection and the first understanding of sliding detection lies in: in the first understanding of sliding detection, when it is determined that the service mode of the audio to be processed is the non-music mode, sliding detection can continue, and sliding detection stops after it is determined that the service mode of the audio to be processed is the music mode; while in the second understanding of sliding detection, whether it is determined that the service mode of the audio to be processed is the non-music mode or the service mode of the audio to be processed is the non-music mode, sliding detection can continue.

[0161] As Figure 9b shown in the decision filtering processing example, the audio sliding window allows 15 audio blocks to be accommodated. When the audio sliding window slides to the first position, the number of active audio blocks ( Figure 9b In the figure, the active audio blocks are represented by rectangular blocks filled with lines, and the inactive audio blocks are represented by rectangular blocks filled with pure white) among the audio blocks currently included in the audio sliding window is 5, which is equal to the audio block number threshold of 5. It can be determined that the service mode of the audio to be processed is the non-music mode, and sliding detection can continue; when the audio sliding window slides to the second position, the number of active audio blocks among the audio blocks currently included in the audio sliding window is 6, which is greater than the audio block number threshold of 5. It can be determined that the service mode of the audio to be processed is the music mode, and sliding detection can continue; when the audio sliding window slides to the third position, the audio sliding window is in the window loading state.

[0162] It is particularly noteworthy that the audio to be processed mentioned in the embodiments of the present application is real-time audio. Therefore, the audio to be processed mentioned in the embodiments of the present application can be the entire audio generated in an audio-video communication scenario, or can be a part of the entire audio generated in an audio-video communication scenario. Further, the service mode of the audio to be processed mentioned in the embodiments of the present application being the music mode (or non-music mode) can be understood as: the service mode of the entire audio generated in the audio-video communication scenario is the music mode (or non-music mode); or, the service mode of a part of the entire audio generated in the audio-video communication scenario is the music mode. For example Figure 9a or Figure 9b in, the audio to be processed can be the audio composed of the audio frames included in the audio sliding window slid to the second position, and the service mode of this audio to be processed is the music mode.

[0163] S711, perform service processing on the audio to be processed according to the determined service mode.

[0164] After determining the service mode of the audio to be processed, service processing can be performed on the audio to be processed according to the determined service mode. When the service mode is the music mode, performing service processing on the audio to be processed according to the determined service mode can include: performing noise reduction processing on the audio to be processed according to the low noise reduction processing rule corresponding to the music mode; or, outputting a music mode selection interface, and in response to a music mode selection operation input in the music mode selection interface, performing noise reduction processing on the audio to be processed according to the low noise reduction processing rule corresponding to the music mode. When the service mode is the non-music mode, performing service processing on the audio to be processed according to the determined service mode can include: performing noise reduction processing on the audio to be processed according to the high noise reduction processing rule corresponding to the non-music mode; or, outputting a non-music mode selection interface, and in response to a non-music mode selection operation input in the non-music mode selection interface, performing noise reduction processing on the audio to be processed according to the high noise reduction processing rule corresponding to the non-music mode.

[0165] Among them, the noise reduction degree indicated by the low noise reduction processing rule is lower than the noise reduction degree indicated by the high noise reduction processing rule. Performing noise reduction processing on the audio to be processed according to the low noise reduction processing rule corresponding to the music mode (i.e., low noise reduction processing) means: weakening the noise reduction degree so that the music played in the audio-video communication scenario is retained; or, removing other noise contents in the audio except for the voice content and music content so that the music played in the audio-video communication scenario is retained (which can also be called steady-state noise reduction). Performing noise reduction processing on the audio to be processed according to the high noise reduction processing rule corresponding to the non-music mode (i.e., high noise reduction processing) means: suppressing the music content in the audio as noise.

[0166] Figure 10aAn exemplary music mode selection interface 101 is shown. The music mode selection interface 101 may include music mode prompt information 1011, such as Figure 10a "Music detected. You can turn on the music mode for a great experience" as shown, and may also include a music mode activation option 1012, such as Figure 10a "Activate" as shown. The music mode selection operation input in the music mode selection interface refers to the operation of triggering the music mode activation option 1012. Figure 10b An exemplary non-music mode selection interface 102 is shown. The non-music mode selection interface 102 may include non-music mode prompt information 1021, such as Figure 10b "Music not detected. You can turn on the non-music mode for a great experience" as shown, and may also include a non-music mode activation option 1022, such as Figure 10b "Activate" as shown. The non-music mode selection operation input in the non-music mode selection interface refers to the operation of triggering the non-music mode activation option 1022. Both the music mode selection interface 101 and the non-music mode selection interface 102 may be window interfaces in the audio-video communication interface 10. It is easy to think that for a solution that directly performs service processing without outputting a service mode selection interface (i.e., the above-mentioned music mode selection interface or non-music mode selection interface), the service processing in the music mode and the service processing in the non-music mode are automatically performed. In this way, the service mode for audio service processing can be automatically adapted and switched in the audio-video communication scenario, improving the audio effect in the audio-video communication scenario; for a solution that outputs a service mode selection interface, not only can the audio effect in the audio-video communication scenario be improved, but also the participants in the audio-video communication scenario can control whether to switch the service mode through the service mode selection interface, which can also improve the interaction experience in the audio-video communication scenario. Moreover, the service mode selection interface can be automatically closed after being displayed for a period of time or manually closed. The embodiments of the present application do not limit this.

[0167] It should also be noted that in the actual business processing process, for the first understanding of the above sliding detection, the sliding detection stops after determining that the business mode of the audio to be processed is the music mode. When the music playback in the audio-video communication scenario stops, the participating objects in the audio-video communication scenario can actively turn off the music mode and enter the non-music mode, so that the audio to be processed can be denoised according to the high denoising processing rules corresponding to the non-music mode. For the second understanding of the above sliding detection, the sliding detection can continue after determining that the business mode of the audio to be processed is the music mode. When the music playback in the audio-video communication scenario stops, it can be determined to enter the non-music mode. When music is played again in the audio-video communication scenario, it can be determined to enter the music mode again. In this way, it can automatically and flexibly switch between the music mode and the non-music mode according to the actual music playback situation in the audio-video communication scenario.

[0168] In the embodiments of the present application, the main function of the statistical filtering process is to count the probability of music frames appearing in an audio block. If the probability is high (that is, the number of music frames appearing is greater than the audio frame number threshold), the audio block is activated. If the probability is low (that is, the number of music frames appearing is less than or equal to the audio frame number threshold), the audio block is not activated. The main function of the decision filtering process is to count the probability of the audio blocks in the activated state within an audio sliding window forward from the current moment. If the probability is high (that is, the number of audio blocks in the activated state is greater than the audio block number threshold), a prompt is reported or directly enter the music mode. If the probability is low (that is, the number of audio blocks in the activated state is less than or equal to the audio block number threshold), a prompt is reported or directly enter the non-music mode. Through the mutual cooperation among the short filtering algorithm strategy, the long filtering algorithm strategy, and the sliding window strategy, the decision accuracy of the business mode of the audio to be processed can be improved, thereby improving the accuracy of music scene detection.

[0169] In summary Figures 3 - 10b Regarding the relevant content of the embodiments shown Figure 11 The audio processing method provided by the embodiments of the present application can be summarized as the

[0170] (1) Collect the audio to be processed in the audio-video communication scenario. For any audio frame (i.e., the target audio frame) in the audio to be processed, if the sampling rate of the original time-domain signal of the target audio frame is greater than the target sampling rate, the frequency-domain separation process can be performed on the original time-domain signal of the target audio frame, and the separated low-frequency time-domain signal with the sampling rate of the target sampling rate is used as the target time-domain signal and input into the classification and recognition module. If the sampling rate of the original time-domain signal of the target audio frame is less than or equal to the target sampling rate, the original time-domain signal itself is a low-frequency time-domain signal, and the original time-domain signal can be directly used as the target time-domain signal and input into the classification and recognition module.

[0171] (2) In the classification and recognition module, the target time-domain signal can be subjected to frequency-domain conversion processing to obtain the frequency-domain signal of the target audio frame. The frequency-domain signal is subjected to frequency-band separation processing to obtain multiple frequency-band signals. Then, a classification model (including a feature encoding layer (Encode-DNN), a feature learning layer (RNN), and a feature decoding layer (Decode-DNN)) can be called to classify and recognize the multiple frequency-band signals to obtain a classification result. After that, the classification result can be input into the threshold decision sub-module (Threshold module) to determine the frame type of the target audio frame. After determining the frame type of the audio frames in the audio to be processed, the frame type of the audio frames in the audio to be processed is input into the music mode decision module.

[0172] (3) In the music mode decision module, the audio frames in the audio to be processed are divided into multiple audio blocks. The main function of the statistical filtering process (i.e., short filtering process) in the music mode decision module is to statistically calculate the probability of music frames appearing within an audio block. If the probability is high (i.e., the number of music frames appearing is greater than the audio frame number threshold), then the audio block is activated; if the probability is low (i.e., the number of music frames appearing is less than or equal to the audio frame number threshold), then the audio block is not activated. The main function of the decision filtering process (i.e., long filtering process) in the music mode decision module is to statistically calculate the probability of the audio blocks in an audio sliding window that are in the activated state from the current moment forward. If the probability is high (i.e., the number of activated audio blocks is greater than the audio block number threshold), then a prompt is reported or the service mode of the audio to be processed is directly determined as the music mode; if the probability is low (i.e., the number of activated audio blocks is less than or equal to the audio block number threshold), then a prompt is reported or the service mode of the audio to be processed is directly determined as the non-music mode. After determining the service mode of the audio to be processed, the service mode of the audio to be processed belongs to the service processing module.

[0173] (4) In the service processing module, low noise reduction processing can be performed on the audio to be processed in the music mode, and high noise reduction processing can be performed on the audio to be processed in the non-music mode.

[0174] (5) After performing service processing on the audio to be processed, the processed audio can be encoded, and the encoded audio stream is sent to the audio receiving end. The embodiments of the present application can perform music scene detection efficiently and accurately in an audio-video communication scenario.

[0175] The above has elaborated in detail the system and method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.

[0176] Please refer to Figure 12 , Figure 12It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application. The audio processing device can be arranged in the computer device provided by the embodiment of the present application, and the computer device can be the audio generation end or the server mentioned above. Figure 12 The audio processing device shown can be a computer program (including program code) running in a computer device, and this audio processing device can be used to execute Figure 3 , Figure 5 or Figure 7 part or all of the steps in the method embodiments shown. Please refer to Figure 12 , and this audio processing device can include the following units:

[0177] An acquisition unit 1201, configured to acquire multiple band signals of a target audio frame in the audio to be processed;

[0178] A processing unit 1202, configured to perform feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals;

[0179] The processing unit 1202 is further configured to perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector;

[0180] The processing unit 1202 is further configured to perform feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determine the frame type of the target audio frame according to the classification result;

[0181] The processing unit 1202 is further configured to perform filtering processing on the audio to be processed based on the frame type of the target audio frame to obtain a service mode of the audio to be processed, and perform service processing on the audio to be processed according to the determined service mode; the filtering processing includes statistical filtering processing and decision filtering processing.

[0182] In one implementation manner, when the acquisition unit 1201 is configured to acquire multiple band signals of a target audio frame in the audio to be processed, it is specifically configured to perform the following steps:

[0183] Acquire the original time-domain signal of the target audio frame in the audio to be processed, and acquire the target time-domain signal from the original time-domain signal according to the sampling information, and the sampling rate of the target time-domain signal satisfies the signal analysis conditions corresponding to the sampling information;

[0184] Perform frequency-domain conversion processing on the target time-domain signal to obtain the frequency-domain signal of the target audio frame;

[0185] Perform frequency band division processing on the frequency-domain signal to obtain multiple band signals.

[0186] In one implementation, when the obtaining unit 1201 is used to obtain a target time-domain signal from an original time-domain signal according to sampling information, it is specifically used to perform the following steps:

[0187] Obtain the sampling rate when the original time-domain signal is sampled;

[0188] If the sampling rate of the original time-domain signal meets the signal analysis condition, determine the original time-domain signal as the target time-domain signal;

[0189] If the sampling rate of the original time-domain signal does not meet the signal analysis condition, perform frequency separation processing on the original time-domain signal based on the sampling information to obtain the target time-domain signal.

[0190] In one implementation, the number of band signals is M, where M is an integer greater than 1; the feature encoding process is performed by the feature encoding layer of the classification model. The feature encoding layer includes M feature encoding units, and the M feature encoding units correspond one-to-one to the M band signals. The M band signals are sequentially input into the corresponding feature encoding units in the order of the magnitudes of the band information corresponding to the M band signals. When the processing unit 1202 is used to perform feature encoding on each of the multiple band signals to obtain the feature vector corresponding to each of the multiple band signals, it is specifically used to perform the following steps:

[0191] Call the m-th feature encoding unit among the M feature encoding units to perform feature encoding on the m-th band signal corresponding to the m-th feature encoding unit to obtain the feature vector corresponding to the m-th band signal, where m is a positive integer less than or equal to M.

[0192] In one implementation, the feature learning process is performed by the feature learning layer of the classification model. The feature learning layer includes N feature learning units; the number of feature vectors is M, and the M feature vectors are divided into N vector groups, and each vector group includes at least two feature vectors. Both N and M are integers greater than 1, and N is less than or equal to M. When the processing unit 1202 is used to perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector, it is specifically used to perform the following steps:

[0193] The feature learning layer calls the first feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the first vector group among the N vector groups to obtain the feature learning result of the first feature learning unit;

[0194] The feature learning layer calls the second feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the second vector group among the N vector groups and the feature learning result of the first feature learning unit to obtain the feature learning result of the second feature learning unit;

[0195] Among them, the number of classification indication vectors is P, and the P classification indication vectors are determined according to the feature learning results of the last P feature learning units among the N feature learning units, where P is a positive integer less than or equal to N.

[0196] In one implementation, the number of classification indication vectors is P, where P is a positive integer; the feature decoding process is executed by the feature decoding layer of the classification model. The feature decoding layer includes Y feature decoding units, and different feature decoding units among the Y feature decoding units correspond to different frame types, where Y is an integer greater than 1; when the processing unit 1202 is used to perform feature decoding on the classification indication vectors to obtain the classification result of the target audio frame, it is specifically used to perform the following steps:

[0197] Call the y-th feature decoding unit among the Y feature decoding units to perform dimensionality reduction processing on the P classification indication vectors to obtain the dimensionality reduction result of the y-th feature decoding unit; and perform regression processing on the dimensionality reduction result of the y-th feature decoding unit to obtain the probability that the target audio frame belongs to the frame type corresponding to the y-th feature decoding unit, where y is a positive integer less than or equal to Y.

[0198] Among them, the classification result includes: the probability that the target audio frame belongs to the frame type corresponding to each of the Y feature decoding units.

[0199] In one implementation, the audio frames in the audio to be processed are divided into multiple audio blocks, and each audio block includes at least two audio frames; the audio sliding window slides in the audio to be processed with one audio block as the sliding step length, and the audio sliding window allows at least two audio blocks to be accommodated; when the processing unit 1202 is used to perform filtering processing on the audio to be processed based on the frame type of the target audio frame to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0200] Based on the frame type of the target audio frame, perform statistical filtering processing on the parent audio block of the target audio frame to determine the state of the parent audio block, where the parent audio block refers to the audio block containing the target audio frame;

[0201] If the audio sliding window slides to the position containing the parent audio block, perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed.

[0202] In one implementation, when the processing unit 1202 is used to perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0203] Statistically analyze the states of each audio block currently included in the audio sliding window, where the parent audio block is within the audio sliding window;

[0204] If the result of the state statistics indicates that the number of target audio blocks among each audio block currently included in the audio sliding window meets the audio block quantity condition, determine that the service mode of the audio to be processed is the music mode, where the target audio block refers to the audio block in the activated state among each audio block currently included in the audio sliding window;

[0205] If the result of the state statistics indicates that the number of target audio blocks among each audio block currently included in the audio sliding window does not meet the audio block quantity condition, determine that the service mode of the audio to be processed is the non - music mode.

[0206] In one implementation, the processing unit 1202, when used to perform statistical filtering processing on the parent audio block of the target audio frame based on the frame type of the target audio frame and determine the state of the parent audio block, specifically is used to perform the following steps:

[0207] Statistically analyze the frame types of each audio frame in the parent audio block;

[0208] If the result of the frame type statistics indicates that the number of audio frames belonging to the target frame type among each audio frame included in the parent audio block meets the audio frame quantity condition, determine that the state of the parent audio block is the activated state.

[0209] In one implementation, when the service mode is the music mode, the processing unit 1202, when used to perform service processing on the audio to be processed according to the determined service mode, specifically is used to perform the following steps:

[0210] Perform noise reduction processing on the audio to be processed according to the low - noise reduction processing rule corresponding to the music mode; or, output a music mode selection interface, and in response to the music mode selection operation input in the music mode selection interface, perform noise reduction processing on the audio to be processed according to the low - noise reduction processing rule corresponding to the music mode;

[0211] When the service mode is the non - music mode, the processing unit 1202, when used to perform service processing on the audio to be processed according to the determined service mode, specifically is used to perform the following steps: Perform noise reduction processing on the audio to be processed according to the high - noise reduction processing rule corresponding to the non - music mode.

[0212] In one implementation, the classification result includes the probability that the target audio frame belongs to each of the Y frame types, where Y is an integer greater than 1; different frame types among the Y frame types correspond to different probability ranges; the processing unit 1202, when used to determine the frame type of the target audio frame according to the classification result, specifically is used to perform the following steps:

[0213] Determine the audio scene type of the audio to be processed;

[0214] According to the audio scene type, determine the weight of each of the Y frame types;

[0215] According to the weights of each of the Y frame types, perform a weighted sum processing on the probabilities of the corresponding frame types among the Y frame types to obtain a weighted average probability;

[0216] According to the probability range to which the weighted average probability belongs, determine the frame type of the target audio frame.

[0217] According to another embodiment of the present application, Figure 12 Each unit in the audio processing device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the audio processing device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0218] According to another embodiment of the present application, it can be achieved by running, on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), a computer program (including program code) capable of executing the steps involved in the corresponding method embodiments shown in Figure 3 , Figure 5 or Figure 7 to construct the audio processing device shown in Figure 12 and to implement the audio processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.

[0219] In the embodiments of the present application, after obtaining multiple band signals of a target audio frame in the audio to be processed, by performing feature encoding on the multiple band signals, performing feature learning on the feature vectors obtained by the feature encoding, and performing feature decoding on the classification indication vectors obtained by the classification learning, the accuracy of the classification results obtained by the feature decoding can be improved, the accuracy of the frame type of the target audio frame determined according to the classification results can be improved, and thus the accuracy of music scene detection can be improved. Moreover, based on the frame type of the target audio frame, performing statistical filtering processing and decision filtering processing on the audio to be processed can further improve the accuracy of music scene detection.

[0220] Based on the above system, method, and apparatus embodiments, the embodiments of the present application provide a computer device, which may be the aforementioned audio generation end or server. Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. Figure 13 The computer device shown at least includes a processor 1301, an input interface 1302, an output interface 1303, and a computer-readable storage medium 1304. Among them, the processor 1301, the input interface 1302, the output interface 1303, and the computer-readable storage medium 1304 can be connected through a bus or other means.

[0221] The input interface 1302 can be used to obtain the audio to be processed and obtain the target audio frame in the audio to be processed. The output interface 1303 can be used to send the encoded audio stream to the audio receiving end.

[0222] The computer-readable storage medium 1304 can be stored in the memory of the computer device. The computer-readable storage medium 1304 is used to store a computer program, and the computer program includes computer instructions. The processor 1301 is used to execute the program instructions stored in the computer-readable storage medium 1304. The processor 1301 (or CPU (Central Processing Unit, central processing unit)) is the computing core and control core of the computer device, and is adapted to implement one or more computer instructions, and is specifically adapted to load and execute one or more computer instructions to implement the corresponding method flow or corresponding function.

[0223] The embodiments of the present application also provide a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device, and of course can also include the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. And, one or more computer instructions suitable for being loaded and executed by the processor are also stored in this storage space. These computer instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0224] One or more computer instructions stored in the computer-readable storage medium 1304 can be loaded and executed by the processor 1301 to implement the corresponding steps of the audio processing method related to Figure 3 , Figure 5 or Figure 7 shown above. In a specific implementation, the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform the following steps:

[0225] Obtain multiple band signals of a target audio frame in the audio to be processed;

[0226] Perform feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals;

[0227] Perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector;

[0228] Perform feature decoding on the classification indication vector to obtain the classification result of the target audio frame, and determine the frame type of the target audio frame according to the classification result;

[0229] Based on the frame type of the target audio frame, perform filtering processing on the audio to be processed to obtain the service mode of the audio to be processed, and perform service processing on the audio to be processed according to the determined service mode; the filtering processing includes statistical filtering processing and decision filtering processing.

[0230] In one implementation, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to obtain multiple band signals of a target audio frame in the audio to be processed, it is specifically used to perform the following steps:

[0231] Obtain the original time-domain signal of the target audio frame in the audio to be processed, and obtain the target time-domain signal from the original time-domain signal according to the sampling information, where the sampling rate of the target time-domain signal meets the signal analysis conditions corresponding to the sampling information;

[0232] Perform frequency-domain conversion processing on the target time-domain signal to obtain the frequency-domain signal of the target audio frame;

[0233] Perform frequency band division processing on the frequency-domain signal to obtain multiple frequency band signals.

[0234] In one implementation, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to obtain the target time-domain signal from the original time-domain signal according to the sampling information, it is specifically used to perform the following steps:

[0235] Obtain the sampling rate when sampling the original time-domain signal;

[0236] If the sampling rate of the original time-domain signal meets the signal analysis conditions, determine the original time-domain signal as the target time-domain signal;

[0237] If the sampling rate of the original time-domain signal does not meet the signal analysis conditions, perform frequency separation processing on the original time-domain signal based on the sampling information to obtain the target time-domain signal.

[0238] In one implementation, the number of frequency band signals is M, where M is an integer greater than 1; the feature encoding process is executed by the feature encoding layer of the classification model. The feature encoding layer includes M feature encoding units, and the M feature encoding units correspond one-to-one with the M frequency band signals. The M frequency band signals are input into the corresponding feature encoding units in the order of the magnitudes of the frequency band information corresponding to the M frequency band signals respectively; when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform feature encoding on each of the multiple frequency band signals to obtain the feature vector corresponding to each of the multiple frequency band signals, it is specifically used to perform the following steps:

[0239] Call the m-th feature encoding unit among the M feature encoding units to perform feature encoding on the m-th frequency band signal corresponding to the m-th feature encoding unit to obtain the feature vector corresponding to the m-th frequency band signal, where m is a positive integer less than or equal to M.

[0240] In one implementation, the feature learning process is performed by the feature learning layer of the classification model. The feature learning layer includes N feature learning units; the number of feature vectors is M. The M feature vectors are divided into N vector groups, and each vector group includes at least two feature vectors. Both N and M are integers greater than 1, and N is less than or equal to M. When the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform classification learning on the multiple feature vectors obtained by feature encoding to obtain the classification indication vectors, the following steps are specifically used:

[0241] Call the first feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the first vector group among the N vector groups to obtain the feature learning result of the first feature learning unit;

[0242] The feature learning layer calls the second feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the second vector group among the N vector groups and the feature learning result of the first feature learning unit to obtain the feature learning result of the second feature learning unit;

[0243] Among them, the number of classification indication vectors is P. The P classification indication vectors are determined according to the feature learning results of the last P feature learning units among the N feature learning units. P is a positive integer less than or equal to N.

[0244] In one implementation, the number of classification indication vectors is P, and P is a positive integer; the feature decoding process is performed by the feature decoding layer of the classification model. The feature decoding layer includes Y feature decoding units. Different feature decoding units among the Y feature decoding units correspond to different frame types. Y is an integer greater than 1. When the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform feature decoding on the classification indication vectors to obtain the classification result of the target audio frame, the following steps are specifically used:

[0245] Call the y-th feature decoding unit among the Y feature decoding units to perform dimensionality reduction processing on the P classification indication vectors to obtain the dimensionality reduction result of the y-th feature decoding unit; and perform regression processing on the dimensionality reduction result of the y-th feature decoding unit to obtain the probability that the target audio frame belongs to the frame type corresponding to the y-th feature decoding unit, where y is a positive integer less than or equal to Y;

[0246] Among them, the classification result includes: the probability that the target audio frame belongs to the frame type corresponding to each feature decoding unit among the Y feature decoding units.

[0247] In one implementation, the audio frames in the audio to be processed are divided into multiple audio blocks, and each audio block includes at least two audio frames; an audio sliding window slides in the audio to be processed with one audio block as the sliding step, and the audio sliding window allows at least two audio blocks to be accommodated; when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform filtering processing on the audio to be processed based on the frame type of the target audio frame to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0248] Based on the frame type of the target audio frame, perform statistical filtering processing on the parent audio block of the target audio frame to determine the state of the parent audio block, where the parent audio block refers to the audio block containing the target audio frame;

[0249] If the audio sliding window slides to a position containing the parent audio block, perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed.

[0250] In one implementation, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform decision filtering processing on the states of the respective audio blocks included in the audio sliding window to obtain the service mode of the audio to be processed, it is specifically used to perform the following steps:

[0251] Statistically count the states of the respective audio blocks currently included in the audio sliding window, and the parent audio block is within the audio sliding window;

[0252] If the state statistical result indicates that the number of target audio blocks in the respective audio blocks currently included in the audio sliding window meets the audio block number condition, determine that the service mode of the audio to be processed is the music mode, where the target audio block refers to the audio block in the respective audio blocks currently included in the audio sliding window that is in the active state;

[0253] If the state statistical result indicates that the number of target audio blocks in the respective audio blocks currently included in the audio sliding window does not meet the audio block number condition, determine that the service mode of the audio to be processed is the non-music mode.

[0254] In one implementation, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform statistical filtering processing on the parent audio block of the target audio frame based on the frame type of the target audio frame to determine the state of the parent audio block, it is specifically used to perform the following steps:

[0255] Statistically count the frame types of the respective audio frames in the parent audio block;

[0256] If the statistical result of the frame type indicates that the number of audio frames belonging to the target frame type among the respective audio frames included in the parent audio block satisfies the audio frame number condition, determine the state of the parent audio block as the active state.

[0257] In one implementation, when the service mode is the music mode, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform service processing on the to-be-processed audio according to the determined service mode, it is specifically used to perform the following steps:

[0258] Perform noise reduction processing on the to-be-processed audio according to the low noise reduction processing rule corresponding to the music mode; or, output a music mode selection interface, and in response to the music mode selection operation input in the music mode selection interface, perform noise reduction processing on the to-be-processed audio according to the low noise reduction processing rule corresponding to the music mode;

[0259] When the service mode is a non-music mode, when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to perform service processing on the to-be-processed audio according to the determined service mode, it is specifically used to perform the following steps: Perform noise reduction processing on the to-be-processed audio according to the high noise reduction processing rule corresponding to the non-music mode.

[0260] In one implementation, the classification result includes the probability that the target audio frame belongs to each of the Y frame types, where Y is an integer greater than 1; different frame types among the Y frame types correspond to different probability ranges; when the computer instructions in the computer-readable storage medium 1304 are loaded and executed by the processor 1301 to determine the frame type of the target audio frame according to the classification result, it is specifically used to perform the following steps:

[0261] Determine the audio scene type of the to-be-processed audio;

[0262] According to the audio scene type, determine the weight of each of the Y frame types;

[0263] According to the weight of each of the Y frame types, perform weighted summation processing on the probabilities of the corresponding frame types among the Y frame types to obtain a weighted average probability;

[0264] According to the probability range to which the weighted average probability belongs, determine the frame type of the target audio frame.

[0265] In the embodiments of the present application, after obtaining multiple band signals of a target audio frame in the audio to be processed, by performing feature encoding on the multiple band signals, performing feature learning on the feature vectors obtained by the feature encoding, and performing feature decoding on the classification indication vectors obtained by the classification learning, the accuracy of the classification result obtained by the feature decoding can be improved, and the accuracy of the frame type of the target audio frame determined according to the classification result can be improved, so that the accuracy of music scene detection can be improved. Moreover, based on the frame type of the target audio frame, performing statistical filtering processing and decision filtering processing on the audio to be processed can further improve the accuracy of music scene detection.

[0266] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio processing method provided in the above various optional manners.

[0267] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An audio processing method, characterized in that, the method includes: Obtaining multiple band signals of a target audio frame in the audio to be processed; the audio frames in the audio to be processed are divided into multiple audio blocks, and each audio block includes at least two audio frames; an audio sliding window slides in the audio to be processed with one audio block as the sliding step, and the audio sliding window allows accommodation of at least two audio blocks; Performing feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals; Performing classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector; Performing feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determining the frame type of the target audio frame according to the classification result; Based on the frame type of the target audio frame, performing statistical filtering processing on the parent audio block of the target audio frame to determine the state of the parent audio block, where the parent audio block refers to the audio block containing the target audio frame; If the audio sliding window slides to a position containing the parent audio block, then performing statistics on the states of the respective audio blocks currently contained in the audio sliding window; If the state statistics result indicates that the number of target audio blocks in the respective audio blocks currently contained in the audio sliding window satisfies the audio block number condition, then determining that the service mode of the audio to be processed is the music mode, where the target audio block refers to the audio block in the respective audio blocks currently contained in the audio sliding window that is in the active state; If the state statistics result indicates that the number of target audio blocks in the respective audio blocks currently contained in the audio sliding window does not satisfy the audio block number condition, then determining that the service mode of the audio to be processed is the non-music mode; Performing service processing on the audio to be processed according to the determined service mode.

2. The method according to claim 1, characterized in that, the obtaining of the multiple band signals of the target audio frame in the audio to be processed includes: Obtaining the original time-domain signal of the target audio frame in the audio to be processed, and obtaining a target time-domain signal from the original time-domain signal according to the sampling information, where the sampling rate of the target time-domain signal satisfies the signal analysis condition corresponding to the sampling information; Performing frequency-domain conversion processing on the target time-domain signal to obtain the frequency-domain signal of the target audio frame; Performing frequency band division processing on the frequency-domain signal to obtain multiple band signals.

3. The method according to claim 2, characterized in that, the obtaining of the target time-domain signal from the original time-domain signal according to the sampling information includes: Obtaining the sampling rate when the original time-domain signal is sampled; If the sampling rate of the original time-domain signal satisfies the signal analysis condition, then determining the original time-domain signal as the target time-domain signal; If the sampling rate of the original time-domain signal does not satisfy the signal analysis condition, then performing frequency separation processing on the original time-domain signal based on the sampling information to obtain the target time-domain signal.

4. The method according to claim 1, characterized in that, The number of the band signals is M, where M is an integer greater than 1; the feature encoding process is performed by a feature encoding layer of a classification model. The feature encoding layer includes M feature encoding units, and the M feature encoding units correspond to the M band signals one by one. The M band signals are input into the corresponding feature encoding units in sequence according to the magnitude order of the band information corresponding to the M band signals respectively; Performing feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals includes: Invoking the m-th feature encoding unit among the M feature encoding units to perform feature encoding on the m-th band signal corresponding to the m-th feature encoding unit, so as to obtain a feature vector corresponding to the m-th band signal, where m is a positive integer less than or equal to M.

5. The method according to claim 1, wherein, the feature learning process is performed by a feature learning layer of a classification model. The feature learning layer includes N feature learning units; the number of the feature vectors is M, and the M feature vectors are divided into N vector groups, each vector group includes at least two feature vectors. Both N and M are integers greater than 1, and N is less than or equal to M; Performing classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector includes: Invoking the first feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the first vector group among the N vector groups, so as to obtain a feature learning result of the first feature learning unit; The feature learning layer invokes the second feature learning unit among the N feature learning units to perform feature learning on the feature vectors in the second vector group among the N vector groups and the feature learning result of the first feature learning unit, so as to obtain a feature learning result of the second feature learning unit; wherein, the number of the classification indication vectors is P, and the P classification indication vectors are determined according to the feature learning results of the last P feature learning units among the N feature learning units, where P is a positive integer less than or equal to N.

6. The method according to claim 1, wherein, the number of the classification indication vectors is P, where P is a positive integer; the feature decoding process is performed by a feature decoding layer of a classification model. The feature decoding layer includes Y feature decoding units, and different feature decoding units among the Y feature decoding units correspond to different frame types, where Y is an integer greater than 1; Performing feature decoding on the classification indication vector to obtain a classification result of the target audio frame includes: Invoking the y-th feature decoding unit among the Y feature decoding units to perform dimensionality reduction processing on the P classification indication vectors, so as to obtain a dimensionality reduction result of the y-th feature decoding unit; and performing regression processing on the dimensionality reduction result of the y-th feature decoding unit to obtain a probability that the target audio frame belongs to the frame type corresponding to the y-th feature decoding unit, where y is a positive integer less than or equal to Y. Among them, the classification result includes: the probability that the target audio frame belongs to the frame type corresponding to each of the Y feature decoding units.

7. The method according to claim 1, wherein, the performing statistical filtering processing on the parent audio block of the target audio frame based on the frame type of the target audio frame to determine the state of the parent audio block includes: performing statistics on the frame types of each audio frame in the parent audio block; if the frame type statistics result indicates that the number of audio frames belonging to the target frame type among each audio frame included in the parent audio block satisfies the audio frame number condition, determining the state of the parent audio block as the active state.

8. The method according to claim 1, wherein, when the service mode is the music mode, the performing service processing on the audio to be processed according to the determined service mode includes: performing noise reduction processing on the audio to be processed according to the low noise reduction processing rule corresponding to the music mode; or, outputting a music mode selection interface, and in response to a music mode selection operation input in the music mode selection interface, performing noise reduction processing on the audio to be processed according to the low noise reduction processing rule corresponding to the music mode; when the service mode is a non - music mode, the performing service processing on the audio to be processed according to the determined service mode includes: performing noise reduction processing on the audio to be processed according to the high noise reduction processing rule corresponding to the non - music mode.

9. The method according to claim 1, wherein, the classification result includes the probability that the target audio frame belongs to each of the Y frame types, Y is an integer greater than 1; different frame types among the Y frame types correspond to different probability ranges; the determining the frame type of the target audio frame according to the classification result includes: determining the audio scene type of the audio to be processed; determining the weight of each of the Y frame types according to the audio scene type; performing weighted summation processing on the probabilities of the corresponding frame types among the Y frame types according to the weights of each of the Y frame types to obtain a weighted average probability; determining the frame type of the target audio frame according to the probability range to which the weighted average probability belongs.

10. An audio processing device, wherein, the audio processing device includes: an acquisition unit, configured to acquire multiple band signals of a target audio frame in the audio to be processed; the audio frames in the audio to be processed are divided into multiple audio blocks, each audio block includes at least two audio frames; an audio sliding window slides in the audio to be processed with one audio block as a sliding step, and the audio sliding window allows at least two audio blocks to be accommodated; a processing unit, configured to perform feature encoding on each of the multiple band signals to obtain a feature vector corresponding to each of the multiple band signals; the processing unit is further configured to perform classification learning on the multiple feature vectors obtained by feature encoding to obtain a classification indication vector; The processing unit is further configured to perform feature decoding on the classification indication vector to obtain a classification result of the target audio frame, and determine a frame type of the target audio frame according to the classification result; The processing unit is further configured to perform statistical filtering processing on a parent audio block of the target audio frame based on the frame type of the target audio frame to determine a state of the parent audio block, where the parent audio block refers to an audio block including the target audio frame; if the audio sliding window slides to a position including the parent audio block, perform statistics on states of each audio block currently included in the audio sliding window; if a state statistics result indicates that the number of target audio blocks in each audio block currently included in the audio sliding window meets an audio block number condition, determine that a service mode of the to-be-processed audio is a music mode, where the target audio block refers to an audio block in an active state among each audio block currently included in the audio sliding window; if the state statistics result indicates that the number of target audio blocks in each audio block currently included in the audio sliding window does not meet the audio block number condition, determine that the service mode of the to-be-processed audio is a non-music mode; The processing unit is further configured to perform service processing on the to-be-processed audio according to the determined service mode.

11. A computer device, characterized in that the computer device includes: a processor, adapted to implement a computer program; a computer-readable storage medium storing a computer program, where the computer program is adapted to be loaded and executed by the processor to perform the audio processing method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, where the computer program is adapted to be loaded and executed by a processor to perform the audio processing method according to any one of claims 1-9.

13. A computer program product, characterized in that the computer program product includes computer instructions, and when the computer instructions are executed by a processor, the audio processing method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Voice endpoint detection method and system based on long short-term memory network

    CN112967739A

  • Audio scene detection method, device and equipment and storage medium

    CN113257276A