Audio recognition method and multi-task audio recognition model training method

CN116913286BActive Publication Date: 2026-09-11TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311013736.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2026-09-11
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

[0003]然而,由于诸多原因,如音频中同时存在多种关联违规内容,违规内容有时比较晦涩或者因音频中正常播放声音、如背景音乐(BGM)的干扰,当前的音频违规内容检测效果不佳

Benefits of technology

[0044]本技术方案提供了音频识别的方法,构建多任务音频识别模型,多任务音频识别模型具有共享编码器,通过上下文关系和音频特点等关联信息,形成音频编码特征向量,并行经过分类子模型和编码子模型,同时得到分类识别结果和语音内容识结果,进而基于音频类别与语音内容的识别结果识别音频中的低俗、色情等违规内容。本发明方案可有效提高音频识别准确率,节省音频识别计算资源成本和人工审核成本;音频识别结果可进一步作为对用户或黑产恶意上传的违规音频作出警告或封号处理的依据,为音频内容安全检测提供技术保障。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116913286B_ABST
    Figure CN116913286B_ABST
Patent Text Reader

Abstract

The application discloses a multi-task audio recognition method, comprising receiving an audio signal; performing endpoint processing on the audio signal to obtain an effective audio segment; extracting an acoustic feature vector of the effective audio segment; inputting the acoustic feature vector of the effective audio segment into a trained multi-task audio recognition model to obtain an audio classification recognition result and a speech content recognition result; and identifying illegal content of the audio according to the audio classification recognition result and the speech content recognition result. The application can be used for quickly identifying the existence of illegal content such as pornography and vulgarity in a context or situation, improving the audio content safety detection accuracy, and effectively reducing the audio recognition calculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multimedia content processing, specifically to an audio recognition method and a multi-task audio recognition model training method. Additionally, this application also relates to related electronic devices and storage media. Background Technology

[0002] With the rapid development of short video and live streaming industries on the internet, a massive amount of audio content or video content containing audio has been generated. Some users, in order to attract traffic or vent their emotions, include pornographic, vulgar, or potentially rule-breaking audio content in short videos or live streams. Therefore, it is necessary to detect and prevent violations of audio content generated by short videos and live streams.

[0003] However, due to various reasons, such as the simultaneous presence of multiple related violations in the audio, the violation content sometimes being obscure, or interference from normally played sounds in the audio, such as background music (BGM), the current detection effect of audio violation content is not good.

[0004] Therefore, in the process of realizing this invention, the inventors found that the prior art requires the establishment of multiple independent algorithm models, but the computing resources consumed by multiple independent algorithm models are multiplied, and it is impossible to quickly and accurately identify illegal audio content.

[0005] The background description is provided for the purpose of understanding the relevant technologies in this field and is not intended as an admission of prior art. Summary of the Invention

[0006] Therefore, the embodiments of the present invention aim to provide an audio recognition method, a multi-task audio recognition model training method, and related electronic devices and computer storage media. The audio recognition scheme of the embodiments of the present invention can improve recognition accuracy and reduce audio recognition computational costs.

[0007] In a first aspect, embodiments of the present invention provide an audio recognition method, comprising:

[0008] Receive audio signals;

[0009] Endpoint processing is performed on the audio signal to obtain valid audio segments;

[0010] Extract the acoustic feature vector of the effective speech segment;

[0011] The acoustic feature vectors of the effective speech segments are input into a trained multi-task audio recognition model. The acoustic feature vectors are processed by the shared encoder of the multi-task audio recognition model to obtain encoded feature vectors. These encoded feature vectors are then processed by a first classification sub-model of the multi-task audio recognition model to obtain audio classification and recognition results. Furthermore, the encoded feature vectors are processed by a parallel second decoder sub-model of the multi-task audio recognition model to obtain speech content recognition results.

[0012] Based on the audio classification and recognition results and the speech content recognition results, it is determined whether the audio meets the preset violation conditions.

[0013] In some embodiments of the present invention, the shared encoder includes a bottleneck layer that supports input audio of arbitrary length having the same dimension, the bottleneck layer being composed of multiple neural network layers.

[0014] In some embodiments of the present invention, the first classification sub-model includes a linear projection layer and a classifier corresponding to the number of audio classifications.

[0015] In some embodiments of the present invention, the second decoder sub-model includes a CTC decoder for obtaining multiple candidate speech content recognition results and an attention decoder for scoring the multiple candidate speech content recognition results.

[0016] In some embodiments of the present invention, the decoder includes multiple Transformer decoding layers, each with the same dimension as the shared encoder.

[0017] In some embodiments of the present invention, the encoded feature vector is processed by a parallel second decoder sub-model of the multi-task audio recognition model to obtain a speech content recognition result, including:

[0018] The CTC decoder calculates the candidate results and scores for speech content recognition, and outputs the candidate results in descending order of scores.

[0019] The candidate results are re-scored using the attention decoder; and,

[0020] The candidate result with the highest score is output as the speech content recognition result.

[0021] In some embodiments of the present invention, extracting the acoustic feature vector of the effective speech segment includes:

[0022] Extract one or more acoustic features of a specified dimension from the effective speech segment after short-time Fourier transform; and,

[0023] The acoustic feature vector is composed based on one or more of the aforementioned acoustic features.

[0024] In some embodiments of the present invention, receiving the audio signal includes:

[0025] The audio signal is received from an audio channel in an audio / video source, the audio / video source including audio / video files and / or live stream links; and / or,

[0026] The audio signal is preprocessed; the preprocessing of the audio signal includes performing one or more preprocessing operations on the audio signal, including specified encoding format conversion, normalization and pre-emphasis.

[0027] In some embodiments of the present invention, endpoint processing is performed on the acquired audio signal to obtain the effective audio segment, including:

[0028] Determine one or more characterization information of the audio signal;

[0029] The effective audio segment is extracted from the audio signal based on one or more characterization information, wherein the effective segment includes a silent segment and / or a noise segment, and the characterization information includes one or more of amplitude, energy, zero-crossing rate, and fundamental frequency.

[0030] In some embodiments of the present invention, the acoustic feature vector may be downsampled before being processed by the shared encoder.

[0031] In some embodiments of the present invention, the multi-task audio recognition model further includes: a language detection decoder for identifying the language of the audio; and / or a context recognizer for identifying the environment in which the audio is located.

[0032] Secondly, a multi-task audio recognition model training method is proposed, including:

[0033] Extract the acoustic feature vectors from the training audio;

[0034] The acoustic feature vector of the training audio is input into the multi-task audio recognition model to be trained, wherein the multi-task audio recognition model includes a shared encoder, a first classification sub-model and a second encoder sub-model, wherein the first classification sub-model includes a classifier and the second encoder sub-model includes one or more decoders, wherein the first classification sub-model and the second encoder sub-model are set in parallel.

[0035] The first classification sub-model and the second encoder sub-model of the multi-task audio recognition model respectively output the audio classification recognition training results and the speech content recognition training results.

[0036] The parameters of the multi-task audio recognition model are iteratively updated based on the audio classification and recognition training results and the speech content recognition training results until a preset iteration termination condition is reached, so as to obtain a trained multi-task audio recognition model.

[0037] In some embodiments of the present invention, the target loss function of the multi-task audio recognition model may include a weighted first classification sub-loss function and one or more second decoder sub-loss functions.

[0038] In some embodiments of the present invention, the second decoding sub-loss function may include a CTC decoding sub-loss function and an attention decoding sub-loss function, wherein the sum of the weight values ​​of the first classification sub-loss function, the CTC decoding sub-loss function and the attention decoding sub-loss function is 1.

[0039] In some embodiments of the present invention, the method further includes: before being input into the multi-task audio recognition model to be trained, performing data augmentation processing on the acoustic feature vector of the training audio signal, wherein the data augmentation processing includes adding noise and / or reverberation to the time-domain signal and / or frequency-domain signal of the training audio signal.

[0040] In some embodiments of the present invention, the second encoder sub-model includes a CTC decoder and an attention decoder.

[0041] In some embodiments of the present invention, the outputs of both the CTC decoder and the attention decoder are connected to a linear layer.

[0042] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein, when the program is executed by a processor, it implements an audio recognition method of any embodiment of the present invention, or a multi-task audio recognition model training method of any embodiment of the present invention.

[0043] Fourthly, embodiments of the present invention provide an electronic device, including: a processor and a memory storing a computer program, wherein the processor is configured to execute an audio recognition method of any embodiment of the present invention and a multi-task audio recognition model training method of any embodiment of the present invention when running the computer program.

[0044] This technical solution provides an audio recognition method that constructs a multi-task audio recognition model. This model features a shared encoder, which generates an audio encoded feature vector through contextual relationships and audio characteristics. This vector is then passed in parallel through a classification sub-model and an encoding sub-model, simultaneously yielding classification and speech content recognition results. Based on these results, vulgar, pornographic, and other illegal content in the audio can be identified. This invention effectively improves audio recognition accuracy, saves computational resources and manual review costs, and provides a basis for warning or banning users or maliciously uploaded audio content, thus offering technical support for audio content security detection.

[0045] In contrast, in some solutions known to the inventors, multiple independent algorithm models need to be established for audio detection, such as an algorithm model for panting, an algorithm model for distinguishing between songs and speech audio types, etc. Although multiple independent tasks and their models can meet the audio detection requirements, the computational resources consumed by multiple independent algorithm models are multiplied, and they do not consider the correlation between audio tasks, so they cannot achieve fast and accurate identification of audio violations.

[0046] Other optional features and technical effects of the embodiments of the present invention are partly described below and partly apparent from reading this document. Attached Figure Description

[0047] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The elements shown are not limited to the scale shown in the drawings, and the same or similar reference numerals in the drawings denote the same or similar elements, wherein:

[0048] Figure 1 One example flowchart of an audio recognition method according to an embodiment of the present invention is shown;

[0049] Figure 2 This is a second example flowchart of the audio recognition method according to an embodiment of the present invention;

[0050] Figure 3 This is the third example flowchart of the audio recognition method according to an embodiment of the present invention;

[0051] Figure 4 This diagram illustrates one of the structural schematics of a multi-task audio recognition model in the audio recognition method of this invention.

[0052] Figure 5 This is the fourth example flowchart of the audio recognition method according to an embodiment of the present invention;

[0053] Figure 6 This is the second schematic diagram of the multi-task audio recognition model structure in the audio recognition method of this invention.

[0054] Figure 7 This diagram illustrates an example flowchart of a multi-task audio recognition model training method according to an embodiment of the present invention.

[0055] Figure 8 An exemplary structural diagram of an audio recognition device according to an embodiment of the present invention is shown;

[0056] Figure 9 An exemplary structural diagram of a multi-task audio recognition model training apparatus according to an embodiment of the present invention is shown;

[0057] Figure 10 An exemplary structural diagram of an electronic device capable of implementing the method according to an embodiment of the present invention is shown. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.

[0059] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0060] In embodiments of the present invention, "network" has the conventional meaning in the field of machine learning, such as neural network NN, deep neural network DNN, convolutional neural network CNN, recurrent neural network RNN, Transformer, Conformer, other machine learning or deep learning networks, or combinations or modifications thereof. In some embodiments, the relevant content based on Transformer can be referred to in the paper "Attention Is All You Need" published in 2017 by Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, AN, ... & Polosukhin, I. from *Advances in Neural Information Processing Systems*; the relevant content based on Conformer can be referred to in the conference paper "Conformer: Convolution-augmented Transformer for Speech Recognition" published in 2020 by Gulati, A., Ahuja, V., Misra, H., & Narayanan, S. from *International Conference on Learning Representations*.

[0061] In embodiments of the present invention, the acoustic model may primarily employ a DNN model, such as neural network structures like CNN, CRNN, TDNN, LSTM, and RNN.

[0062] In embodiments of the present invention, the end-to-end speech recognition method mainly includes a connection-time classification model (CTC), a recurrent neural network converter model (RNN-T), an attention-based sequence-to-sequence (Attention-based Seq2Seq) model, Transformer and Conformer models, etc. The connection-time classification model (CTC) and its corresponding decoder aim to directly associate speech with corresponding text, achieving temporal problem classification. The attention mechanism and its corresponding decoder focus on information more critical to the current task from a large amount of input information, reducing attention to other information and even filtering out irrelevant information, thus solving the information overload problem and improving the efficiency and accuracy of task processing.

[0063] In this embodiment of the invention, "model" has the conventional meaning in the field of machine learning. For example, a model can be a machine learning or deep learning model, such as a machine learning or deep learning model that includes or is composed of the above-mentioned network.

[0064] In this embodiment of the invention, "loss function" and "loss value" have their conventional meanings in the field of machine learning.

[0065] In this embodiment of the invention, the processing of "audio" includes identifying videos and short videos with sound, as well as audio works recorded by one or more people with different vocal expressions and recording formats, including but not limited to songs, music, audiobooks, audio recordings, rap programs, etc., and performing necessary audio and video technical processing.

[0066] This invention provides an audio recognition method and apparatus, a related training method and system for a multi-task audio recognition model, as well as a storage medium and electronic device. The method, system, apparatus / model can be implemented using one or more computers. In some embodiments, the system, apparatus / model can be implemented by software, hardware, or a combination of both. In some embodiments, the electronic device or computer can be implemented by the computer described herein or other electronic devices capable of performing the corresponding functions.

[0067] This invention can be applied to the backend content security management of products such as live streaming or audio-related applications. By identifying the audio content of real-time audio and video streams from user live streams or uploaded audio and video content, it can take action against users or cybercriminals uploading malicious and illegal audio content, issuing warnings or banning accounts. This reduces the cost of manpower for security review and improves the efficiency of security review, protecting the business security of the product line and preventing the product from causing adverse social impacts and regulatory notices. This technical solution, by combining audio characteristics and voice content recognition, can accurately identify vulgar, pornographic, and other illegal audio and video content.

[0068] Therefore, as Figure 1 As shown, in one exemplary embodiment, an audio recognition method is provided, comprising:

[0069] S110: Receives audio signals.

[0070] In some embodiments of the present invention, receiving the audio signal may include receiving the audio signal from an audio channel in an audio / video source, wherein the audio / video source includes audio / video files and / or live stream links.

[0071] Specifically, whether it's uploaded audio / video files or live audio / video streams provided as live stream links, they can all be downloaded or retrieved in real time. Because video and audio encoding / decoding principles are different and they reside in the same signal channel, the corresponding audio signal can be separated from the video signal containing sound.

[0072] In some embodiments of the present invention, the audio signal is preprocessed; the preprocessing of the audio signal includes performing one or more preprocessing operations on the audio signal, including specified encoding format conversion, normalization and pre-emphasis.

[0073] Specifically, since the obtained audio signals may have different encoding formats, the audio signal data must first be transcoded according to a specified encoding type during preprocessing. For example, depending on the encoding method, audio signal encoding can be achieved through waveform encoding, parametric encoding, and hybrid encoding to form a unified audio code. For ease of computation, such as in subsequent endpoint detection to determine a fixed threshold, normalization preprocessing is usually required for the audio signal. Pre-emphasis processing is a prerequisite for audio signal processing, used to enhance the high-frequency components in the audio signal, especially for the application scenarios of this invention, for the security detection of some illegal audio, such as moaning. Audio is produced through the human vocal system, starting from the lungs, with airflow passing through the vocal cords, causing periodic vibrations, and then passing through the pharynx, oral cavity, lips, and tongue to form the final sound. It is generally dominated by low-frequency components, but high-frequency components contain more information, thus improving the completeness of the audio signal representation. In addition, other preprocessing operations, including but not limited to framing and windowing, can be used, depending on the needs of the audio signal and subsequent audio feature extraction.

[0074] S120: Perform endpoint processing on the acquired audio signal to obtain a valid audio segment.

[0075] In some embodiments of the present invention, endpoint processing is performed on the acquired audio signal to obtain the effective audio segment, such as... Figure 2 As shown, it includes:

[0076] S121: Determine one or more representational information of the audio signal.

[0077] In some embodiments of the present invention, the characterization information includes one or more of amplitude, energy, zero-crossing rate, and fundamental frequency.

[0078] The fundamental frequency has many uses, including detecting speech noise, detecting special sounds, gender discrimination, speaker recognition, and parameter adaptation.

[0079] S122: Extract valid audio segments from an audio signal based on one or more characterization information.

[0080] In some embodiments of the present invention, a silent segment and / or a noise segment are determined and extracted based on one or more characterization information.

[0081] Specifically, calculating representation information and judging and detecting effective speech is to remove silent segments and noise segments, extract effective speech segments, and reduce the impact of silent segments and noise segments on the recognition results.

[0082] S130: Extract the acoustic feature vectors of valid speech segments.

[0083] In some embodiments of the present invention, the acoustic feature vector of the effective speech segment is extracted, such as... Figure 3 As shown, it includes:

[0084] S131: Extract one or more acoustic features of a specified dimension from a valid speech segment after short-time Fourier transform.

[0085] The acoustic features include Mel frequency cepstral coefficients (MFCC) and / or filter bank spectrum (Fbank).

[0086] Specifically, since the conventional Fourier transform can only reflect the characteristics of audio signals in the frequency domain and cannot analyze the signals in the time domain, this embodiment of the invention employs a Short-Time Fourier Transform (STFT) to process effective speech segments in order to link the time and frequency domains. Essentially, it is a windowed Fourier transform. For example, 80-dimensional FBank features are extracted from effective speech segments.

[0087] S132: An acoustic feature vector is formed based on one or more of the aforementioned acoustic features.

[0088] In the embodiments of this application, these acoustic features, such as Mel-frequency cepstral coefficients (MFCC) and / or filter bank spectra (FBank), can be used to form an acoustic feature vector. The manner in which the feature vector is constructed can be determined based on known techniques.

[0089] S140: Input the acoustic feature vectors of the effective speech segments into the trained multi-task audio recognition model.

[0090] In some embodiments of the present invention, the acoustic feature vector input to a trained multi-task audio recognition model is processed by the shared encoder of the multi-task audio recognition model to obtain an encoded feature vector; and,

[0091] The encoded feature vector will be processed by the first classification sub-model of the multi-task audio recognition model to obtain the audio classification and recognition result, and the encoded feature vector will be processed by the parallel second decoder sub-model of the multi-task audio recognition model to obtain the speech content recognition result.

[0092] In some embodiments of the present invention, the acoustic feature vector is downsampled before being processed by the shared encoder.

[0093] Specifically, downsampling the acoustic feature vector includes passing the feature vector through a convolutional layer and outputting an acoustic feature vector with a sampling rate lower than a preset proportion of the original sampling rate. For example, before inputting to the shared encoder, the feature vector is first passed through a convolutional layer for downsampling, reducing the sampling rate to 1 / 4 of the original.

[0094] In some embodiments of the present invention, the shared encoder includes a bottleneck layer that supports input audio of arbitrary length with the same dimension, the bottleneck layer being composed of multiple neural network layers.

[0095] Specifically, the shared encoder comprises a multi-layered neural network encoding layer, which includes one or more of Transformer, Conformer, CNN, and RNN, and the encoding layer includes a bottleneck layer. Specifically, this embodiment of the invention uses a Transformer neural network as an example. Transformer is a sequence generation neural network based on a sequence-to-sequence (seq2seq) structure, and it employs attention neurons. Compared to RNN, the advantage of the attention mechanism used by Transformer is that the training process is parallel, making it more suitable for training in large-scale distributed clusters; compared to CNN, the advantage of the attention mechanism used by Transformer is that it allows for viewing the global data.

[0096] Specifically, the number of encoding layers in the neural network of the shared encoder and the parameter values ​​of the neural network are determined according to the encoding requirements. For example, using Transformer, the shared encoder consists of 12 Transformer encoding layers, with each Transformer block having an attention dimension of 256 and a feedforward dimension of 2048. Therefore, the acoustic feature vector is input to the shared encoder for encoding, including receiving acoustic feature vectors of arbitrary length and of the same dimension for encoding. Thus, the multi-task audio recognition method of this invention achieves audio category recognition and speech recognition by sharing audio feature extraction and encoder, essentially representing the sharing of audio features.

[0097] For reference Figure 4 In some embodiments of the present invention, the first classification sub-model includes a linear projection layer and a classifier corresponding to the number of audio classifications.

[0098] It can be understood that a linear projection layer is essentially a fully connected layer, used to connect each node to all nodes in the previous layer, thus combining the features extracted earlier. In some specific implementations, the dimension of the output vector of the linear projection layer is determined based on the number of audio classification categories; and the output vector of the linear projection layer is input into the classifier to output audio category labels.

[0099] In some embodiments of the present invention, the second decoder sub-model includes a CTC decoder for obtaining multiple candidate speech content recognition results and an attention decoder for scoring the multiple candidate speech content recognition results.

[0100] It is understandable that the goal of the connection-time classification model CTC and its corresponding CTC decoder is to directly associate speech with corresponding text to achieve temporal problem classification. The attention mechanism and its corresponding attention decoder focus on the information that is more critical to the current task from a large amount of input information, reduce attention to other information, and even filter out irrelevant information, which can solve the problem of information overload and improve the efficiency and accuracy of task processing.

[0101] In some embodiments of the present invention, the decoder includes multiple Transformer decoding layers, each with the same dimension as the shared encoder.

[0102] Specifically, the attention decoder includes a specified number of decoding layers corresponding to the neural network in the shared encoder; and the dimension of each of the decoding layers is the same as the dimension of the acoustic feature vector encoded by the shared encoder.

[0103] In some embodiments of the present invention, the encoded feature vector is processed by a parallel second decoder sub-model of the multi-task audio recognition model to obtain the speech content recognition result, such as... Figure 5 As shown, it includes:

[0104] S141: The CTC decoder calculates the candidate results and scores for speech content recognition, and outputs the candidate results in descending order of scores.

[0105] S142: The attention decoder re-scores the candidate results.

[0106] S143: Output the candidate result with the highest score as the speech content recognition result.

[0107] Specifically, the linear projection layer for audio classification outputs a vector with the same dimension as the number of categories, which is then passed through a classifier to output audio category labels. Speech content recognition is performed by a decoder using a two-step decoding method. First, a CTC decoder obtains multiple candidate results, and finally, an attention decoder re-scores these candidate results. The attention decoder contains six Transformer decoding layers, each with the same dimension as the shared encoder. This two-step decoding approach achieves accurate speech content recognition. For each audio frame, audio category labels and speech recognition results are output simultaneously, enabling multi-task audio recognition.

[0108] In some embodiments of the present invention, the multi-task audio recognition model, such as Figure 6 As shown, it also includes: a language detection decoder for identifying the language of the audio; and / or, a context recognizer for identifying the environment in which the audio is located.

[0109] Specifically, based on the solution of this invention, various additional recognition technologies can be added, such as language detection, and a corresponding language decoder can be connected according to the language of the speech to achieve multilingual speech recognition; it is also possible to combine audio category recognition and speech recognition results to determine the context or environment in which the user is speaking.

[0110] S150: Confirm whether the audio meets the preset violation conditions based on the audio classification and recognition results and the speech content recognition results.

[0111] In some embodiments of the present invention, before performing audio recognition, at least one preset violation condition can be preset to determine whether the audio violates regulations. The audio features used to confirm whether the audio violates regulations include, but are not limited to, the audio containing / involving hate speech and discriminatory content, violent and threatening content, pornographic content, illegal gambling information, false and fraudulent information, terrorist information, spam advertising information, personal attacks, and politically sensitive information, etc., without limitation. The preset violation condition can consist of a violation condition confirming the audio classification and recognition result and a violation condition confirming the voice content recognition result. Specifically, the form of the preset violation condition includes, but is not limited to, classification rules, thresholds, or ratio scores, for example: the received audio is identified and classified as gambling information; the received audio is identified as containing pornographic language content; the received audio is identified as containing violent and threatening words exceeding a ratio threshold, etc.

[0112] Those skilled in the art should understand that they can pre-set pre-defined violation conditions that meet their expectations based on the specific application scenario of the audio recognition method disclosed in this application and their own needs. The content and composition of the pre-defined violation conditions mentioned above are only examples for reference and are not intended to limit the technical solutions described in this application.

[0113] This technical solution provides an audio recognition method that constructs a multi-task audio recognition model. This model features a shared encoder, which generates an audio encoded feature vector through contextual relationships and audio characteristics. This vector is then passed in parallel through a classification sub-model and an encoding sub-model, simultaneously yielding classification and speech content recognition results. Based on these results, vulgar, pornographic, and other illegal content in the audio can be identified. This invention effectively improves audio recognition accuracy, saves computational resources and manual review costs, and provides a basis for warning or banning users or maliciously uploaded audio content, thus offering technical support for audio content security detection.

[0114] In contrast, some known solutions require the development of multiple independent algorithm models for audio detection, such as models for detecting moaning or for differentiating between songs and spoken audio. While multiple independent tasks and their models can meet the audio detection requirements, the computational resources consumed by these models are significantly increased, and they do not consider the correlation between audio tasks, thus failing to achieve rapid and accurate identification of inappropriate audio content.

[0115] like Figure 7 As shown, in one exemplary embodiment, a multi-task audio recognition model training method is provided. The multi-task audio recognition model training method of this embodiment includes:

[0116] S210: Extract the acoustic feature vector of the training audio.

[0117] In some embodiments of the present invention, the method further includes: before being input into the multi-task audio recognition model to be trained, performing data augmentation processing on the acoustic feature vector of the training audio signal, wherein the data augmentation processing includes adding noise and / or reverberation to the time-domain signal and / or frequency-domain signal of the training audio signal.

[0118] S220: Input the acoustic feature vector of the training audio into the multi-task audio recognition model to be trained.

[0119] The multi-task audio recognition model includes a shared encoder, a first classification sub-model, and a second encoder sub-model. The first classification sub-model includes a classifier, and the second encoder sub-model includes one or more decoders. The first classification sub-model and the second encoder sub-model are configured in parallel.

[0120] In some embodiments of the present invention, the second encoder sub-model includes a CTC decoder and an attention decoder.

[0121] In some embodiments of the present invention, the outputs of both the CTC decoder and the attention decoder are connected to a linear layer.

[0122] The purpose of connecting the outputs of the CTC decoder and the attention decoder to the linear layer is to ensure that the dimension of the output speech content recognition training result is the same as the size of the speech content recognition dictionary.

[0123] In some embodiments of the present invention, the target loss function of the multi-task audio recognition model includes a weighted first classification sub-loss function and one or more second decoder sub-loss functions.

[0124] In some embodiments of the present invention, the second decoding sub-loss function includes a CTC decoding sub-loss function and an attention decoding sub-loss function, wherein the sum of the weights of the first classification sub-loss function, the CTC decoding sub-loss function, and the attention decoding sub-loss function is 1.

[0125] Specifically, the objective loss function of the multi-task audio recognition model can be constructed using the Adam optimizer. The objective loss function for training the multi-task audio recognition model is L = αLoss. Classifier +βLoss CTC +(1-α-β)Loss Attention Among them, Loss Classifier Loss CTC Loss Attention These are the classifier loss function, the CTC decoder loss function, and the attention decoder loss function, respectively; and α and β are hyperparameters used to adjust the weights of the classifier loss function and the CTC decoder loss function.

[0126] S230: The first classification sub-model and the second encoder sub-model of the multi-task audio recognition model output the audio classification recognition training results and the speech content recognition training results, respectively.

[0127] S240: Iteratively update the parameters of the multi-task audio recognition model based on the training results of audio classification and recognition and speech content recognition until the preset iteration termination condition is reached, so as to obtain the trained multi-task audio recognition model.

[0128] In embodiments of the present invention, such as Figure 8 The diagram illustrates a multi-task audio recognition device 800, which can incorporate features of any embodiment of the multi-task audio recognition method. Figure 8 In this embodiment, the multi-task audio recognition device 800 includes:

[0129] The first module 801 is configured to perform endpoint processing on the received audio signal to obtain a valid audio segment.

[0130] The second module 802 is configured to acquire valid speech segments of audio.

[0131] The third module 803 is configured to extract acoustic feature vectors of valid speech segments.

[0132] The fourth module 804 is configured to input the acoustic feature vectors of valid speech segments into a trained multi-task audio recognition model.

[0133] The fifth module 805 is configured to confirm whether the audio meets the preset violation conditions based on the recognition results output by the audio recognition model.

[0134] The acoustic feature vector is processed by the shared encoder of the multi-task audio recognition model to obtain an encoded feature vector, the encoded feature vector is processed by the first classification sub-model of the multi-task audio recognition model to obtain an audio classification recognition result, and the encoded feature vector is processed by the parallel second decoder sub-model of the multi-task audio recognition model to obtain a speech content recognition result.

[0135] It should be understood that the audio recognition device 800 of one embodiment of this specification can also perform... Figures 1 to 6 The method features executed by the audio recognition device (or equipment) and the implementation of the audio recognition device (or equipment) in Figures 1 to 6 The functionality of the example shown will not be elaborated upon here.

[0136] In embodiments of the present invention, such as Figure 9 As shown, a multi-task audio recognition model training device 900 can be combined with the features of any embodiment of the multi-task audio recognition model training method. Figure 9 In this embodiment, the multi-task audio recognition model training device 900 includes:

[0137] The first module 901 is configured to extract acoustic feature vectors from the training audio.

[0138] The second module 902 inputs the acoustic feature vectors of the training audio into the multi-task audio recognition model to be trained.

[0139] The multi-task audio recognition model includes a shared encoder, a first classification sub-model, and a second encoder sub-model. The first classification sub-model includes a classifier, and the second encoder sub-model includes one or more decoders. The first classification sub-model and the second encoder sub-model are configured in parallel.

[0140] The third module 903 outputs the audio classification and recognition training results and the speech content recognition training results from the first classification sub-model and the second encoder sub-model of the multi-task audio recognition model, respectively.

[0141] The fourth module 904 iteratively updates the parameters of the multi-task audio recognition model based on the training results of audio classification and recognition and speech content recognition until the preset iteration termination condition is reached, so as to obtain the trained multi-task audio recognition model.

[0142] It should be understood that the multi-task audio recognition model training device 900 of the embodiments of this specification can also perform... Figure 8 The method features of the training device (or apparatus) for multi-task audio recognition model are described, and the training device (or apparatus) for multi-task audio recognition model is implemented in... Figure 8 The functionality of the example shown will not be elaborated upon here.

[0143] In this embodiment of the invention, an electronic device is provided, including: a processor and a memory storing a computer program, wherein the processor is configured to execute, when running the computer program, a multi-task audio recognition method of any embodiment and a multi-task audio recognition model training method of any embodiment.

[0144] Figure 10 The diagram illustrates a method for implementing embodiments of the present invention or an electronic device 1000 for implementing embodiments of the present invention. In some embodiments, it may include more or fewer electronic devices than illustrated. In some embodiments, it may be implemented using a single or multiple electronic devices. In some embodiments, it may be implemented using cloud-based or distributed electronic devices.

[0145] like Figure 10 As shown, the electronic device 1000 includes a processor 1001, which can perform various appropriate operations and processes based on programs and / or data stored in read-only memory (ROM) 1002 or programs and / or data loaded from storage portion 1008 into random access memory (RAM) 1003. The processor 1001 may be a multi-core processor or may contain multiple processors. In some embodiments, the processor 1001 may include a general-purpose main processor and one or more special coprocessors, such as a central processing unit (CPU), graphics processing unit (GPU), neural network processor (NPU), digital signal processor (DSP), etc. Various programs and data required for the operation of the electronic device 1000 are also stored in RAM 1003. The processor 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0146] The processor and memory described above are used together to execute programs stored in the memory. When the program is executed by a computer, it can implement the methods, steps, or functions described in the above embodiments.

[0147] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, touchscreen, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 1010 as needed so that computer programs read from it can be installed into storage section 1008 as needed. Figure 10 The diagram only shows a portion of the components and does not imply that the computer system 1000 only includes... Figure 10 The components shown.

[0148] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, smartphone, personal computer, laptop computer, in-vehicle human-machine interface device, personal digital assistant, media player, navigation device, game console, tablet computer, wearable device, smart TV, Internet of Things system, smart home, industrial computer, server, or a combination thereof.

[0149] Although not shown, in this embodiment of the invention, a storage medium is provided, the storage medium storing a computer program configured to be executed at runtime to perform the multi-task audio recognition method of any embodiment and the multi-task audio recognition model training method of any embodiment.

[0150] Storage media in embodiments of the present invention include articles that are permanent and non-permanent, removable and non-removable, capable of storing information by any method or technology. Examples of storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0151] The methods, programs, systems, apparatuses, etc., in embodiments of the present invention can be executed or implemented in one or more networked computers, or practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be performed by remote processing devices connected via a communication network.

[0152] Those skilled in the art will understand that the embodiments described in this specification can be provided as methods, systems, or computer program products. Therefore, those skilled in the art will realize that the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware, or a combination of both.

[0153] Unless explicitly stated otherwise, the actions or steps of the methods and procedures described in the embodiments of the present invention do not necessarily have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0154] This document describes several embodiments of the present invention; however, for the sake of brevity, the descriptions of the embodiments are not exhaustive, and identical or similar features or parts between the embodiments may be omitted. In this document, "one embodiment," "some embodiments," "example," "specific example," or "some examples" refers to embodiments applicable to at least one, but not all, of the present invention. The above terms do not necessarily refer to the same embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of the different embodiments or examples.

[0155] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are merely examples of the best mode for implementing the systems and methods. Those skilled in the art will understand that various changes can be made to the embodiments of the systems and methods described herein without departing from the spirit and scope of the invention as defined in the appended claims when implementing the systems and / or methods.

Claims

1. An audio recognition method, characterized in that, include: Receive audio signals; Endpoint processing is performed on the audio signal to obtain valid audio segments; Extract the acoustic feature vector of the effective audio segment; The acoustic feature vector of the effective audio segment is input into a trained multi-task audio recognition model. The acoustic feature vector is processed by the shared encoder of the multi-task audio recognition model to obtain an encoded feature vector. The encoded feature vector is processed by the first classification sub-model of the multi-task audio recognition model to obtain an audio classification recognition result. The audio classification recognition result includes an audio category label. The encoded feature vector is processed by the parallel second decoder sub-model of the multi-task audio recognition model to obtain a speech content recognition result. Based on the audio classification and recognition results and the speech content recognition results, it is determined whether the audio meets the preset violation conditions.

2. The audio recognition method according to claim 1, characterized in that, The shared encoder includes a bottleneck layer that supports input audio of arbitrary length with the same dimension, the bottleneck layer being composed of multiple neural network layers; The first classification sub-model includes a linear projection layer and a classifier corresponding to the number of audio categories; The second decoder sub-model includes a connection-time classification model (CTC) decoder for obtaining multiple candidate speech content recognition results and an attention decoder for scoring the multiple candidate speech content recognition results.

3. The audio recognition method according to claim 2, characterized in that, The encoded feature vector is processed by the parallel second decoder sub-model of the multi-task audio recognition model to obtain the speech content recognition result, including: The CTC decoder is used to calculate candidate results and scores for speech content recognition, and the candidate results are output in descending order of scores. The candidate results are re-scored using the attention decoder; and, The candidate result with the highest score is output as the speech content recognition result.

4. The audio recognition method according to claim 1, characterized in that, Endpoint processing is performed on the acquired audio signal to obtain the valid audio segment, including: Determine one or more characterization information of the audio signal; and, The effective audio segment is extracted from the audio signal based on one or more characterization information, the effective audio segment including a silent segment and / or a noise segment, and the characterization information including one or more of amplitude, energy, zero-crossing rate and fundamental frequency.

5. The audio recognition method according to any one of claims 1 to 4, characterized in that, The multi-task audio recognition model further includes: a language detection decoder for identifying the language of the audio; and / or a context recognizer for identifying the environment in which the audio is located.

6. A method for training an audio recognition model, characterized in that, include: Extract the acoustic feature vectors from the training audio; The acoustic feature vector of the training audio is input into the multi-task audio recognition model to be trained, wherein the multi-task audio recognition model includes a shared encoder, a first classification sub-model and a second encoder sub-model, wherein the first classification sub-model includes a classifier and the second encoder sub-model includes one or more decoders, wherein the first classification sub-model and the second encoder sub-model are set in parallel. The first classification sub-model and the second encoder sub-model of the multi-task audio recognition model output audio classification recognition training results and speech content recognition training results, respectively, wherein the audio classification recognition training results include audio category labels; The parameters of the multi-task audio recognition model are iteratively updated based on the audio classification and recognition training results and the speech content recognition training results until a preset iteration termination condition is reached, so as to obtain a trained multi-task audio recognition model.

7. The audio recognition model training method according to claim 6, characterized in that, The target loss function of the multi-task audio recognition model includes a weighted first classifier sub-loss function and one or more second decoder sub-loss functions.

8. The audio recognition model training method according to claim 7, characterized in that, The second decoder sub-loss function includes a CTC decoder sub-loss function and an attention decoder sub-loss function, wherein the sum of the weights of the first classification sub-loss function, the CTC decoder sub-loss function, and the attention decoder sub-loss function is 1.

9. The audio recognition model training method according to any one of claims 6 to 8, characterized in that, Also includes: Before being input into the multi-task audio recognition model to be trained, the acoustic feature vector of the training audio is subjected to data augmentation processing, which includes adding noise and / or reverberation to the time-domain signal and / or frequency-domain signal of the training audio. The second encoder sub-model includes a CTC decoder and an attention decoder; The outputs of both the CTC decoder and the attention decoder are connected to a linear layer.

10. An electronic device, characterized in that, It includes a processor and a memory storing a computer program, the processor being configured to perform the method of any one of claims 1 to 9 when running the computer program.

11. A storage medium, characterized in that, The storage medium stores a computer program configured to execute the method of any one of claims 1 to 9 when run.

Citation Information

Patent Citations

  • Speech recognition method and device, medium and computing equipment

    CN115064153A

  • Audio sensitive content detection method, computer equipment and computer program product

    CN115148211A