Cross-language audio detection method and device and related equipment
Patent Information
- Application Number
- CN202411999833.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-13
Smart Images

Figure CN120148477A_ABST
Abstract
Description
Background Art
[0002] With the development of multimedia technology, in the User Generated Content (UGC) on the Internet, the audio content produced or shared by users is uploaded directly to the Internet, and it is inevitable that there will be illegal or bad content, especially when it comes to files in non-Chinese (foreign) languages, the possibility of missing bad voice detection is greater.
[0003] Currently, there are more than 4,000 languages in the world, and the voice files or voice streams to be detected are often of unknown languages. Therefore, how to provide a method or model that can cover multi-language speech detection at the same time, rather than developing a separate speech detection model for each language, is a technical problem that needs to be urgently solved in this field.
[0004] In addition, for foreign multilingual speech, manual review faces challenges of low efficiency, difficulty and high cost. Even if multilingual speech is transcribed into the language of the corresponding country, there is a lack of sensitive word library or semantic recognition detection means.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The present disclosure provides a cross-language audio detection method, apparatus and related equipment, which at least to a certain extent overcome the technical problem that cross-language audio detection in the related art is time-consuming and labor-intensive.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a cross-language audio detection method is provided, the method comprising: obtaining audio to be detected; performing language recognition on the audio to be detected; if the audio to be detected is target audio of a target language, inputting the target audio into a target language audio recognition model, and performing semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information; if the audio to be detected is non-target audio of a non-target language, inputting the non-target audio into a cross-language audio recognition model, and performing semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information; inputting the target audio and / or the non-target audio for which no sensitive information is recognized into a pre-trained spectrum information detection model, and outputting an audio detection result of whether the target audio and / or the non-target audio is abnormal audio.
[0009] In some embodiments, the audio to be detected includes: at least one audio segment; wherein, performing language identification on the audio to be detected includes: extracting a target audio segment from the audio to be detected, wherein the target audio segment includes: one or more audio segments; performing language identification on the target audio segment; if the language identification result of the target audio segment is the target language, determining that the audio to be detected is the target audio in the target language; if the language identification result of the target audio segment is a non-target language, determining that the audio to be detected is the non-target audio in the non-target language.
[0010] In some embodiments, inputting the target audio into a target language audio recognition model and performing semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information includes: inputting the target audio into the target language audio recognition model and outputting the audio recognition result of the target language; inputting the audio recognition result of the target language into the semantic detection model of the target language and outputting the semantic analysis result; determining whether the target audio contains sensitive information according to the semantic analysis result.
[0011] In some embodiments, the cross-language audio recognition model is an audio recognition model obtained by fine-tuning a trained large language model and capable of recognizing audio in multiple languages.
[0012] In some embodiments, the cross-language audio recognition model is an audio recognition model obtained by fine-tuning a trained large language model and capable of recognizing audio in multiple languages; wherein, inputting the non-target audio into the cross-language audio recognition model and performing semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information includes: inputting the non-target audio into the cross-language audio recognition model and outputting the audio recognition result in a single language; inputting the audio recognition result in a single language output by the cross-language audio recognition model into the semantic detection model of the corresponding language and outputting the semantic analysis result; determining whether the non-target audio contains sensitive information according to the semantic analysis result.
[0013] In some embodiments, obtaining the audio to be detected includes: in response to an upload request or sharing request of an audio file, obtaining the audio to be detected.
[0014] According to another aspect of the present disclosure, there is also provided a cross-lingual audio detection device, which includes: an audio acquisition module for acquiring the audio to be detected; a language identification module for identifying the language of the audio to be detected; a target language audio identification module for, if the audio to be detected is a target audio in the target language, inputting the target audio into a target language audio identification model and performing semantic analysis on the audio identification result of the target language audio identification model to determine whether the target audio contains sensitive information; a cross-lingual audio identification module for, if the audio to be detected is a non-target audio in a non-target language, inputting the non-target audio into a cross-lingual audio identification model and performing semantic analysis on the audio identification result of the cross-lingual audio identification model to determine whether the non-target audio contains sensitive information; a spectrogram information detection module for inputting the target audio and / or the non-target audio that does not identify sensitive information into a pre-trained spectrogram information detection model and outputting an audio detection result as to whether the target audio and / or the non-target audio is an abnormal audio.
[0015] According to another aspect of the present disclosure, there is also provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the cross-lingual audio detection method according to any one of the above by executing the executable instructions.
[0016] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the cross-lingual audio detection method according to any one of the above is implemented.
[0017] According to another aspect of the present disclosure, there is also provided a computer program product, including: a computer program or instruction, and when the computer program or instruction is executed by a processor, the cross-lingual audio detection method according to any one of the above is implemented.
[0018] In the embodiments of the present disclosure, a cross - language audio detection method, apparatus and related equipment, after obtaining the audio to be detected, first perform language identification on the audio to be detected. If the audio to be detected is the target audio in the target language, the target audio is input into the target - language audio recognition model, and semantic analysis is performed on the audio recognition result of the target - language audio recognition model to determine whether the target audio contains sensitive information; if the audio to be detected is a non - target audio in a non - target language, the non - target audio is input into the cross - language audio recognition model, and semantic analysis is performed on the audio recognition result of the cross - language audio recognition model to determine whether the non - target audio contains sensitive information; finally, the target audio and / or non - target audio that do not identify sensitive information are input into a pre - trained spectrogram information detection model, and an audio detection result of whether the target audio and / or non - target audio is abnormal audio is output.
[0019] Through the embodiments of the present disclosure, it is possible to quickly and accurately detect whether cross - language audio is abnormal audio, especially suitable for audio detection scenarios with high real - time requirements.
[0020] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0022] Figure 1 Shows a schematic diagram of an application system architecture in an embodiment of the present disclosure;
[0023] Figure 2 Shows a flowchart of a cross - language audio detection method in an embodiment of the present disclosure;
[0024] Figure 3 Shows a flowchart of a language identification in an embodiment of the present disclosure;
[0025] Figure 4 Shows a flowchart of an audio sensitive information identification in an embodiment of the present disclosure;
[0026] Figure 5 Shows another flowchart of an audio sensitive information identification in an embodiment of the present disclosure;
[0027] Figure 6 Shows a schematic diagram of a cross - language audio detection system architecture in an embodiment of the present disclosure;
[0028] Figure 7 Shows the implementation flowchart of a cross - language audio detection method in an embodiment of the present disclosure;
[0029] Figure 8 Shows the schematic diagram of a cross - language audio detection device in an embodiment of the present disclosure;
[0030] Figure 9 Shows the structural block diagram of an electronic device in an embodiment of the present disclosure. Detailed implementation manners
[0031] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.
[0032] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0033] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are explained as follows:
[0034] LAM: Large Audio Model, a large - scale audio model, which refers to a machine - learning model with a large number of parameters and a complex structure and can complete audio tasks in multi - language and multi - task scenarios.
[0035] WER: Word Error Rate, a word error rate, which is an important indicator for evaluating the performance of a speech recognition algorithm or system and is used to evaluate the error rate between the predicted text and the standard text.
[0036] LID: Language Identification, which refers to the technology of automatically identifying and judging the speech signal of a speaker by a computer system to obtain the language type corresponding to the speech.
[0037] The following will explain in detail the specific implementation manners of the embodiments of the present disclosure with reference to the accompanying drawings.
[0038] Figure 1The figure shows a schematic diagram of an exemplary application system architecture to which the cross - language audio detection method in the embodiments of the present disclosure can be applied. As Figure 1 shown, the system architecture may include a terminal device 101, a network 102, and a server 103.
[0039] The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103, and can be a wired network or a wireless network.
[0040] Optionally, the above - mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network). In some embodiments, technologies and / or formats including hypertext markup language (HTML), extensible markup language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as secure socket layer (SSL), transport layer security (TLS), virtual private network (VPN), Internet protocol security (IPSec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above - mentioned data communication technologies.
[0041] The terminal device 101 can be various electronic devices, including but not limited to smartphones, tablets, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.
[0042] Optionally, the clients of the application programs installed in different terminal devices 101 are the same, or are clients of the same type of application programs based on different operating systems. Depending on the different terminal platforms, the specific form of the client of the application program can also be different. For example, the client of the application program can be a mobile client, a PC client, etc.
[0043] Server 103 can be a server that provides various services, such as a background management server that supports the devices operated by users using terminal device 101. The background management server can analyze and process data such as received requests, and feedback the processing results to the terminal device.
[0044] Optionally, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0045] Those skilled in the art can know that Figure 1 the numbers of the terminal devices, networks, and servers in
[0046] are only illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. The embodiments of the present disclosure do not limit this.
[0047] Under the above system architecture, an audio detection method across languages is provided in the embodiments of the present disclosure, and this method can be executed by any electronic device with computing and processing capabilities.
[0048] Figure 2 The flowchart of an audio detection method across languages in the embodiments of the present disclosure is shown. As Figure 2 shown, the audio detection method across languages provided in the embodiments of the present disclosure includes the following steps:
[0049] S202, obtain the audio to be detected.
[0050] It should be noted that the audio to be detected in the embodiments of the present disclosure can be any kind of audio, which can be an audio file or real-time transmitted audio stream data, can be human speech, or can be other audio. It should be noted that the audio detected in the embodiments of the present disclosure can be audio with relatively high real-time requirements or audio with relatively low real-time requirements, can be all audio, or can be sampled audio. The audio detection method provided in the embodiments of the present disclosure can be applied but is not limited to various multimedia applications or social applications in business scenarios such as real-time, non-real-time full inspection or sampling inspection, and intelligent review of files and audio streams, such as detection or review in live broadcast, audio and video conferencing, games, audio and video website / APP / miniprogram dial testing, network disk file service, instant messaging service, etc.
[0051] In one embodiment, the audio to be detected can be obtained through the following steps: in response to an upload request or sharing request of an audio file, obtain the audio to be detected.
[0052] S204, perform language identification on the audio to be detected.
[0053] It should be noted that when performing abnormal audio identification (such as illegal audio identification), it is usually necessary to convert the audio into the corresponding text, and then detect whether there are sensitive information such as sensitive words in the corresponding text; however, for language texts in different languages, the sensitive words to be detected are also different. Therefore, in order to accurately detect whether the audio is abnormal audio (such as whether it contains sensitive information), it is necessary to first perform language identification on the audio to be detected.
[0054] S206, if the audio to be detected is the target audio in the target language, input the target audio into the target language audio recognition model, and perform semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information.
[0055] It should be noted that since the audio to be detected is in an unknown language and there are many existing language types, how to quickly detect the language of the audio to be detected has a great impact on subsequent abnormal audio identification. In the embodiments of the present disclosure, one or more languages with relatively high probabilities are predefined as the target languages, and then it is detected whether the audio to be detected is the audio in the target language. Only when the audio to be detected is not the audio in the target language, cross-language detection is performed, which can greatly improve the efficiency of language identification.
[0056] In one embodiment, Chinese can be configured as the target language, and then it is detected whether the audio to be detected is Chinese audio or non-Chinese audio, and the non-Chinese audio is input into the cross-language audio recognition model for recognition.
[0057] S208. If the audio to be detected is a non-target audio in a non-target language, input the non-target audio into a cross-lingual audio recognition model, and perform semantic analysis on the audio recognition result of the cross-lingual audio recognition model to determine whether the non-target audio contains sensitive information.
[0058] It should be noted that the above cross-lingual audio recognition model is a pre-trained model capable of multi-lingual audio recognition, and the output of this model is a single-language audio. In specific implementation, it is difficult to create corresponding sensitive information libraries for all languages. Therefore, the cross-lingual audio recognition model in the embodiments of the present disclosure can be an audio recognition model with multi-lingual translation function, that is, for any audio in any language, this model can output an audio in the same language (such as English). For non-target audio identified as a non-target language, only a sensitive information library in a certain language needs to be configured.
[0059] S210. Input the target audio and / or non-target audio that do not identify sensitive information into a pre-trained spectrogram information detection model, and output the audio detection result of whether the target audio and / or non-target audio is abnormal audio.
[0060] It should be noted that for some illegal abnormal audio, there is no language, and its corresponding text cannot be recognized through speech recognition technology. Therefore, it is impossible to determine whether the audio is abnormal audio by detecting whether the audio contains sensitive information by creating a sensitive information library. To detect this type of audio, in the embodiments of the present disclosure, a model capable of detecting audio information is trained through machine learning, and the spectrogram information is detected to determine whether the audio to be detected is illegal abnormal audio.
[0061] In some embodiments, the audio to be detected includes: at least one audio segment; wherein, performing language recognition on the audio to be detected includes:
[0062] S302. Extract a target audio segment from the audio to be detected, where the target audio segment includes: one or more audio segments;
[0063] S304. Perform language recognition on the target audio segment;
[0064] S306. If the language recognition result of the target audio segment is the target language, determine that the audio to be detected is the target audio in the target language;
[0065] S308. If the language recognition result of the target audio segment is a non-target language, determine that the audio to be detected is a non-target audio in a non-target language.
[0066] It should be noted that for language identification of a piece of audio, it is often sufficient to identify based on a certain segment in the audio. However, to avoid the situation where there are some audio segments in a target language audio that are in other languages (for example, there is an English sentence in a Chinese audio), resulting in the identified language based on some segments not being the language of the audio, thus, multiple audio segments in a piece of audio can be selected for language identification. For example, three segments are extracted from a piece of audio to be detected. The languages of two segments are Chinese and the language of one segment is English, then it is determined that this piece of audio is a Chinese audio segment.
[0067] Through the above embodiments, extracting one or more audio segments from the audio to be detected for language identification can greatly improve the efficiency of language identification.
[0068] In some embodiments, the target audio is input into a target language audio recognition model, and semantic analysis is performed on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information, including:
[0069] S402, input the target audio into the target language audio recognition model, and output the audio recognition result of the target language;
[0070] S404, input the audio recognition result of the target language into the semantic detection model of the target language, and output the semantic analysis result;
[0071] S406, determine whether the target audio contains sensitive information according to the semantic analysis result.
[0072] It should be noted that the above target language audio recognition model is a model that can recognize the audio of the target language and is pre-trained through machine learning. If the audio to be detected is the target audio of the target language, it is input into the target language audio recognition model, and the audio recognition result of the target language (that is, the text corresponding to the audio of the target language) is output. The audio recognition result of the target language is input into the semantic detection model of the target language, and the semantic analysis result is output. Then, it is determined whether the target audio contains sensitive information according to the semantic analysis result. For example, after inputting a Chinese audio into the Chinese audio recognition model, the output is Chinese text. The Chinese text is segmented and matched with a pre-created Chinese sensitive word library to determine whether the text corresponding to the Chinese audio contains sensitive words. In one embodiment, the above semantic detection model of the target language can be pre-trained through machine learning.
[0073] In some embodiments, the above cross-language audio recognition model is an audio recognition model that can recognize the audio of multiple languages and is obtained by fine-tuning a trained large language model.
[0074] In some embodiments, a non-target audio is input into a cross-lingual audio recognition model, and semantic analysis is performed on the audio recognition result of the cross-lingual audio recognition model to determine whether the non-target audio contains sensitive information, including:
[0075] S502, input the non-target audio into the cross-lingual audio recognition model, and output the audio recognition result in a single language;
[0076] S504, input the audio recognition result in a single language output by the cross-lingual audio recognition model into the semantic detection model in the corresponding language, and output the semantic analysis result;
[0077] S506, determine whether the non-target audio contains sensitive information according to the semantic analysis result.
[0078] It should be noted that the single language output by the above cross-lingual audio recognition model can be a pre-defined language, which is usually a relatively common language, for example, English.
[0079] In one embodiment, the above cross-lingual audio recognition model is an audio recognition large model based on a speech large model that can recognize audio in multiple languages. If the audio to be detected is a non-target audio in a non-target language, it is input into the cross-lingual audio recognition model, and the audio recognition result in a single language (which can be a pre-defined and relatively common language, such as English) is output, and then it is input into the semantic detection model in the corresponding language, and the semantic analysis result is output. Furthermore, it is determined whether the target audio contains sensitive information according to the semantic analysis result. For example, when the non-target audio is a Japanese audio, the Japanese audio is input into the cross-lingual audio recognition model, and the output is an English text. After the English text is tokenized, it is matched with a pre-created English sensitive word library to determine whether the Japanese audio contains sensitive words. In one embodiment, if the language output by the cross-lingual audio recognition model is a certain language, a corresponding semantic detection model can be trained through machine learning.
[0080] Through the above embodiments, by using the cross-lingual audio recognition model obtained by fine-tuning the speech large model, the audio in any language is converted into a pre-defined single language, and only a sensitive word library for the single language needs to be created, instead of creating corresponding sensitive word libraries for all languages, which greatly reduces the complexity of the audio recognition system.
[0081] Next, in combination with Figure 6 and Figure 7 to describe in detail the cross-lingual audio detection method provided in the embodiments of the present disclosure. As shown in Figure 6 and Figure 7 shown, the implementation process is as follows:
[0082] 1) If the audio to be detected (such as speech) is in an unknown language, first extract small audio segments from the sampled audio for quick language identification. The output of the language identification is the language ID.
[0083] 2) If the language ID is Chinese, send the audio to be detected into a Chinese speech recognition engine for speech transcription, and then send it to the semantic detection of sensitive information for detection. If there is a violation, directly output the result.
[0084] 3) If no abnormality is found, send the speech to be detected into a special moaning sound detection model for detection. If there is a violation, directly output the result.
[0085] 4) If the result recognized by the language recognition model is non-Chinese (output the language ID of other countries' languages), then send the speech to be detected into a multilingual speech translation model based on a speech large model. And output the translation result of a single language text (such as English). The English text is then sent to the semantic detection model of the English text for discrimination, and finally return the classification result (such as "malicious promotion", etc.) and the discrimination result (such as "release", "normal", "suspected", etc.).
[0086] 5) For foreign language semantic analysis, if no abnormality is found, send the speech to be detected into a special moaning sound detection model for detection. If there is a violation, directly output the result.
[0087] It should be noted that for language identification, the existing speech large models have poor accuracy and long time consumption for Chinese identification. For example, the word error rate (WER) of openai / whisper-large-v2 for Chinese identification is 14.7, while the WER of openai / whisper-tiny for Chinese identification is as high as 40. The transcription speed for a 13-minute audio reaches 4 minutes, which cannot meet the real-time detection production environment. If directly using a multilingual speech large model to identify Chinese and foreign languages simultaneously, the identification effect of Chinese may be very poor, seriously affecting the accuracy of the detection result.
[0088] Therefore, in the previous steps of the method, first perform language identification and do not use all the audio. Only extract a small amount of audio for language identification to improve the review efficiency. If the language identification is Chinese, then use a Chinese speech recognition model with a low word error rate and high maturity for speech transcription later. After language identification, according to the language ID result, if it is the ID of a foreign language, then send it to the transcription / translation based on the speech large model. Avoid first performing Chinese speech recognition, and then performing secondary speech recognition and speech translation for foreign languages after the recognition as semantic garbled characters.
[0089] In addition, it should be noted that when implementing multilingual detection based on a speech large model, the language ID is sent as a parameter to the speech large model, and the speech large model does not need to predict the language by itself, which can greatly improve the recognition accuracy of the speech large model. To implement this function, the implementation method provided by the embodiments of the present disclosure is as follows:
[0090] 1) Use the large model for joint training of multiple tasks and multiple languages, and use a set of training data to simultaneously perform multilingual speech translation tasks.
[0091] 2) Use the large model to train multilingual speech translation, so that the input speech in any language can output text in a single language (such as English text) after passing through a large model.
[0092] 3) Perform model fine-tuning and incremental training on special languages that are key concerns.
[0093] 4) Considering the large number of parameters and slow inference speed of the large model, based on ctranslate2, convert and accelerate the trained speech large model.
[0094] Taking the audio file review scenario of a certain cloud disk as an example, when a user shares a file or a file group with other users, it will go through an audio detection and review system, as well as a manual review system. After the review is passed, the user sharing can take effect; otherwise, it will be blocked. The real-time requirement is that the end-to-end detection of audio below 100M by the entire detection system takes within 180 seconds.
[0095] In the audio detection system, the language recognition model can be implemented based on a large model or a small model. When implemented based on a large model, it can be based on Facebook's MMS (Massive Multilingual Speech) LID series models, and different models from 126 languages to 4017 languages can be selected according to business requirements.
[0096] The Chinese speech recognition small model can be incrementally trained based on the WeNet open-source pre-trained model, and the word error rate (WER) after training is 5.12. The recognized text is segmented and sent to the semantic detection module, and the semantic detection module will use a separate semantic model and a sensitive word library to make semantic judgments. By introducing methods such as homophone expansion, special character processing, sentiment analysis, scene discrimination, large model text classification, and policy determination, it is possible to return detection results within seconds for ultra-long texts.
[0097] The audio to be detected passes through the LID model, and the ID information corresponding to the language (such as the ISO code of the International Organization for Standardization) can be output. According to the ISO code of the corresponding language, this ISO code is used as a parameter in the inference method of the speech large model translation model to complete the translation process.
[0098] On the speech large model translation model, incremental training and fine-tuning can be performed based on models such as the Fast Whisper series (an optimized version of OpenAI's Whisper model, which aims to improve the speed and efficiency of audio transcription and speech recognition tasks) or the Facebook / wav2vec2-xls-r-300m model (which can be translated as "Facebook / Waveform to Vector 2 - Extended Language Representation - 300 million parameters" model. Here, "Waveform to Vector 2" corresponds to "Wav2Vec2", which is a pre-trained model for speech recognition; "Extended Language Representation" corresponds to "XLS-R", indicating that the model has extended language representation capabilities; and "300 million parameters" represents the size of the model). The data for incremental training is trained using negative samples marked by manual review in the detection system. All audio in unknown languages is translated into English, and then the English text is input into the English sensitive word library or English NLP detection. The voice formats supported include aac, ac3, amr, ape, flac, m4a, mp3, ogg, wav, wma, etc.
[0099] When implementing a spectrogram information detection model (such as a gasping sound recognition model) based on a small model, the Panns-cnn10 and Ecapa-tdnn models can be used for training. The positive and negative samples and data for training are as follows: The training results are loss value loss: 0.05732; accuracy accuracy: 0.98355. In the negative samples, audiobooks, phone recordings, pure music instrument pieces, etc. are added to make the model more accurate and robust. Among them, Panns-CNN10 is a model based on convolutional neural network (CNN), specifically for environmental sound recognition. Its name comes from "PANNs" (PyTorch Audio NeuralNetworks) and "CNN10", where "CNN10" indicates that 10 layers of convolutional neural network are used in the model. ECAPA-TDNN is the abbreviation of "eXtended Channel Attention Pooling for Time-Delay Neural Network", which is a deep learning model for speaker recognition.
[0100] Compared with existing audio recognition solutions, the audio recognition method provided in the embodiments of the present disclosure realizes language identification based on speech segments and can quickly identify languages; based on a cross-language audio recognition model (an audio recognition model implemented based on a speech large model), it can solve the problems of multi-language prediction, multi-language translation, and multi-language bad information detection. Through the audio recognition method provided in the embodiments of the present disclosure, the following technical effects can be achieved but are not limited to:
[0101] 1) The most time-consuming part of the entire detection is speech recognition. For cross-languages, only one inference of the speech recognition (or translation) model is required, which improves the detection efficiency and accuracy, especially for services with real-time or quasi-real-time review requirements.
[0102] 2) Due to the limitations of the sensitive word library or text detection algorithm for languages, all languages are translated into a certain language, and then only the semantics of this language is detected for violations. For audio mixed with multiple languages, only one large model for a large speech translation task (instead of multiple large models) is used to solve the difficulty that there is no corresponding sensitive word library for most languages, reducing the system complexity and the difficulty of technical implementation, and further reducing the difficulty of technical development, product implementation and operation of multi-language bad audio detection.
[0103] 3) It can simultaneously perform abnormal audio recognition for multiple languages including Chinese and non-Chinese (such as recognizing moaning sounds without language), support cross-language audio recognition, and can cover various possible requirements for detecting illegal audio (such as foreign language audio violations, moaning sounds, Chinese audio violations, etc.). It solves the difficulty that it is difficult to accurately detect foreign language speech in machine review and manual review in business requirements.
[0104] It should be noted that in the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations. In the embodiments of the present disclosure, various types of data such as personal identity data, operation data, and behavior data related to individuals, customers, and groups have all been authorized.
[0105] Based on the same inventive concept, an audio detection device for cross-languages is also provided in the embodiments of the present disclosure, as described in the following embodiments. Since the principle of solving problems in the device embodiments is similar to that of the above method embodiments, the implementation of the device embodiments can refer to the implementation of the above method embodiments, and the repeated parts will not be described again.
[0106] Figure 8 The schematic diagram of an audio detection device for cross-languages in the embodiments of the present disclosure is shown, as Figure 8 shown, the device includes: an audio acquisition module 801, a language identification module 802, a target language audio identification module 803, a cross-language audio identification module 804, and a spectrogram information detection module 805.
[0107] Among them, the audio acquisition module 801 is used to acquire the audio to be detected; the language recognition module 802 is used to recognize the language of the audio to be detected; the target language audio recognition module 803 is used to, if the audio to be detected is the target audio in the target language, input the target audio into the target language audio recognition model, and perform semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information; the cross-language audio recognition module 804 is used to, if the audio to be detected is the non-target audio in the non-target language, input the non-target audio into the cross-language audio recognition model, and perform semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information; the spectrogram information detection module 805 is used to input the target audio and / or non-target audio that does not recognize sensitive information into the pre-trained spectrogram information detection model, and output the audio detection result of whether the target audio and / or non-target audio is abnormal audio.
[0108] In some embodiments, the audio to be detected includes: at least one audio segment; among them, the language recognition module 802 is further used to: extract the target audio segment from the audio to be detected, where the target audio segment includes: one or more audio segments; recognize the language of the target audio segment; if the language recognition result of the target audio segment is the target language, determine that the audio to be detected is the target audio in the target language; if the language recognition result of the target audio segment is the non-target language, determine that the audio to be detected is the non-target audio in the non-target language.
[0109] In some embodiments, the target language audio recognition module 803 is further used to: input the target audio into the target language audio recognition model, and output the audio recognition result in the target language; input the audio recognition result in the target language into the semantic detection model in the target language, and output the semantic analysis result; according to the semantic analysis result, determine whether the target audio contains sensitive information; if the target audio contains sensitive information, determine that the target audio is abnormal audio that meets the preset abnormal conditions; if the target audio does not contain sensitive information, determine that the target audio is abnormal audio that meets the preset abnormal conditions.
[0110] In some embodiments, the cross-language audio recognition model is an audio recognition model obtained by fine-tuning a trained large language model and can recognize audio in multiple languages.
[0111] In some embodiments, the cross-lingual audio recognition model is an audio recognition model obtained by fine-tuning a trained large language model and capable of recognizing audio in multiple languages. Among them, the cross-lingual audio recognition module 804 is further configured to: input non-target audio into the cross-lingual audio recognition model and output an audio recognition result in a single language; input the audio recognition result in a single language output by the cross-lingual audio recognition model into the semantic detection model in the corresponding language and output a semantic analysis result; determine whether the non-target audio contains sensitive information according to the semantic analysis result; if the non-target audio contains sensitive information, determine that the non-target audio is an abnormal audio that meets the preset abnormal conditions; if the non-target audio does not contain sensitive information, determine that the non-target audio is an abnormal audio that meets the preset abnormal conditions.
[0112] In some embodiments, the above audio acquisition module 801 is further configured to: in response to an upload request or a sharing request of an audio file, acquire the audio to be detected.
[0113] It should be noted here that the examples and application scenarios implemented by each module in the above device embodiments are the same as the corresponding steps in the method embodiments, but are not limited to the content disclosed in the above method embodiments. It should be noted that the above modules, as part of the device, can be executed in a computer system such as a set of computer executable instructions.
[0114] Those skilled in the art of the present technical field can understand that various aspects of the present disclosure can be specifically implemented in the following forms, that is: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module" or "system" here.
[0115] Based on the same inventive concept, an electronic device is further provided in the embodiments of the present disclosure. The electronic device includes: a processor; and a memory for storing executable instructions of the processor. Among them, the processor is configured to execute the XXXXXXXXXXXX method of any one of the above via executing the executable instructions. Since the principle of solving problems in this electronic device embodiment is similar to that in the above method embodiment, the implementation of this electronic device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be described again.
[0116] Next, refer to Figure 9 to describe the electronic device 900 according to this embodiment of the present disclosure. Figure 9 The shown electronic device 900 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0117] As Figure 9As shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one of the above-mentioned processing units 910, at least one of the above-mentioned storage units 920, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).
[0118] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of the present specification above. For example, the processing unit 910 may execute the following steps of the above method embodiment: obtaining the audio to be detected; performing language identification on the audio to be detected; if the audio to be detected is the target audio in the target language, inputting the target audio into the target language audio recognition model, and performing semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information; if the audio to be detected is the non-target audio in the non-target language, inputting the non-target audio into the cross-language audio recognition model, and performing semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information; inputting the target audio and / or non-target audio that does not recognize sensitive information into a pre-trained spectrogram information detection model, and outputting an audio detection result indicating whether the target audio and / or non-target audio is abnormal audio.
[0119] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202, and may further include a read-only storage unit (ROM) 9203.
[0120] The storage unit 920 may further include a program / utilities 9204 having a set (at least one) of program modules 9205. Such program modules 9205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0121] The bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0122] The electronic device 900 can also communicate with one or more external devices 940 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 950. Moreover, the electronic device 900 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 960. As shown in the figure, the network adapter 960 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0123] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0124] Based on the same inventive concept, the embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the cross-language audio detection method in any one of the above. Since the principle of solving problems in the embodiment of this computer-readable storage medium is similar to that in the above method embodiment, the implementation of the embodiment of this computer-readable storage medium can refer to the implementation of the above method embodiment, and the repeated parts will not be described again.
[0125] More specific examples of the computer-readable storage medium in the present disclosure can include but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0126] In the present disclosure, a computer-readable storage medium may include a data signal included in a baseband or propagated as part of a carrier wave, in which a readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0127] Optionally, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0128] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0129] Based on the same inventive concept, an embodiment of the present disclosure also provides a computer program product, including: a computer program or instruction, which when executed by a processor implements the cross-lingual audio detection method in any one of the above method embodiments. Since the principle of solving the problem in this computer program product embodiment is similar to that of the above method embodiments, the implementation of this computer program product embodiment may refer to the implementation of the above method embodiments, and the repeated parts will not be elaborated.
[0130] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0131] In addition, although the steps of the methods in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0132] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0133] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A cross-language audio detection method, characterized in that: include: Get the audio to be detected; Performing language recognition on the audio to be detected; If the audio to be detected is a target audio of a target language, the target audio is input into a target language audio recognition model, and a semantic analysis is performed on an audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information; If the audio to be detected is non-target audio of a non-target language, inputting the non-target audio into a cross-language audio recognition model, and performing semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information; The target audio and / or the non-target audio in which sensitive information is not identified are input into a pre-trained spectrum information detection model, and an audio detection result of whether the target audio and / or the non-target audio is abnormal audio is output.
2. The cross-language audio detection method according to claim 1, characterized in that: The audio to be detected includes: at least one audio segment; The step of performing language recognition on the audio to be detected includes: Extracting a target audio segment from the audio to be detected, wherein the target audio segment includes: one or more audio segments; Performing language recognition on the target audio segment; If the language recognition result of the target audio segment is the target language, determining that the audio to be detected is the target audio of the target language; If the language recognition result of the target audio segment is a non-target language, the audio to be detected is determined to be a non-target audio of the non-target language.
3. The cross-language audio detection method according to claim 1, characterized in that: The step of inputting the target audio into a target language audio recognition model and performing semantic analysis on an audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information includes: Inputting the target audio into a target language audio recognition model, and outputting an audio recognition result of the target language; Inputting the audio recognition result of the target language into the semantic detection model of the target language, and outputting the semantic analysis result; According to the semantic analysis result, determine whether the target audio contains sensitive information.
4. The cross-language audio detection method according to claim 1, characterized in that: The cross-language audio recognition model is an audio recognition model that is obtained by fine-tuning a trained large language model and can recognize audio in multiple languages.
5. The cross-language audio detection method according to claim 4, characterized in that: The cross-language audio recognition model is an audio recognition model that is obtained by fine-tuning a trained large language model and can recognize audio in multiple languages; The non-target audio is input into a cross-language audio recognition model, and a semantic analysis is performed on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information, including: Inputting the non-target audio into a cross-language audio recognition model, and outputting an audio recognition result in a single language; Inputting the audio recognition result of a single language output by the cross-language audio recognition model into a semantic detection model of the corresponding language, and outputting a semantic analysis result; According to the semantic analysis result, determine whether the non-target audio contains sensitive information.
6. The cross-language audio detection method according to any one of claims 1 to 5, characterized in that: Get the audio to be detected, including: In response to an upload request or a share request for an audio file, the audio to be detected is obtained.
7. A cross-language audio detection device, characterized in that: include: An audio acquisition module, used to acquire the audio to be detected; A language recognition module, used to perform language recognition on the audio to be detected; A target language audio recognition module, for inputting the target audio into a target language audio recognition model if the audio to be detected is a target audio of the target language, and performing semantic analysis on the audio recognition result of the target language audio recognition model to determine whether the target audio contains sensitive information; A cross-language audio recognition module, for inputting the non-target audio into a cross-language audio recognition model if the audio to be detected is a non-target audio in a non-target language, and performing semantic analysis on the audio recognition result of the cross-language audio recognition model to determine whether the non-target audio contains sensitive information; The audio spectrum information detection module is used to input the target audio and / or the non-target audio that has not been identified with sensitive information into a pre-trained audio spectrum information detection model, and output an audio detection result of whether the target audio and / or the non-target audio is abnormal audio.
8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the cross-language audio detection method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-language audio detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the cross-language audio detection method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Speech recognition method, device, terminal and storage medium
CN111261144A
Illegal audio detection method and device, electronic equipment and storage medium
CN116705031A
Audio recognition method and multi-task audio recognition model training method
CN116913286A
Audio recognition method and device, equipment and medium
CN118098200A