Illegal audio detection method, device, electronic device and computer-readable storage medium

By sampling the spectrum feature of the audio content and matching the phoneme string, the target phoneme sequence is generated and converted into text, the problem of large errors in audio to text in the prior art is solved, and the accuracy of violation detection is improved.

CN115440194BActive Publication Date: 2025-05-13BEIJING KNOWNSEC INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211064103.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-05-13
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

The prior art converts audio content into text for illegal word detection, with large errors, resulting in low accuracy of the audit results.

Method used

By sampling and drawing frames of the detected audio, multiple consecutive spectrum features are obtained and converted into phoneme strings. The previous and subsequent text matching is performed based on the sequential relationship between spectrum features, the target phoneme sequence is obtained, and it is converted into target text for violation word detection.

Benefits of technology

Reduce errors when converting audio content into text, improve the accuracy of audit results, and make audio violation detection more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440194B_ABST
    Figure CN115440194B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention proposes a method, device, electronic device and computer-readable storage medium for detecting illegal audio, which belongs to the field of speech recognition technology. The method includes: sampling and framing the audio to be detected to obtain multiple continuous spectral features, converting each spectral feature into at least one phoneme string, and based on the sequential relationship between the spectral features, matching all the phoneme strings of all the spectral features with the context to obtain a target phoneme sequence, so that all the phoneme strings in the target phoneme sequence have common language characteristics, thereby converting the target phoneme sequence into a target text, and detecting illegal words on the target text to determine whether the audio to be detected is illegal, which can greatly improve the accuracy of illegal audio detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a method, device, electronic device and computer-readable storage medium for detecting illegal audio. Background Art

[0002] With the widespread popularity of audio and video on Internet platforms such as entertainment social platforms (for example, mainstream self-media platforms, news websites, and various audio social platforms), the security of audio and video content has received special attention.

[0003] The supervision and review of audio content is an important part of ensuring the safety of audio content. Currently, the supervision and review methods of audio content mainly include manual review and machine review. Manual review, that is, manual review of audio content for violations, is costly and inefficient. Machine review mainly converts the audio content to be reviewed into text through a hybrid language model, and detects illegal words in combination with a vocabulary. However, the method of converting the audio content to be reviewed into text through a hybrid language model has a large error, resulting in low accuracy of the review results. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, electronic device and computer-readable storage medium for detecting illegal audio, which can reduce the error when converting the audio content to be reviewed into text and improve the accuracy of the review results.

[0005] In order to achieve the above purpose, the technical solution adopted by the embodiment of the present invention is as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for detecting illegal audio, the method comprising:

[0007] The audio to be detected is sampled and framed to obtain multiple continuous spectrum features;

[0008] For each of the spectral features, convert the spectral features into phonemes to obtain at least one phoneme string;

[0009] Based on the sequential relationship between the spectral features, context matching is performed on all the phoneme strings of all the spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string of each of the spectral features;

[0010] The target phoneme sequence is converted into a target text, and illegal word detection is performed on the target text to determine whether the audio to be detected violates the rules.

[0011] Furthermore, the step of performing context matching on all phoneme strings of all the spectral features based on the sequential relationship between the spectral features to obtain a target phoneme sequence includes:

[0012] Taking each phoneme string of the first spectrum feature as a starting point of the phoneme sequence, and taking the spectrum features of all the spectrum features of the audio to be detected except the first spectrum feature as matching spectrum features;

[0013] Using the preset language matching model, each phoneme string of each matching spectral feature is matched with the phoneme sequence composed of the phoneme strings of all the previous spectral features to obtain the target phoneme sequence.

[0014] Furthermore, the step of using a preset language matching model to match each phoneme string of each matching spectral feature with a phoneme sequence composed of phoneme strings of all spectral features of the preceding sequence to obtain a target phoneme sequence includes:

[0015] For each matching spectral feature, each phoneme string of the matching spectral feature is matched with a phoneme sequence composed of phoneme strings of all spectral features before the matching spectral feature, and a matching degree is calculated to obtain multiple matching sequences and a matching degree of each matching sequence;

[0016] According to the matching degree, the best matching sequence is determined from all matching sequences as a phoneme sequence composed of the phoneme strings of the matching spectral features and all the spectral features of the preceding sequence.

[0017] Further, the phoneme string includes an empty phoneme string indicating that the phoneme string is empty;

[0018] The step of converting the target phoneme sequence into a target text comprises:

[0019] Deduplication of adjacent and repeated phoneme strings in the target phoneme sequence, and removal of empty phoneme strings in the target phoneme sequence, to obtain a target phoneme sequence after deduplication;

[0020] The target phoneme sequence after deduplication is converted into text to obtain a target text.

[0021] Furthermore, the step of detecting illegal words on the target text to determine whether the audio to be detected is illegal includes:

[0022] Detect whether a phrase or a character in a preset illegal word library appears in the target text. If so, determine that the audio to be detected is illegal.

[0023] Furthermore, before the step of performing frame sampling on the audio to be detected to obtain a plurality of continuous frequency spectrum features, the method further includes:

[0024] Convert the audio file to be detected into an audio file in wav format to obtain the audio file to be detected.

[0025] Furthermore, before the step of converting the audio file to be detected into an audio file in wav format to obtain the audio to be detected, the method further includes:

[0026] Perform audio extraction on the video file to be detected to obtain the audio file to be detected.

[0027] In a second aspect, an embodiment of the present invention provides an illegal audio detection device, the device comprising a feature extraction module, a phoneme conversion module, a matching module and a detection module:

[0028] The feature extraction module is used to sample and extract frames of the audio to be detected to obtain multiple continuous frequency spectrum features;

[0029] The phoneme conversion module is used to convert the spectrum feature into phonemes for each spectrum feature to obtain at least one phoneme string;

[0030] The matching module is used to perform context matching on all the phoneme strings of all the spectral features based on the sequential relationship between the spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string for each of the spectral features;

[0031] The detection module is used to convert the target phoneme sequence into a target text, and perform illegal word detection on the target text to determine whether the audio to be detected violates the rules.

[0032] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the illegal audio detection method as described in the first aspect.

[0033] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the illegal audio detection method as described in the first aspect is implemented.

[0034] The illegal audio detection method, device, electronic device and computer-readable storage medium provided by the embodiments of the present invention obtain multiple continuous spectral features of the audio to be detected, convert each spectral feature into at least one phoneme string, and then perform context matching on the phoneme strings of all spectral features according to the sequential relationship between the spectral features to obtain a target phoneme sequence. All phoneme strings in the target phoneme sequence have common language characteristics, which can greatly reduce the error of the target phoneme sequence, and then convert the target phoneme sequence into a target text, and use the target text to detect illegal words, so that the audio violation detection can be more accurate.

[0035] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 A block diagram of an illegal audio detection system provided by an embodiment of the present invention is shown.

[0038] Figure 2 One of the flow charts of the illegal audio detection method provided by an embodiment of the present invention is shown.

[0039] Figure 3 Shows Figure 2 Schematic diagram of the process of some sub-steps of step S16.

[0040] Figure 4 Shows Figure 3 Schematic diagram of the process of some sub-steps of step S162.

[0041] Figure 5 Shows Figure 2 Schematic diagram of the process of some sub-steps of step S18.

[0042] Figure 6 The second flowchart of the illegal audio detection method provided by the embodiment of the present invention is shown.

[0043] Figure 7 A block diagram of an illegal audio detection device provided by an embodiment of the present invention is shown.

[0044] Figure 8 A block diagram of an electronic device provided by an embodiment of the present invention is shown.

[0045] Figure numerals: 100 - illegal audio detection system; 110 - client; 120 - electronic device; 130 - illegal audio detection device; 140 - feature extraction module; 150 - phoneme conversion module; 160 - matching module; 170 - detection module. DETAILED DESCRIPTION

[0046] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0047] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0048] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0049] The supervision and review of audio content is an important part of ensuring the safety of audio content. Currently, the supervision and review methods of audio content mainly include manual review and machine review. Manual review, that is, manual review of audio content for violations, is costly and inefficient. Machine review mainly converts the audio content to be reviewed into text through a hybrid language model, and detects illegal words in combination with a vocabulary.

[0050] However, the existing machine review method is difficult to deploy and has high technical barriers. After converting the audio content to be reviewed into a phoneme string through a hybrid language model, the phoneme string is directly converted into text, resulting in large errors in audio-to-text conversion and low accuracy of the review results.

[0051] Based on the above considerations, an embodiment of the present invention provides a method for detecting illegal audio, which can reduce the error when converting the audio content to be reviewed into text and improve the accuracy of the review result. The illegal audio detection method is introduced below.

[0052] The illegal audio detection method provided by the embodiment of the present invention can be applied to Figure 1In the illustrated illegal audio detection system 100 , the system includes an electronic device 120 and a plurality of clients 110 , and the plurality of clients 110 can be communicatively connected with the electronic device 120 via a network.

[0053] The electronic device 120 may be a server or a terminal. The client 110 includes but is not limited to: a mobile phone, a tablet, an iPad, a personal computer, a notebook computer, a portable wearable device and other smart devices.

[0054] The user sends the recorded or produced audio content or video content to the electronic device through the client 110 .

[0055] The electronic device 120 is used to receive audio content or video content sent by each client 110, extract the video content to obtain audio content, and use the illegal audio detection method provided by the embodiment of the present invention to perform illegal detection on the audio content sent by the client 110 and the audio content extracted from the video content.

[0056] In one embodiment, reference Figure 2 The embodiment of the present invention provides a method for detecting illegal audio. In this embodiment, the method for detecting illegal audio is applied to Figure 1 Taking the electronic device 120 as an example, the illegal audio detection method may include the following steps.

[0057] S12, sampling and decimating the audio to be detected to obtain a plurality of continuous frequency spectrum features.

[0058] The audio to be detected can be an audio file sent by the client to the electronic device, or it can be an audio file extracted from a video file sent by the client to the electronic device, that is, the audio file to be detected is obtained by extracting audio from the video file. The spectrum feature is the audio feature of the audio to be detected. A spectrum feature is the spectrum feature of a segment of the audio to be detected. All spectrum features are connected in sequence to form the overall spectrum feature of the audio to be detected.

[0059] S14, for each spectral feature, convert the spectral feature into phonemes to obtain at least one phoneme string.

[0060] A phoneme string is a phoneme of a possible text string of a spectrum feature. After the spectrum feature is converted into phonemes, at least one phoneme string is obtained, that is, there are as many possible text strings corresponding to each spectrum feature as there are phoneme strings.

[0061] S16, based on the sequential relationship between the spectral features, context matching is performed on all the phoneme strings of all the spectral features to obtain a target phoneme sequence.

[0062] The target phoneme sequence includes a phoneme string for each spectral feature.

[0063] S18, converting the target phoneme sequence into a target text, and performing illegal word detection on the target text to determine whether the audio to be detected violates the rules.

[0064] In one embodiment, each spectral feature can be converted to phoneme using a pre-trained acoustic model to obtain at least one possible phoneme string for each spectral feature, and only one of the phoneme strings is a correct phoneme string. The acoustic model is used to convert the spectral feature into a phoneme string in combination with a dictionary.

[0065] Exemplarily, when the electronic device receives an audio file or a video file sent by any client, the audio file or the audio file extracted from the video file is processed into the audio to be detected (if the format is correct, the audio file or the extracted audio file is the audio to be detected). The audio is actually a superposition of a series of sine waves of different frequencies and phases, so sampling the audio to be detected can obtain the spectrum characteristics, and the sampling result is divided into multiple continuous spectrum characteristics in a frame extraction manner.

[0066] After obtaining all the spectral features of the audio to be detected, each spectral feature is input into the pre-trained acoustic model to perform phoneme conversion on the spectral features to obtain at least one phoneme string corresponding to each audio feature. Then, based on the order relationship between the spectral features, all the phoneme strings of all the spectral features are matched with the context to obtain the target phoneme sequence. The target phoneme sequence is converted into the target text, and the target text is detected for illegal words to determine whether the audio to be detected is illegal.

[0067] Compared with the existing machine review method of converting audio content into phoneme strings and then directly converting them into text, the above-mentioned illegal audio detection method provided by the embodiment of the present invention, after obtaining the phoneme string of each spectral feature, performs context matching based on all phoneme strings of all spectral features to determine the target phoneme sequence, so that all phoneme strings in the target phoneme sequence have common language characteristics, which can greatly reduce the error of the target phoneme sequence, and thus make the audio violation detection more accurate.

[0068] The method of performing context matching on all phoneme strings to obtain the target phoneme sequence can be flexibly set. For example, the target phoneme sequence can be obtained according to a preset rule, or the target phoneme sequence can be obtained by machine learning, which is not specifically limited in this embodiment. Figure 3 , the above step S16 may include the following steps.

[0069] S161, taking each phoneme string of the first spectrum feature as a starting point of the phoneme sequence, and taking the spectrum features of all the spectrum features of the audio to be detected except the first spectrum feature as matching spectrum features.

[0070] S162, using a preset language matching model, matching each phoneme string of each matching spectral feature with a phoneme sequence composed of phoneme strings of all preceding spectral features to obtain a target phoneme sequence.

[0071] For example, suppose the audio to be detected has three spectral features P1, P2 and P3, and the phoneme string of P1 is P 11 , the phoneme string of P2 includes P 21 and P 22 , the phoneme string of P3 includes P 31 and P 32 .P 11 As the starting point of the phoneme sequence, P2 and P3 are both matching spectral features. 11 Do the matching and get the phoneme sequence {P 11 , P 21} and {P 11 , P 22 For the spectral feature P3, the phoneme string of P3 is matched with the phoneme sequence composed of the phoneme strings of P2 and P1 to obtain the phoneme sequence {P 11 , P 21 , P 31}、{P 11 , P 21 , P 32}、{P 11 , P 22 , P 31} and {P 11 , P 22 , P 32}, the optimal phoneme sequence among the four phoneme sequences is the target phoneme sequence.

[0072] Furthermore, in order to reduce the amount of calculation when there are multiple spectral features, the language matching model in step S162 can be obtained by training a hidden Markov model (HMM). During training, each audio in the training sample is converted into multiple continuous spectral features, and each spectral feature has a corresponding phoneme string. The spectral feature of each audio in the training sample is used as the iterative input, and the phoneme string corresponding to the spectral feature is used as the target output to train the HMM model to obtain a mature language matching model.

[0073] The calculation mechanism of the HMM model is: each state depends on a finite number of previous states, that is, if the current state is n+1, it depends on the previous n states to generate a random sequence when the n+1 state is reached. The language matching model obtained through training is more in line with the context (previous and next context) characteristics of language habits. When the matching spectral features are known, the language matching model can find the best matching phoneme string in the phoneme string of the current matching spectral feature and the phoneme sequence composed of the phoneme string of the previous spectral feature.

[0074] In S162, the language matching model may introduce a matching degree in the matching process of each matching spectral feature, so as to select the best phoneme sequence in the matching of each matching spectral feature, and only match the best phoneme sequence with the next matching spectral feature, so as to reduce the amount of calculation. Figure 4 , step S162 may include the following sub-steps.

[0075] S1621, for each matching spectral feature, match each phoneme string of the matching spectral feature with a phoneme sequence composed of phoneme strings of all spectral features before the matching spectral feature, and calculate the matching degree to obtain multiple matching sequences and the matching degree of each matching sequence.

[0076] S1622: According to the matching degree, determine the best matching sequence from all matching sequences as the phoneme sequence composed of the phoneme strings of the matching spectral features and all the preceding spectral features.

[0077] It should be understood that when the matching spectrum feature is the last spectrum feature of the audio to be detected, the optimal matching sequence corresponding to the matching spectrum feature is the target phoneme sequence.

[0078] For a complete audio to be detected, in most cases there are some time periods without speech, that is, there is a gap between two segments of speech. At this time, the phoneme strings of some spectral features of the audio to be detected may be empty, so there may be empty phoneme strings in all the phoneme strings of the target phoneme sequence. At the same time, the audio segments corresponding to adjacent spectral features may have repeated speech content, which also makes it possible for repeated phoneme strings to exist in the target phoneme sequence. Both empty phoneme strings and repeated phoneme strings will affect the target text, thereby affecting the accuracy of violation detection.

[0079] In order to improve the impact of empty phoneme strings and repeated phoneme strings on the target text, in a real-time manner, deduplication and deletion of empty phoneme strings are introduced during the process of converting the target phoneme sequence into the target text in step S18. Figure 5 , step S18 may include the following sub-steps.

[0080] S181, deduplicate adjacent and repeated phoneme strings in the target phoneme sequence, and remove empty phoneme strings in the target phoneme sequence to obtain a target phoneme sequence after deduplication.

[0081] S182, converting the target phoneme sequence after deleting duplicates into text to obtain a target text.

[0082] The accuracy of the target text can be improved through the above steps S181 and S182.

[0083] After obtaining the target file, the method of detecting the target text for illegal words can be flexibly set. For example, the target text can be identified to determine whether there are illegal words, or machine learning can be used to detect the target text. In one embodiment, a preset illegal word library can be used as an auxiliary to perform illegal detection. Referring to the figure, step S18 can also include step S183.

[0084] S183, detecting whether a phrase or a character in a preset illegal vocabulary appears in the target text, and if so, determining that the audio to be detected is illegal.

[0085] The target text may be first segmented to obtain phrases and characters, and then all the phrases and characters obtained by the segmentation may be matched with the phrases and characters in the violation vocabulary. As long as one character or one phrase is successfully matched, the detected audio is in violation.

[0086] In order to improve the detection accuracy, in one embodiment, referring to Figure 6 Before step S12, the illegal audio detection method provided by the embodiment of the present invention also includes step S11.

[0087] S11, converting the audio file to be detected into an audio file in wav format to obtain the audio file to be detected.

[0088] Through the above step S11, the compressed audio file (for example, MP3 format) is converted into an original WAV file, and the audio file is restored as much as possible to help improve the detection accuracy.

[0089] In one embodiment, the acoustic model and the language matching model can be combined into a CTC converter. At this time, the detection process of the audio to be detected includes: converting the audio to be detected into WAV format, sampling and extracting the converted audio to be detected to obtain spectral features, and passing the obtained multiple continuous spectral features into the CTC converter. The CTC converter uses the acoustic model to convert the acoustic features into high-order phoneme strings, and then uses the language matching model to perform softmax prediction regression processing on the phoneme strings of all spectral features to obtain the optimal target phoneme sequence; the target phoneme sequence is deduplicated and the empty phoneme strings are deleted and converted into the target text, and then the target text is detected with the help of the illegal word library table. If illegal text is detected, the audio to be detected is determined to be illegal audio.

[0090] Among them, the CTC converter can be expressed as, CTC(x)=softmax(W·Encoder(x)).

[0091] For example, if the target phoneme sequence is %%dd%%e%ee%%pp%%x%x, where % represents an empty phoneme string and dd is a repeated d, the target phoneme sequence after deduplication and removal of the empty phoneme string is 'deepxx'.

[0092] The illegal audio detection method provided by the embodiment of the present invention mainly uses the acoustic model and the language matching model to obtain the target phoneme sequence of the frequency spectrum to be detected, and all the phoneme strings of the target phoneme sequence have language characteristics, which can greatly reduce the error of the target phoneme sequence, thereby making the illegal audio detection more accurate. It realizes the identification of illegal and sensitive audio in a fast, low-cost, and artificial intelligence manner, effectively improving the accuracy and feasibility of illegal audio.

[0093] Based on the concept of the above-mentioned illegal audio detection method, in one implementation, referring to Figure 7 The embodiment of the present invention further provides a violation audio detection device 130, which can be applied to Figure 1 In the electronic device 120, the illegal audio detection device 130 may include a feature extraction module 140, a phoneme conversion module 150, a matching module 160 and a detection module 170.

[0094] The feature extraction module 140 is used to sample and extract frames of the audio to be detected to obtain a plurality of continuous frequency spectrum features.

[0095] The phoneme conversion module 150 is used to convert the spectrum feature into phonemes for each spectrum feature to obtain at least one phoneme string.

[0096] The matching module 160 is used to perform context matching on all phoneme strings of all spectral features based on the sequential relationship between the spectral features to obtain a target phoneme sequence.

[0097] The target phoneme sequence includes a phoneme string for each spectral feature.

[0098] The detection module 170 is used to convert the target phoneme sequence into a target text and perform illegal word detection on the target text to determine whether the audio to be detected violates the rules.

[0099] In the above-mentioned illegal audio detection device 130, after the feature extraction module 140 obtains multiple continuous spectral features of the audio to be detected, the phoneme conversion module 150 converts each spectral feature into at least one phoneme string, and then the matching module 160 performs context matching on the phoneme strings of all spectral features according to the sequential relationship between the spectral features to obtain a target phoneme sequence. All phoneme strings in the target phoneme sequence have common language characteristics, which can greatly reduce the error of the target phoneme sequence. Then the detection module 170 converts the target phoneme sequence into a target text, and uses the target text to detect illegal words, thereby making the audio violation detection more accurate.

[0100] For the specific definition of the illegal audio detection device 130, please refer to the definition of the illegal audio detection method above, which will not be repeated here. Each module in the above-mentioned illegal audio detection device 130 can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device 120 in the form of hardware, or can be stored in the memory of the electronic device 120 in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0101] In one embodiment, an electronic device 120 is provided, which may be a server, and its internal structure diagram may be as shown in the figure. The electronic device 120 includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the electronic device 120 is used to provide computing and control capabilities. The memory of the electronic device 120 includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device 120 is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, an operator network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, the illegal audio detection method improved as in the above-mentioned embodiment is implemented.

[0102] Figure 8The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the electronic device 120 to which the solution of the present invention is applied. The specific electronic device 120 may include Figure 8 More or fewer components may be shown, or certain components may be combined, or may have a different arrangement of components.

[0103] In one embodiment, the illegal audio detection device provided by the present invention can be implemented in the form of a computer program. The computer program can be used in Figure 8 The memory of the electronic device 120 may store various program modules constituting the illegal audio detection device 130, for example, Figure 7 The feature extraction module 140, the phoneme conversion module 150, the matching module 160 and the detection module 170 are shown. The computer program composed of various program modules enables the processor to execute the steps of the illegal audio detection method described in this specification.

[0104] For example, Figure 8 The electronic device 120 shown may be Figure 7 The feature extraction module 140 in the illegal audio detection device 130 shown in the figure performs step S12. The electronic device 120 can perform step S14 through the phoneme conversion module 150. The electronic device 120 can perform step S16 through the matching module 160. The electronic device 120 can perform step S18 through the detection module 170.

[0105] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program: sampling and framing the audio to be detected to obtain a plurality of continuous spectral features; for each spectral feature, converting the spectral feature into phonemes to obtain at least one phoneme string; based on the sequential relationship between the spectral features, performing context matching on all phoneme strings of all spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string for each of the spectral features; converting the target phoneme sequence into a target text, and performing illegal word detection on the target text to determine whether the audio to be detected is illegal.

[0106] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: sampling and framing the audio to be detected to obtain a plurality of continuous spectral features; for each spectral feature, converting the spectral feature into phonemes to obtain at least one phoneme string; based on the sequential relationship between the spectral features, context matching is performed on all phoneme strings of all spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string for each of the spectral features; converting the target phoneme sequence into a target text, and performing illegal word detection on the target text to determine whether the audio to be detected violates the rules.

[0107] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0108] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0109] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0110] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting illegal audio, characterized in that: The method comprises: The audio to be detected is sampled and framed to obtain multiple continuous spectrum features; For each of the spectral features, the spectral features are converted into phonemes to obtain at least one phoneme string; wherein one of the phoneme strings is a phoneme of a text string of the spectral feature; Based on the sequential relationship between the spectral features, context matching is performed on all the phoneme strings of all the spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string of each of the spectral features; Convert the target phoneme sequence into a target text, and perform illegal word detection on the target text to determine whether the audio to be detected violates the rules; The step of performing context matching on all the phoneme strings of all the spectral features based on the sequential relationship between the spectral features to obtain a target phoneme sequence includes: Taking each phoneme string of the first spectrum feature as a starting point of the phoneme sequence, and taking the spectrum features of all the spectrum features of the audio to be detected except the first spectrum feature as matching spectrum features; Using the preset language matching model, each phoneme string of each matching spectral feature is matched with the phoneme sequence composed of the phoneme strings of all the preceding spectral features to obtain the target phoneme sequence.

2. The illegal audio detection method according to claim 1, characterized in that: The step of using a preset language matching model to match each phoneme string of each matching spectral feature with a phoneme sequence composed of phoneme strings of all spectral features of the preceding sequence to obtain a target phoneme sequence includes: For each matching spectral feature, each phoneme string of the matching spectral feature is matched with a phoneme sequence composed of phoneme strings of all spectral features before the matching spectral feature, and a matching degree is calculated to obtain multiple matching sequences and a matching degree of each matching sequence; According to the matching degree, the best matching sequence is determined from all matching sequences as a phoneme sequence composed of the phoneme strings of the matching spectral features and all the spectral features of the preceding sequence.

3. The illegal audio detection method according to claim 1 or 2, characterized in that: The phoneme string includes an empty phoneme string indicating that the phoneme string is empty; The step of converting the target phoneme sequence into a target text comprises: Deduplication of adjacent and repeated phoneme strings in the target phoneme sequence, and removal of empty phoneme strings in the target phoneme sequence, to obtain a target phoneme sequence after deduplication; The target phoneme sequence after deduplication is converted into text to obtain a target text.

4. The illegal audio detection method according to claim 1 or 2, characterized in that: The step of detecting illegal words on the target text to determine whether the audio to be detected is illegal includes: Detect whether a phrase or a character in a preset illegal word library appears in the target text, and if so, determine that the audio to be detected is illegal.

5. The illegal audio detection method according to claim 1 or 2, characterized in that: Before the step of performing frame sampling on the audio to be detected to obtain a plurality of continuous frequency spectrum features, the method further comprises: Convert the audio file to be detected into an audio file in wav format to obtain the audio file to be detected.

6. The illegal audio detection method according to claim 5, characterized in that: Before the step of converting the audio file to be detected into an audio file in wav format to obtain the audio to be detected, the method further includes: Perform audio extraction on the video file to be detected to obtain the audio file to be detected.

7. A device for detecting illegal audio, characterized in that: The device comprises a feature extraction module, a phoneme conversion module, a matching module and a detection module: The feature extraction module is used to sample and extract frames of the audio to be detected to obtain multiple continuous frequency spectrum features; The phoneme conversion module is used to convert the spectrum feature into phonemes for each spectrum feature to obtain at least one phoneme string; wherein one phoneme string is a phoneme of a text string of the spectrum feature; The matching module is used to perform context matching on all the phoneme strings of all the spectral features based on the sequential relationship between the spectral features to obtain a target phoneme sequence; wherein the target phoneme sequence includes a phoneme string for each of the spectral features; The detection module is used to convert the target phoneme sequence into a target text, and perform illegal word detection on the target text to determine whether the audio to be detected violates the rules; The matching module is further used for: Taking each phoneme string of the first spectrum feature as a starting point of the phoneme sequence, and taking the spectrum features of all the spectrum features of the audio to be detected except the first spectrum feature as matching spectrum features; Using the preset language matching model, each phoneme string of each matching spectral feature is matched with the phoneme sequence composed of the phoneme strings of all the preceding spectral features to obtain the target phoneme sequence.

8. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the illegal audio detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the illegal audio detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Voice recognition method, voice recognition system, computer equipment and computer readable storage medium

    CN107871499A

  • Method and device for detecting violation terms

    CN109508402A