Audio recognition method, related apparatus and storage medium
Patent Information
- Application Number
- CN202310721455.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-06-16
AI Technical Summary
[0003]目前,现有的语音检测技术一般是使用语音较单一的语音特征对语音检测模型进行训练,然后通过训练后的语音检测模型基于较单一的语音特征对深度伪造的语音进行语音检测,这样会使得语音检测模型学习到的知识局限性较大,或使得语音特征包含一些不重要的信息,造成语音检测模型学习偏差,从而导致语音检测模型对语音检测的准确性较低
[0032]在一种可能的设计中,上述芯片系统还包括通信接口,用于输入和/或输出信息。
Smart Images

Figure CN117153149B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically to an audio recognition method, related apparatus, and storage medium, wherein the related apparatus includes an audio recognition device, a computing device (also referred to as a computer device or computer), a computer program product, and a chip system, etc. Background Technology
[0002] With the development of technology, voice is being used more and more widely. Voice carries human language and speaker identity information, and may be used to imitate the voice of the target speaker to deceive human or machine hearing. Therefore, voice detection technology can be used to detect deepfake voice.
[0003] Currently, existing speech detection technologies generally use relatively simple speech features to train speech detection models, and then use the trained speech detection models to detect deepfake speech based on these simple speech features. This results in the speech detection models having limited knowledge or the speech features containing some unimportant information, causing learning bias in the speech detection models and thus leading to low accuracy in speech detection. Summary of the Invention
[0004] This application provides an audio recognition method, related device, and storage medium that can improve the accuracy of audio forgery detection, resulting in a low false recognition rate and higher recognition sensitivity.
[0005] In a first aspect, embodiments of this application provide an audio recognition method, the method comprising:
[0006] Obtain the audio to be processed;
[0007] The audio to be processed is frequency domain transformed to obtain the target frequency domain acoustic features of the audio to be processed, and the audio to be processed is feature extracted to obtain the target audio attribute features of the audio to be processed.
[0008] The target frequency domain acoustic features and the target audio attribute features are fused to obtain the target audio features;
[0009] The target audio features are identified to obtain an identification result, which is used to indicate whether the audio to be processed is fake.
[0010] In one embodiment, before identifying the target audio features using an audio recognition model, the method further includes:
[0011] Obtain the audio training set;
[0012] Extract the frequency domain acoustic features and audio attribute features of each audio training sample in the audio training set;
[0013] A feature set is obtained, which includes multiple audio features, each of which is obtained by fusing the frequency domain acoustic features and audio attribute features of the same audio training sample;
[0014] The audio recognition model is trained based on the feature set to obtain the trained audio recognition model.
[0015] In one implementation, training the audio recognition model based on the feature set to obtain the trained audio recognition model includes:
[0016] Obtain the true labels of the audio training samples;
[0017] Identify each audio feature in the feature set to obtain the predicted label of each audio training sample in the audio training set;
[0018] Obtain the difference information between the real label and the predicted label;
[0019] Based on the difference information, the parameters of the audio recognition model are adjusted to obtain the trained audio recognition model.
[0020] In one implementation, the step of extracting features from the audio to be processed to obtain the target audio attribute features of the audio to be processed includes:
[0021] The audio to be processed is subjected to multi-layer convolutional feature extraction through a pre-trained speech model to obtain multi-layer attribute features.
[0022] The multi-layer attribute features are weighted and fused to obtain the target audio attribute features of the audio to be processed.
[0023] Secondly, embodiments of this application provide an audio recognition device that implements the audio recognition method corresponding to the first aspect described above. The function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above function, and the modules can be software and / or hardware.
[0024] In one embodiment, the audio recognition device includes:
[0025] The input / output module is configured to acquire the audio to be processed;
[0026] The processing module is configured to perform frequency domain transformation on the audio to be processed to obtain the target frequency domain acoustic features of the audio to be processed, and to extract features from the audio to be processed to obtain the target audio attribute features of the audio to be processed; to fuse the target frequency domain acoustic features and the target audio attribute features to obtain the target audio features; and to identify the target audio features to obtain the identification result, which is used to indicate whether the audio to be processed is fake.
[0027] Thirdly, embodiments of this application provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio recognition method provided in the first aspect.
[0028] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the audio recognition method as described in the first aspect.
[0029] Fifthly, embodiments of this application provide a computer program product including instructions, the computer program product including program instructions, which, when run on a computer or processor, cause the computer to execute the audio recognition method provided in the first aspect.
[0030] In a sixth aspect, embodiments of this application provide a chip including a processor coupled to a transceiver of a terminal device, for executing the audio recognition method provided in the first aspect of embodiments of this application.
[0031] In a seventh aspect, embodiments of this application provide a chip system including a processor for supporting a terminal device in implementing the functions involved in the first aspect above, such as generating or processing information involved in the audio recognition method provided in the first aspect above.
[0032] In one possible design, the aforementioned chip system also includes a communication interface for inputting and / or outputting information.
[0033] In one possible design, the aforementioned chip system also includes a memory for storing program instructions and data necessary for the terminal device. The chip system can be composed of chips or may include chips and other discrete components.
[0034] Compared to existing technologies, in this embodiment, the acquired audio to be processed is first transformed in the frequency domain to obtain target frequency domain acoustic features with audio frequency domain representation, and features are extracted from the audio to be processed to obtain target audio attribute features. Then, the target frequency domain acoustic features and the target audio attribute features are fused to obtain the final target audio features used for identification. Since the target frequency domain acoustic features can represent the audio frequency domain features of the audio to be processed, and the target audio attribute features can represent multi-dimensional audio information, which can contain richer audio features, the target audio features obtained after fusing the target frequency domain acoustic features and the target audio attribute features will contain a richer and more comprehensive audio representation. Therefore, compared to the simple identification based on a single audio feature in related technologies, the embodiments of this application, when the audio recognition model identifies the target audio feature, will obtain richer and more comprehensive audio features because the target audio feature has the characteristic of containing richer and more comprehensive audio representation. Thus, the embodiments of this application can accurately determine whether the audio to be processed is fake by identifying the target audio feature containing richer and more comprehensive audio representation, thereby improving the accuracy of detecting forgery traces in audio, resulting in a low false recognition rate and higher recognition sensitivity. Attached Figure Description
[0035] The objectives, features, and advantages of the embodiments of this application will become readily understood by referring to the accompanying drawings and the detailed description of the embodiments. Wherein:
[0036] Figure 1 This is a schematic diagram of an audio recognition system based on the audio recognition method in this application.
[0037] Figure 2 This is a schematic flowchart of an audio recognition method according to an embodiment of this application;
[0038] Figure 3 This is a schematic diagram of audio frame segmentation for an audio recognition method according to an embodiment of this application;
[0039] Figure 4 This is a schematic diagram of another audio framing method for audio recognition according to an embodiment of this application;
[0040] Figure 5 This is a schematic diagram of an audio recognition model training method according to an embodiment of this application;
[0041] Figure 6 This is a schematic diagram illustrating the application of an audio recognition model of the audio recognition method according to an embodiment of this application;
[0042] Figure 7 This is a schematic diagram of an optimized audio recognition model for an audio recognition method according to an embodiment of this application;
[0043] Figure 8 This is another schematic flowchart of the audio recognition method according to an embodiment of this application;
[0044] Figure 9 This is another schematic flowchart of the audio recognition method according to an embodiment of this application;
[0045] Figure 10 This is another schematic flowchart of the audio recognition method according to an embodiment of this application;
[0046] Figure 11 This is a schematic diagram of the structure of an audio recognition device according to an embodiment of this application;
[0047] Figure 12 This is a schematic diagram of the structure of a computing device according to an embodiment of this application;
[0048] Figure 13 This is a schematic diagram of the structure of a mobile phone in one embodiment of this application;
[0049] Figure 14 This is a schematic diagram of a server structure in one embodiment of this application.
[0050] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0051] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects (e.g., the first confidence interval and the second confidence interval represent different confidence intervals, and so on), and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules appearing in the embodiments of this application is merely a logical division; in actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not performed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through interfaces, indirect couplings between modules, or electrical or other similar forms of communication connections, none of which are limited in the embodiments of this application. Moreover, the modules or sub-modules described as separate components may or may not be physically separate, may or may not be physical modules, or may be distributed across multiple circuit modules. Some or all of these modules can be selected according to actual needs to achieve the purpose of the embodiments of this application.
[0052] This application also provides an audio recognition method, related apparatus, and storage medium, which can be applied to an audio recognition system capable of detecting audio forgery in scenarios. The audio recognition system may include an audio recognition device, which can at least be used to recognize audio to be processed to determine whether the audio is forged. The audio recognition device may be an application that identifies the audio to be processed based on the audio to be processed to obtain a recognition result determining whether the audio is forged, or a server or terminal device that has installed an application that identifies the audio to be processed based on the audio to be processed to obtain a recognition result determining whether the audio is forged. The application may be, for example, an audio recognition model. The audio recognition device may also be a server or terminal device that has deployed an audio recognition model. This application example uses a server that has deployed an audio recognition model to identify the audio to be processed and obtain a recognition result determining whether the audio is forged. For a terminal device that has deployed an audio recognition model to identify the audio to be processed, the method can be referenced to the method for identifying the audio to be processed on the server, and will not be elaborated further.
[0053] The solutions provided in this application involve technologies such as Artificial Intelligence (AI) and Machine Learning (ML), and are specifically illustrated through the following embodiments:
[0054] AI, or Artificial Intelligence, refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, Artificial Intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine capable of reacting in a manner similar to human intelligence. Artificial Intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0055] AI technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0056] In existing technologies, audio recognition typically uses relatively simple audio features to train a model, and then the trained model performs recognition based on these simple audio features. This results in a limited range of knowledge learned by the model, or the audio features may contain some unimportant information, causing learning bias in the model and leading to a low accuracy rate for audio recognition after training.
[0057] Compared to existing technologies, in this embodiment, long-duration audio can be segmented into frames and multiple thresholds can be set to detect the audio. Furthermore, the audio features extracted from frequency domain acoustic features and pre-trained speech models are fused to obtain audio features containing richer and more comprehensive audio representations. Based on these richer and more comprehensive audio features, the audio recognition model can accurately determine whether the audio to be processed is forged. This can improve the overall performance of the audio recognition model to a certain extent, enhance the accuracy of detecting forgery traces in audio, and result in a lower false recognition rate and higher recognition sensitivity.
[0058] In some implementations, taking the audio recognition device integrated into server 10 as an example, refer to... Figure 1 The audio recognition method provided in this application embodiment can be based on Figure 1The illustration shows an implementation of an audio recognition system. This audio recognition system may include a server 10 and a terminal device 20. The server 10 may include an audio recognition device, and the server 10 may deploy an audio recognition program or an audio recognition model, such as an audio recognition model trained using machine learning methods.
[0059] Taking a telephone fraud prevention application scenario as an example, server 10 can receive the audio of a call sent by terminal device 20, using it as the audio to be processed. Server 10 can then extract the target frequency domain acoustic features of the audio to be processed, and extract the target audio attribute features through a pre-trained speech model. Server 10 can then fuse the target frequency domain acoustic features and target audio attribute features to obtain the final target audio features used for identification, which contain a richer and more comprehensive audio representation. Finally, server 10 can use an audio recognition model to identify the target audio features and determine whether the audio to be processed is fraudulent, i.e., whether the call audio is fraudulent. Because identification can be performed using target audio features containing a richer and more comprehensive audio representation to accurately determine whether the audio to be processed is fraudulent, the accuracy of detecting fraudulent traces in audio can be improved, resulting in a lower false recognition rate and higher recognition sensitivity. If the call audio is fraudulent, server 10 can send a prompt message to terminal device 20 so that terminal device 20 can display the prompt message to remind the user that the current call is a telephone fraud, avoiding unnecessary losses for the user and improving call security.
[0060] It should be noted that the server involved in the embodiments of this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0061] The terminal devices involved in the embodiments of this application can be devices that provide voice and / or data connectivity to users, handheld devices with wireless connectivity, or other processing devices connected to a wireless modem. Examples include mobile phones (or "cellular" phones) and computers with mobile terminals, such as portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with a wireless access network. Examples include Personal Communication Service (PCS) phones, cordless phones, Session Initialization Protocol (SIP) phones, Wireless Local Loop (WLL) stations, Personal Digital Assistants (PDAs), and other devices.
[0062] Reference Figure 2 , Figure 2 This is a flowchart illustrating an audio recognition method provided in an embodiment of this application. The method can be executed by an audio recognition device and can be applied to audio recognition scenarios requiring authentication. It involves acquiring the target frequency domain acoustic features and target audio attribute features of the audio to be processed, fusing these features to obtain target audio features containing a richer and more comprehensive audio representation, and then performing recognition based on these target audio features to obtain the recognition result. The audio recognition method may include steps S101 to S104:
[0063] Step 101: The audio recognition device acquires the audio to be processed.
[0064] The audio to be processed can be audio that needs to be identified, to determine whether the audio is fake, or whether the speaker's voice contained in the audio is fake, etc. The audio to be processed can include the speaker's voice (which can be simply referred to as human voice). The audio to be processed can be audio that has undergone preprocessing such as human voice detection, noise reduction, filtering or framing, or audio that has not been preprocessed, and the specifics are not limited here.
[0065] The methods for obtaining the audio to be processed may include, but are not limited to, the following: the audio recognition device obtains the audio to be recognized from a database used for storing audio, and obtains the audio to be processed; or, the audio recognition device can receive audio recorded by the user's speech sent by the terminal device, and obtain the audio to be processed; or, the audio recognition device can crawl the audio to be recognized from the Internet using web crawling technology, and obtain the audio to be processed, etc.
[0066] Since excessively long audio clips can affect recognition accuracy, to improve accuracy and avoid the negative impact of excessive audio length, the audio recognition device can preprocess the audio to be recognized, such as by segmenting it into frames, and use the preprocessed audio as the audio to be processed. In some embodiments, the audio recognition device acquires the audio to be processed, which may include:
[0067] S11. Obtain the original audio.
[0068] S12. If the length of the original audio is greater than the preset length threshold, the original audio is divided into frames to obtain multiple audio frames, and these multiple audio frames are used as audio to be processed.
[0069] S13. If the length of the original audio is less than or equal to the preset length threshold, then the original audio is used as the audio to be processed.
[0070] Specifically, the audio recognition device can obtain raw audio from a database or receive raw audio recorded by the user from a terminal device; the audio recognition device determines whether the length of the raw audio is greater than a preset length threshold, which can be flexibly set according to actual needs and is not limited here.
[0071] If the length of the original audio exceeds the preset length threshold, it indicates that the original audio is too long and belongs to long-duration audio. At this time, the audio recognition device can perform frame-segmentation processing on the original audio to obtain multiple audio frames. These multiple audio frames are used as audio to be processed. By performing frame-segmentation processing on long-duration audio, accurate and objective recognition of long-duration audio can be achieved in the future, thereby accurately determining whether the audio to be processed is fake.
[0072] The framing method can be set according to actual needs, for example, Figure 3 As shown, long audio segments can be extracted using a sliding window method with a preset step size to obtain multiple audio frames. The preset step size can be greater than 0 and less than the window length of the sliding window. For example, as... Figure 4 As shown, long audio segments can be extracted using a sliding window method with a step size of 0, resulting in multiple audio frames.
[0073] If the length of the original audio is less than or equal to the preset length threshold, it means that the original audio is short-time audio. In this case, the audio recognition device does not need to perform frame-by-frame processing on the original audio and can directly use the original audio as the audio to be processed.
[0074] It should be noted that, to reduce audio noise interference, the audio recognition device can also perform preprocessing such as noise reduction and filtering on the original audio to remove noise. When it is necessary to determine whether the speaker's voice in the audio is fake, the audio recognition device can also detect whether there is a human voice in the original audio. If no human voice is detected, the device can either not perform further recognition on the original audio or output a prompt message indicating that no human voice is present. If human voice is detected in the original audio, the audio recognition device can further determine whether the length of the original audio exceeds a preset length threshold. If the length of the original audio exceeds the preset length threshold, the device can perform frame segmentation on the original audio.
[0075] Step 102: The audio recognition device performs frequency domain transformation on the audio to be processed to obtain the target frequency domain acoustic features of the audio to be processed, and performs feature extraction on the audio to be processed to obtain the target audio attribute features of the audio to be processed.
[0076] After obtaining the audio to be processed, the audio recognition device can perform a series of signal processing steps, such as frequency domain transformation, to convert the audio from the time domain to the frequency domain. This allows the device to obtain the target frequency domain acoustic features of the audio in the frequency domain dimension. These target frequency domain acoustic features can be stored in the form of a feature matrix, a list, or other formats, which are not limited here. When the audio to be processed includes multiple audio frames, the audio recognition device can perform frequency domain transformation on each audio frame separately and obtain the target frequency domain acoustic features corresponding to each audio frame in the frequency domain dimension.
[0077] The target frequency domain acoustic features can be used to characterize the timbre and pitch of audio in the frequency domain dimension. The frequency domain acoustic features can be constant Q cepstral coefficients (CQCC), Mel-Frequency cepstral coefficients (MFCC), or Linear Frequency Cepstral Coefficients (LFCC), etc.
[0078] Furthermore, the audio recognition device can extract features from the audio to be processed to obtain the target audio attribute features. When the audio to be processed includes multiple audio frames, the audio recognition device can extract features from each audio frame separately to obtain the target audio attribute features corresponding to each audio frame. These target audio attribute features can be used to characterize multi-dimensional audio information such as semantics, timbre, and pitch. These target audio attribute features can be stored in the form of a feature matrix, a list, or other forms, without limitation here.
[0079] After obtaining the audio to be processed, the audio recognition device can extract features from the audio to be processed using a pre-trained speech model to obtain the target audio attribute features of the audio to be processed. The pre-trained speech model can be a trained wav2vec model or other speech models, which are not limited here.
[0080] Because pre-trained speech models have high processing efficiency and can improve the accuracy of feature extraction after training, they can be used to extract target audio attribute features of the audio to be processed in order to improve the accuracy and efficiency of target audio attribute feature acquisition. In some implementations, the audio recognition device extracts features from the audio to be processed to obtain the target audio attribute features, which may include:
[0081] S21. By using a pre-trained speech model to extract features layer by layer from the audio to be processed through multiple convolutional layers, multi-layer attribute features are obtained.
[0082] S22. Perform weighted fusion of multi-layer attribute features to obtain the target audio attribute features of the audio to be processed.
[0083] Specifically, by using a pre-trained speech model to extract features layer by layer from the audio to be processed through multiple convolutional layers, multi-layer attribute features are obtained. These multi-layer attribute features can be weighted and fused to obtain the target audio attribute features of the audio to be processed, or the last layer of attribute features can be directly used as the target audio attribute features of the audio to be processed. Based on the prior knowledge of audio feature extraction learned by the pre-trained speech model after training, not only is the accuracy of target audio attribute feature acquisition improved, but the obtained target audio attribute features can also better represent the audio.
[0084] Step 103: The audio recognition device fuses the target frequency domain acoustic features and the target audio attribute features to obtain the target audio features.
[0085] After obtaining the target frequency domain acoustic features and target audio attribute features of the audio to be processed, the audio recognition device can fuse these features to obtain richer and more comprehensive target audio features. When the audio to be processed includes multiple audio frames, the target frequency domain acoustic features and target audio attribute features corresponding to each audio frame can be obtained. In this case, the target frequency domain acoustic features and target audio attribute features corresponding to each audio frame can be concatenated along the feature depth dimension to obtain richer and more comprehensive target audio features.
[0086] One feature fusion method involves concatenating the frequency domain acoustic features and audio attribute features of the audio to be processed along the feature depth dimension to obtain the target audio features. For example, M×N dimensional frequency domain acoustic features and R×N audio attribute features can be concatenated along the feature depth dimension to obtain (M+R)×N target audio features. This makes the audio representation of the target audio features richer and more comprehensive, capable of fully representing all audio attributes. N represents the number of frames in the audio to be processed, M represents the feature depth of the frequency domain acoustic features, and R represents the feature depth of the audio attribute features.
[0087] Step 104: The audio recognition device identifies the target audio features and obtains the recognition result.
[0088] The identification results are used to indicate whether the audio to be processed is fake.
[0089] In some implementations, the identification result may further include a authenticity label and a confidence level. The authenticity label indicates whether the audio to be processed is forged. For example, if the authenticity label is 0, it indicates that the audio to be processed is not forged; if the authenticity label is 1, it indicates that the audio to be processed is forged. When it is necessary to detect whether the speaker's voice contained in the audio to be processed is the target user's real voice, the authenticity label is used to indicate whether the audio to be processed is the target user's real voice. For example, if the authenticity label is 0, it indicates that the audio to be processed is the target user's real voice; if the authenticity label is 1, it indicates that the audio to be processed is not the user's real voice. The confidence level may include the confidence level corresponding to the audio to be processed being forged and the confidence level corresponding to the audio to be processed not being forged.
[0090] Because the target audio features are integrated with the target frequency domain acoustic features and the target audio attribute features for identification, the audio representation of the target audio features is richer and more comprehensive, and can fully represent the audio attributes. Therefore, by identifying the target audio features containing richer and more comprehensive audio representations, it is possible to accurately determine whether the audio to be processed is fake, and improve the accuracy of detecting forgery traces in audio.
[0091] For telephone fraud prevention applications, audio recognition devices can identify the audio of a call. When the identification result indicates that the audio is a fake speaker's voice (i.e., fake audio), a prompt message can be output to remind the user that the current call is a telephone scam, allowing the user to stop the loss in time and avoid unnecessary losses, thus improving the security of the call.
[0092] Because audio recognition models have high processing efficiency and can learn prior knowledge of audio recognition after training, thus improving the accuracy of audio recognition, audio recognition devices can use trained audio recognition models to identify target audio features to obtain a first confidence level for the audio to be processed. The audio recognition device can then determine whether the audio to be processed is forged based on this first confidence level. In some embodiments, identifying target audio features to obtain a recognition result may include:
[0093] S41. Using an audio recognition model, identify the target audio features to obtain the first confidence level of the audio to be processed.
[0094] The first confidence level can be used to characterize the similarity between the audio to be processed and real audio (e.g., the real speaking voice of the target user). The audio recognition model can be a deep learning model with audio anti-spoofing capabilities. For example, the audio recognition model can be an audio anti-spoofing model, which can be a lookup-based convolutional neural network (LCNN) or a voiceprint model (ECAPA-TDNN), etc.
[0095] Before using an audio recognition model to identify target audio features, the audio recognition model can be pre-trained. To improve the robustness and generalization of the audio recognition model training, the audio recognition device can pre-train the audio recognition model using rich and comprehensive audio features, enabling the audio recognition model to learn richer and more comprehensive prior knowledge of audio features. In some embodiments, before using the audio recognition model to identify target audio features, the audio recognition model can be trained first, including:
[0096] a1. Obtain the audio training set.
[0097] The audio recognition device can obtain multiple audio training samples from a database to obtain an audio training set, or receive multiple audio training samples sent by a terminal device to obtain an audio training set, or crawl multiple audio training samples from the Internet to obtain an audio training set, etc. The audio training samples are real audio, such as the real voice of the target user.
[0098] a2. Extract the frequency domain acoustic features and audio attribute features of each audio training sample in the audio training set.
[0099] The audio recognition device can perform a series of signal processing steps, such as Fourier transform, on each audio training sample in the frequency training set to transform the audio training samples from the time domain to the frequency domain. This allows for the extraction of frequency domain acoustic features from the audio training samples. These frequency domain acoustic features can be stored in the form of a feature matrix, a list, or other formats, without limitation here. These frequency domain acoustic features can be used to characterize the timbre and pitch of the audio in the frequency domain. The frequency domain acoustic features can be CQCC, MFCC, or LFCC, etc.
[0100] Furthermore, the audio recognition device can extract audio attribute features of each audio training sample in the audio training set through the pre-trained speech model wav2vec. The audio attribute features can be used to characterize multi-dimensional audio information such as semantics, timbre, and pitch. These audio attribute features can be stored in the form of a feature matrix, a list, or other forms, without limitation here.
[0101] a3. Obtain the feature set, which includes multiple audio features. Each audio feature is obtained by fusing the frequency domain acoustic features and audio attribute features of the same audio training sample.
[0102] Audio recognition devices can acquire feature sets, which include multiple audio features. For example, after obtaining the frequency domain acoustic features and audio attribute features of the audio training samples, each audio feature can be fused based on the frequency domain acoustic features and audio attribute features of the same audio training sample, thereby obtaining richer and more comprehensive audio features.
[0103] a4. Train the audio recognition model based on the feature set to obtain the trained audio recognition model.
[0104] Audio recognition devices can train audio recognition models, such as LCNN, based on audio features to obtain trained audio recognition models. For example, for audio training samples from users A, B, and C, the frequency domain acoustic features and audio attribute features of user A's audio training samples can be fused to obtain user A's audio features; the frequency domain acoustic features and audio attribute features of user B's audio training samples can be fused to obtain user B's audio features; and the frequency domain acoustic features and audio attribute features of user C's audio training samples can be fused to obtain user C's audio features. Then, the audio recognition model learns the audio features of users A, B, and C respectively, so that it can accurately identify whether a certain audio is a forged audio from user A, user B, or user C. By training the audio recognition model through the fusion of frequency domain acoustic features and audio attribute features, the audio recognition model can learn richer and more comprehensive prior knowledge of audio features, improving its robustness and generalization ability.
[0105] Optionally, in some embodiments of this application, during the training of the audio recognition model, in order to improve the reliability of the audio recognition model training, the parameters of the audio recognition model can be adjusted by minimizing the difference between the real label and the predicted label, so that the prediction result of the audio recognition model can approach the real label. Specifically, the audio recognition model is trained based on the feature set to obtain the trained audio recognition model, including:
[0106] a41. Obtain the true labels of the audio training samples.
[0107] The real label can be the label of the speaker's voice corresponding to the audio training sample in the feature set. For example, if the audio training sample contains the real voice of user A, the real label is 0; if the audio training sample contains the fake voice of user A, the real label is 1.
[0108] a42. Identify each audio feature in the feature set to obtain the predicted label of each audio training sample in the audio training set.
[0109] The predicted label can be the confidence level of the speaker's voice predicted by the audio recognition model from the audio training samples, and the value of the predicted label can be from 0 to 1 (inclusive).
[0110] a43. Obtain the difference information between the real label and the predicted label; adjust the parameters of the audio recognition model based on the difference information to obtain the trained audio recognition model.
[0111] The difference information can be whether the real label and the predicted label are consistent or inconsistent, or the difference between the real label and the predicted label.
[0112] During the training of the audio recognition model, the audio recognition device can obtain the pre-set real labels of the audio training samples in the feature set, and identify (i.e. predict) each audio feature in the feature set through the audio recognition model to obtain the predicted labels (e.g., confidence) of each audio training sample in the audio training set. Then, the audio recognition device can compare the real labels with the predicted labels to obtain the difference information (e.g., difference) between the real labels and the predicted labels. At this time, a loss function can be constructed based on the difference information, and the parameters of the audio recognition model can be adjusted based on the loss function until the loss converges, that is, when the condition of minimizing the difference information is met, the training stops, and the trained audio recognition model is obtained.
[0113] This embodiment trains an audio recognition model by fusing audio features obtained from frequency domain acoustic features and audio attribute features. On the one hand, this allows the audio recognition model to have certain prior knowledge, preventing it from collapsing to unimportant features. On the other hand, it improves the representation performance of the audio recognition model, more comprehensively representing audio attributes, enabling the audio recognition model to learn more comprehensive knowledge, avoiding large-scale biases, and thus improving the robustness and generalization of the model. When encountering unknown audio, the audio recognition model can still maintain a high detection accuracy.
[0114] For example, such as Figure 5 As shown, taking the audio recognition model as an example of the audio authentication model, after obtaining the audio training set, the audio recognition device extracts the frequency domain acoustic features and audio attribute features of each audio training sample in the audio training set. The frequency domain acoustic features and audio attribute features based on the same audio training sample are fused to obtain multiple audio features. The multiple audio features constitute a feature set. The audio authentication model is trained based on the feature set to obtain the trained audio authentication model, which improves the robustness and generalization of the audio authentication model.
[0115] It should be noted that audio training samples can be short audio clips with a length less than or equal to a preset length threshold, or long audio clips with a length greater than the preset length threshold. The preset length threshold can be flexibly set according to actual needs and is not limited here. When the audio training sample is a long audio clip, it can be segmented into frames to ensure that the length of all audio clips in the training sample is less than or equal to the preset length threshold. Each audio training sample is assigned a corresponding label, which can be used to characterize whether the audio training sample represents the real voice of the target user (i.e., the speaker).
[0116] S42. Obtain the identification result based on the first confidence level.
[0117] After the audio recognition device identifies the target audio features and obtains the first confidence level through the trained audio recognition model, it can generate the recognition result of the audio to be processed based on the first confidence level. For example, the generated recognition result includes the confidence level and may also include true / false labels.
[0118] For example, such as Figure 6 As shown, taking an audio recognition model as an example of an audio counterfeiting detection model, the audio recognition device can perform frame-by-frame processing on the original audio to obtain audio to be processed containing multiple audio frames. It extracts the target frequency domain acoustic features corresponding to each audio frame in the audio to be processed in the frequency domain dimension, and extracts the target audio attribute features for each audio frame pair in the audio to be processed through a pre-trained speech model. Then, it fuses the target frequency domain acoustic features and the target audio attribute features to obtain the target audio features. These target audio features are then input into the trained audio counterfeiting detection model, which performs multi-threshold judgments to identify the audio to be processed and obtain a result determining whether the audio to be processed is counterfeit. By using the audio counterfeiting detection model based on target audio features containing richer and more comprehensive audio representations, it can accurately determine whether the audio to be processed is counterfeit, improving the accuracy of detecting counterfeit traces in audio, resulting in a low false recognition rate and higher recognition sensitivity.
[0119] When the audio to be processed includes multiple audio frames, multiple threshold judgments can be used to detect fake audio. This avoids simply selecting the highest confidence level from multiple confidence levels as the final confidence level of the audio to be processed, which would increase the false alarm rate and reduce the fault tolerance rate, thus improving the accuracy of audio fake audio detection. In some implementations, the first confidence level includes the second confidence level of each audio frame. The recognition result obtained based on the first confidence level may include:
[0120] b1. Determine the threshold range to which the second confidence level of each audio frame belongs.
[0121] b2. Determine the recognition result based on the threshold range to which the second confidence level of each audio frame belongs. The recognition result is used to indicate whether there are audio frames in the audio to be processed that meet the preset forgery conditions.
[0122] When the audio to be processed includes multiple audio frames, the audio recognition device can perform anti-counterfeiting (which can be called reasoning) based on the logic of multi-threshold judgment of multiple audio frames, so as to achieve accurate and objective detection of the audio. The multiple audio frames can be obtained by dividing the audio to be processed into frames.
[0123] Specifically, after obtaining the second confidence scores of multiple audio frames, the audio recognition device can determine the threshold interval to which the second confidence score of each audio frame belongs. The threshold interval may include multiple intervals, and the value range of each threshold interval does not overlap. The multiple threshold intervals constitute a continuous threshold range. The threshold interval can be flexibly set according to actual needs and is not limited here.
[0124] Then, the audio recognition device can determine whether the audio to be processed is fake audio based on the threshold range to which the second confidence score of each audio frame belongs, thereby obtaining the recognition result of the audio to be processed. The confidence score can range from [0, 1]. Determining whether the audio to be processed is fake audio through multi-threshold judgment, i.e., determining whether the audio to be processed is real or fake, improves the accuracy of audio fakeness detection. This avoids the problem of simply selecting the highest confidence score from multiple confidence scores as the final confidence score of the audio to be processed, which would increase the false alarm rate and reduce the fault tolerance rate, making accurate audio fakeness detection impossible.
[0125] In determining the recognition result based on the threshold range to which the second confidence level of each audio frame belongs, the large number of audio frames makes the process of authentication through multiple threshold judgments prone to confusion. Therefore, a combination of multiple threshold judgments and list updates can be used, with the updated list providing a more intuitive representation for audio authentication. The intuitiveness of list storage enhances readability and facilitates the determination of the recognition result based on the threshold range to which the second confidence level of each audio frame belongs, thus improving the efficiency of audio authentication. In some implementations, determining the recognition result based on the threshold range to which the second confidence level of each audio frame belongs may include:
[0126] b21. If there is at least one first audio frame among multiple audio frames with a second confidence level not less than the target confidence level threshold, then the first audio frame is determined to meet the preset forgery conditions, and the first target audio frame is updated to the first list.
[0127] The target confidence threshold can be flexibly set according to actual needs and is not limited here. For example, the target confidence threshold can be 0.5. After obtaining the second confidence scores corresponding to multiple audio frames, the audio recognition device can determine whether there is at least one first audio frame among the multiple audio frames with a second confidence score not less than the target confidence threshold.
[0128] If there is at least one first audio frame among multiple audio frames with a second confidence level not less than the target confidence level threshold, then the first audio frame is determined to meet the preset forgery conditions, that is, the first audio frame is a forged audio frame, and the first audio frame is updated to the first list, which can be a forgery list fake_list.
[0129] b22. If there is at least one second audio frame among multiple audio frames with a second confidence level less than the target confidence level threshold, determine that the second audio frame does not meet the preset forgery conditions, and update the second target audio frame to the second list.
[0130] If at least one second audio frame among multiple audio frames has a second confidence level less than the target confidence level threshold, it is determined that the second audio frame does not meet the preset forgery conditions, that is, the second audio frame is not a forged audio frame, and the second target audio frame is updated to the second list, which can be the real list real_list.
[0131] b23. Based on the first list and the second list, determine whether the audio to be processed meets the preset forgery conditions and obtain the recognition result.
[0132] The audio recognition device can determine whether the audio to be processed meets the preset forgery conditions based on a first list and a second list, and obtain the recognition result. For example, if the first list is empty, it is determined that the audio to be processed does not meet the preset forgery conditions; if the second list is empty, it is determined that the audio to be processed meets the preset forgery conditions.
[0133] For example, when the audio recognition model outputs a score greater than or equal to 0.5 for a certain audio frame, the audio frame and its corresponding score can be stored in the first list, `fake_list`. When the audio recognition model outputs a score less than 0.5 for a certain audio frame, the audio frame and its corresponding score can be stored in the second list, `real_list`. After judging and storing multiple audio frames, if `fake_list` is empty, the audio to be processed is determined to be real audio (`real`), and the confidence score is the minimum value in `real_list`, `min(real_list)`. If `real_list` is empty, the audio to be processed is determined to be fake audio (`fake`), and the confidence score is the maximum value in `fake_list`, `max(fake_list)`. The determination of whether the audio to be processed is fake audio is achieved by comparing the confidence scores of multiple audio frames with the target confidence threshold, i.e., judging whether the audio to be processed is real or fake.
[0134] By using multiple thresholds for authentication, the false alarm rate and fault tolerance rate can be improved by avoiding the problem of simply selecting the highest confidence level from multiple confidence levels as the final confidence level of the audio to be processed, which would increase the false alarm rate and reduce the fault tolerance rate. Therefore, the accuracy of audio false detection can be improved. Furthermore, by combining multiple thresholds with a more intuitive list update method, and performing audio false detection based on the updated list, the readability is improved due to the intuitiveness of list storage. This allows for the quick statistical analysis and determination of the threshold range to which the second confidence level of each audio frame belongs, achieving the goal of rapid false detection and improving the efficiency of audio false detection.
[0135] Since the first list has already been updated with first audio frames that meet the preset forgery conditions, and the second list has already been updated with second audio frames that meet the preset forgery conditions, information such as the confidence level and number of audio frames in the first and second lists can be statistically analyzed to accurately determine whether the audio to be processed meets the preset forgery conditions based on the statistical results. In some embodiments, determining whether the audio to be processed meets the preset forgery conditions based on the first and second lists may include:
[0136] b231. Get the first total number of audio frames in the first list.
[0137] b232. Based on the first total, and according to at least one of the average confidence level, the target number, and the second total of audio frames in the second list, determine whether the audio to be processed meets the preset forgery conditions. The average confidence level is the average confidence level of each audio frame in the first list, and the target number is the target number of audio frames in the first list whose confidence level belongs to the preset confidence level range.
[0138] The target number may include a first number of audio frames in the first list whose confidence level is in the first confidence level interval, a second number of audio frames whose confidence level is in the second confidence level interval, and a third number of audio frames whose confidence level is in the third confidence level interval, etc. The first confidence level interval, the second confidence level interval, and the third confidence level interval constitute a range of confidence level values that are less than the target confidence level threshold, and they do not overlap with each other.
[0139] Specifically, if the first list and the second list are not empty, the audio recognition device can count the first total number of audio frames in the first list (fake_count) and the second total number of audio frames in the second list (real_count).
[0140] The audio recognition device can also calculate the average confidence score of the audio frames in the first list: sample_score = np.mean(fake_list).
[0141] And count the first number of audio frames in the first list whose confidence level is within the first confidence interval: fakescore_count1 = sum(i>threshold1 for i in fake_list), where i>threshold1 is the first confidence interval.
[0142] And count the second number of audio frames in the first list whose confidence level is in the second confidence interval: fakescore_count2 = sum(i>=threshold2 for i in fake_list), where i>=threshold2 is the second confidence interval.
[0143] And count the third number of audio frames in the first list whose confidence level is in the third confidence level interval: fakescore_count3 = sum(i>=threshold3 for i in fake_list), where i>=threshold3 is the third confidence level interval.
[0144] The first confidence interval, the second confidence interval, and the third confidence interval constitute the range of confidence values less than the target confidence threshold, and they do not overlap. The specific values of the third confidence interval, the first confidence interval, the second confidence interval, and the third confidence interval can be flexibly set according to actual needs, and are not limited here.
[0145] In the process of multi-threshold judgment, in order to fully and reasonably utilize the first total, average confidence, target quantity, and second total for accurate counterfeiting detection, based on the first total, and according to at least one of the average confidence, target quantity, and the second total of audio frames in the second list, it is determined whether the audio to be processed meets the preset counterfeiting conditions, including at least one of the following:
[0146] c1. If the average confidence level is greater than the first threshold and the number of targets is greater than or equal to half of the first total, then the audio to be processed is determined to meet the preset forgery conditions.
[0147] Specifically, if the average confidence score sample_score is greater than the first threshold (e.g., 0.5), and the second quantity fakescore_count2 is greater than or equal to half of the first total quantity (0.5 * fake_count), and the first total quantity fake_count is greater than 1, i.e., sample_score > 0.5 and fakescore_count2 >= (0.5 * fake_count) and fake_count > 1, then the audio to be processed is determined to meet the preset forgery conditions; otherwise, the audio to be processed is determined not to meet the preset forgery conditions.
[0148] c2. If the first total quantity is greater than or equal to the first preset multiple of the second total quantity, then the audio to be processed is determined to meet the preset forgery conditions.
[0149] Specifically, if the first total fake_count is greater than or equal to the first preset multiple (e.g., 3 times) of the second total real_count, i.e., fake_count>=(3*real_count), then the audio to be processed is determined to meet the preset forgery conditions, and the first preset multiple is greater than 1; otherwise, the audio to be processed is determined to not meet the preset forgery conditions.
[0150] c3. If the number of targets is greater than or equal to the second threshold, and the first total is greater than the third threshold, then the audio to be processed is determined to meet the preset forgery conditions, the third threshold is greater than the second threshold, and the second threshold is greater than the first threshold; or,
[0151] Specifically, if the third quantity fakescore_count3 is greater than or equal to the second threshold (e.g., a value of 2), and the first total quantity fake_count is greater than the third threshold (e.g., a value of 3), i.e., fakescore_count3>=2) and fake_count>3, then the audio to be processed is determined to meet the preset forgery conditions; otherwise, the audio to be processed is determined to not meet the preset forgery conditions.
[0152] c3. If the target quantity is greater than or equal to the second preset multiple of the first total quantity, and the first total quantity is greater than the fourth threshold, then the audio to be processed is determined to meet the preset forgery conditions, the second preset multiple is less than the first preset multiple, and the fourth threshold is greater than the third threshold.
[0153] Specifically, if the first quantity fakescore_count1 is greater than or equal to the second preset multiple (e.g., 0.8) of the first total quantity fake_count, and the first total quantity fake_count is greater than the fourth threshold (e.g., 5), i.e., fakescore_count1 >= (0.8 * fake_count) and fake_count > 5, then the audio to be processed is determined to meet the preset forgery conditions; otherwise, the audio to be processed is determined not to meet the preset forgery conditions. The fourth threshold is greater than the third threshold, the third threshold is greater than the second threshold, the second threshold is greater than the first threshold, and the second preset multiple is less than the first preset multiple.
[0154] The first threshold, second threshold, third threshold, fourth threshold, first preset multiple, and second preset multiple can all be flexibly set according to actual needs, and are not limited here.
[0155] In summary, during the inference process of the audio recognition model, when the audio to be processed is divided into multiple audio frames, two score lists can be created: a real list (real_list) and a fake list (fake_list). When the audio recognition model outputs a score greater than or equal to 0.5 for a certain audio frame, the audio frame and its corresponding score can be stored in the first list (fake_list). When the audio recognition model outputs a score less than 0.5 for a certain audio frame, the audio frame and its corresponding score can be stored in the second list (real_list). After judging and storing multiple audio frames, ...
[0156] 1) If fake_list is empty, then the audio to be processed is determined to be real, and the confidence level is min(real_list);
[0157] 2) If real_list is empty, the audio to be processed is determined to be fake, and the confidence level is max(fake_list);
[0158] 3) If both fake_list and real_list are not empty, then the following metrics are calculated:
[0159] sample_score=np.mean(fake_list);
[0160] fakescore_count1=sum(i>threshold 1 for i in fake_list);
[0161] fakescore_count2=sum(i>=threshold 2 for i in fake_list);
[0162] fakescore_count3=sum(i>=threshold 3for i in fake_list);
[0163] Furthermore, a list is classified as fake if any of the following conditions exist, with a confidence level of max(fake_list); otherwise, it is classified as real with a confidence level of min(real_list):
[0164] 3.1)sample_score>0.5and fakescore_count2>=(0.5*fake_count)and fake_count>1;
[0165] 3.2)fake_count>=(3*real_count);
[0166] 3.3) fakescore_count3>=2)and fake_count>3;
[0167] 3.4) fakescore_count1>=(0.8*fake_count)and fake_count>5.
[0168] This application embodiment can perform frame-by-frame processing on long-duration audio and combine it with multi-threshold judgment logic for fake detection. Therefore, it can achieve accurate and objective detection of long-duration audio. Compared with simply selecting the highest confidence level from multiple confidence levels as the final confidence level of the audio, it reduces the false alarm rate. In particular, it greatly improves the fault tolerance rate on open set data (untrained) and real network data (Internet). This allows the audio recognition model to still accurately detect fake audio when encountering long audio, thus improving the accuracy of audio fake detection.
[0169] After obtaining the identification result of whether the audio to be processed is forged, the audio synthesis model can be trained using the identified non-forged audio. In some embodiments, the audio recognition method further includes identifying the target audio features and obtaining the identification result:
[0170] d1. When the audio to be processed does not meet the preset forgery conditions, the audio to be processed is used as a training sample for the audio synthesis model.
[0171] d2. Train the audio synthesis model using training samples to obtain the trained audio synthesis model.
[0172] In audio synthesis applications, when an audio recognition device, through an audio recognition model, determines that the audio to be processed does not meet preset forgery conditions, it indicates that the audio is not forged, such as a user's real voice. In this case, the audio to be processed can be used as a training sample for the audio synthesis model. The model is trained using this training sample to obtain a trained audio synthesis model. For example, the prompt information from the training sample can be input into the audio synthesis model, which generates audio. The generated audio is then compared with the training sample, and the parameters of the audio synthesis model are adjusted based on the differences between the two. Training stops when the differences between the two are minimized, ensuring that the trained audio synthesis model can accurately generate the required audio. This embodiment identifies the audio to be processed as genuine and inputs it into the audio synthesis model for training, thereby obtaining reliable audio for precise model training and improving the accuracy of the audio synthesis model in generating the required audio.
[0173] After training the audio synthesis model, the training results can be fed back to the audio recognition model for optimization. In some implementations, the audio synthesis model is trained using training samples. After obtaining the trained audio synthesis model, the audio recognition method further includes:
[0174] e1. Generate the target audio for the target user using the trained audio synthesis model.
[0175] e2. Obtain the similarity between the target audio and the real audio corresponding to the target user.
[0176] e3. Adjust the parameters of the trained audio recognition model based on the similarity to obtain the optimized audio recognition model.
[0177] After training the audio synthesis model, the target audio corresponding to the target user can be generated using the trained audio synthesis model. Then, the similarity between the target audio and the real audio corresponding to the target user is obtained to determine whether the target audio and the real audio come from the same user. Based on the similarity, the audio synthesis model can know the synthesis effect. The parameters of the trained audio recognition model are adjusted according to the similarity, that is, the parameters of the trained audio recognition model are adjusted based on the synthesis effect to optimize the audio recognition model. This allows for the acquisition of higher quality audio for model optimization, enabling the optimized audio recognition model to more accurately identify fake audio.
[0178] For example, such as Figure 7 As shown, taking an audio recognition model as an audio authentication model as an example, in the audio authentication stage, the trained audio authentication model can first authenticate the audio to be recognized. If the audio is fake, such as a forged voice of the target user, it can be discarded; if the audio is real, such as the real voice of the target user, it can proceed to the stage of training the audio synthesis model. At this time, the audio to be recognized can be used as training samples to train the audio synthesis model, resulting in a trained audio synthesis model. Then, in the stage of optimizing the audio authentication model, the trained audio synthesis model can generate target audio corresponding to the target user, obtain the similarity between the target audio and the real audio of the target user, and optimize the trained audio authentication model based on the similarity, resulting in an optimized audio authentication model. Since the authentication result of the audio authentication model can be used to provide samples for the training of the audio synthesis model, and the target audio can be generated by the trained audio synthesis model, and the generation effect of the target audio can be fed back to the audio authentication model to optimize it, the accuracy of the optimized audio authentication model in audio authentication can be improved.
[0179] For example, such as Figure 8 As shown, the audio recognition device and the audio synthesis device can interact, including steps S31 to S38, wherein...
[0180] S31. The audio recognition device can acquire the audio to be recognized.
[0181] S32. The audio recognition device can extract the target frequency domain acoustic features and target audio attribute features of the audio to be recognized, and fuse the target frequency domain acoustic features and target audio attribute features to obtain the target audio features.
[0182] S33. The audio recognition device performs multi-threshold recognition on the audio to be recognized using the trained audio recognition model. If the audio is identified as fake, such as a fake voice of the target user, the audio can be discarded.
[0183] S34. If the audio is identified as genuine (i.e., real audio), such as the target user's real voice, the real audio can be sent to the audio synthesis device.
[0184] S35. The audio synthesis device can use real audio as training samples to train the audio synthesis model, thereby obtaining the trained audio synthesis model. The trained audio synthesis model can then be used to generate the target audio corresponding to the target user.
[0185] S36. The audio synthesis device can obtain the similarity between the target audio and the target user's real audio.
[0186] S37. The audio synthesis device can send the similarity between the target audio and the real audio to the audio recognition device.
[0187] S38. The audio recognition device can optimize the trained audio recognition model based on similarity to obtain an optimized audio recognition model. Through the interaction between the audio synthesis device and the audio recognition device, the recognition results of the audio recognition model are used to provide samples for the audio synthesis model to train. The trained audio synthesis model generates target audio, and the generation effect of the target audio is fed back to the audio recognition model to optimize it. This can improve the accuracy of the optimized audio recognition model in detecting forgery traces in audio.
[0188] In this embodiment, the acquired audio to be processed is first transformed in the frequency domain to obtain target frequency domain acoustic features with audio frequency domain representation, and feature extraction is performed on the audio to be processed to obtain target audio attribute features. Then, the target frequency domain acoustic features and the target audio attribute features are fused to obtain the final target audio features used for identification. Since the target frequency domain acoustic features can represent the audio frequency domain features of the audio to be processed, and the target audio attribute features can represent multi-dimensional audio information, which can contain richer audio features, the target audio features obtained after fusing the target frequency domain acoustic features and the target audio attribute features will contain a richer and more comprehensive audio representation. Therefore, compared to the simple identification based on a single audio feature in related technologies, the embodiments of this application, when the audio recognition model identifies the target audio feature, will obtain richer and more comprehensive audio features because the target audio feature has the characteristic of containing richer and more comprehensive audio representation. Thus, the embodiments of this application can accurately determine whether the audio to be processed is fake by identifying the target audio feature containing richer and more comprehensive audio representation, thereby improving the accuracy of detecting forgery traces in audio, resulting in a low false recognition rate and higher recognition sensitivity.
[0189] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.
[0190] This embodiment uses an audio recognition device integrated into a server as an example, applied to a telephone fraud prevention scenario. When terminal device A and terminal device B are in a call, after receiving the audio of the call sent by terminal device A, terminal device B can request its corresponding server to recognize the audio. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a schematic flowchart illustrating the audio recognition method provided in an embodiment of this application. The method flow may include:
[0191] S11. Terminal device A sends the audio of the call to terminal device B.
[0192] Terminal device A can establish a call connection with terminal device B. After receiving the audio recorded by the user based on the current call, terminal device A can send the call audio to terminal device B.
[0193] S12. Terminal device B sends the audio of the call to the server.
[0194] After receiving the call audio from terminal device A, terminal device B can forward the call audio to the server so that the server can identify the call audio.
[0195] S13. The server extracts the frequency domain acoustic features and audio attribute features of the audio and fuses them to obtain the target audio features.
[0196] The server can perform frequency domain transformation on the audio of a call and extract the frequency domain acoustic features (such as the CQCC feature matrix) of the audio in the frequency domain dimension. It can also extract features from the audio of the call through a pre-trained speech model (such as wav2vec) to obtain the audio attribute features (such as the audio feature matrix). Then, the frequency domain acoustic features and audio attribute features can be fused to obtain target audio features with richer and more comprehensive audio representation.
[0197] S14. The server uses the trained audio recognition model to perform multi-threshold recognition on the audio based on the target audio features.
[0198] The server can use a trained audio recognition model (such as LCNN) to determine the confidence level of a call's audio as the target user's real voice based on the target audio features. Based on this confidence level, the server can then determine whether the call's audio is a forged audio impersonating the target user. For example, the server can determine the threshold range to which the audio's confidence level falls, and based on this threshold range, determine whether the call's audio is a forged audio impersonating a particular user.
[0199] S15. If the audio of the call is fake, the server sends a prompt message to terminal device B.
[0200] S16. Terminal device B displays a prompt message.
[0201] If the audio of the call is not fake, terminal device B will continue the current call. If the audio is fake, the server can generate a notification message indicating that the audio is fake and send the message to terminal device B. Terminal device B can display the notification message to remind the user that the current call is a scam. In this case, terminal device B can also automatically terminate the current call or report the scam, allowing the user to stop the loss in time and avoid unnecessary losses, thus improving call security.
[0202] It should be noted that the server can also act as a firewall between terminal device A and terminal device B. Terminal device A can establish a call connection with terminal device B. After receiving the audio recorded by the user based on the current call, terminal device A can first send the call audio to the server. The server extracts the frequency domain acoustic features and audio attribute features of the audio and fuses them to obtain the target audio features. Based on the target audio features, the trained audio recognition model performs multi-threshold recognition on the audio. If the call audio is spoofed, the server sends a prompt message to terminal device B and can also block the call connection between terminal device A and terminal device B. If the call audio is not spoofed, the server can send the call audio to terminal device B.
[0203] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0204] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.
[0205] This embodiment uses an audio recognition device integrated into a terminal device as an example, applied to a telephone fraud prevention scenario. Terminal device A and terminal device B are making a call. After receiving the audio of the call sent by terminal device A, terminal device B can recognize the audio. Please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic flowchart illustrating the audio recognition method provided in an embodiment of this application. The method flow may include:
[0206] S21. Terminal device A sends the audio of the call to terminal device B.
[0207] Terminal device A can establish a call connection with terminal device B. After receiving the audio recorded by the user based on the current call, terminal device A can send the call audio to terminal device B.
[0208] S22. Terminal device B extracts the frequency domain acoustic features and audio attribute features of the audio and fuses them to obtain the target audio features.
[0209] Terminal device B can perform frequency domain conversion on the audio of a call and extract the frequency domain acoustic features of the audio in the frequency domain dimension. It can also extract features from the audio of the call through a pre-trained speech model to obtain the audio attribute features of the audio. Then, the frequency domain acoustic features and audio attribute features can be fused to obtain target audio features with richer and more comprehensive audio representation.
[0210] S23. Terminal device B uses the trained audio recognition model to perform multi-threshold recognition of audio based on the target audio features.
[0211] Terminal device B can use a trained audio recognition model to determine the confidence level of a call's audio as the target user's real voice based on the target audio features. Based on this confidence level, it can then determine whether the call's audio is a forged audio imitating the target user. For example, terminal device B can determine the threshold range to which the audio's confidence level falls, and based on this threshold range, determine whether the call's audio is a forged audio imitating a particular user.
[0212] S24. If the audio of the call is fake, terminal device B displays a prompt message.
[0213] If the audio of the call is not fake, terminal device B will continue the current call. If the audio of the call is fake, terminal device B can display a prompt message indicating that the audio is fake to remind the user that the current call is a phone scam. In this case, terminal device B can also automatically terminate the current call or report the scam, allowing the user to stop the loss in time and avoid unnecessary losses, thus improving the security of the call.
[0214] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0215] The above describes an audio recognition method in the embodiments of this application. The following describes the audio recognition device (e.g., a server) that performs the above audio recognition method.
[0216] See Figure 11 ,like Figure 11 The diagram illustrates the structure of an audio recognition device, which can be applied to servers in audio recognition scenarios requiring authentication. It acquires the target frequency domain acoustic features and target audio attribute features of the audio to be processed, and fuses these features to obtain target audio features containing richer and more comprehensive audio representations. Recognition is then performed based on these target audio features to obtain the recognition result. The audio recognition device in this embodiment can achieve the functionality described above. Figure 2 The steps of the audio recognition method executed in the corresponding embodiments are described above. The functions implemented by the audio recognition device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, which can be software and / or hardware. The audio recognition device 60 may include an input / output module 601 and a processing module 602, wherein the functional implementation of the input / output module 601 and the processing module 602 can be referred to... Figure 2 The operations performed in the corresponding embodiments will not be described in detail here.
[0217] Input / output module 601 is configured to acquire audio to be processed;
[0218] The processing module 602 is configured to perform frequency domain transformation on the audio to be processed to obtain the target frequency domain acoustic features of the audio to be processed, and to extract features from the audio to be processed to obtain the target audio attribute features of the audio to be processed; to fuse the target frequency domain acoustic features and the target audio attribute features to obtain the target audio features; and to identify the target audio features to obtain an identification result, wherein the identification result is used to indicate whether the audio to be processed is forged.
[0219] In some implementations, the processing module 602 is configured to identify the target audio features using an audio recognition model to obtain a first confidence level of the audio to be processed; and to obtain a recognition result based on the first confidence level.
[0220] In some embodiments, the audio recognition device 60 further includes:
[0221] The first acquisition module is configured to acquire the audio training set;
[0222] The extraction module is configured to extract the frequency domain acoustic features and audio attribute features of each audio training sample in the audio training set;
[0223] The second acquisition module is configured to acquire a feature set, which includes multiple audio features. Each audio feature is obtained by fusing the frequency domain acoustic features and audio attribute features of the same audio training sample.
[0224] The first training module is configured to train the audio recognition model based on the feature set, thereby obtaining the trained audio recognition model.
[0225] In some implementations, the first training module is configured to: acquire the true labels of the audio training samples; identify each audio feature in the feature set to obtain the predicted labels of each audio training sample in the audio training set; acquire the difference information between the true labels and the predicted labels; and adjust the parameters of the audio recognition model based on the difference information to obtain the trained audio recognition model.
[0226] In some implementations, the audio to be processed includes multiple audio frames, the first confidence level includes the second confidence level of each audio frame, and the processing module 602 is configured to determine the threshold range to which the second confidence level of each audio frame belongs; and determine the recognition result based on the threshold range to which the second confidence level of each audio frame belongs, the recognition result being used to indicate whether there are audio frames in the audio to be processed that meet the preset forgery conditions.
[0227] In some embodiments, the processing module 602 is configured to: if there is at least one first audio frame among multiple audio frames with a second confidence level not less than a target confidence level threshold, determine that the first audio frame meets the preset forgery conditions and update the first labeled audio frame to a first list; if there is at least one second audio frame among multiple audio frames with a second confidence level less than the target confidence level threshold, determine that the second audio frame does not meet the preset forgery conditions and update the second labeled audio frame to a second list; and determine whether the audio to be processed meets the preset forgery conditions based on the first list and the second list to obtain the recognition result.
[0228] In some implementations, the processing module 602 is configured to obtain a first total number of audio frames in a first list; based on the first total number and at least one of the average confidence level, the target number, and the second total number of audio frames in a second list, determine whether the audio to be processed meets a preset forgery condition, wherein the average confidence level is calculated as the average confidence level of each audio frame in the first list, and the target number is the target number of audio frames in the first list whose confidence level belongs to a preset confidence level range.
[0229] In some implementations, the processing module 602 is configured to determine that the audio to be processed meets preset forgery conditions if the average confidence level is greater than a first threshold and the number of targets is greater than or equal to half of the first total; if the first total is greater than or equal to a first preset multiple of the second total, the audio to be processed meets preset forgery conditions; if the number of targets is greater than or equal to a second threshold and the first total is greater than a third threshold, the audio to be processed meets preset forgery conditions, wherein the third threshold is greater than the second threshold and the second threshold is greater than the first threshold; or, if the number of targets is greater than or equal to a second preset multiple of the first total and the first total is greater than a fourth threshold, the audio to be processed meets preset forgery conditions, wherein the second preset multiple is less than the first preset multiple and the fourth threshold is greater than the third threshold.
[0230] In some implementations, the input / output module 601 is configured to acquire the original audio; if the length of the original audio is greater than a preset length threshold, the original audio is segmented into frames to obtain multiple audio frames, and the multiple audio frames are used as audio to be processed; if the length of the original audio is less than or equal to the preset length threshold, the original audio is used as audio to be processed.
[0231] In some implementations, the processing module 602 is configured to perform layer-by-layer feature extraction on the audio to be processed through a pre-trained speech model using multiple convolutional layers to obtain multi-layer attribute features; and to perform weighted fusion of the multi-layer attribute features to obtain the target audio attribute features of the audio to be processed.
[0232] In some embodiments, the audio recognition device 60 further includes:
[0233] The sample acquisition module is configured to use the audio to be processed as a training sample for the audio synthesis model when the audio to be processed does not meet the preset forgery conditions.
[0234] The second training module is configured to train the audio synthesis model using training samples to obtain the trained audio synthesis model.
[0235] In some embodiments, the audio recognition device 60 further includes:
[0236] The generation module is configured to generate target audio for the target user using a trained audio synthesis model.
[0237] The similarity acquisition module is configured to obtain the similarity between the target audio and the real audio corresponding to the target user.
[0238] The optimization module is configured to adjust the parameters of the trained audio recognition model based on similarity, resulting in an optimized audio recognition model.
[0239] In this embodiment, the acquired audio to be processed is first transformed in the frequency domain to obtain target frequency domain acoustic features with audio frequency domain representation, and feature extraction is performed on the audio to be processed to obtain target audio attribute features. Then, the target frequency domain acoustic features and the target audio attribute features are fused to obtain the final target audio features used for identification. Since the target frequency domain acoustic features can represent the audio frequency domain features of the audio to be processed, and the target audio attribute features can represent multi-dimensional audio information, which can contain richer audio features, the target audio features obtained after fusing the target frequency domain acoustic features and the target audio attribute features will contain a richer and more comprehensive audio representation. Therefore, compared to the simple identification based on a single audio feature in related technologies, the embodiments of this application, when the audio recognition model identifies the target audio feature, will obtain richer and more comprehensive audio features because the target audio feature has the characteristic of containing richer and more comprehensive audio representation. Thus, the embodiments of this application can accurately determine whether the audio to be processed is fake by identifying the target audio feature containing richer and more comprehensive audio representation, thereby improving the accuracy of detecting forgery traces in audio, resulting in a low false recognition rate and higher recognition sensitivity.
[0240] The audio recognition device 60 in this application embodiment has been described above from the perspective of modular functional entities. The audio recognition device in this application embodiment will be described below from the perspective of hardware processing.
[0241] It should be noted that, Figure 11The physical device corresponding to the input / output module 601 shown can be a transceiver, radio frequency circuit, communication module, and input / output (I / O) interface, etc., and the physical device corresponding to the processing module 602 can be a processor.
[0242] Figure 11 The devices shown can all have the following characteristics: Figure 12 The structure shown, when Figure 11 The audio recognition device 60 shown has, for example: Figure 12 When the structure shown is used, Figure 12 The processor and transceiver in the device can perform the same or similar functions as the input / output module 601 and processing module 602 provided in the aforementioned device embodiments. Figure 13 The memory stores the computer programs that the processor needs to call when executing the above audio recognition method.
[0243] This application also provides a terminal device, such as... Figure 13 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal device can be any terminal device including mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, etc. Taking a mobile phone as an example:
[0244] Figure 13 This diagram illustrates a partial structure of a mobile phone related to the terminal device provided in the embodiments of this application. (Reference) Figure 13 The mobile phone includes components such as a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art will understand that... Figure 13 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0245] The following is combined Figure 13 A detailed introduction to each component of a mobile phone:
[0246] The RF circuit 1010 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 1080; additionally, it transmits uplink data to the base station. Typically, the RF circuit 1010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the RF circuit 1010 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).
[0247] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0248] The input unit 1030 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1031), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 1080, and can also receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 may also include other input devices 1032. Specifically, other input devices 1032 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0249] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041, which may optionally be configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar display. Further, a touch panel 1031 may cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it transmits the information to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 based on the type of touch event. Although in Figure 13 In this embodiment, the touch panel 1031 and the display panel 1041 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.
[0250] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1041 according to the ambient light level, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0251] The audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the mobile phone. The audio circuit 1060 converts the received audio data into electrical signals and transmits them to the speaker 1061, where the speaker 1061 converts them into sound signals for output. On the other hand, the microphone 1062 converts the collected sound signals into electrical signals, which are then received by the audio circuit 1060, converted into audio data, and then processed by the processor 1080 before being transmitted via the RF circuit 1010 to, for example, another mobile phone, or the audio data can be output to the memory 1020 for further processing.
[0252] Wi-Fi is a short-range wireless transmission technology. Through the Wi-Fi module 1070, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 13 The Wi-Fi module 1070 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.
[0253] The processor 1080 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 1020 and calls data stored in the memory 1020 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 1080.
[0254] The mobile phone also includes a power supply 1090 (such as a battery) that supplies power to various components. Optionally, the power supply can be logically connected to the processor 1080 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0255] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0256] In this embodiment of the application, the processor 1080 included in the mobile phone also has the function of controlling the execution of the audio recognition method process performed by the audio recognition device.
[0257] This application also provides a server; please refer to [link / reference]. Figure 14 , Figure 14 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. The memory 1132 and storage media 1130 may be temporary or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server. Furthermore, the CPU 1122 may be configured to communicate with the storage media 1130 and execute the series of instruction operations in the storage media 1130 on the server 1100.
[0258] Server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.
[0259] The steps performed by the server in the above embodiments can be based on this Figure 14 The structure of server 1100 shown. For example, as in the above embodiment, by Figure 11 The steps performed by the audio recognition device 60 shown can be based on this Figure 14 The server architecture is shown. For example, the central processing unit 1122 performs the following operations by calling instructions from memory 1132:
[0260] The audio to be processed is obtained through the input / output interface 1158;
[0261] The central processing unit 1122 performs frequency domain transformation on the audio to be processed to obtain the target frequency domain acoustic features of the audio to be processed, and performs feature extraction on the audio to be processed to obtain the target audio attribute features of the audio to be processed; it fuses the target frequency domain acoustic features and the target audio attribute features to obtain the target audio features; and it identifies the target audio features to obtain an identification result used to indicate whether the audio to be processed is fake.
[0262] The recognition results can also be sent through the input / output interface 1158.
[0263] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0264] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0265] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0266] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0267] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0268] According to one aspect of this application, a computer-readable storage medium is provided, which includes instructions that, when executed on a computer, cause the computer to perform the audio recognition method of this application as described above.
[0269] According to one aspect of this application, a computing device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio recognition method of this application. The computing device may be the aforementioned terminal device or server.
[0270] According to one aspect of this application, a chip is provided, which includes a processor coupled to a transceiver of a terminal device for executing the audio recognition method provided in the embodiments of this application. Embodiments of this application also provide a chip system including a processor for supporting the terminal device in implementing the functions involved in the aforementioned audio recognition method. In one possible design, the chip system further includes a communication interface for inputting and / or outputting information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the terminal device. The chip system can be composed of a chip or may include chips and other discrete devices.
[0271] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions (i.e., program instructions) stored in a computer-readable storage medium. A processor of a computer device (which may be referred to as a computing device or computer) reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.
[0272] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0273] A computer program product includes one or more computer instructions. When a computer program is loaded and executed on a computer, it generates, in whole or in part, the processes or functions provided according to the embodiments of this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0274] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.
Claims
1. An audio recognition method, characterized in that, The method includes: Obtain the audio to be processed; The audio to be processed is frequency domain transformed to obtain the target frequency domain acoustic features of the audio to be processed, and the audio to be processed is feature extracted to obtain the target audio attribute features of the audio to be processed. The target frequency domain acoustic features and the target audio attribute features are fused to obtain the target audio features; The target audio features are identified using an audio recognition model to obtain a recognition result, which is used to indicate whether the audio to be processed is fake. The process of identifying the target audio features using an audio recognition model and obtaining the recognition result further includes: When the audio to be processed does not meet the preset forgery conditions, the audio to be processed is used as a training sample for the audio synthesis model; the audio synthesis model is trained using the training sample to obtain the trained audio synthesis model; the target audio corresponding to the target user is generated using the trained audio synthesis model; the similarity between the target audio and the real audio corresponding to the target user is obtained; the parameters of the trained audio recognition model are adjusted according to the similarity to obtain the optimized audio recognition model.
2. The audio recognition method according to claim 1, characterized in that, The process of identifying the target audio features to obtain the identification result includes: The target audio features are identified using an audio recognition model to obtain the first confidence level of the audio to be processed. The identification result is obtained based on the first confidence level.
3. The audio recognition method according to claim 2, characterized in that, Before identifying the target audio features using the audio recognition model, the method further includes: Obtain the audio training set; Extract the frequency domain acoustic features and audio attribute features of each audio training sample in the audio training set; A feature set is obtained, which includes multiple audio features, each of which is obtained by fusing the frequency domain acoustic features and audio attribute features of the same audio training sample; The audio recognition model is trained based on the feature set to obtain the trained audio recognition model.
4. The audio recognition method according to claim 3, characterized in that, The step of training the audio recognition model based on the feature set to obtain the trained audio recognition model includes: Obtain the true labels of the audio training samples; Each audio feature in the feature set is identified to obtain the predicted label of each audio training sample in the audio training set; Obtain the difference information between the real label and the predicted label; Based on the difference information, the parameters of the audio recognition model are adjusted to obtain the trained audio recognition model.
5. The audio recognition method according to claim 2, characterized in that, The audio to be processed includes multiple audio frames, the first confidence level includes the second confidence level of each audio frame, and obtaining the recognition result based on the first confidence level includes: Determine the threshold range to which the second confidence score of each audio frame belongs; The identification result is determined based on the threshold range to which the second confidence level of each audio frame belongs. The identification result is used to indicate whether there are audio frames in the audio to be processed that meet the preset forgery conditions.
6. The audio recognition method according to claim 5, characterized in that, Determining the recognition result based on the threshold interval to which the second confidence level of each audio frame belongs includes: If there is at least one first audio frame among the plurality of audio frames with a second confidence level not less than the target confidence level threshold, then the first audio frame is determined to meet the preset forgery condition, and the first audio frame is updated to the first list; If at least one second audio frame among the plurality of audio frames has a second confidence level less than the target confidence threshold, it is determined that the second audio frame does not meet the preset forgery conditions, and the second audio frame is updated to the second list; Based on the first list and the second list, determine whether the audio to be processed meets the preset forgery conditions, and obtain the recognition result.
7. The audio recognition method according to claim 6, characterized in that, The step of determining whether the audio to be processed meets the preset forgery conditions based on the first list and the second list includes: Obtain the first total number of audio frames in the first list; Based on the first total, and according to at least one of the average confidence level, the target number, and the second total number of audio frames in the second list, it is determined whether the audio to be processed meets the preset forgery conditions. The average confidence level is the average confidence level of each audio frame in the first list, and the target number is the target number of audio frames in the first list whose confidence level belongs to a preset confidence level range.
8. The audio recognition method according to claim 7, characterized in that, The determination of whether the audio to be processed meets the preset forgery conditions based on the first total amount and at least one of the average confidence level, the target number, and the second total amount of audio frames in the second list includes at least one of the following: If the average confidence level is greater than the first threshold and the number of targets is greater than or equal to half of the first total, then the audio to be processed is determined to meet the preset forgery conditions. If the first total is greater than or equal to a first preset multiple of the second total, then the audio to be processed is determined to meet the preset forgery conditions; If the target quantity is greater than or equal to the second threshold, and the first total quantity is greater than the third threshold, then the audio to be processed is determined to meet the preset forgery conditions, wherein the third threshold is greater than the second threshold, and the second threshold is greater than the first threshold; or, If the target quantity is greater than or equal to a second preset multiple of the first total quantity, and the first total quantity is greater than a fourth threshold, then the audio to be processed is determined to meet the preset forgery conditions, where the second preset multiple is less than the first preset multiple, and the fourth threshold is greater than the third threshold.
9. The audio recognition method according to claim 1, characterized in that, The acquisition of the audio to be processed includes: Obtain the raw audio; If the length of the original audio is greater than a preset length threshold, the original audio is divided into frames to obtain multiple audio frames, and the multiple audio frames are used as audio to be processed. If the length of the original audio is less than or equal to a preset length threshold, then the original audio is used as the audio to be processed.
10. The audio recognition method according to claim 1, characterized in that, The step of extracting features from the audio to be processed to obtain the target audio attribute features of the audio to be processed includes: The audio to be processed is subjected to multi-layer convolutional feature extraction through a pre-trained speech model to obtain multi-layer attribute features. The multi-layer attribute features are weighted and fused to obtain the target audio attribute features of the audio to be processed.
11. An audio recognition device, characterized in that, include: The input / output module is configured to acquire the audio to be processed; The processing module is configured to perform frequency domain transformation on the audio to be processed to obtain the target frequency domain acoustic features of the audio to be processed, and to extract features from the audio to be processed to obtain the target audio attribute features of the audio to be processed; to fuse the target frequency domain acoustic features and the target audio attribute features to obtain target audio features; to identify the target audio features through an audio recognition model to obtain a recognition result, the recognition result being used to indicate whether the audio to be processed is forged; wherein, after identifying the target audio features through the audio recognition model and obtaining the recognition result, the module further includes: when the audio to be processed does not meet the preset forgery conditions, using the audio to be processed as a training sample for an audio synthesis model; training the audio synthesis model through the training sample to obtain a trained audio synthesis model; generating target audio corresponding to the target user through the trained audio synthesis model; obtaining the similarity between the target audio and the real audio corresponding to the target user; and adjusting the parameters of the trained audio recognition model according to the similarity to obtain an optimized audio recognition model.
12. A computing device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, It includes instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 10.
14. A computer program product comprising instructions, the computer program product including program instructions that, when executed on a computer or processor, cause the computer or processor to perform the method as claimed in any one of claims 1 to 10.
15. A chip system, characterized in that, The chip system includes: A communication interface used for inputting and / or outputting information; A processor for executing a computer-executable program, causing a device having the chip system mounted to perform the method as described in any one of claims 1 to 10.