Audio authentication method, related device and storage medium
By extracting the timing feature and calculating the similarity of fake audio and real audio, the problems of high error rejection rate and low detection rate of audio pseudo-identification methods in the prior art under the data distribution differences are solved, and a higher accuracy of fake verification is achieved.
Patent Information
- Application Number
- CN202510543174.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The difference in data distribution between the existing audio pseudo-identification method in the business field and the model training stage leads to a high error rejection rate and a poor detection rate, which makes it impossible to effectively identify forged audio.
By obtaining multiple fake audio and real audio, timing feature extraction is performed, partitioned into different sets, and the similarity parameters of the audio to be identified and the fake audio and real audio timing feature set are calculated. If the first similarity parameter is greater than the second similarity parameter, the audio category is determined to be a fake category.
The accuracy of audio fake identification is improved, and the accuracy of using rich fake feature information is improved.
Smart Images

Figure CN120340501A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and more specifically, to an audio forgery detection method, a related device, and a storage medium. Background Art
[0002] With the rapid development of 5G technology, voice deep forgery related technologies, such as text-to-speech (TTS) and voice conversion (VC), have become increasingly mature and have been widely applied in fields such as medical rehabilitation (e.g., "reconstructing" the voice of patients with aphonia) and entertainment (e.g., funny videos). While meeting people's daily needs, they also bring many security risks. In response to this security risk, many scholars have proposed related voice forgery detection methods. However, due to the difference in data distribution between the business site and the model training stage, the false rejection rate is relatively high and the detection rate is poor. Therefore, there is an urgent need for a new audio forgery detection method to improve the accuracy of audio forgery detection. Summary of the Invention
[0003] The embodiments of the present application provide an audio forgery detection method, a related device, and a storage medium, which can improve the accuracy of audio forgery detection.
[0004] In a first aspect, the embodiments of the present application provide an audio forgery detection method, which includes:
[0005] Obtain a plurality of forged audios and a plurality of genuine audios;
[0006] Extract temporal features from the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios;
[0007] Respectively divide the plurality of forged audio temporal features and the plurality of genuine audio temporal features into different sets to obtain a plurality of forged audio temporal feature sets and a plurality of genuine audio temporal feature sets;
[0008] When a to-be-recognized audio is obtained, calculate a first similarity parameter between the to-be-recognized audio and the plurality of forged audio temporal feature sets, and a second similarity parameter between the to-be-recognized audio and the plurality of genuine audio temporal feature sets respectively;
[0009] If the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the to-be-recognized audio is a forged category.
[0010] In one embodiment, the extracting of the temporal features of the multiple forged audios and the multiple genuine audios to obtain multiple forged audio temporal features corresponding to the multiple forged audios and multiple genuine audio temporal features corresponding to the multiple genuine audios includes:
[0011] Split the forged audio into multiple forged sub-audios sorted in time sequence;
[0012] Extract frequency domain features for each forged sub-audio respectively to obtain first forged frequency domain features corresponding to each forged sub-audio;
[0013] Determine the forged audio temporal features corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio.
[0014] In one embodiment, the determining of the forged audio temporal features corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio includes:
[0015] Determine each forged sub-audio as a target sub-audio respectively;
[0016] Determine the feature difference between the first forged frequency domain feature of the first forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio, and obtain the adjacent frequency domain change features of each target sub-audio;
[0017] Correspondingly splice the multiple first forged frequency domain features and the multiple adjacent frequency domain change features corresponding to the forged audio to obtain multiple second forged frequency domain features of the forged audio;
[0018] Determine the forged audio temporal features corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio.
[0019] In one embodiment, the determining of the forged audio temporal features corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio includes:
[0020] Determine the feature difference between the first forged frequency domain feature of the second forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the alternating frequency domain change feature of the target sub-audio, and obtain the alternating frequency domain change features of each target sub-audio;
[0021] Correspondingly splice the multiple second forged frequency domain features and the multiple phase-interval frequency domain change features corresponding to the forged audio to obtain multiple third forged frequency domain features of the forged audio;
[0022] Determine the forged audio timing feature corresponding to the forged audio based on the multiple third forged frequency domain features of the forged audio.
[0023] In one embodiment, the calculating the first similarity parameter between the audio to be recognized and multiple sets of forged audio timing features and the second similarity parameter between the audio to be recognized and multiple sets of real audio timing features includes:
[0024] Determine the average value of the multiple forged audio timing features in the set of forged audio timing features as the set audio feature of the set of forged audio timing features;
[0025] Calculate the audio feature similarity between the audio feature of the audio to be recognized and the multiple set audio features of the multiple sets of forged audio timing features to obtain the multiple audio feature similarities of the multiple sets of forged audio timing features;
[0026] Determine the maximum value of the multiple audio feature similarities of the multiple sets of forged audio timing features as the first similarity parameter.
[0027] In one embodiment, the extracting the timing features of the multiple forged audios and the multiple real audios to obtain the multiple forged audio timing features corresponding to the multiple forged audios and the multiple real audio timing features corresponding to the multiple real audios includes:
[0028] Perform frame division and windowing operations on the forged audio to obtain each frame of audio signal;
[0029] Perform fast Fourier transform on each frame of audio signal to obtain the amplitude spectrum in the linear frequency domain;
[0030] Use a linear frequency filter bank to filter the amplitude spectrum in the linear frequency domain to obtain the energy of each filter in the filter bank;
[0031] Based on the energy of each filter, perform logarithmic operation and discrete cosine transform to obtain the forged audio timing feature of the forged audio.
[0032] In a second aspect, an audio forgery detection device provided by an embodiment of the present application has a function of implementing the audio forgery detection method provided in the first aspect corresponding thereto. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware.
[0033] In one embodiment, the audio anti-counterfeiting device includes:
[0034] An acquisition module, configured to acquire a plurality of forged audios and a plurality of genuine audios;
[0035] A feature extraction module, configured to perform temporal feature extraction on the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios;
[0036] A partitioning module, configured to partition the plurality of forged audio temporal features and the plurality of genuine audio temporal features into different sets respectively, to obtain a plurality of forged audio temporal feature sets and a plurality of genuine audio temporal feature sets;
[0037] A calculation module, configured to calculate a first similarity parameter between the audio to be recognized and the plurality of forged audio temporal feature sets, and a second similarity parameter between the audio to be recognized and the plurality of genuine audio temporal feature sets respectively when the audio to be recognized is acquired;
[0038] A determination module, configured to determine that the audio category of the audio to be recognized is a forged category if the first similarity parameter is greater than the second similarity parameter.
[0039] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the audio anti-counterfeiting method as described in the first aspect.
[0040] In a fourth aspect, an embodiment of the present application provides a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the audio anti-counterfeiting method as described in the first aspect when executing the computer program.
[0041] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor coupled to a transceiver of a terminal device, for executing the technical solution provided in the first aspect of the embodiments of the present application.
[0042] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor for supporting a terminal device to implement the functions involved in the above first aspect, for example, generating or processing the information involved in the image processing method provided in the above first aspect.
[0043] In a possible design, the above chip system further includes a memory, which is used to store program instructions and data necessary for the terminal. The chip system may be composed of chips or may include chips and other discrete devices.
[0044] In a seventh aspect, an embodiment of the present application provides a computer program product including instructions. When the computer program product runs on a computer, the computer is caused to execute the audio anti-forgery method provided in the first aspect above.
[0045] Compared with the prior art, in an embodiment of the present application, multiple forged audios and multiple authentic audios are obtained; temporal features are extracted from the multiple forged audios and the multiple authentic audios to obtain multiple forged audio temporal features corresponding to the multiple forged audios and multiple authentic audio temporal features corresponding to the multiple authentic audios; the multiple forged audio temporal features and the multiple authentic audio temporal features are respectively divided into different sets to obtain multiple forged audio temporal feature sets and multiple authentic audio temporal feature sets; when an audio to be recognized is obtained, a first similarity parameter between the audio to be recognized and the multiple forged audio temporal feature sets and a second similarity parameter between the audio to be recognized and the multiple authentic audio temporal feature sets are respectively calculated; if the first similarity parameter is greater than the second similarity parameter, it is determined that the audio category of the audio to be recognized is a forged category. In the present application, a first similarity parameter between the audio to be recognized and a forged audio temporal feature set including multiple forged audio temporal features is calculated, and then a second similarity parameter between the audio to be recognized and an authentic audio temporal feature set including multiple authentic audio temporal features is calculated. Whether the audio category of the audio to be recognized is a forged category is determined according to the first similarity parameter and the second similarity parameter, which can utilize rich forged feature information and improve the accuracy of audio anti-forgery. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The objectives, features, and advantages of the embodiments of the present application will become easy to understand by referring to the detailed description of the embodiments of the present application with reference to the accompanying drawings. Among them:
[0047] Figure 1 It is a schematic diagram of an audio anti-forgery system for the audio anti-forgery method in an embodiment of the present application;
[0048] Figure 2 It is a schematic flowchart of the audio anti-forgery method in an embodiment of the present application;
[0049] Figure 3 It is a schematic structural diagram of an audio anti-forgery device in an embodiment of the present application;
[0050] Figure 4 It is a schematic structural diagram of a computing device in an embodiment of the present application;
[0051] Figure 5 It is a schematic structural diagram of a mobile phone in an embodiment of the present application;
[0052] Figure 6 It is a schematic structural diagram of a server in an embodiment of the present application.
[0053] In the accompanying drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners
[0054] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and the above drawings of the embodiments of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules does not necessarily limit to the clearly listed steps or modules, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. The division of modules in the embodiments of the present application is only a logical division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling and communication connection between modules can be in an electrical or other similar form, which are not limited in the embodiments of the present application. And the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.
[0055] The embodiments of the present application also provide an audio anti-counterfeiting method, a related device and a storage medium, which can be applied to an audio anti-counterfeiting system. The audio anti-counterfeiting system may include an audio anti-counterfeiting device, and the audio anti-counterfeiting device can be integrally deployed or separately deployed. The audio anti-counterfeiting device is at least used to obtain a plurality of forged audios and a plurality of genuine audios; extract temporal features from the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios; respectively divide the plurality of forged audio temporal features and the plurality of genuine audio temporal features into different sets to obtain a plurality of forged audio temporal feature sets and a plurality of genuine audio temporal feature sets; when an audio to be recognized is obtained, calculate a first similarity parameter between the audio to be recognized and the plurality of forged audio temporal feature sets, and a second similarity parameter between the audio to be recognized and the plurality of genuine audio temporal feature sets; if the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the audio to be recognized is a forged category.
[0056] The solution provided by the embodiments of this application involves technologies such as Artificial Intelligence (AI) and Machine Learning (ML). The specific description is as follows through the following embodiments:
[0057] Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0058] AI technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0059] Due to the difference in the data distribution at the business site and the data distribution in the model training stage, the existing audio forensics methods may have a high false rejection rate and a poor detection rate.
[0060] Compared with the prior art, in the embodiments of this application, multiple forged audios and multiple genuine audios are obtained; temporal features are extracted from the multiple forged audios and multiple genuine audios to obtain multiple forged audio temporal features corresponding to the multiple forged audios and multiple genuine audio temporal features corresponding to the multiple genuine audios; the multiple forged audio temporal features and the multiple genuine audio temporal features are respectively divided into different sets to obtain multiple forged audio temporal feature sets and multiple genuine audio temporal feature sets; when a to-be-identified audio is obtained, a first similarity parameter between the to-be-identified audio and the multiple forged audio temporal feature sets and a second similarity parameter between the to-be-identified audio and the multiple genuine audio temporal feature sets are respectively calculated; if the first similarity parameter is greater than the second similarity parameter, it is determined that the audio category of the to-be-identified audio is the forged category. In this application, the first similarity parameter between the to-be-identified audio and the forged audio temporal feature set containing multiple forged audio temporal features is calculated, and then the second similarity parameter between the to-be-identified audio and the genuine audio temporal feature set containing multiple genuine audio temporal features is calculated. According to the first similarity parameter and the second similarity parameter, it is determined whether the audio category of the to-be-identified audio is the forged category, which can utilize rich forged feature information and improve the accuracy of audio forensics.
[0061] In some embodiments, with reference to Figure 1 , the audio anti-forgery method provided by the embodiments of the present application can be implemented based on Figure 1 an audio anti-forgery system shown. The audio anti-forgery system may include an electronic device 100 and a memory 200. The electronic device 100 may be a server or a terminal device.
[0062] It should be noted that the server involved in the embodiments of the present application may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0063] The terminal device involved in the embodiments of the present application may be a device that provides voice and / or data connectivity to users, a handheld device with wireless connection capabilities, or other processing devices connected to a wireless modem. For example, a mobile phone (or a "cellular" phone) and a computer with a mobile terminal. For example, it may be a portable, pocket-sized, handheld, computer-integrated or vehicle-mounted mobile device that exchanges voice and / or data with a radio access network. For example, a Personal Communication Service (PCS) phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA), etc.
[0064] With reference to Figure 2 , Figure 2 is a schematic flowchart of an audio anti-forgery method provided by the embodiments of the present application. This method can be executed by an audio anti-forgery device. The method includes steps 101-105:
[0065] Step 101, obtain a plurality of forged audios and a plurality of genuine audios.
[0066] In the embodiments of the present application, the category of the forged audio is the forged category, and the category of the real audio is the real category. The real audio can be a naturally recorded audio. For example, the real audio is a recording of a person speaking naturally, that is, natural speech. The forged audio can be synthetically generated. The audio type of the forged audio can be forged type 1, forged type 2, etc., which can be set according to specific circumstances. Different forged types can represent forged voices in different fields, such as the mathematics field, the literature field, the finance field, etc., which can be set according to specific circumstances.
[0067] In a specific embodiment, a real audio is obtained; the real audio is segmented to obtain speech segments of the real audio; a random splicing process is performed on the speech segments of the real audio to obtain a spliced speech, and the spliced speech is determined as the forged audio.
[0068] To improve the randomness of segmentation, the real audio can be divided into multiple segments according to different durations. Therefore, in some embodiments, the process of segmenting the real audio to obtain the speech segments of the real audio includes: segmenting the real audio according to a first length to obtain speech segments of the first length of the real audio; segmenting the real audio according to a second length to obtain speech segments of the second length of the real audio. Among them, the speech segments of the real audio include the speech segments of the first length of the real audio and the speech segments of the second length of the real audio. Among them, the first length and the second length can be preset. For example, the first length can be 0.2 seconds, and the second length can be 0.4 seconds.
[0069] To improve the randomness of splicing, in some embodiments, the process of performing a random splicing process on the speech segments to obtain a spliced speech includes the following steps: obtaining a first segment set and a second segment set, where the first segment set includes speech segments of the first length of the real audio, and the second segment set includes speech segments of the second length of the real audio; randomly selecting a speech segment from the first segment set as the first speech segment; randomly selecting a speech segment from the second segment set as the second speech segment; splicing the first speech segment and the second speech segment to obtain a spliced speech.
[0070] Specifically, in some embodiments, randomly selecting a speech segment from the first segment set as the first speech segment can be randomly selecting N speech segments from the first segment set as the first speech segment. Randomly selecting a speech segment from the second segment set as the second speech segment can be randomly selecting N speech segments from the second segment set as the second speech segment. Among them, N is a positive integer.
[0071] Step 102: Extract the temporal features of multiple forged audio and multiple genuine audio to obtain multiple forged audio temporal features corresponding to the multiple forged audio and multiple genuine audio temporal features corresponding to the multiple genuine audio.
[0072] In an embodiment of the present application, in a specific embodiment, the multiple forged audio and the multiple genuine audio are respectively input into a first preset feature extraction model for temporal feature extraction, to obtain multiple forged audio temporal features corresponding to the multiple forged audio and multiple genuine audio temporal features corresponding to the multiple genuine audio. Among them, the first preset feature extraction model can be a recurrent neural network (RNN) or a Transformer model. The input of the first preset feature extraction model is serialized acoustic features or the original audio waveform. The recurrent connection of the hidden layer is used to capture the dependencies in the time series, and LSTM or GRU variants are used to solve the long-term dependence problem. For example, LSTM / GRU networks directly model the serialized acoustic features and are used for speaker recognition or speech emotion analysis.
[0073] In another specific embodiment, the extraction of the temporal features of the multiple forged audio and the multiple genuine audio to obtain multiple forged audio temporal features corresponding to the multiple forged audio and multiple genuine audio temporal features corresponding to the multiple genuine audio includes:
[0074] (1) Split the forged audio into multiple forged sub-audio sorted in time sequence.
[0075] Specifically, the forged audio is split into multiple forged sub-audio with equal durations sorted in time sequence. The number of the multiple forged sub-audio can be set according to specific circumstances. For example, if the duration of the forged audio is 8s, the forged audio is split into 8 forged sub-audio with a duration of 1s each.
[0076] (2) Respectively extract the frequency domain features of each forged sub-audio to obtain the first forged frequency domain features corresponding to each forged sub-audio.
[0077] Among them, the first forged frequency domain feature is the LFCC (Linear Frequency Cepstral Coefficients) feature. In other embodiments, the first forged frequency domain feature may be the MFCC (Mel Frequency Cepstral Coefficients) feature. MFCC (Mel Frequency Cepstral Coefficients) is a commonly used feature extraction method in speech processing, and its parameter settings will directly affect the performance of the features and the effect of subsequent audio forgery detection tasks. The LFCC feature is a feature extraction method for speech signal processing. The core difference from the traditional MFCC (Mel Frequency Cepstral Coefficients) lies in the way of dividing the frequency scale. MFCC is based on the Mel frequency scale (simulating the human ear's auditory characteristics), while LFCC directly uses the linear frequency scale and is suitable for scenarios that require linear frequency resolution (such as audio analysis, acoustic feature modeling, etc.).
[0078] In a specific embodiment, frequency domain feature extraction is performed on each forged sub-audio to obtain the first forged frequency domain feature corresponding to each forged sub-audio, including: performing framing and windowing operations on the forged sub-audio to obtain each frame of audio signal. Performing a fast Fourier transform on each frame of audio signal to obtain the amplitude spectrum in the linear frequency domain. Using a linear frequency filter bank to filter the amplitude spectrum in the linear frequency domain to obtain the energy of each filter in the filter bank. Performing logarithmic operation and discrete cosine transform based on the energy of each filter to obtain the first forged frequency domain feature corresponding to the forged sub-audio.
[0079] (3) Determine the forged audio timing feature corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio.
[0080] In a specific embodiment, the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio are determined as the forged audio timing feature corresponding to the forged audio.
[0081] For example, the duration of the forged audio is 8s. The forged audio is split into 8 forged sub-audios with a duration of 1s each, and each first forged frequency domain feature contains 5 coefficients. The multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio are represented by matrix P1. Matrix P1 contains 8 columns and 5 rows. Matrix P1 is an 8*5 matrix, and matrix P1 is as follows,
[0082]
[0083] In another specific embodiment, determining the multiple forged audio timing features corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio includes:
[0084] (1) Determine each forged sub-audio as the target sub-audio.
[0085] (2) Determine the feature difference between the first forged frequency domain feature of the first forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio, and obtain the adjacent frequency domain change feature of each target sub-audio.
[0086] Wherein, when there is no forged sub-audio after the target sub-audio, the first forged frequency domain feature of the target sub-audio is determined as the adjacent frequency domain change feature of the target sub-audio.
[0087] For example, represent the multiple adjacent frequency domain change features corresponding to the multiple forged sub-audios of the forged audio using matrix P2. Matrix P2 contains 8 columns and 5 rows, and matrix P2 is an 8*5 matrix. Matrix P2 is shown as follows.
[0088]
[0089] (3) Concatenate the multiple first forged frequency domain features and the multiple adjacent frequency domain change features corresponding to the forged audio to obtain the multiple second forged frequency domain features of the forged audio.
[0090] Specifically, record the multiple second forged frequency domain features of the forged audio as matrix P5. Matrix P5 contains 8 columns and 10 rows, and matrix P2 is an 8*10 matrix. Then P5 is shown as follows.
[0091]
[0092] (4) Determine the forged audio time series feature corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio.
[0093] In a specific embodiment, determine the multiple second forged frequency domain features of the forged audio as the forged audio time series feature corresponding to the forged audio.
[0094] In another specific embodiment, determining the forged audio time series feature corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio includes:
[0095] (1) Determine the feature difference between the first forged frequency domain feature of the second forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the interphase frequency domain change feature of the target sub-audio, and obtain the interphase frequency domain change feature of each target sub-audio.
[0096] Wherein, when there is no second forged sub-audio after the target sub-audio, the first forged frequency domain feature of the target sub-audio is determined as the adjacent frequency domain change feature.
[0097] For example, represent the multiple adjacent frequency domain change features corresponding to multiple forged sub-audios of the forged audio using matrix P3. Matrix P3 contains 8 columns and 5 rows, and matrix P3 is an 8*5 matrix. Matrix P3 is as follows:
[0098]
[0099] (2) Concatenate the multiple second forged frequency domain features and multiple alternating frequency domain change features corresponding to the forged audio to obtain multiple third forged frequency domain features of the forged audio;
[0100] Specifically, record the multiple third forged frequency domain features of the forged audio as matrix P6. Matrix P6 contains 8 columns and 15 rows, and matrix P6 is an 8*15 matrix. Then P6 is as follows:
[0101]
[0102] (3) Determine the forged audio time series features corresponding to the forged audio based on the multiple third forged frequency domain features of the forged audio.
[0103] In a specific embodiment, determine the multiple third forged frequency domain features of the forged audio as the forged audio time series features corresponding to the forged audio.
[0104] In another specific embodiment, determining the forged audio time series features corresponding to the forged audio based on the multiple third forged frequency domain features of the forged audio includes:
[0105] Input the multiple third forged frequency domain features of the forged audio into a second preset feature extraction model for feature extraction to obtain the forged audio time series features corresponding to the forged audio. The second preset feature extraction model can be a recurrent neural network (RNN) or a Transformer model.
[0106] Specifically, obtain a second preset feature extraction model and a classification model. The second preset feature extraction model is used to receive audio time series features and input the corrected audio time series features. The classification model is used to classify according to the corrected audio time series features. Obtain multiple audio samples and corresponding sample labels. The multiple audio samples include multiple forged audios and multiple real audios. The sample labels are forged categories and forged types, real categories and real types. For example, the sample label of audio sample 1 is: forged category, forged type 1. It means that audio sample 1 belongs to the forged category, and the forged type is specifically forged type 1.
[0107] Iteratively train a second preset feature extraction model and a classification model based on multiple audio samples and corresponding sample labels to obtain the trained second preset feature extraction model and classification model. Input the multiple third forged frequency domain features of the forged audio into the trained second preset feature extraction model for feature extraction to obtain the forged audio time series features corresponding to the forged audio.
[0108] Based on the same principle, in the embodiments of the present application, perform time series feature extraction on multiple said forged audios and multiple said real audios to obtain multiple forged audio time series features corresponding to the multiple said forged audios and multiple real audio time series features corresponding to the multiple said real audios, further including:
[0109] (1) Split the real audio into multiple real sub-audios sorted in time sequence.
[0110] (2) Perform frequency domain feature extraction on each real sub-audio respectively to obtain the first real frequency domain feature corresponding to each real sub-audio.
[0111] Wherein, the first real frequency domain feature is the LFCC (Linear Frequency Cepstral Coefficients) feature.
[0112] In a specific embodiment, performing frequency domain feature extraction on each real sub-audio respectively to obtain the first real frequency domain feature corresponding to each real sub-audio includes: performing frame splitting and windowing operations on the real sub-audio to obtain each frame of audio signal. Perform a fast Fourier transform on each frame of audio signal to obtain the amplitude spectrum in the linear frequency domain. Use a linear frequency filter bank to filter the amplitude spectrum in the linear frequency domain to obtain the energy of each filter in the filter bank. Based on the energy of each filter, perform logarithmic operation and discrete cosine transform to obtain the first real frequency domain feature corresponding to the real sub-audio.
[0113] (3) Determine the real audio time series feature corresponding to the real audio based on the multiple first real frequency domain features corresponding to the multiple real sub-audios of the real audio.
[0114] In a specific embodiment, determine the multiple first real frequency domain features corresponding to the multiple real sub-audios of the real audio as the real audio time series feature corresponding to the real audio.
[0115] In another specific embodiment, determining the multiple real audio time series features corresponding to the real audio based on the multiple first real frequency domain features corresponding to the multiple real sub-audios of the real audio includes:
[0116] (1) Determine each real sub-audio as the target sub-audio respectively.
[0117] (2) Determine the feature difference between the first true frequency domain feature of the first true sub-audio after the target sub-audio and the first true frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio, and obtain the adjacent frequency domain change feature of each target sub-audio.
[0118] Among them, when there is no true sub-audio after the target sub-audio, determine the first true frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio.
[0119] (3) Concatenate the multiple first true frequency domain features and the multiple adjacent frequency domain change features corresponding to the true audio to obtain multiple second true frequency domain features of the true audio.
[0120] (4) Determine the true audio time series feature corresponding to the true audio based on the multiple second true frequency domain features of the true audio.
[0121] In a specific embodiment, determine the multiple second true frequency domain features of the true audio as the true audio time series feature corresponding to the true audio.
[0122] In another specific embodiment, determining the true audio time series feature corresponding to the true audio based on the multiple second true frequency domain features of the true audio includes:
[0123] (1) Determine the feature difference between the first true frequency domain feature of the second true sub-audio after the target sub-audio and the first true frequency domain feature of the target sub-audio as the interphase frequency domain change feature of the target sub-audio, and obtain the interphase frequency domain change feature of each target sub-audio.
[0124] Among them, when there is no second true sub-audio after the target sub-audio, determine the first true frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio.
[0125] (2) Concatenate the multiple second true frequency domain features and the multiple interphase frequency domain change features corresponding to the true audio to obtain multiple third true frequency domain features of the true audio;
[0126] (3) Determine the true audio time series feature corresponding to the true audio based on the multiple third true frequency domain features of the true audio.
[0127] In a specific embodiment, determine the multiple third true frequency domain features of the true audio as the true audio time series feature corresponding to the true audio.
[0128] In another specific embodiment, determining the true audio time series feature corresponding to the true audio based on the multiple third true frequency domain features of the true audio includes:
[0129] Input multiple third true frequency domain features of the real audio into the second preset feature extraction model for feature extraction to obtain the real audio time series features corresponding to the real audio. The second preset feature extraction model can be a recurrent neural network (RNN) or a Transformer model.
[0130] Specifically, obtain the second preset feature extraction model and the classification model. The second preset feature extraction model is used to receive the audio time series features and input the corrected audio time series features. The classification model is used to classify according to the corrected audio time series features. Obtain multiple audio samples and their corresponding sample labels. The multiple audio samples include multiple real audios and multiple forged audios. The sample labels are the real category and the real type, the real category and the real type. For example, the sample label of audio sample 1 is: real category, real type 1. It means that audio sample 1 belongs to the real category, and the real type is specifically real type 1.
[0131] Iteratively train the second preset feature extraction model and the classification model based on the multiple audio samples and their corresponding sample labels to obtain the trained second preset feature extraction model and classification model. Input multiple third true frequency domain features of the real audio into the trained second preset feature extraction model for feature extraction to obtain the real audio time series features corresponding to the real audio.
[0132] Step 103, respectively divide multiple real audio time series features and multiple real audio time series features into different sets to obtain multiple real audio time series feature sets and multiple real audio time series feature sets.
[0133] In a specific embodiment, respectively divide multiple forged audio time series features and multiple real audio time series features into different sets to obtain multiple forged audio time series feature sets and multiple real audio time series feature sets, including: obtain the forged type of the forged audio corresponding to the forged audio time series features, and put the forged audio time series features of the same forged type into the same set as a forged audio time series feature set to obtain multiple forged audio time series feature sets. For example, in the forged audio time series feature set corresponding to the forged type 1, the forged types of the forged audios corresponding to each forged audio time series feature are all forged type 1. In the forged audio time series feature set corresponding to the forged type 2, the forged types of the forged audios corresponding to each forged audio time series feature are all forged type 2.
[0134] Furthermore, obtain the real type of the real audio corresponding to the real audio time series features, and put the real audio time series features of the same real type into the same set as a real audio time series feature set to obtain multiple real audio time series feature sets.
[0135] In another specific embodiment, a plurality of forged audio temporal features and a plurality of real audio temporal features are respectively divided into different sets, obtaining a plurality of forged audio temporal feature sets and a plurality of real audio temporal feature sets, including: clustering the plurality of forged audio temporal features to obtain a plurality of feature clustering clusters; determining one feature clustering cluster as one forged audio temporal feature set, obtaining a plurality of forged audio temporal feature sets.
[0136] Further, clustering the plurality of real audio temporal features to obtain a plurality of feature clustering clusters; determining one feature clustering cluster as one real audio temporal feature set, obtaining a plurality of real audio temporal feature sets.
[0137] Step 104, when the audio to be recognized is obtained, calculate the first similarity parameter between the audio to be recognized and the plurality of forged audio temporal feature sets, and the second similarity parameter between the audio to be recognized and the plurality of real audio temporal feature sets respectively.
[0138] In the embodiment of the present application, calculating the first similarity parameter between the audio to be recognized and the plurality of forged audio temporal feature sets, and the second similarity parameter between the audio to be recognized and the plurality of real audio temporal feature sets respectively, includes:
[0139] (1) Determine the average value of the plurality of forged audio temporal features in the forged audio temporal feature set as the set audio feature of the forged audio temporal feature set.
[0140] (2) Calculate the audio feature similarity between the audio feature of the audio to be recognized and the plurality of set audio features of the plurality of forged audio temporal feature sets, obtaining the plurality of audio feature similarities of the plurality of forged audio temporal feature sets.
[0141] In the embodiment of the present application, denote the set audio feature as m ij , where i ∈ {0, 1} represents the true or false category, i = 0 represents the true category, i = 1 represents the forged category, and j represents the serial number of the set corresponding to the set audio feature.
[0142] In the embodiment of the present application, calculate the audio feature similarity between the audio feature of the audio to be recognized and the set audio feature of the forged audio temporal feature set, obtaining the audio feature similarity of the forged audio temporal feature set. Among them, the audio feature similarity can be a measurement method such as cosine distance, Euclidean distance, etc.
[0143] In the embodiment of the present application, the audio feature similarity is the Euclidean distance. For the audio x to be recognized, extract the features of the audio x to be recognized, obtaining the feature f(x) of the audio x to be recognized, and calculate the Euclidean distance between the feature f(x) of the audio x to be recognized and each set audio feature m ij The Euclidean distance is used to measure the audio x to be recognized and the set audio feature mij The similarity between them. The feature f(x) of the audio x to be recognized and the audio features m of each set ij The Euclidean distance is calculated by the following formula
[0144]
[0145] where d(f(x), m ij ) represents the Euclidean distance between the feature f(x) of the audio x to be recognized and the audio feature m of the set audio, that is, the audio feature similarity ij
[0146] (3) Determine the maximum value of the audio feature similarities of multiple forged audio temporal feature sets as the first similarity parameter.
[0147] For multiple set audio features of multiple forged audio temporal feature sets of forged categories, the first similarity parameter d1 is calculated as follows:
[0148] d1 = max d(f(x), m ij ), i = 1, j ∈ {1, …, K}.
[0149] where d1 represents the first similarity parameter, and d(f(x), m ij ) represents the audio feature similarity between the feature f(x) of the audio x to be recognized and the audio feature m of the set audio ij
[0150] In the embodiments of the present application, calculating the first similarity parameter between the audio to be recognized and multiple forged audio temporal feature sets, and the second similarity parameter between the audio to be recognized and multiple real audio temporal feature sets respectively includes:
[0151] (1) Determine the average value of multiple real audio temporal features in the real audio temporal feature set as the set audio feature of the real audio temporal feature set;
[0152] (2) Calculate the audio feature similarity between the audio feature of the audio to be recognized and the multiple set audio features of multiple real audio temporal feature sets to obtain multiple audio feature similarities of multiple real audio temporal feature sets.
[0153] (3) Weighted sum the multiple audio feature similarities of multiple real audio temporal feature sets based on the set weights of each real audio temporal feature set to obtain the second similarity parameter.
[0154] In a specific embodiment, obtain the feature quantity parameter of the true audio time-series features in the set of true audio time-series features, obtain the signal-to-noise ratio parameter of the true audio corresponding to the set of true audio time-series features, and determine the set weight γ of the set of true audio time-series features based on the feature quantity parameter and the signal-to-noise ratio parameter of the set of true audio time-series features j The set weight corresponding to the j-th set of true audio time-series features is γ j The signal-to-noise ratio (SNR) is the ratio of the audio signal power to the noise power, with the unit of decibel (dB). The higher the SNR, the smaller the impact of noise (such as current noise, background noise) on the sound quality
[0155] When determining the set weight of the set of true audio time-series features, the feature quantity parameter and the signal-to-noise ratio parameter can be first normalized and mapped to the interval [0, 1], and then the comprehensive weight is calculated through a preset weight distribution formula. Specifically, let the feature quantity parameter be N (reflecting the richness of the feature dimension), the signal-to-noise ratio parameter be SNR (reflecting the signal purity), the normalized feature quantity be N j and the normalized signal-to-noise ratio be S j Then the set weight γ j can be expressed as γ j = a * N j + b * S j , where both a and b are within [0, 1] and can be set according to specific circumstances
[0156] Suppose that the number of true audio time-series features in one set of true audio time-series features is 10, the feature quantity parameter is 10, and the SNR of this audio is 25 dB. After normalization, the feature quantity parameter is 0.5, and the normalized signal-to-noise ratio is 0.5. If a = 0.6 and b = 0.4 are set, the j-th set weight γ j = 0.5
[0157] For the multiple set audio features of the multiple sets of true audio time-series features of the true category, the second similarity parameter d2 is calculated as follows
[0158]
[0159] where γ j represents the set weight corresponding to the j-th set of true audio time-series features
[0160] In the embodiment of the present application, obtain the hardware parameters of the electronic device, determine the maximum number of processes and the maximum processing quantity of each process based on the GPU hardware parameters of the electronic device. Establish multiple processes based on the maximum number of processes and the maximum processing quantity of each process, and control each process to process the corresponding quantity of true audio
[0161] Step 105, if the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the audio to be recognized is a forged category.
[0162] Further, if the first similarity parameter is less than the second similarity parameter, determine that the audio category of the audio to be recognized is a genuine category.
[0163] Compared with the prior art, in the embodiments of the present application, multiple forged audios and multiple genuine audios are obtained; temporal features are extracted from the multiple forged audios and the multiple genuine audios to obtain multiple forged audio temporal features corresponding to the multiple forged audios and multiple genuine audio temporal features corresponding to the multiple genuine audios; the multiple forged audio temporal features and the multiple genuine audio temporal features are respectively divided into different sets to obtain multiple forged audio temporal feature sets and multiple genuine audio temporal feature sets; when the audio to be recognized is obtained, the first similarity parameter between the audio to be recognized and the multiple forged audio temporal feature sets and the second similarity parameter between the audio to be recognized and the multiple genuine audio temporal feature sets are respectively calculated; if the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the audio to be recognized is a forged category. In the present application, the first similarity parameter between the audio to be recognized and the forged audio temporal feature set containing multiple forged audio temporal features is calculated, and then the second similarity parameter between the audio to be recognized and the genuine audio temporal feature set containing multiple genuine audio temporal features is calculated, and the audio category of the audio to be recognized is determined to be a forged category according to the first similarity parameter and the second similarity parameter, which can utilize rich forged feature information and improve the accuracy of audio forgery detection.
[0164] Refer to Figure 3 ,such as Figure 3 the structural schematic diagram of an audio forgery detection device shown. The audio forgery detection device in the embodiments of the present application can implement the steps corresponding to the audio forgery detection method executed in the corresponding embodiments in the above Figure 2 The functions implemented by the audio forgery detection device can be realized by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. For the function implementation of the acquisition module 601, the feature extraction module 602, the division module 603, the calculation module 604, and the determination module 605 of the audio forgery detection device 60, reference can be made to Figure 2 the operations executed in the corresponding embodiments, which will not be elaborated here.
[0165] The audio forgery detection device includes:
[0166] An acquisition module 601, configured to acquire multiple forged audios and multiple genuine audios;
[0167] The feature extraction module 602 is configured to perform temporal feature extraction on the multiple forged audios and the multiple genuine audios, so as to obtain multiple forged audio temporal features corresponding to the multiple forged audios and multiple genuine audio temporal features corresponding to the multiple genuine audios;
[0168] The partitioning module 603 is configured to partition the multiple forged audio temporal features and the multiple genuine audio temporal features into different sets respectively, so as to obtain multiple forged audio temporal feature sets and multiple genuine audio temporal feature sets;
[0169] The calculation module 604 is configured to, when the audio to be recognized is obtained, calculate a first similarity parameter between the audio to be recognized and the multiple forged audio temporal feature sets, and a second similarity parameter between the audio to be recognized and the multiple genuine audio temporal feature sets respectively;
[0170] The determination module 605 is configured to determine that the audio category of the audio to be recognized is the forged category if the first similarity parameter is greater than the second similarity parameter.
[0171] The audio forgery detection device 60 in the embodiments of the present application has been described above from the perspective of modular functional entities. Next, the audio forgery detection device in the embodiments of the present application will be described from the perspective of hardware processing.
[0172] Figure 3 The devices shown can all have the structure as Figure 4 shown. When Figure 3 the audio forgery detection device 60 shown has the structure as Figure 4 shown, Figure 4 the processor and transceiver in Figure 4 can implement the same or similar functions as the modules provided in the device embodiments corresponding to the device, and
[0173] The embodiments of the present application also provide a terminal device. As Figure 5 shown, for the sake of convenience of description, only parts related to the embodiments of the present application are shown. For those specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) terminal device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0174] Figure 5 What is shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided in the embodiments of the present application. Refer to Figure 5, the mobile phone includes components such as a Radio Frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, sensors 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art can understand that Figure 5 the mobile phone structure shown in
[0175] does not limit the mobile phone and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 5 The following specifically introduces each component of the mobile phone:
[0176] The RF circuit 1010 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1080 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 1010 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0177] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1020 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0178] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 can include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect the touch operations of the user on or near it (such as the operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1031), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1031 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive the commands sent by the processor 1080 and execute them. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1031. In addition to the touch panel 1031, the input unit 1030 can also include other input devices 1032. Specifically, the other input devices 1032 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, joysticks, etc.
[0179] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 5 , the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and input functions of the mobile phone, but in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.
[0180] The mobile phone may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.
[0181] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent to another mobile phone, for example, via the RF circuit 1010, or the audio data is output to the memory 1020 for further processing.
[0182] Wi-Fi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the Wi-Fi module 1070. It provides users with wireless broadband Internet access. Although Figure 5 the Wi-Fi module 1070 is shown, it can be understood that it does not belong to the essential components of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.
[0183] The processor 1080 is the control center of the mobile phone. It connects various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1020, and by calling data stored in the memory 1020, it performs various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.
[0184] The mobile phone further includes a power source 1090 (such as a battery) that powers each component. Optionally, the power source can be logically connected to the processor 1080 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system.
[0185] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0186] In the embodiment of the present application, the processor 1080 included in the mobile phone also has the function of controlling the execution of the above audio anti-counterfeiting method flow performed by the audio anti-counterfeiting device.
[0187] The embodiment of the present application also provides a server. Please refer to Figure 6 , Figure 6It is a schematic diagram of a server structure provided by an embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (full English name: central processing units, English abbreviation: CPU) 1122 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 for storing application programs 1142 or data 1144 (for example, one or more mass storage devices). Among them, the memory 1132 and the storage media 1130 may be transient storage or persistent storage. The programs stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1122 may be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the server 1100.
[0188] The server 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and so on.
[0189] The steps performed by the server in the above embodiments may be based on the Figure 6 structure of the server 1100 shown. For example, for example, the steps performed by the Figure 6 audio forgery detection device 60 shown may be based on the Figure 6 server structure shown. For example, the central processing unit 1122 performs the following operations by calling the instructions in the memory 1132:
[0190] Obtain a plurality of forged audios and a plurality of genuine audios;
[0191] Extract temporal features from the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios;
[0192] Respectively divide the plurality of forged audio temporal features and the plurality of genuine audio temporal features into different sets to obtain a plurality of forged audio temporal feature sets and a plurality of genuine audio temporal feature sets;
[0193] When the audio to be recognized is obtained, calculate the first similarity parameter between the audio to be recognized and multiple sets of forged audio temporal features, and the second similarity parameter between the audio to be recognized and multiple sets of real audio temporal features respectively;
[0194] If the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the audio to be recognized is the forged category.
[0195] In one embodiment, the step of performing temporal feature extraction on the multiple forged audios and the multiple real audios to obtain multiple sets of forged audio temporal features corresponding to the multiple forged audios and multiple sets of real audio temporal features corresponding to the multiple real audios includes:
[0196] Split the forged audio into multiple forged sub-audios sorted in time sequence;
[0197] Perform frequency domain feature extraction on each forged sub-audio respectively to obtain the first forged frequency domain feature corresponding to each forged sub-audio;
[0198] Determine the forged audio temporal feature corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio.
[0199] In one embodiment, the step of determining the forged audio temporal feature corresponding to the forged audio based on the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio includes:
[0200] Determine each forged sub-audio as the target sub-audio respectively;
[0201] Determine the feature difference between the first forged frequency domain feature of the first forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio, and obtain the adjacent frequency domain change feature of each target sub-audio;
[0202] Concatenate the multiple first forged frequency domain features corresponding to the multiple forged sub-audios of the forged audio and the multiple adjacent frequency domain change features correspondingly to obtain the multiple second forged frequency domain features of the forged audio;
[0203] Determine the forged audio temporal feature corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio.
[0204] In one embodiment, the step of determining the forged audio temporal feature corresponding to the forged audio based on the multiple second forged frequency domain features of the forged audio includes:
[0205] Determine the feature difference between the first forged frequency-domain feature of the second forged sub-audio after the target sub-audio and the first forged frequency-domain feature of the target sub-audio as the alternating frequency-domain change feature of the target sub-audio, and obtain the alternating frequency-domain change feature of each target sub-audio;
[0206] Correspondingly splice the multiple second forged frequency-domain features and the multiple alternating frequency-domain change features corresponding to the forged audio to obtain multiple third forged frequency-domain features of the forged audio;
[0207] Determine the forged audio timing feature corresponding to the forged audio based on the multiple third forged frequency-domain features of the forged audio.
[0208] In one embodiment, the calculating the first similarity parameter between the audio to be recognized and multiple sets of forged audio timing features, and the second similarity parameter between the audio to be recognized and multiple sets of real audio timing features includes:
[0209] Determine the average value of the multiple forged audio timing features in the set of forged audio timing features as the set audio feature of the set of forged audio timing features;
[0210] Calculate the audio feature similarity between the audio feature of the audio to be recognized and the multiple set audio features of the multiple sets of forged audio timing features to obtain the multiple audio feature similarities of the multiple sets of forged audio timing features;
[0211] Determine the maximum value of the multiple audio feature similarities of the multiple sets of forged audio timing features as the first similarity parameter.
[0212] In one embodiment, the extracting the timing features of the multiple forged audios and the multiple real audios to obtain the multiple forged audio timing features corresponding to the multiple forged audios and the multiple real audio timing features corresponding to the multiple real audios includes:
[0213] Perform frame division and windowing operations on the forged audio to obtain each frame of audio signal;
[0214] Perform fast Fourier transform on each frame of audio signal to obtain the amplitude spectrum in the linear frequency domain;
[0215] Use a linear frequency filter bank to filter the amplitude spectrum in the linear frequency domain to obtain the energy of each filter in the filter bank;
[0216] Perform logarithmic operation and discrete cosine transform based on the energy of each filter to obtain the forged audio timing feature of the forged audio.
[0217] In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0218] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0219] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces, and the indirect coupling or communication connection of devices or modules may be in electrical, mechanical, or other forms.
[0220] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0221] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules. If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0222] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementation manners.
[0223] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0224] A computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, the procedures or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server, a data center, etc. that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0225] The technical solutions provided by the embodiments of the present application have been introduced in detail above. In the embodiments of the present application, specific examples are used to illustrate the principles and implementation manners of the embodiments of the present application. The descriptions of the above embodiments are only used to help understand the methods and the core ideas of the embodiments of the present application. At the same time, for those of ordinary skill in the art, according to the ideas of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the embodiments of the present application.
Claims
1. An audio anti-counterfeiting method, characterized in that, The audio forgery detection method includes: Obtaining a plurality of forged audios and a plurality of genuine audios; Performing temporal feature extraction on the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios; Respectively dividing the plurality of forged audio temporal features and the plurality of genuine audio temporal features into different sets to obtain a plurality of forged audio temporal feature sets and a plurality of genuine audio temporal feature sets; When a to-be-identified audio is obtained, respectively calculating a first similarity parameter between the to-be-identified audio and the plurality of forged audio temporal feature sets, and a second similarity parameter between the to-be-identified audio and the plurality of genuine audio temporal feature sets; If the first similarity parameter is greater than the second similarity parameter, determining that the audio category of the to-be-identified audio is a forged category.
2. The audio anti-counterfeiting method according to claim 1, wherein, The performing temporal feature extraction on the plurality of forged audios and the plurality of genuine audios to obtain a plurality of forged audio temporal features corresponding to the plurality of forged audios and a plurality of genuine audio temporal features corresponding to the plurality of genuine audios includes: Splitting the forged audio into a plurality of forged sub-audios sorted in time sequence; Performing frequency domain feature extraction on each forged sub-audio to obtain a first forged frequency domain feature corresponding to each forged sub-audio; Determining the forged audio temporal feature corresponding to the forged audio based on the plurality of first forged frequency domain features corresponding to the plurality of forged sub-audios of the forged audio.
3. The audio forgery detection method according to claim 2, wherein The determining the forged audio temporal feature corresponding to the forged audio based on the plurality of first forged frequency domain features corresponding to the plurality of forged sub-audios of the forged audio includes: Respectively determining each forged sub-audio as a target sub-audio; Determining the feature difference between the first forged frequency domain feature of the first forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the adjacent frequency domain change feature of the target sub-audio, and obtaining the adjacent frequency domain change feature of each target sub-audio; Correspondingly splicing the plurality of first forged frequency domain features and the plurality of adjacent frequency domain change features corresponding to the forged audio to obtain a plurality of second forged frequency domain features of the forged audio; Determining the forged audio temporal feature corresponding to the forged audio based on the plurality of second forged frequency domain features of the forged audio.
4. The audio forgery detection method according to claim 3, wherein The determining the forged audio temporal feature corresponding to the forged audio based on the plurality of second forged frequency domain features of the forged audio includes: Determining the feature difference between the first forged frequency domain feature of the second forged sub-audio after the target sub-audio and the first forged frequency domain feature of the target sub-audio as the alternate frequency domain change feature of the target sub-audio, and obtaining the alternate frequency domain change feature of each target sub-audio; Correspondingly splicing the plurality of second forged frequency domain features and the plurality of alternate frequency domain change features corresponding to the forged audio to obtain a plurality of third forged frequency domain features of the forged audio; Determine the forged audio time series feature corresponding to the forged audio based on the multiple third forged frequency domain features of the forged audio.
5. The audio forgery detection method according to claim 1, characterized in that The calculating the first similarity parameter between the audio to be identified and multiple forged audio time series feature sets, and the second similarity parameter between the audio to be identified and multiple real audio time series feature sets includes: Determine the average value of multiple forged audio time series features in the forged audio time series feature set as the set audio feature of the forged audio time series feature set; Calculate the audio feature similarity between the audio feature of the audio to be identified and multiple set audio features of multiple forged audio time series feature sets to obtain multiple audio feature similarities of multiple forged audio time series feature sets; Determine the maximum value of multiple audio feature similarities of multiple forged audio time series feature sets as the first similarity parameter.
6. The audio forgery detection method according to claim 1, wherein The extracting time series features from multiple forged audios and multiple real audios to obtain multiple forged audio time series features corresponding to multiple forged audios and multiple real audio time series features corresponding to multiple real audios includes: Perform frame division and windowing operations on the forged audio to obtain each frame of audio signal; Perform fast Fourier transform on each frame of audio signal to obtain the amplitude spectrum in the linear frequency domain; Use a linear frequency filter bank to filter the amplitude spectrum in the linear frequency domain to obtain the energy of each filter in the filter bank; Perform logarithmic operation and discrete cosine transform based on the energy of each filter to obtain the forged audio time series feature of the forged audio.
7. An audio anti-counterfeiting device, characterized in that, The audio forgery detection device includes: An acquisition module configured to acquire multiple forged audios and multiple real audios; A feature extraction module configured to extract time series features from multiple forged audios and multiple real audios to obtain multiple forged audio time series features corresponding to multiple forged audios and multiple real audio time series features corresponding to multiple real audios; A division module configured to respectively divide multiple forged audio time series features and multiple real audio time series features into different sets to obtain multiple forged audio time series feature sets and multiple real audio time series feature sets; A calculation module configured to, when an audio to be identified is acquired, calculate the first similarity parameter between the audio to be identified and multiple forged audio time series feature sets, and the second similarity parameter between the audio to be identified and multiple real audio time series feature sets; A determination module configured to, if the first similarity parameter is greater than the second similarity parameter, determine that the audio category of the audio to be identified is a forged category.
8. A computing device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1-6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1-6.
10. A computer program product comprising instructions, the computer program product including program instructions which, when run on a computer or a processor, cause the computer or the processor to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Counterfeit voice recognition method and device, electronic equipment and storage medium
CN113257255A
Audio authenticity detection method, related device and storage medium
CN116030831A
Voice authentic identification method and related equipment
CN119360887A
Deep counterfeit audio identification method based on hybrid enhancement and contrast loss function
CN119418715A
Method, device, and computer-readable storage medium for recognizing fake speech
WO2021135454A1