Audio verification method and device, audio forensics method and device
By introducing ultrasonic features into audio verification, the security threat posed by deepfake audio technology is addressed, enabling high-precision, transferable verification of audio data authenticity and ensuring audio quality and detection accuracy.
Patent Information
- Application Number
- CN202211504136.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-11-28
AI Technical Summary
Existing technologies are insufficient to effectively defend against the security threats posed by deepfake audio technology, especially AI-based audio tampering methods. Furthermore, existing defense methods suffer from low detection accuracy, poor portability, high cost, or environmental limitations.
By introducing ultrasound as a credibility factor, ultrasound features are extracted from the audio to be verified, and the authenticity of the speech data is verified based on the correlation between ultrasound features and speech data. The frequency range of ultrasound does not overlap with that of low-frequency speech, thus avoiding the impact on speech quality. Detection is performed by combining Doppler frequency shift, acoustic nonlinearity, and time-of-flight features.
It improves the accuracy and transferability of audio detection, can detect various tampering methods, ensures the authenticity and quality of audio data, and reduces the false positive rate of forged audio.
Smart Images

Figure CN116030820B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology, specifically to an audio verification method and apparatus, and an audio forensics method and apparatus. Background Technology
[0002] Speech corpus information is a large, information-rich, and intuitive component in judicial evidence collection and reliable evidence preservation, but it is also the most vulnerable link. There are numerous and inexpensive methods to tamper with speech corpus information, and deepfake audio technology based on artificial intelligence further exacerbates the security threats faced by speech corpus information. Summary of the Invention
[0003] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide an audio verification method and apparatus, and an audio forensics method and apparatus.
[0004] In a first aspect, one embodiment of this application provides an audio verification method, comprising: acquiring an audio to be verified; if the audio to be verified contains a first ultrasonic wave, determining the ultrasonic feature corresponding to the first ultrasonic wave; extracting speech data contained in the audio to be verified; and verifying whether the speech data is speech data expressed by a target object in a target scene based on the correlation between the ultrasonic feature and the speech data.
[0005] In conjunction with the first aspect, in some implementations of the first aspect, before determining the ultrasonic feature corresponding to the first ultrasonic wave, the method further includes: if the audio to be verified contains the first ultrasonic wave, then determining the time-spectrum diagram corresponding to the first ultrasonic wave; based on the time-spectrum diagram corresponding to the first ultrasonic wave, determining the slope change trend corresponding to the first ultrasonic wave; obtaining the slope change trend corresponding to the second ultrasonic wave, wherein the second ultrasonic wave is the ultrasonic wave emitted towards the target object when the target object emits sound in the target scene in advance; determining whether the slope change trend corresponding to the first ultrasonic wave and the slope change trend corresponding to the second ultrasonic wave are the same; the method further includes, when the slope change trend corresponding to the first ultrasonic wave and the slope change trend corresponding to the second ultrasonic wave are the same, determining the ultrasonic feature corresponding to the first ultrasonic wave.
[0006] In conjunction with the first aspect, in some implementations of the first aspect, the audio verification method further includes: if the slope change trend corresponding to the first ultrasonic wave and the slope change trend corresponding to the second ultrasonic wave are different, then it is determined that the voice data is not the voice data expressed by the target object in the target scene.
[0007] In conjunction with the first aspect, in some implementations of the first aspect, determining the slope change trend of the first ultrasonic wave based on the time-spectrum graph corresponding to the first ultrasonic wave includes: performing binarization processing on the time-spectrum graph corresponding to the first ultrasonic wave based on the amplitude at each moment in the time-spectrum graph corresponding to the first ultrasonic wave; calculating the slope between adjacent moments with the same value in the binarized time-spectrum graph; and determining the slope change trend of the first ultrasonic wave based on the slope between adjacent moments with the same value.
[0008] In conjunction with the first aspect, in some implementations of the first aspect, the audio verification method further includes: determining the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data respectively; for each of the multiple spectral channels, determining the similarity between the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data; and determining the correlation between the ultrasonic feature and the speech data based on the similarity corresponding to each of the multiple spectral channels.
[0009] In conjunction with the first aspect, in certain implementations of the first aspect, determining the time-spectrum map corresponding to the ultrasonic features and the time-spectrum map corresponding to the speech data respectively includes: performing speech endpoint detection on the speech data to obtain speech endpoint detection results; using the speech endpoint detection results to segment the ultrasonic features and speech data respectively to obtain the ultrasonic features of the speech segment and the speech data of the speech segment; determining the time-spectrum map corresponding to the ultrasonic features of the speech segment as the time-spectrum map corresponding to the ultrasonic features; and determining the time-spectrum map corresponding to the speech data of the speech segment as the time-spectrum map corresponding to the speech data.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, the ultrasonic features include Doppler frequency shift features; or, the ultrasonic features include Doppler frequency shift features, and the ultrasonic features also include acoustic nonlinear features and / or time-of-flight features.
[0011] Secondly, one embodiment of this application provides an audio forensics method, including: transmitting ultrasonic waves to a target object located in a target scene; when the target object makes a sound in the target scene, collecting the voice data expressed by the target object and the ultrasonic waves reflected by the target object to obtain credible audio data corresponding to the target object.
[0012] Thirdly, one embodiment of this application provides an audio verification device, comprising: an acquisition module for acquiring audio to be verified; a determination module for determining the ultrasonic feature corresponding to the first ultrasonic wave if the audio to be verified contains a first ultrasonic wave; an extraction module for extracting speech data contained in the audio to be verified; and a verification module for verifying whether the speech data is speech data expressed by a target object in a target scene based on the correlation between the ultrasonic feature and the speech data.
[0013] Fourthly, one embodiment of this application provides an audio evidence collection device, including: a transmitting module for transmitting ultrasonic waves to a target object located in a target scene; and a collecting module for collecting voice data expressed by the target object and ultrasonic waves reflected by the target object when the target object makes a sound in the target scene, thereby obtaining credible audio data corresponding to the target object.
[0014] Fifthly, one embodiment of this application provides a computer-readable storage medium storing a computer program for performing the methods described in the first and second aspects.
[0015] In a sixth aspect, one embodiment of this application provides an electronic device, the electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being configured to perform the methods described in the first and second aspects.
[0016] The audio verification method provided in this application has the following beneficial effects:
[0017] First, although ultrasound is inaudible to the human ear, it possesses high perceptual accuracy. This application introduces a reliability factor—ultrasound—and detects whether the speech data in the audio to be verified has been tampered with based on the corresponding ultrasonic features. This improves the detection accuracy of the audio to be verified. Furthermore, the frequency range of ultrasound does not overlap with the frequency range of low-frequency speech, so it will not affect the quality of the speech data in the audio to be verified. Second, the audio verification method in this application has strong transferability, enabling the detection of tampering methods that cause a mismatch between ultrasonic features and speech data, thereby resisting other unknown tampering methods. Attached Figure Description
[0018] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0019] Figure 1a The diagram shown is an application scenario applicable to an embodiment of this application.
[0020] Figure 1b The diagram shown illustrates another application scenario applicable to the embodiments of this application.
[0021] Figure 2 The diagram shown is a flowchart of an audio verification method provided in an exemplary embodiment of this application.
[0022] Figure 3The diagram shown is a flowchart illustrating the process of determining the ultrasonic features corresponding to a first ultrasonic wave, as provided in an exemplary embodiment of this application.
[0023] Figure 4 The diagram shown is a flowchart illustrating the process of determining the slope change trend corresponding to the first ultrasound wave according to an exemplary embodiment of this application.
[0024] Figure 5 The diagram shown is a flowchart of an audio verification method provided in another exemplary embodiment of this application.
[0025] Figure 6 The diagram shown is a flowchart illustrating the determination of the time spectrum diagram provided in an exemplary embodiment of this application.
[0026] Figure 7 The diagram shown is a flowchart of an audio forensics method provided in an exemplary embodiment of this application.
[0027] Figure 8 The diagram shown is a system architecture diagram of audio forensics and verification provided in an exemplary embodiment of this application.
[0028] Figure 9 The diagram shown is a structural schematic of an audio verification device provided in an exemplary embodiment of this application.
[0029] Figure 10 The diagram shown is a schematic representation of the audio evidence collection device provided in an exemplary embodiment of this application.
[0030] Figure 11 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] Application Overview
[0033] Data acquisition: The process of receiving analog signals through sensors, decoding them into digital signals, and encoding the digital signals into a specific data format.
[0034] Ultrasound: Mechanical waves with frequencies higher than the upper limit of human hearing (i.e., greater than 18kHz).
[0035] Copy-move tampering: Copying a speech segment from the same target object to the location of another speech segment in order to hide or modify the semantics recorded in the entire audio file.
[0036] Audio deepfake: Using artificial intelligence technology to generate the voice of a target object.
[0037] With the development of technology, deepfake audio technology based on artificial intelligence has exacerbated the security threats faced by audio corpora. Existing data-level analysis-based defenses lag behind rapidly evolving tampering techniques. Furthermore, blockchain technology aims to solve trust-related application problems. Product traceability, transaction monitoring, and investigation all involve the trust barrier between off-chain and on-chain data, making it a crucial battleground for ensuring trustworthiness at the source. Once false data is uploaded to the blockchain with the purpose of falsifying the production and transportation processes of genuine goods or fabricating investigative evidence, it will pollute the on-chain data and pose new challenges to the credibility and informational value of the information uploaded to the blockchain.
[0038] Descriptions of people, things, and places off-chain typically rely on data collection from offline IoT devices. This collected data is then uploaded to the blockchain for comprehensive profiling, forming a tangible proof of existence for the on-chain abstract model. Audio data, with its core corpus information, can directly record factual events. The on-chain abstract model, in different application scenarios, possesses certain financial attributes (pledged goods), legal attributes (factual evidence), and value attributes (commodities and raw materials). This on-chain data, combined with specific business models, can generate relevant business value. While on-chain trusted transfer technology is relatively mature at present, the process of transferring information (especially corpus information) from the off-chain physical world to its actual on-chain representation still faces many unresolved issues.
[0039] Related methods for encrypting corpus information include audio watermarking technology, environmental feature extraction technology, similarity detection technology, audio deepfake detection technology, and blockchain technology.
[0040] Audio watermarking technology encrypts audio files by adding a digital key to the acquired audio file. The disadvantages of this scheme are: (1) the added digital key will significantly reduce the quality and intelligibility of the audio file; (2) the digital key may be forged, which may cause the real audio file to be mistaken for a forgery.
[0041] Environmental feature extraction technology uses the environmental features contained in the audio recording process as a basis, and further verifies whether the audio has been tampered with by detecting the continuity of the environmental features. The disadvantages of this approach are: (1) it is susceptible to interference from environmental noise; (2) its application environment is limited. For example, power grid frequency detection technology is only suitable for indoor environments with strong power grid harmonics and cannot be applied to open outdoor scenarios.
[0042] Similarity detection technology prevents copy-and-move tampering by detecting the similarity between corpora from the same source. The disadvantages of this approach are: (1) low detection accuracy; (2) when there are too many audio files, the time and computational cost required for detection increase exponentially.
[0043] Audio deepfake detection technology distinguishes audio from natural human speech by mining the features introduced during the audio deepfake process. The disadvantages of this approach are: (1) it requires a large number of fake audio files for deepfake audio for training; (2) it has poor transferability, and can only be applied to detect audio deepfake techniques contained in the training dataset, and cannot defend against unseen audio deepfake techniques; (3) it is outdated and cannot defend against new audio deepfake techniques that may emerge in the future.
[0044] Blockchain technology ensures the authenticity of audio recordings by encrypting the recording time and location information and uploading it to the blockchain, generating a blockchain key that is embedded in the audio. The disadvantages of this approach are: (1) the audio's time and location information can be forged on a mobile phone, thus compromising its authenticity; and (2) the authenticity of the audio data cannot be guaranteed.
[0045] In summary, this application proposes an audio verification method. The method involves acquiring an audio file to be verified, determining the ultrasonic features corresponding to the first ultrasonic wave if the audio file contains a first ultrasonic wave, extracting the speech data contained in the audio file, and verifying whether the speech data represents the speech data expressed by the target object in the target scene based on the correlation between the ultrasonic features and the speech data. Firstly, although ultrasonic waves are inaudible to the human ear, they possess high perceptual accuracy. This application introduces a credibility factor—ultrasonic waves—and detects whether the speech data in the audio file to be verified has been tampered with based on the ultrasonic features corresponding to the ultrasonic waves. This improves the detection accuracy of the audio file to be verified. Furthermore, the frequency range of ultrasonic waves does not overlap with the frequency range of low-frequency speech, thus not affecting the quality of the speech data in the audio file to be verified. Secondly, the audio verification method in this application has strong transferability; any tampering methods that cause a mismatch between ultrasonic features and speech data can be detected, and it can also resist other unknown tampering methods.
[0046] Exemplary scenario
[0047] During the data collection process, additional, difficult-to-forge, and non-repudiable credibility factors are introduced into the collected corpus information using hardware such as sensors. These credibility factors serve as the basis for preventing data tampering, thereby effectively raising the data security threshold and increasing the cost of malicious activity for perpetrators. The value of this application lies in increasing the cost of malicious activity to prevent it from occurring, thus ensuring the correct correspondence between the off-chain physical world and on-chain information.
[0048] Figure 1a The diagram illustrates an application scenario applicable to an embodiment of this application. The application scenario mentioned in this embodiment includes a mobile terminal 11, a server 12, and a speaker 13, with a communication connection between the mobile terminal 11 and the server 12. Specifically, during the audio evidence collection stage, when the target object makes a sound in the target scene, the speaker 13 emits modulated ultrasonic waves towards the target object. The mobile terminal 11 uses its built-in microphone to collect the voice information of the target object and the ultrasonic waves reflected by the target object, and sends the collected voice information to the server 12 for authenticity verification, i.e., verifying whether the voice information has been tampered with. If it has not been tampered with, the collected voice information is uploaded to the blockchain to ensure the correct correspondence between the off-chain physical world and the on-chain information.
[0049] In another feasible scenario, trusted data acquisition functionality is integrated into a pre-developed trusted acquisition application (APP). This trusted acquisition APP can be deployed on mobile terminals, such as smartphones, without requiring external devices or hardware modifications. Specifically, for example... Figure 1b As shown, this use case includes edge devices, which are equipped with a trusted data collection app and other applications. For example, the other applications are mobile banking apps, and the edge device is a mobile phone. When a user conducts business using the mobile banking app, they may request audio data collection. The mobile banking app redirects to the trusted data collection app, which, based on the user's request, emits ultrasonic waves towards the target object and simultaneously records the target object's audio data. Based on the ultrasonic signals contained in the audio data, the audio data is verified. If the verified audio data is the voice data expressed by the target object in the target scenario, the audio data is uploaded to the blockchain for evidence storage. Afterwards, the trusted data collection app sends the evidence code returned by the blockchain to the mobile banking app. This completes the audio data collection, verification, and blockchain uploading process for the target object. If the mobile banking app needs to use the audio collection results, it sends the evidence code to the blockchain and receives the audio collection results for the target object from the blockchain.
[0050] Exemplary methods
[0051] Figure 2The diagram shown is a flowchart illustrating an exemplary embodiment of the audio verification method provided in this application. Figure 2 As shown in the embodiments of this application, the audio verification method includes the following steps.
[0052] Step 210: Obtain the audio to be verified.
[0053] For example, in response to an audio verification request, audio to be verified is obtained. The audio verification request is used to verify whether the audio to be verified contains speech data expressed by a target object in a target scene, the target scene including emitted ultrasonic waves, and the audio data obtained during the acquisition of the target object's speech data includes ultrasonic waves reflected by the target object.
[0054] Furthermore, the audio to be verified may contain real speech data of the target object and speech data of non-target object, or it may contain altered speech data of the target object and speech data of non-target object. Of course, the audio to be verified may not contain speech data of non-target object.
[0055] Step S220: If the audio to be verified contains a first ultrasonic wave, then determine the ultrasonic feature corresponding to the first ultrasonic wave.
[0056] Step S230: Extract the speech data contained in the audio to be verified.
[0057] For example, a high-pass filter is used to filter low-frequency speech in the audio to be verified, and the filtering result is obtained. The filtering result is checked to see if it contains high-frequency ultrasound components. If it does, the audio to be verified is considered to contain the first ultrasound wave; if it does not, the audio to be verified is considered not to contain the first ultrasound wave.
[0058] Furthermore, if the audio to be verified contains a first ultrasonic wave, the ultrasonic features corresponding to the first ultrasonic wave are determined, and the voice data to be verified contained in the audio to be verified is extracted based on the audio to be verified.
[0059] Step S240: Based on the correlation between ultrasound features and speech data, verify whether the speech data is the speech data expressed by the target object in the target scene.
[0060] Specifically, when the target object emits sound in the target scene, ultrasonic waves are emitted into the target scene. These ultrasonic waves can capture the vocal tract motion characteristics of the target object when it emits sound. If the target object's speech data has not been tampered with, the ultrasonic features obtained from the demodulation of the audio to be verified have a high correlation with the target object's speech data. If the target object's speech data has been tampered with, the ultrasonic features obtained from the demodulation of the audio to be verified will not match the features in the target object's speech data.
[0061] For example, an equivalent threshold can be set. If the correlation between the ultrasound features and the speech data is greater than the equivalent threshold, the speech data extracted from the audio data to be verified can be considered as the speech data expressed by the target object in the target scene; otherwise, the speech data contained in the audio to be verified is considered to have been tampered with.
[0062] In this application embodiment, firstly, although ultrasound is inaudible to the human ear, it possesses high perceptual accuracy. This application introduces a credibility factor—ultrasound—and detects whether the speech data in the audio to be verified has been tampered with based on the ultrasound features corresponding to the ultrasound. This improves the detection accuracy of the audio to be verified. Furthermore, the frequency range of ultrasound does not overlap with the frequency range of the speech data expressed by the target object, thus not affecting the quality of the speech data in the audio to be verified. Secondly, the audio verification method in this application has strong transferability, enabling the detection of tampering methods that result in a mismatch between ultrasound features and speech data, thereby resisting other unknown tampering methods.
[0063] Figure 3 The diagram shown is a schematic flowchart illustrating the process of determining the ultrasonic feature corresponding to a first ultrasonic wave according to an exemplary embodiment of this application. Figure 2 Extending from the illustrated embodiment Figure 3 The illustrated embodiment will be described in detail below. Figure 3 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0064] like Figure 3 As shown in the embodiment of this application, before determining the ultrasonic features corresponding to the first ultrasonic wave, the following steps are further included.
[0065] Step S310: Determine whether the audio to be verified contains the first ultrasonic wave.
[0066] For example, if the judgment result of step S310 is yes, then step S320 is executed, that is, the time spectrum diagram corresponding to the first ultrasound is determined; if the judgment result of step S310 is no, then step S370 is executed, that is, the voice data is determined to be the voice data expressed by the target object in the target scene.
[0067] Specifically, a high-pass filter is used to filter out low-frequency speech information in the audio to be verified, obtaining the high-frequency first ultrasonic wave. A Fourier transform is then performed on the first ultrasonic wave to obtain its corresponding time-spectrum.
[0068] In this application, it is assumed that the scene in which the audio to be verified is generated includes ultrasonic waves reflected by the target object. Therefore, if the audio to be verified does not contain the first ultrasonic wave, it is considered that the voice data of the target object in the audio to be verified has been tampered with, that is, the voice data in the audio to be verified is not the voice data expressed by the target object in the target scene.
[0069] Step S330: Based on the time-spectrum diagram corresponding to the first ultrasound, determine the slope change trend corresponding to the first ultrasound.
[0070] Step S340: Obtain the slope change trend corresponding to the second ultrasonic wave. The second ultrasonic wave is the ultrasonic wave emitted towards the target object when the target object emits sound in the target scene beforehand.
[0071] Step S350: Determine whether the slope change trend corresponding to the first ultrasound and the slope change trend corresponding to the second ultrasound are the same.
[0072] Specifically, to ensure the accuracy of the verification, when the target object emits sound in the target scene, the ultrasonic wave emitted towards the target object is an ultrasonic wave of a specific varying frequency. That is, the second ultrasonic wave is an ultrasonic wave of a specific varying frequency, rather than an ultrasonic wave of a constant frequency. This is because ultrasonic waves of constant frequency are low-cost to use maliciously and are not easily detected.
[0073] Furthermore, when the second ultrasonic wave is emitted towards the target object, the ultrasonic frequency of the second ultrasonic wave is known, and consequently, the slope trend of the second ultrasonic wave's frequency change is known. By comparing the slope trend of the first ultrasonic wave's frequency change, it can be determined whether the speech data in the audio to be verified is the speech data expressed by the target object in the target scene. For example, if the ultrasonic frequency of the second ultrasonic wave changes sinusoidally, then the slope trend of the second ultrasonic wave's frequency change is cosine.
[0074] For example, if the judgment result of step S350 is yes, then step S360 is executed, that is, the ultrasonic feature corresponding to the first ultrasonic wave is determined; if the judgment result of step S360 is no, then step S370 is executed, that is, the voice data is determined to be the voice data expressed by the target object in the target scene.
[0075] In other words, if the slope change trend corresponding to the first ultrasound is the same as that corresponding to the second ultrasound, the ultrasound characteristics of the first ultrasound can be further determined so as to further verify the audio to be verified. This dual verification improves the accuracy of the verification. If the slope change trends corresponding to the first ultrasound and the second ultrasound are different, it can be considered that the speech data in the audio data to be verified has been tampered with.
[0076] In this embodiment, the audio to be verified is first screened based on the slope change trends corresponding to the first and second ultrasound waves to determine whether the voice data about the target object in the audio has been tampered with. This method has low computational complexity, thus improving the verification speed of the audio.
[0077] Figure 4The diagram shown is a schematic flowchart illustrating the process of determining the slope change trend corresponding to the first ultrasonic wave, as provided in an exemplary embodiment of this application. Figure 3 Extending from the illustrated embodiment Figure 4 The illustrated embodiment will be described in detail below. Figure 4 The illustrated embodiments and Figure 3 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0078] like Figure 4 As shown in the embodiment of this application, the slope change trend of the first ultrasonic wave is determined based on the time-frequency spectrum corresponding to the first ultrasonic wave, including the following steps.
[0079] Step S410: Based on the amplitude values at each moment in the time spectrum diagram corresponding to the first ultrasonic wave, perform binarization processing on the time spectrum diagram corresponding to the first ultrasonic wave.
[0080] Specifically, the time-spectrum diagram of the first ultrasonic wave is binarized based on the amplitude at each moment in the time-spectrum. For example, a high energy threshold is set, points greater than the high energy threshold are marked as 1, and the remaining points are marked as 0, to obtain the binarized time-spectrum diagram.
[0081] Step S420: Calculate the slope between adjacent times of the same value in the binarized time spectrum.
[0082] Step S430: Based on the slope between adjacent moments with the same value, determine the slope change trend corresponding to the first ultrasound.
[0083] Following the example in step S410, calculate the slope between adjacent time points with a value of 1 in the binarized time-spectrum graph. Accordingly, if points greater than the high energy threshold are marked as 0 during binarization, calculate the slope between adjacent time points with a value of 0 in the binarized time-spectrum graph.
[0084] Continue Figure 3 In the example, if the slope of the second ultrasonic wave frequency changes in a cosine manner, and if the speech data about the target object in the audio to be verified has not been tampered with, then the slope of the first ultrasonic wave frequency should also change in a cosine manner.
[0085] In this embodiment of the application, binarizing the time spectrum of the first ultrasonic wave can more conveniently and accurately determine the slope change trend of the ultrasonic frequency of the first ultrasonic wave, so as to perform reliable verification of the audio to be verified.
[0086] Figure 5 The diagram shown is a flowchart illustrating an audio verification method provided in another exemplary embodiment of this application. Figure 2Extending from the illustrated embodiment Figure 5 The illustrated embodiment will be described in detail below. Figure 5 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0087] like Figure 5 As shown in the embodiments of this application, the audio verification method further includes the following steps.
[0088] Step S510: Determine the time-spectrum diagram corresponding to the ultrasound features and the time-spectrum diagram corresponding to the speech data, respectively.
[0089] Step S520: For each of the multiple spectral channels, determine the similarity between the time-spectral graph corresponding to the ultrasound feature and the time-spectral graph corresponding to the speech data.
[0090] For example, a low-pass filter is used to filter the speech data to remove high-frequency components. Fourier transforms are then performed on the ultrasonic features and the filtered speech data to obtain the time-spectrum diagrams corresponding to the ultrasonic features and the speech data. Furthermore, embodiments of this application can utilize a detection algorithm to determine the similarity between the time-spectrum diagrams corresponding to the ultrasonic features and the speech data, or a trained detection model can be used to determine the similarity between the time-spectrum diagrams corresponding to the ultrasonic features and the speech data.
[0091] Below, we will take a trained detection model as an example to explain in detail how to determine similarity.
[0092] For example, during the training phase of the detection model, the time-spectrum graphs corresponding to the ultrasound features and the time-spectrum graphs of the speech data corresponding to the ultrasound features are concatenated by spectral channels to form positive samples. Further, the correspondence between the time-spectrum graphs of the ultrasound features and the speech data is shuffled, and then concatenated by spectral channels to form negative samples. The positive and negative samples are divided into a training set and a validation set according to a certain ratio. For example, the ratio of positive samples to negative samples is 4:1.
[0093] During the training phase of the detection model, the spliced spectrogram corresponding to each spectral channel is used as input, and the matching result is used as output, which is a natural number between 0 and 1. The higher the output value, the stronger the correlation between the ultrasound features and the speech data. The model is trained using the cross-entropy loss function and the backpropagation algorithm, and the performance of the model is verified using a validation set.
[0094] Furthermore, in the application stage of the detection model, the time-spectrum diagrams corresponding to the ultrasound features and the time-spectrum diagrams corresponding to the speech data are spliced together by spectral channels and then input into the detection model to obtain similarity comparison results.
[0095] Step S530: Based on the similarity of each of the multiple spectral channels, determine the correlation between ultrasound features and speech data.
[0096] For example, if the similarity result meets the equivalence judgment condition, then the ultrasound features and the speech data are considered to be strongly correlated, that is, the speech data in the audio to be verified is considered to be the speech data expressed by the target object in the target scene.
[0097] In this embodiment of the application, the high-frequency band of ultrasound carries relevant information of low-frequency voice data. Based on this, the correlation between the time spectrum diagram corresponding to the ultrasound features and the time spectrum corresponding to the voice data is detected to ensure the authenticity of the audio to be verified and to detect whether it has been forged before the audio to be verified is uploaded to the blockchain.
[0098] Figure 6 The diagram shown is a flowchart illustrating the determination of the time-frequency spectrum according to an exemplary embodiment of this application. Figure 5 Based on the illustrated embodiment, the following extensions are made: Figure 6 The illustrated embodiment will be described in detail below. Figure 6 The illustrated embodiments and Figure 5 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0099] like Figure 6 As shown in the embodiments of this application, determining the time-spectrum diagram corresponding to the ultrasonic features and the time-spectrum diagram corresponding to the speech data includes the following steps.
[0100] Step S610: Perform voice endpoint detection on the voice data to obtain the voice endpoint detection result.
[0101] Step S620: The ultrasonic features and speech data are segmented using the speech endpoint detection results to obtain the ultrasonic features of the speech segment and the speech data of the speech segment.
[0102] Step S630: Determine the time-spectrum diagram corresponding to the ultrasonic features of the speech segment as the time-spectrum diagram corresponding to the ultrasonic features, and determine the time-spectrum diagram corresponding to the speech data of the speech segment as the time-spectrum diagram corresponding to the speech data.
[0103] For example, the spectral entropy method can be used to perform endpoint detection on speech data to obtain endpoint detection results, or the dual threshold method can be used to perform endpoint detection on speech data. The embodiments of this application do not limit the method of endpoint detection.
[0104] Furthermore, based on the speech endpoint detection results, the ultrasonic features and speech data are segmented, specifically into ultrasonic features of speech segments and ultrasonic features of non-speech segments, as well as speech data of speech segments and speech data of non-speech segments. Moreover, to reduce the computational load of detection, similarity comparison can be performed only on the ultrasonic features and speech data of the speech segments.
[0105] In another implementation, after determining the ultrasonic features of the speech segment and the ultrasonic features of the non-speech segment, as well as the speech data of the speech segment and the speech data of the non-speech segment, using the speech endpoint detection results, filtering of the ultrasonic features of the non-speech segment and the speech data of the non-speech segment may not be performed. Instead, a similarity comparison is made between the ultrasonic features of the speech segment and the speech data of the speech segment, and a similarity comparison is made between the ultrasonic features of the non-speech segment and the speech data of the non-speech segment. In this case, if the ultrasonic features of the non-speech segment show related vocal tract motion features, it can be proven that the audio to be detected has been tampered with.
[0106] In this embodiment of the application, the ultrasonic features and speech data are divided into speech segments and non-speech segments, which enables more accurate, refined and targeted detection of the audio to be verified.
[0107] In an exemplary embodiment of this application, the ultrasonic features include Doppler frequency shift features; or, the ultrasonic features include Doppler frequency shift features, and the ultrasonic features also include acoustic nonlinear features and / or time-of-flight features.
[0108] That is, this application can use Doppler frequency shift as an ultrasonic feature to determine its correlation with speech data; it can also use Doppler frequency shift features, acoustic nonlinear features, time-of-flight features, etc. as ultrasonic features to determine its correlation with speech data; or it can use Doppler frequency shift features and acoustic nonlinear features, or Doppler frequency shift features and time-of-flight features as ultrasonic features to determine its correlation with speech data.
[0109] For example, Doppler frequency shift characteristics, acoustic nonlinear characteristics, and time-of-flight characteristics are used as ultrasonic features. First, the audio to be verified is filtered at low frequencies to obtain the first ultrasonic wave. For example, the acoustic nonlinear characteristics, Doppler frequency shift characteristics, and time-of-flight characteristics are demodulated to the low-frequency band by self-demodulation. The self-demodulation method is shown in formula (1).
[0110]
[0111] In formula (1), U(t) represents the first ultrasonic wave, and F(t) represents the acoustic nonlinearity, Doppler frequency shift, and time-of-flight characteristics contained in the first ultrasonic wave.
[0112] In this embodiment, various acoustic effects are employed to improve the accuracy of audio verification. Furthermore, by comparing the correlation between the acoustic nonlinear characteristics, Doppler frequency shift characteristics, time-of-flight characteristics, and low-frequency speech data carried on the first ultrasonic wave, tampering of the speech data can be detected, enabling the detection of any tampering methods that cause a mismatch between ultrasonic features and speech data.
[0113] The audio verification method corresponds to the audio forensics method. For the same target object, if the audio data collected in the audio forensics method has not been tampered with, then the audio data is the same as the audio to be verified obtained in the audio verification method; otherwise, the audio data is different from the audio to be verified obtained in the audio verification method.
[0114] Figure 7 The diagram shown is a flowchart illustrating an audio forensics method provided in an exemplary embodiment of this application. Figure 7 As shown in the embodiments of this application, the audio forensics method includes the following steps.
[0115] Step S710: Emit ultrasonic waves to the target object located in the target scene.
[0116] Step S720: When the target object makes a sound in the target scene, collect the speech data expressed by the target object and the ultrasonic waves reflected by the target object to obtain the reliable audio data corresponding to the target object.
[0117] For example, the modulation method of the emitted ultrasonic wave is shown in the following formula (2).
[0118]
[0119] In formula (2), Let B be the frequency of the emitted ultrasonic wave, B be the bandwidth of the emitted ultrasonic wave, τ be the period of frequency variation of the emitted ultrasonic wave, and F be the frequency of the emitted ultrasonic wave. bias For the center frequency, A u This indicates the intensity of the emitted ultrasonic wave.
[0120] In this embodiment, by emitting ultrasonic waves during audio acquisition, the high-frequency band of the received audio carries information related to low-frequency speech, ensuring the authenticity of the audio acquisition. Even if the acquired audio is forged before being uploaded to the blockchain, its authenticity can be determined by comparing the correlation between the information carried by the high-frequency band and the low-frequency speech data.
[0121] Figure 8 The diagram shown is a system architecture diagram of audio forensics and audio verification provided in an exemplary embodiment of this application. Figure 8 As shown, the system includes a data acquisition module 810, a distortion detection module 820, an ultrasonic feature demodulation module 830, and a tamper detection module 840.
[0122] The data acquisition module uses the phone's built-in speaker to emit modulated ultrasonic waves and uses the built-in microphone to collect the speech information of the target subject and the reflected ultrasonic waves. The distortion detection module detects the continuity of the ultrasonic wave frequencies in the received speech information; if discontinuities are found, the speech information is considered tampered with. If the ultrasonic wave frequencies are continuous, the signal enters the ultrasonic feature demodulation module. This module demodulates the ultrasonic waves in the speech information to obtain the ultrasonic features produced by the ultrasonic effect. For example, the ultrasonic features include Doppler frequency shift features; or, the ultrasonic features include Doppler frequency shift features, and also include acoustic nonlinear features and / or time-of-flight features. The tampering detection module preprocesses the ultrasonic features and low-frequency speech information and uses a detection model to compare their similarity. If the similarity exceeds a certain threshold, the speech information is considered not tampered with.
[0123] This application ensures the authenticity of the offline product chain without relying on any peripherals or hardware modifications. Firstly, it utilizes various acoustic effects to explore the correlation between audible sound waves (low-frequency speech) and ultrasound. Secondly, it uses ultrasound as a reliable factor to prevent audio tampering; this reliable factor is imperceptible to the human ear, will not interfere with surrounding users, and will not reduce audio quality or intelligibility.
[0124] The above text combined Figures 2 to 8 The method embodiments of this application are described in detail below, in conjunction with... Figure 9 and Figure 10 The present application provides a detailed description of the apparatus embodiments. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be found in the foregoing method embodiments.
[0125] Figure 9 The diagram shown is a structural schematic of an audio verification device provided in an exemplary embodiment of this application. Figure 9 As shown, the audio verification device 90 provided in this application embodiment includes:
[0126] Module 910 is used to acquire the audio to be verified;
[0127] The determination module 920 is used to determine the ultrasonic features corresponding to the first ultrasonic wave if the audio to be verified contains the first ultrasonic wave.
[0128] Extraction module 930 is used to extract the voice data contained in the audio to be verified;
[0129] The verification module 940 is used to verify whether the speech data is the speech data expressed by the target object in the target scene based on the correlation between the ultrasound features and the speech data.
[0130] In one embodiment of this application, the determining module 920 is further configured to, before determining the ultrasonic feature corresponding to the first ultrasonic wave, further include: if the audio to be verified contains the first ultrasonic wave, then determining the time-spectrum diagram corresponding to the first ultrasonic wave; based on the time-spectrum diagram corresponding to the first ultrasonic wave, determining the slope change trend corresponding to the first ultrasonic wave; obtaining the slope change trend corresponding to the second ultrasonic wave, wherein the second ultrasonic wave is the ultrasonic wave emitted towards the target object when the target object emits sound in the target scene in advance; determining whether the slope change trend corresponding to the first ultrasonic wave and the slope change trend corresponding to the second ultrasonic wave are the same; the method further includes, when the slope change trend corresponding to the first ultrasonic wave and the slope change trend corresponding to the second ultrasonic wave are the same, determining the ultrasonic feature corresponding to the first ultrasonic wave.
[0131] In one embodiment of this application, the determining module 920 is further configured to determine that the voice data is not the voice data expressed by the target object in the target scene if the slope change trend corresponding to the first ultrasound and the slope change trend corresponding to the second ultrasound are different.
[0132] In one embodiment of this application, the determining module 920 is further configured to: perform binarization processing on the time spectrum diagram corresponding to the first ultrasonic wave based on the amplitude at each moment in the time spectrum diagram corresponding to the first ultrasonic wave; calculate the slope between adjacent moments with the same value in the time spectrum diagram after binarization; and determine the slope change trend corresponding to the first ultrasonic wave based on the slope between adjacent moments with the same value.
[0133] In one embodiment of this application, the verification module 940 is further configured to: determine the time-spectrum diagram corresponding to the ultrasound feature and the time-spectrum diagram corresponding to the speech data respectively; determine the similarity between the time-spectrum diagram corresponding to the ultrasound feature and the time-spectrum diagram corresponding to the speech data for each of the multiple spectrum channels; and determine the correlation between the ultrasound feature and the speech data based on the similarity corresponding to each of the multiple spectrum channels.
[0134] In one embodiment of this application, the verification module 940 is further configured to: perform voice endpoint detection on the voice data to obtain voice endpoint detection results; segment the ultrasonic features and voice data using the voice endpoint detection results to obtain the ultrasonic features of the voice segment and the voice data of the voice segment; determine the time-spectrum diagram corresponding to the ultrasonic features of the voice segment as the time-spectrum diagram corresponding to the ultrasonic features; and determine the time-spectrum diagram corresponding to the voice data of the voice segment as the time-spectrum diagram corresponding to the voice data.
[0135] In one embodiment of this application, the ultrasonic features include Doppler frequency shift features; or, the ultrasonic features include Doppler frequency shift features, and the ultrasonic features also include acoustic nonlinear features and / or time-of-flight features.
[0136] Figure 10The diagram shown is a structural schematic of an audio evidence collection device provided in an exemplary embodiment of this application. Figure 10 As shown, the audio forensics device 100 provided in this application embodiment includes:
[0137] The transmitting module 1010 is used to transmit ultrasonic waves to a target object located in the target scene;
[0138] The acquisition module 1020 is used to acquire the speech data expressed by the target object and the ultrasonic waves reflected by the target object when the target object emits a sound in the target scene, so as to obtain the reliable audio data corresponding to the target object.
[0139] Below, for reference Figure 11 This describes an electronic device according to embodiments of the present application. Figure 11 The diagram shown is a structural schematic of an electronic device provided in an exemplary embodiment of this application.
[0140] like Figure 11 As shown, the electronic device 110 includes one or more processors 1101 and memory 1102.
[0141] The processor 1101 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 110 to perform desired functions.
[0142] The memory 1102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1101 may execute the program instructions to implement the methods of the various embodiments of this application described above and / or other desired functions. Various content, such as audio to be verified, first ultrasound, second ultrasound, and ultrasound features, may also be stored in the computer-readable storage medium.
[0143] In one example, the electronic device 110 may also include an input device 1103 and an output device 1104, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0144] The input device 1103 may include, for example, a keyboard, a mouse, etc.
[0145] The output device 1104 can output various information to the outside, including audio to be verified, first ultrasound, second ultrasound, ultrasound characteristics, etc. The output device 1104 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0146] Of course, for the sake of simplicity, Figure 11 Only some of the components of the electronic device 110 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 110 may include any other suitable components depending on the specific application.
[0147] In addition to the methods and apparatus described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods described above according to various embodiments of this application.
[0148] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0149] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods described above according to various embodiments of this application.
[0150] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0151] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0152] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0153] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0154] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0155] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. An audio verification method, comprising: Get the audio to be verified; If the audio to be verified contains a first ultrasonic wave, then the ultrasonic feature corresponding to the first ultrasonic wave is determined. Extract the speech data contained in the audio to be verified; Determine the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data respectively; For each of the multiple spectral channels, determine the similarity between the time-spectral graph corresponding to the ultrasonic feature and the time-spectral graph corresponding to the speech data; Based on the similarity of the multiple spectral channels, the correlation between the ultrasound features and the speech data is determined; Based on the correlation between the ultrasonic features and the speech data, it is verified whether the speech data is the speech data expressed by the target object in the target scene.
2. The method according to claim 1, further comprising, before determining the ultrasonic feature corresponding to the first ultrasonic wave: If the audio to be verified contains the first ultrasonic wave, then determine the time-spectrum diagram corresponding to the first ultrasonic wave; Based on the time-spectrum diagram corresponding to the first ultrasonic wave, the slope change trend corresponding to the first ultrasonic wave is determined. Obtain the slope change trend corresponding to the second ultrasonic wave, where the second ultrasonic wave is the ultrasonic wave emitted towards the target object when the target object emits a sound in the target scene in advance; Determine whether the slope change trend corresponding to the first ultrasound and the slope change trend corresponding to the second ultrasound are the same; The method further includes determining the ultrasonic feature corresponding to the first ultrasonic wave when the slope change trend corresponding to the first ultrasonic wave is the same as the slope change trend corresponding to the second ultrasonic wave.
3. The method according to claim 2, further comprising: If the slope change trend corresponding to the first ultrasound is different from the slope change trend corresponding to the second ultrasound, then it is determined that the voice data is not the voice data expressed by the target object in the target scene.
4. The method according to claim 2, wherein determining the slope change trend of the first ultrasonic wave based on the time-spectrum diagram corresponding to the first ultrasonic wave includes: Based on the amplitude values at each moment in the time-spectrum diagram corresponding to the first ultrasonic wave, the time-spectrum diagram corresponding to the first ultrasonic wave is binarized. Calculate the slope between adjacent time points of the same value in the binarized time spectrum; Based on the slope between adjacent moments of the same value, the slope change trend corresponding to the first ultrasound is determined.
5. The method according to any one of claims 1 to 4, wherein determining the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data respectively comprises: Perform voice endpoint detection on the voice data to obtain voice endpoint detection results; The ultrasonic features and the speech data are segmented using the speech endpoint detection results to obtain the ultrasonic features of the speech segment and the speech data of the speech segment. The time-spectrum diagram corresponding to the ultrasonic features of the speech segment is determined as the time-spectrum diagram corresponding to the ultrasonic features; The time-spectrum diagram corresponding to the speech data of the speech segment is determined as the time-spectrum diagram corresponding to the speech data.
6. The method according to any one of claims 1 to 4, The ultrasound features include Doppler frequency shift features; or... The ultrasonic features include the Doppler frequency shift features, and the ultrasonic features also include acoustic nonlinear features and / or time-of-flight features.
7. An audio verification device, comprising: The acquisition module is used to acquire the audio to be verified. The determination module is used to determine the ultrasonic features corresponding to the first ultrasonic wave if the audio to be verified contains a first ultrasonic wave. The extraction module is used to extract the voice data contained in the audio to be verified; The verification module is used to determine the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data respectively; for each of the multiple spectral channels, determine the similarity between the time-spectrum diagram corresponding to the ultrasonic feature and the time-spectrum diagram corresponding to the speech data; based on the similarity corresponding to each of the multiple spectral channels, determine the correlation between the ultrasonic feature and the speech data; and based on the correlation between the ultrasonic feature and the speech data, verify whether the speech data is the speech data expressed by the target object in the target scene.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for performing the method described in any one of claims 1 to 6.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to perform the method described in any one of claims 1 to 6.