Audio signal authenticity verification method and device, equipment and medium
By generating a set of adversarial samples through a generative adversarial network and combining acoustic and non-acoustic features to construct a multi-dimensional feature vector, the problem of identifying voice cloning attacks in existing audio authenticity verification technology is solved, and the robustness and detection effect of the model are improved.
Patent Information
- Application Number
- CN202510826683.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing audio authenticity verification technology is difficult to effectively identify voice cloning attacks, lacks adversarial sample detection capabilities and multimodal feature fusion judgment mechanisms, and is unable to cope with complex and intelligent voice cloning attacks.
By collecting original audio and corresponding text to generate an original audio-text dataset, a generative adversarial network is used to generate an adversarial sample set, which is then input into the audio detection model for joint training together with the original audio-text dataset to extract acoustic and non-acoustic features, construct a multi-dimensional feature vector, generate anomaly indicators and perform graded response operations.
The robustness of the audio detection model has been improved, the ability to identify voice cloning attacks has been enhanced, effective identification and response to new voice cloning attacks have been achieved, and the operability of detection results has been improved.
Smart Images

Figure CN120673780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a method, device, equipment and storage medium for verifying the authenticity of an audio signal. Background Art
[0002] In the current technical system for verifying the authenticity of audio information, facing the rapidly evolving voice cloning attack methods, the existing technology still has obvious shortcomings in many aspects.
[0003] In the fintech sector, voice is widely used in scenarios such as remote identity authentication, voice-command transfers, and customer service interactions. Existing audio detection models are typically trained on static datasets and lack adversarial training mechanisms, making them ineffective at identifying voice clones generated through deepfakes. These forged voices are often highly similar to real voices at the acoustic level, making them difficult for conventional models to distinguish, posing a security risk of real user voices being impersonated. Furthermore, current systems generally lack the ability to integrate multimodal information for verification, often relying solely on acoustic features and lacking comprehensive analysis of related information such as text semantics, device behavior, and interaction patterns. This results in blind spots in the system's recognition when facing complex attacks.
[0004] In the healthcare sector, voice interaction is increasingly being used in scenarios such as intelligent diagnosis and treatment, voice-based medical consultations, and remote assistance. Patient voice information can be maliciously synthesized and used to deceive the system through false registration, misappropriation of medical resources, or misdirected system responses. Existing detection methods often focus on standard pronunciation samples and lack the ability to adapt to varying speaking speeds, dialects, and device environments. Furthermore, they cannot dynamically integrate environmental, device, and behavioral information for comprehensive verification, making them susceptible to interference from voice cloning, which can reduce system stability and security.
[0005] Furthermore, most traditional detection systems lack an active sample enhancement mechanism, making it impossible to generate adversarial training samples to improve the detection model's robustness against voice cloning attacks. During acoustic feature extraction, feature fusion, and anomaly indicator determination, most methods lack the collaborative modeling of acoustic and non-acoustic information, reducing their ability to discern subtle differences between authentic and forged audio. Existing systems also generally lack a response and feedback mechanism based on anomaly indicators, making it impossible to implement automated security actions based on detection results, impacting the efficiency of responding to risk events.
[0006] In summary, existing audio authenticity verification technologies have technical gaps in data sample construction, adversarial robustness training, multi-dimensional feature fusion, and response mechanisms. They are unable to cope with increasingly complex and intelligent voice cloning attacks, and there is an urgent need to improve their overall defense capabilities and detection accuracy. Summary of the Invention
[0007] The main purpose of the present invention is to provide a method, device, equipment and storage medium for verifying the authenticity of an audio signal, aiming to solve the technical problems in the prior art such as the difficulty in effectively identifying voice cloning attacks, insufficient adversarial sample detection capabilities, and the lack of a multimodal feature fusion judgment mechanism.
[0008] To achieve the above object, the present invention provides a method for verifying the authenticity of an audio signal, comprising:
[0009] Collect original audio and corresponding text to generate an original audio text dataset;
[0010] Generate a set of adversarial samples based on the original audio text dataset through a generative adversarial network;
[0011] Inputting the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model;
[0012] Acquiring an audio signal to be detected and extracting acoustic features of the audio signal to be detected;
[0013] Acquiring non-acoustic features associated with the audio signal to be detected;
[0014] Constructing a multidimensional feature vector according to the acoustic features and the non-acoustic features;
[0015] Inputting the multidimensional feature vector into an audio detection model to generate an anomaly indicator;
[0016] Based on the abnormal indicators, a graded response operation is performed.
[0017] Furthermore, to achieve the above-mentioned object, the present invention provides an audio signal authenticity verification device, comprising:
[0018] The data acquisition module is used to collect original audio and corresponding text to generate an original audio text dataset;
[0019] An adversarial sample generation module, configured to generate an adversarial sample set based on the original audio text dataset through a generative adversarial network;
[0020] An adversarial training module is used to input the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model;
[0021] An acoustic feature extraction module, configured to obtain an audio signal to be detected and extract acoustic features of the audio signal to be detected;
[0022] a non-acoustic feature extraction module, configured to obtain non-acoustic features associated with the audio signal to be detected;
[0023] A feature fusion module, configured to construct a multi-dimensional feature vector based on the acoustic features and the non-acoustic features;
[0024] an abnormality identification module, configured to input the multidimensional feature vector into an audio detection model to generate an abnormality indicator;
[0025] The response control module is used to perform a hierarchical response operation based on the abnormal indicators.
[0026] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an audio signal authenticity verification program stored in the memory and runnable on the processor. When the audio signal authenticity verification program is executed by the processor, the steps of the audio signal authenticity verification method described above are implemented.
[0027] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an audio signal authenticity verification program is stored. When the audio signal authenticity verification program is executed by a processor, the steps of the audio signal authenticity verification method described above are implemented.
[0028] Beneficial effects: The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. It discloses a method, device, equipment and medium for verifying the authenticity of an audio signal, including: collecting original audio and corresponding text to generate an original audio text dataset, generating an adversarial sample set based on the dataset through an adversarial generative network, inputting the original audio text dataset and the adversarial sample set into an audio detection model for joint training to obtain an adversarially trained audio detection model; obtaining an audio signal to be detected and extracting its acoustic features, obtaining non-acoustic features associated with the audio signal to be detected, constructing a multidimensional feature vector based on the acoustic features and non-acoustic features, inputting the multidimensional feature vector into the audio detection model to generate anomaly indicators, and performing a graded response operation based on the anomaly indicators. The present invention improves the robustness of the audio detection model through adversarial sample training, combines acoustic features and non-acoustic features to construct a multidimensional feature vector to enhance the model's recognition ability of cloning attacks, introduces a graded response mechanism to improve the operability of the detection results, and realizes effective recognition and response to new types of voice cloning attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0030] Figure 1 A schematic diagram of an application environment of an audio signal authenticity verification method according to an embodiment of the present invention;
[0031] Figure 21. A flow chart of an embodiment of a method for verifying the authenticity of an audio signal according to the present invention;
[0032] Figure 3 Schematic diagram of the functional modules of a preferred embodiment of the audio signal authenticity verification device of the present invention;
[0033] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0034] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0035] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0036] The audio signal authenticity verification method provided by the embodiment of the present invention can be applied in the following Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can collect original audio and corresponding text from the user terminal to generate an original audio text dataset, generate an adversarial sample set based on the dataset through an adversarial generative network, input the original audio text dataset and the adversarial sample set into the audio detection model for joint training, and obtain an adversarially trained audio detection model; obtain the audio signal to be detected and extract its acoustic features, obtain non-acoustic features associated with the audio signal to be detected, construct a multidimensional feature vector based on the acoustic features and non-acoustic features, input the multidimensional feature vector into the audio detection model to generate anomaly indicators, and perform a hierarchical response operation based on the anomaly indicators. The present invention improves the robustness of the audio detection model through adversarial sample training, combines acoustic features and non-acoustic features to construct a multidimensional feature vector to enhance the model's recognition ability against cloning attacks, introduces a hierarchical response mechanism to improve the operability of the detection results, and realizes effective recognition and response to new voice cloning attacks. The user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server terminal can be implemented using an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0037] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the audio signal authenticity verification method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0038] like Figure 2 As shown, the audio signal authenticity verification method proposed by the present invention includes the following steps:
[0039] S10, collect original audio and corresponding text to generate an original audio text dataset;
[0040] In this embodiment, collecting voice data and matching it one-to-one with text is an important basis for realizing voice authenticity detection. Original audio refers to natural speech signals that have not been edited or processed, and the sources include telephone calls, voice assistant interactions, customer service recordings, or platform voice input content. The corresponding text is a text record that strictly matches the semantic content of the audio, usually obtained through manual annotation or automatic speech recognition systems (such as ASR models). The collected voice data should cover multiple speakers, multiple languages, multiple emotions, and multiple scenarios to ensure the representativeness and generalization ability of the data; the text data should retain the complete semantic structure and context association to support subsequent semantic consistency adversarial training.
[0041] A matching index relationship is established between the original audio and the corresponding text. This is typically achieved through a tagging system that generates a time-aligned file, such as a CTM (Conversation Time Marked) format, a JSON structured index, or a Kaldi corpus index file. This ensures that the start and end times and speaker information of each audio item are accurately associated with the text content. To prevent label drift, quality control mechanisms should be implemented, including keyword consistency verification, speech transcription accuracy scoring, and language normalization of text content, to ensure semantic and temporal consistency between text and audio.
[0042] In practical applications, raw audio and text datasets require a unified data structure framework. For example, a structured table containing fields such as audio path, speaker ID, timestamp, signal-to-noise ratio parameter, language category, and text content can be used, encapsulated in data formats such as Parquet, CSV, or TFRecord. This dataset serves not only as an input source for adversarial training but also as benchmark data for subsequent model evaluation.
[0043] In one implementation, the voice signal emitted by the user can be received by a recording acquisition terminal deployed at the front end and uploaded to the server in real time; the server transcribes the text content through an automatic speech recognition model, and then a manual review system or a rule comparison system corrects the text deviation to ensure that it is completely consistent with the original voice content.
[0044] In another implementation, the raw audio and text pairs can come from existing customer service call databases, voice interaction logs, or data resources from voice question-and-answer platforms. Using batch processing tools, conversations that meet semantic integrity and sound quality standards are extracted as raw audio-text pairs. To increase data diversity, synthesized speech can also be introduced as an auxiliary audio source and bound to pre-set text templates to construct a hybrid audio-text dataset to enhance the robustness of the training set.
[0045] For low-resource languages or accents, local dialect speech data can be collected and transcribed using dialect-specific ASR models. To handle multi-speaker mixed speech scenarios, speaker separation models (such as TS-VAD) can be used to separate audio tracks and then generate multi-channel annotated samples based on the text content.
[0046] Example: In the healthcare sector, voice interaction recordings between doctors and patients can be collected from online consultation platforms or voice guidance systems. Corresponding text can be generated from system or manual recordings to form a standard medical corpus dataset. This dataset is used to detect fraudulent voice generation that impersonates doctors, such as illegal prescriptions or false recommendations. The authenticity of medical consultations is ensured through semantic consistency analysis between text and audio.
[0047] In the financial field, users' transaction instructions in voice banking services, such as voice transfers and voice reimbursements, can be collected, and voice-text data pairs can be generated based on actual transaction texts. The model can be trained to identify potential forged or tampered audio instructions, thereby preventing transfer fraud or financial tampering operations achieved by malicious voice cloning, and improving financial security protection capabilities.
[0048] This embodiment collects raw audio and corresponding text to construct a dataset containing speech-semantic correspondences. This provides accurate, rich, and diverse basic data support for subsequent adversarial sample generation and detection model training, effectively improving the model's ability to detect voice cloning attack samples. The audio and text alignment structure ensures synergistic constraints at the speech and semantic levels during model learning, helping to improve the stability of adversarial training and the model's generalization capabilities, thereby reducing false positives and missed detections in real-world attack scenarios.
[0049] S20, generating an adversarial sample set based on the original audio text dataset through a generative adversarial network;
[0050] In this embodiment, when constructing a set of adversarial examples using a generative adversarial network, a raw audio-text dataset is first used as input. This dataset contains the original audio signal and its corresponding text content. The original audio is typically acquired by a voice acquisition device and initially cleaned through steps such as format standardization, sampling rate unification, and noise preprocessing. The text portion is obtained through speech transcription or original annotation to ensure strict alignment between speech-text pairs. This dataset can include both natural speech recordings and synthesized speech to improve the model's ability to identify synthetic and forged samples.
[0051] Subsequently, the original audio text dataset is input into the constructed adversarial generative network. The network can be based on the Generative Adversarial Network (GAN) architecture, among which WaveGAN, SpecGAN or DiffGAN architectures are particularly suitable for audio processing. The generator part takes the original audio or its spectrogram as input, and generates adversarial samples in the form of audio by introducing perturbation vectors. The perturbation can be achieved through frequency domain perturbation injection, speech speed adjustment, voice bending, SNR adjustment (for example, controlled below -5dB), etc., to simulate cloning attack behaviors that are difficult for the human ear to detect but effectively interfere with the model. The discriminator judges the authenticity of the input sample based on the training objective, thereby guiding the generator to iteratively optimize the concealment and aggressiveness of the perturbation.
[0052] Furthermore, to enhance the diversity and representativeness of sample generation, various attack configurations can be constructed at the input stage, such as simulated diffusion model cloning, deep voice forgery, obfuscation, or encrypted perturbation. In practical implementation, a collection of adversarial samples can be constructed by integrating multi-stage perturbation mechanisms. This ensures that the resulting generated samples have similar speech properties to the original audio, but have the potential to induce misjudgment by the detection model.
[0053] In a specific implementation, several audio-text pairs with complete annotations are selected as input, and the speech waveform is unified (such as converted to 16kHz mono PCM format) and text normalization (stop words removal, unified punctuation, etc.) are completed through the pre-processing module. DiffGAN is used as the adversarial generative network framework, and its generator accepts the original audio and its text feature encoding, and inputs random perturbation noise generated by Gaussian distribution, and outputs adversarial samples in audio format; the discriminator adopts BiLSTM network structure, performs time series modeling on the input audio clip, and outputs true and false labels for discriminant feedback. In order to enhance the attack diversity of adversarial samples, the perturbation parameter distribution can be dynamically adjusted during the training process, and a gradient alignment strategy can be introduced to enhance the misleading ability of the generated samples to the target detection model; and the audio spectrum loss function (such as Mel spectrum distance) can be used to constrain the samples output by the generator to be close to the real audio at the perceptual level. The generated adversarial samples can be stored according to the attack type and used for weight adjustment or difficult example sampling in subsequent model training.
[0054] Example: In healthcare scenarios, remote doctor consultation systems must verify the authenticity of patient-uploaded voice reports to prevent malicious counterfeit case data from misleading diagnoses. By training a detection model on a set of adversarial examples containing dialect accents, mild noise perturbations, and semantic perturbations, the model's robustness against unstructured medical data can be enhanced.
[0055] In FinTech scenarios, users use voice to perform sensitive operations like transferring funds, and the system must verify in real time whether voice commands are cloned. By introducing high-fidelity adversarial examples constructed using a generative adversarial network during model training, such as white noise injection or voice style imitation, the detection model's ability to identify forged voices can be effectively improved, preventing social engineering attacks driven by speech synthesis and enhancing transaction security.
[0056] This embodiment significantly improves the simulation coverage of voice cloning attacks through an automatic sample construction mechanism based on a generative adversarial network. Compared with static data training methods, the dynamically generated and updated adversarial sample set effectively improves the audio detection model's ability to identify forged audio. In particular, when facing new attack strategies based on diffusion models and frequency domain perturbations, the model can maintain a higher detection rate and enhance its robustness. By generating pseudo-samples that are subjectively similar to real speech but easily confused by the model, the detection system has stronger generalization capabilities and defensive adaptability.
[0057] S30, inputting the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model;
[0058] In this embodiment, before jointly training the original audio-text dataset and the adversarial sample set, both must be standardized to ensure consistency in input data format, sampling rate, data dimension, and label structure. The original audio-text dataset typically consists of natural speech signals from real users and their corresponding text information, with a stable speech duration distribution and a clear semantic structure. The adversarial sample set, on the other hand, is forged speech generated by a generative adversarial network (GAN) that simulates various cloning attack methods. This inherently exhibits greater deceptiveness and structural complexity. When merging these two types of data, it is necessary to ensure that the training set has a balanced distribution of sample types to prevent biased learning during the model training phase.
[0059] The input of joint training includes the audio signal itself and its corresponding label. The label not only identifies the authenticity of the audio source, but can also be refined into a category label of the attack method to support multi-task learning strategies. In the specific operation, the original audio text dataset and the adversarial sample set are first spliced to form a joint training dataset, and the order is shuffled to enhance the random generalization ability of the model. It is then input into the audio detection model, which usually includes a feature extraction layer, a multimodal fusion layer, a time series modeling layer, and a classification prediction layer. Among them, the feature extraction layer can use a convolutional network or a pre-trained audio encoder to extract the spectral domain features and semantic context information of the audio; the multimodal fusion layer is used to model the coupling relationship between the speech signal and its textual semantic features; the time series modeling layer usually introduces LSTM or Transformer structures to capture continuous changes across the time dimension; and the final classification output layer outputs the authenticity judgment result of the input sample based on the fused feature representation.
[0060] To improve the model's robustness against adversarial examples, gradient alignment, adversarial loss, or sample reweighting can be introduced. Gradient alignment mitigates the impact of adversarial perturbations on model parameters. Adversarial loss, such as the maximum perturbation penalty function in adversarial training, can improve the detection margin. Sample reweighting dynamically adjusts the contribution of adversarial examples to the overall loss, making model training more stable and discriminative.
[0061] One implementation method is to first convert all raw audio signals into a Mel-spectrogram representation of uniform length. The corresponding text information is then segmented and embedded to construct a set of audio-text pair vectors. The adversarial sample set undergoes the same spectral transformation and label normalization as the original samples. The two types of data are fused in a set ratio (e.g., 7:3) to form a training dataset. This dataset is then fed into an audio detection model, whose backbone architecture can use a ResNet-34 as the audio channel and a BERT encoder as the text semantic channel. A cross-attention mechanism is applied during the feature fusion stage. During training, a cross-entropy loss is used as the base loss function. An adversarial loss term is introduced for pseudo-labeled adversarial samples. An intra-epoch segmentation strategy is used to gradually increase the training weights of adversarial samples to enhance their influence on the model parameter space. After training, the model parameters with the best performance are selected and saved as the adversarially trained audio detection model.
[0062] Another approach uses a three-stage training process: the first stage uses pre-training on the original audio-text dataset to stabilize model parameters; the second stage introduces adversarial examples for fine-tuning and observes how the model responds to forged speech; the third stage performs joint training on the mixed dataset, using K-fold cross-validation to ensure generalization performance. Regularization losses and feature variance constraints are added to the training process to increase the model's sensitivity to perturbations of sample features.
[0063] Example: In a FinTech business scenario, a bank's intelligent voice customer service system uses voice to identify customers and complete high-risk operations. If an attacker uses a cloned voice to simulate customer instructions, it could seriously compromise account security. Introducing a set of adversarial examples containing synthesized speech, pitch-shifted speech, and spectrally perturbed speech into the joint training phase helps the detection model accurately identify these high-fidelity attack samples, thereby promptly blocking illegal transaction instructions.
[0064] In healthcare scenarios, remote voice consultation platforms receive large amounts of patient voice input and rely on semantic recognition results to generate preliminary diagnosis reports. Tampering or forgery of patient voice recordings through voice imitation can mislead the diagnosis process. By incorporating a joint training mechanism using adversarial examples, the detection model can identify signs of disturbed speech or semantic forgery, effectively avoiding the risk of misdiagnosis and ensuring the security and reliability of remote medical services.
[0065] This embodiment uses a joint training method based on a dataset of original audio and text and a set of adversarial examples. The detection model no longer relies solely on a feature space constructed from static data, but instead dynamically perceives variations in the speech form and semantic expression of forged audio, significantly improving the model's ability to identify voice cloning attacks. In environments where multiple attack strategies coexist, this training mechanism effectively enhances the model's adaptability and ability to recognize abnormal features such as high-frequency perturbations, tempo changes, and frequency offsets, effectively reducing the missed detection rate of forged speech and improving the reliability and stability of the detection model in real-world scenarios.
[0066] S40, obtaining an audio signal to be detected, and extracting acoustic features of the audio signal to be detected;
[0067] In this embodiment, in order to complete the authenticity detection of the acoustic signal, it is necessary to first obtain the audio signal to be detected from the external audio data stream. The audio signal can come from a call recording, voice command, real-time voice input, or a data format encapsulation pushed by an external system. The audio signal to be detected is usually represented by linear pulse code modulation (PCM) or compressed format (such as MP3, AAC). After obtaining the audio signal, its format needs to be standardized, including sampling rate resampling, bit depth unification, frame length setting, and silent segment removal processing to ensure that subsequent acoustic feature extraction is performed at a unified scale.
[0068] The extraction of acoustic features relies on spectral analysis of the audio signal and the speech modeling structure. The most commonly used representations include Mel-Frequency Cepstral Coefficients (MFCC), Mel Filter Bank Energies (MFEs), Linear Prediction Cepstral Coefficients (LPCCs), and their dynamic change rates (such as ΔMFCC and ΔΔMFCC). In this processing flow, a fixed sliding window technique is used to divide the standardized audio signal into several frames. Each frame is then subjected to a Fast Fourier Transform (FFT) and Mel filter, and its frequency domain energy distribution is calculated and mapped into an acoustic parameter vector. These parameters can represent information in the audio, such as vocal tract structure, speaking speed, and speech spectral texture, thus providing an effective foundation for subsequent multidimensional feature construction and anomaly detection.
[0069] In one embodiment, the receiving system is configured to pull audio data in real time from a remote front-end acquisition module based on a Socket stream or HTTP protocol. During the pulling process, noise reduction preprocessing is performed on the data packet, such as applying a Wiener filter or a spectral subtraction algorithm to reduce the impact of background noise. Subsequently, the audio signal is uniformly converted to a 16kHz, 16-bit PCM format and divided into short-time segments with 25ms per frame and a frame shift of 10ms. Each segment is transformed by FFT and Mel filter bank to extract 40-dimensional Mel frequency cepstral coefficients, and append their first-order and second-order dynamic features to form a 120-dimensional acoustic feature vector sequence per frame. The system can be configured with a GPU to parallelize the processing process, accelerate the efficiency of feature extraction, and meet the needs of high-concurrency audio stream processing.
[0070] Another implementation can employ an architecture that decouples the acoustic front-end from the language model, extracting only the MFCC and pitch information combination in low-latency scenarios. This feature is then quantized and compressed to an 8-bit integer representation for optimal edge device deployment. After feature extraction, it is directly fed into a locally deployed lightweight detection module, achieving a 50ms response time.
[0071] Example description: In the field of healthcare, remote voice interactions between doctors and patients are becoming increasingly frequent. The accuracy of recognizing medical voice commands in electronic medical record systems directly affects the safety of patient diagnosis and treatment. In order to prevent malicious impersonation of medical staff's voices to perform illegal operations, in the task of audio authenticity verification, the system must first obtain the current voice command as the audio signal to be detected, and extract its acoustic features to capture the microscopic voice parameters in the voice that reflect individual identity and behavior patterns. Through frame-level spectral modeling, the system can identify speaker characteristics such as timbre, speaking speed, and resonance peaks, and promptly expose potential forged voice clues before subsequent multimodal fusion, thereby ensuring the trustworthiness of the voice entry.
[0072] In the financial sector, users initiate account operation requests through voice interaction. Especially in high-risk instructions such as voice transfers or authorizations, attackers may exploit TTS models to imitate user voices and commit fraud. When performing audio authenticity verification, the system must immediately acquire the user's voice stream and extract acoustic features, modeling the user's individual voice style through parameters such as frequency distribution, fundamental frequency envelope, and dynamic change rate. This acoustic feature sequence can be extracted within tens of milliseconds and serves as input for subsequent anomaly detection, helping to identify unnatural speech characteristics in forged voices, such as frequency domain drift or resonance peak misalignment, thereby ensuring the security and real-time responsiveness of the financial voice command channel.
[0073] This embodiment can significantly enhance the robustness and adaptability of the audio authenticity detection system by obtaining the audio signal to be detected in a standardized manner and extracting acoustic features. The unified audio preprocessing process ensures the quality and structural consistency of the signal, while the acoustic features at the frequency domain level express the stable properties of speech in the time-frequency structure. Combining short-term analysis and spectral modeling, the system can maintain the ability to effectively identify audio anomalies even in the face of voice change attacks, speech rate disturbances or frequency offsets.
[0074] S50, obtaining non-acoustic features associated with the audio signal to be detected;
[0075] In this embodiment, during the audio authenticity verification process, relying solely on acoustic features cannot fully capture the contextual information and behavioral background of the speech generation process. Therefore, non-acoustic features are introduced to enhance overall discrimination capabilities. Non-acoustic features refer to information types that are not directly extracted from the audio signal waveform but are closely related to the audio generation or transmission process. They typically include categories such as semantic features, device source, user interaction trajectory, and environmental context.
[0076] The first type of non-acoustic features comes from the textual information corresponding to the audio. After the audio signal is transcribed into text using an automatic speech recognition model, natural language processing methods can be used to extract semantic features such as keyword density, grammatical structure complexity, and contextual intent vectors. These semantic vectors can reflect the authenticity of the audio content and whether there are any unusual instructions that are inconsistent with the user's identity or usage context.
[0077] The second type of non-acoustic features comes from the device producing the audio. For example, the microphone model, sampling rate configuration, and system audio driver version used in the recording can be extracted through device fingerprinting technology. This information can help identify whether the audio is coming from a genuine user device or whether there is transcoding spoofing or playback attacks.
[0078] The third category of non-acoustic features reflects the user's interaction with the device, such as the user's action path, action rhythm, click event sequence, and sliding behavior that triggered the recording. By analyzing this temporal behavior data, we can determine whether the audio occurs in a context consistent with the user's usage habits.
[0079] The fourth type of non-acoustic features comes from environmental sensor data collected during audio acquisition, such as ambient noise levels, location information, and gyroscope sensor status. Collecting this environmental information can further assist in determining whether the audio generation scene is natural and realistic, which is particularly critical in mobile devices or IoT scenarios.
[0080] The above different types of non-acoustic information will be combined in a structured or embedded representation to form a non-acoustic feature set, which will be input into the subsequent model together with the acoustic features.
[0081] The speech transcription module can be used to perform speech recognition processing on the raw audio, generating text content for subsequent semantic analysis. After the text is generated, semantic embedding models (such as BERT and RoBERTa) can be applied to extract semantic feature vectors. Device information can be queried through the operating system API or collected through the hardware fingerprint module, such as calling the Build.MODEL or AudioManager module on an Android device to read the audio input channel type. User behavior data can be recorded in real time through the log collection system, recording events such as button clicks and recording starts and stops, and encoded using the time series encoding module.
[0082] For environmental information collection, the device's sensor framework can be combined with ambient noise data (e.g., using a microphone to detect background noise levels) and other sensor signals (e.g., GPS and accelerometers). After feature filtering and quantization, an environmental vector is generated. The processing results of all submodules must be uniformly converted into fixed-length vectors and structured together to form the final set of non-acoustic features.
[0083] Example: In the financial sector, when a user enters a voice transaction instruction via a mobile device, the system can extract the voice text after receiving the audio signal and cross-verify it with the device model and location information. If the user logs in to an account overseas but the voice content is in a Chinese dialect, and the operation does not match the account's past transaction history, the system can determine the risk of fraudulent use based on non-acoustic features, thereby preventing forged voice calls from triggering fund transfers.
[0084] In healthcare, when patients report their physiological conditions via voice, the system collects information about the device they are using, their current environment, and interaction history (such as whether the report is within the treatment period and whether it was initiated by a registered patient account). This allows for effective identification of authentic voice reports. Even if the voice content matches the template, if the device generating the report doesn't match the patient's historical device, or if the background noise in the surrounding environment doesn't match the medical setting, an anomaly alert can be triggered, preventing false data from entering the telemedicine system.
[0085] This embodiment introduces non-acoustic features to provide contextual information related to the audio signal, compensating for the traditional acoustic model's inability to recognize content semantics, device behavior, and environmental differences. The introduction of non-acoustic features enhances the model's global ability to discriminate against forged audio. Even when faced with high-quality cloned audio, it can still identify forgery risks based on usage behavior and environmental characteristics, significantly improving the detection model's robustness and generalization performance.
[0086] S60, constructing a multidimensional feature vector based on the acoustic features and the non-acoustic features;
[0087] In this embodiment, acoustic features and non-acoustic features are jointly constructed into a multi-dimensional feature vector of a unified structure in order to achieve collaborative modeling of different modal information in a unified expression space. In the specific implementation process, the acoustic feature vectors from the audio signal are first standardized. The goal of this processing operation is to eliminate the differences in amplitude, speaking speed, etc. between different speech samples, and to adjust the eigenvalues to a unified numerical range through normalization to improve the comparability between different input samples. The standardization operation can be implemented based on the Z-score or Min-Max normalization method.
[0088] Non-acoustic features contain multiple sub-categories of information and need to be processed using different vectorization strategies depending on the data type. For text semantic features, word vector models or sentence vector models can be used for embedding mapping to convert natural language semantics into high-dimensional dense vectors. This process retains contextual relationships and semantic distribution characteristics, making it easier to integrate with acoustic features. For device information, you can first define a structured template containing fields such as manufacturer type, device model, audio sampling format, and then map each field to a sparse vector or discrete value encoding. For user operation behavior sequences, sequence embedding is achieved through sliding window segmentation and time series coding modules to capture the temporal distribution and pattern characteristics of the behavior. Environmental data such as background noise levels, location coordinates, sensor outputs, etc. can be generated into vector expressions through feature screening and continuous value quantization.
[0089] The aforementioned non-acoustic vectors typically have inconsistent dimensional structures, so dimensional alignment is required before fusion, including but not limited to zero-padding, principal component mapping, or projection compression. Finally, the standardized acoustic feature vector is concatenated with all non-acoustic feature subvectors to generate a multidimensional feature vector of uniform length. This vector encompasses information about the audio signal across multiple dimensions, including physical, semantic, behavioral, and environmental dimensions, providing rich input for subsequent model construction.
[0090] Standardized acoustic features can be implemented using the sklearn.preprocessing.StandardScaler module in Python, which takes audio frame-level features as input and generates a uniformly distributed output. Semantic vectors can be generated using open-source language models such as BERT and RoBERTa. Recognized speech and text can be input into the model to obtain context-sensitive sentence vectors. Device feature vectors can be constructed by reading results from operating system APIs. For example, on Android, parameters such as Build.MODEL and AudioSource.MIC can be extracted and mapped into fixed-dimensional vectors.
[0091] User action sequences can be segmented into event segments using a sliding window approach. Time-dependent vectors can then be extracted using a recurrent neural network (such as an LSTM) model. Environmental feature encoding can include noise amplitude quantization, geographic coordinate discretization, and scalar encoding of sensor states. These vectors can be reduced in dimensionality through feature selection and principal component analysis. These vectors are then combined using a concatenation function into a unified multidimensional feature vector structure for input into downstream modules.
[0092] Example: In a financial business scenario, a customer uses voice input to authenticate or transfer funds. After extracting the acoustic features, the system simultaneously obtains the text instructions corresponding to the voice and extracts its semantic embedding. This is then combined with the unique identifier of the device used by the customer and the sequence of operational behaviors to form a multidimensional feature vector. If the voice instruction is highly sensitive at the semantic level, but the device characteristics are inconsistent with historical operation records, or the behavioral sequence is abnormal (such as operating too quickly or from a new device), the fusion vector can be used to identify potential voice impersonation risks.
[0093] In the healthcare sector, patients report their health status via voice recordings on mobile devices. After extracting acoustic features, the system simultaneously analyzes the semantics of the health statement corresponding to the voice recording, the environmental context (e.g., whether the recording was taken at home or in a hospital), the consistency of the device used with historical data, and whether the user was the one operating the device themselves, thereby constructing a multidimensional feature vector. If the system detects logical inconsistencies in the semantic content, abnormal device sources, or sudden changes in environmental characteristics, it can determine that the voice recording is suspected of being forged, preventing false medical conditions from influencing diagnosis and treatment decisions.
[0094] This embodiment fuses acoustic features with multiple types of non-acoustic features to construct a multidimensional feature vector. This not only enriches the model's basis for identifying the authenticity of speech signals, but also effectively improves the system's accuracy in detecting abnormal speech, voice forgery, and cloning attacks. The uniformly encoded multimodal information provides the model with more comprehensive discriminative clues, helping it maintain robustness and stability in noisy environments, on edge devices, or under adversarial conditions, significantly reducing the probability of missed and false positives.
[0095] S70, inputting the multidimensional feature vector into an audio detection model to generate an abnormality indicator;
[0096] In this embodiment, passing the multidimensional feature vector as a unified input to the audio detection model is a key operation link in achieving speech authenticity judgment. The model should have the ability to process different modal information and perform time series analysis. In the specific implementation, the multidimensional feature vector constructed in the previous order is first input into the feature extraction layer of the model, which is used to capture the cross-modal interaction relationship in the vector. The feature extraction layer may include a linear transformation network, a convolution submodule, or a self-attention mechanism to adapt to the potential correlation structure between dimensions such as acoustics, semantics, devices, and behaviors.
[0097] The extracted feature representations then enter the model's time series analysis layer. This layer learns the dynamic changes in features over time and enhances the model's ability to model continuous behavior, speech rhythm, and contextual semantic changes. Typical implementations include bidirectional recurrent neural networks (Bi-RNNs), Transformer timing modules, or simplified variants such as GRUs and TCNs. This processing helps detect unusual discontinuities, sudden changes, or suspicious imitations within the continuous structure of the audio signal.
[0098] After time series encoding, the model's output layer determines the anomaly probability of the sample corresponding to the input feature based on internal weights. In this output layer, the classification head compresses the time series features into a single anomaly confidence score, representing the degree to which the input deviates from the training distribution. To improve readability and operability, the anomaly confidence score is further normalized using a sigmoid mapping function to an anomaly indicator within a specified range, such as a continuous value between 0 and 1. A higher value indicates a higher level of suspicion, providing input for subsequent graded response operations.
[0099] The feature extraction layer can be implemented using a Transformer encoder embedded with a self-attention module, allowing the model to complete cross-modal feature fusion during the encoding phase. When the input dimensions do not match the internal dimensions of the model, fully connected layers can be used for dimensionality up- or down-conversion. During the time series analysis phase, pre-trained temporal modeling modules such as the Transformer architecture in wav2vec2.0 can be used, or a multi-layer bidirectional LSTM structure can be constructed to enhance memory capacity. For resource-constrained scenarios, a lightweight version of the TCN (Time Convolutional Network) can be deployed to reduce inference costs.
[0100] The classification output layer can be configured with a 1-D Sigmoid output structure, interpreting the final output value as anomaly confidence. During model deployment, anomaly indicators are generated using the forward propagation module in TensorFlow or PyTorch environments, and anomaly classification is performed using custom thresholds. Model predictions can be run on edge devices, local gateways, or server clusters, allowing deployment adjustments based on scenario requirements.
[0101] Example: In a financial transaction, a customer uses voice to verify a transaction. The system extracts a multidimensional feature vector from the voice and feeds it into an audio detection model. After time series modeling and classification, an anomaly indicator is generated. When this indicator exceeds a set threshold, indicating that the voice may be synthesized or cloned, the system automatically blocks the transaction and requires the customer to undergo secondary identity verification or switch the call to a human.
[0102] In the healthcare sector, patients upload their medical information via remote voice recording. The system processes the multidimensional feature vectors of the input voice and generates anomaly indicators. If the anomaly indicator value of a particular input is significantly higher than the patient's historical voice samples, the system automatically marks it as a high-risk sample, prompting medical staff to conduct a review to avoid the risk of false reporting or identity theft. This mechanism is particularly important for remote diagnosis and treatment and chronic disease management during the epidemic.
[0103] This embodiment integrates multi-dimensional feature vectors into an audio detection model to generate anomaly indicators reflecting the sample's credibility, achieving a quantitative assessment of speech authenticity by integrating multimodal information. This approach not only enhances the model's ability to identify anomalous samples such as forged speech and voice cloning, but also improves the system's robustness in noisy environments and with diverse device inputs. Anomaly indicators provide clear criteria for subsequent risk intervention, automated processing, and manual review, ensuring high real-time and scalability.
[0104] S80: Execute a hierarchical response operation based on the abnormality indicator.
[0105] In this embodiment, the core of implementing a graded response operation lies in dividing the response paths into different levels based on the value of the anomaly indicator to meet the needs of handling audio authenticity issues at different risk levels. The anomaly indicator, as the output of the previous model, is the quantified result of a multidimensional feature vector analyzed by the audio detection model, typically a continuous value between 0 and 1. This indicator reflects the comprehensive differences between the input audio and known real speech samples in the model, such as distance, similarity, and distribution deviation. Based on this indicator, the response level classification logic can be constructed.
[0106] Response levels are typically categorized into three levels: safe (anomaly indicators below the first threshold), alert (anomaly indicators between the first and second thresholds), and high-risk (anomaly indicators above the second threshold). Different levels correspond to different handling strategies. The classification thresholds can be set based on the distribution statistics of anomaly scores in the training set or dynamically adjusted by system administrators based on business security requirements. The design of the tiered strategy should consider the trade-off between false positive and false negative rates to ensure a balance between security and user experience.
[0107] Once the grading is complete, the system will execute the appropriate response actions. These actions may include general logic judgment, automatic interception, secondary verification requests, and manual review task generation. Interfaces can also trigger security policy updates in external systems or policy reassessment mechanisms in risk control systems. The entire tiered response process should support low-latency processing and retain complete response path logs for subsequent audit analysis.
[0108] Anomaly indicator classification thresholds can be statistically determined from a training dataset using cluster analysis, or static values such as 0.3 and 0.7 can be preset. The system can leverage a rules engine to determine the classification level in real time based on anomaly indicators. Safe-level audio can be directly processed by the business system; alert-level audio can automatically generate secondary verification tasks, such as prompting a text verification code or requiring a re-recording of the audio. High-risk audio triggers a risk interception mechanism, suspending operations related to the audio and forwarding it to a manual review interface or security team.
[0109] In specific implementations, a response scheduling module can be built to receive anomaly indicators and drive the corresponding processing flow. Response actions can return status to the front-end interactive system via a Web API, issue policy update requests to the back-end risk control module, or automatically mark high-risk records in the voice log management system. For time-series scenarios, this module should support an event queue processing mechanism to ensure the sequential and stable nature of response actions.
[0110] Example: In a financial business scenario, a user initiates a transfer operation through a voice command. After the system calculates the abnormality index of the voice input, if the index value is 0.25, the system determines it to be at a safe level and directly sends the transfer instruction to the core business module; if the index value is 0.52, the system marks it as an alert level and automatically triggers a second voice confirmation, requiring the user to repeat the transfer amount and beneficiary information; if the index value is 0.88, the system transfers the voice to a high-risk path, suspends the operation, and generates a review work order and pushes it to the risk control personnel for system review.
[0111] In healthcare, when a doctor remotely receives a patient's voice report, the system determines that the current voice sample is at an alert level based on abnormal indicators. The system automatically compares the voice report with the patient's historical voice data and indicates significant discrepancies. It also prompts the medical staff to reconfirm the symptom description via video call. For high-risk input, the system directly blocks data access and uploads the abnormal voice and device identification to the cloud-based model update module for online training feedback.
[0112] This embodiment implements differentiated processing for different types of voice input by implementing graded response actions based on anomaly indicators. This mechanism not only improves the system's ability to handle potential forged audio, voice cloning, and voice fraud, but also effectively reduces the negative impact of false positives on the user experience. The flexible processing capabilities of graded response enable the system to quickly respond to new attack methods while ensuring the integrity of the operational chain and providing underlying support for business process security.
[0113] The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. A method, device, equipment and medium for verifying the authenticity of an audio signal are disclosed, including: collecting original audio and corresponding text to generate an original audio text dataset, generating an adversarial sample set based on the dataset through an adversarial generative network, inputting the original audio text dataset and the adversarial sample set into an audio detection model for joint training to obtain an adversarially trained audio detection model; obtaining an audio signal to be detected and extracting its acoustic features, obtaining non-acoustic features associated with the audio signal to be detected, constructing a multidimensional feature vector based on the acoustic features and non-acoustic features, inputting the multidimensional feature vector into the audio detection model to generate an anomaly indicator, and performing a graded response operation based on the anomaly indicator. The present invention improves the robustness of the audio detection model through adversarial sample training, combines acoustic features and non-acoustic features to construct a multidimensional feature vector to enhance the model's recognition ability for cloning attacks, introduces a graded response mechanism to improve the operability of the detection results, and achieves effective recognition and response to new types of voice cloning attacks.
[0114] In one embodiment, the above step S20 includes:
[0115] S201, build a generative adversarial network;
[0116] S202, simulating a noise injection attack in the adversarial generative network, adding noise disturbance to the audio signal of the original audio text dataset, and generating a noise-disturbed speech sample;
[0117] S203, simulating a variable speed perturbation attack in the generative adversarial network, performing variable speed perturbation processing on the audio signal in the original audio text dataset, and generating a variable speed perturbation speech sample;
[0118] S204, simulating a frequency domain shift attack in the adversarial generative network, performing frequency domain shift processing on the audio signal in the original audio text dataset, and generating a frequency domain perturbed speech sample;
[0119] S205 , integrating the noise-perturbed speech samples, the speed-varied perturbed speech samples, and the frequency-domain perturbed speech samples to generate an adversarial sample set.
[0120] In this embodiment, building an adversarial generative network is the basis for realizing attack simulation, and a structure consisting of a generator and a discriminator is usually adopted. In the audio scenario, the generator can receive real audio and a perturbation condition vector as input, and output a synthetic speech sample containing perturbation; the discriminator is used to distinguish whether the sample contains detectable forged features, and the two form a dynamic convergence of the perturbation attack strategy through adversarial optimization iteration. The network structure can be selected based on conditional GAN (Conditional Generative Adversarial Network) or adversarial autoencoder (AAE) to ensure that the output speech samples have a high degree of similarity at the perceptual level and are perturbative in the model space.
[0121] In a generative adversarial network, noise injection attacks simulate common real-world background interference and channel contamination. Specific methods include adding Gaussian white noise, pink noise, or simulated fan noise or electromagnetic interference found in call environments to the audio signal. The perturbation intensity can be controlled and parameterized using the signal-to-noise ratio (SNR), typically set to 15dB, 10dB, and 5dB, to assess the model's robustness to varying degrees of noise perturbation.
[0122] Variable-speed perturbation attacks simulate changes in speech rate or playback rhythm adjustments, typically achieved through time-domain compression or expansion techniques. These include spectral time-axis interpolation and resampling after a short-time Fourier transform (STFT), or pitch-preserving variable-speed distortion using the WSOLA (Waveform Similarity Overlap and Add) algorithm. This perturbation method aims to generate samples that remain unchanged in speech recognition but shift in the model distribution space.
[0123] Frequency-domain shift attacks focus on perturbations in the spectral domain. By performing specific frequency shifts or weighted perturbations on Mel-frequency cepstral coefficients (MFCCs), Log-Mel features, or linear spectrum features, the signal remains natural in the auditory space, but the frequency energy distribution deviates from the statistical distribution of the training set. This type of attack can bypass detection models based on spectral similarity.
[0124] The resulting adversarial sample set is constructed by integrating samples subjected to different types of perturbations. This set must ensure sample diversity and representativeness of attack methods, including samples with added noise and samples perturbed by a combination of speed changes and frequency deviations. The integrated samples can be labeled and indexed based on attack type labels, enabling type-specific training strategy adjustments and tiered optimization of attack intensity in subsequent training stages.
[0125] By constructing a generative adversarial network and simulating various types of audio perturbation attacks, this embodiment effectively generates a set of high-quality adversarial samples covering multiple attack methods, providing rich attack scenario input for subsequent detection model training. This approach improves the model's ability to perceive and recognize complex voice cloning attacks in real-world scenarios, significantly enhancing the model's robustness and generalization under different types of adversarial perturbations. In particular, it possesses stronger discrimination and fault tolerance when faced with unforeseen new forged samples, effectively reducing missed detection rates and false alarm rates.
[0126] In one embodiment, the above step S30 includes:
[0127] S301, setting initial network parameters of the audio detection model;
[0128] S302, merging the original audio text dataset and the adversarial sample set to generate a joint training dataset;
[0129] S303, applying a gradient alignment module to adjust the weights of the adversarial samples in the joint training dataset;
[0130] S304, optimizing the network parameters of the audio detection model using the weighted joint training dataset to generate an optimized audio detection model;
[0131] S305, testing the optimized audio detection model on the dialect dataset to generate performance verification results;
[0132] S306: When the performance verification result reaches a preset performance threshold, the optimized audio detection model is output as an adversarially trained audio detection model.
[0133] In this embodiment, setting the initial network parameters of the audio detection model is part of the startup process of the training phase, which involves setting the model architecture and pre-training state. The network structure can be based on a temporal convolutional network (TCN), a bidirectional long short-term memory network (BiLSTM), or a hybrid attention network architecture, and its convolution kernel size, hidden layer dimension, and activation function type are configured. Parameter initialization typically uses Xavier initialization or He initialization to ensure gradient stability during forward propagation.
[0134] The process of merging the original audio-text dataset and the adversarial sample set to form a joint training dataset should adhere to the principles of label consistency and balanced sample distribution. Original data is typically labeled with true annotations, while adversarial samples are labeled with perturbation type or attack intent. During the merging process, a representative ratio of positive and negative samples should be maintained. The training set should support a dynamic sampling mechanism to improve the model's learning ability under weakly supervised attack conditions.
[0135] A gradient alignment module is introduced in joint training to correct the weights of adversarial examples. This mechanism measures the direction and magnitude of the gradients induced by different examples to adjust the contribution of adversarial examples in backpropagation. This can be achieved through gradient projection weighting or sample reweighting based on sensitivity analysis (e.g., a variant of sharpness-aware minimization). This module significantly mitigates the impact of adversarial perturbations on model convergence and improves training stability.
[0136] Optimizing the network parameters of the audio detection model using the weighted joint training dataset follows a standard backpropagation optimization process. Multi-objective optimization can be performed using a cross-entropy loss function, a combination of adversarial loss functions, or a variational discriminant objective. The optimization algorithm typically uses momentum-based adaptive learning rate methods such as Adam, Ranger, or AdaBelief. Dynamic learning rate decay and early stopping can be configured during training to prevent overfitting.
[0137] Testing on dialect datasets is part of model robustness verification. Dialect datasets can cover a variety of low-resource languages, including Cantonese, Wu, and Minnan, and include multi-segment, cross-device, and multi-rate scenarios. Verification involves evaluating the model's recognition accuracy, forgery detection recall, and confidence interval stability in non-mainstream and mixed-language environments.
[0138] When performance verification results reach the preset performance threshold, the optimized model is output. This requires that multiple evaluation dimensions (such as AUC, F1 score, and false alarm rate) exceed the training target values. Performance thresholds can be set based on business requirements for true detection accuracy, security defense capabilities, and false trigger tolerance, ensuring the final output model has broad adaptability and practical defense capabilities.
[0139] This embodiment integrates the original audio with adversarial samples and inputs them into the detection model for joint training. It also introduces a gradient alignment mechanism to adjust the weights of the adversarial samples, enabling the model to effectively perceive the differential characteristics of adversarial perturbations while learning the structure of real speech. After optimized training, the model's recognition ability for multiple types of attack samples is significantly enhanced, and it maintains high robustness and a low missed detection rate, especially under complex speech conditions such as dialects, speed changes, and frequency deviations. The resulting detection model has performance advantages such as strong generalization, low false positive rate, and resistance to unknown forged speech, effectively improving the level of voice content security protection.
[0140] In one embodiment, the above step S40 includes:
[0141] S401, receiving an audio input signal to be detected;
[0142] S402, performing noise reduction processing on the audio input signal to generate a noise-reduced audio signal;
[0143] S403, performing frame processing on the noise-reduced audio signal to generate an audio frame sequence;
[0144] S404, extracting fundamental frequency envelope features of the audio frame sequence to generate a fundamental frequency envelope feature sequence;
[0145] S405: Combining the fundamental frequency envelope feature sequence into an acoustic feature vector.
[0146] In this embodiment, receiving the audio input signal to be detected refers to the system acquiring voice signal data in real time or in batches from the user end or transmission channel. This signal can come from a mobile device microphone, a call recording system, or a front-end voice acquisition module, and is typically encoded in linear PCM or a compressed encoding format (such as OPUS or AAC). The receiving link needs to support different channel formats (mono / stereo), different sampling rates (common examples include 16kHz and 8kHz), and different industry audio protocols, and have a stable data buffering mechanism to ensure the stability of subsequent feature extraction.
[0147] Denoising the audio input signal involves suppressing background noise while preserving the effective speech components, thereby improving feature extraction quality. Denoising algorithms can include spectral subtraction, Wiener filtering, or an end-to-end noise reduction network based on a deep neural network. This process aims to reduce interference from non-speech components such as keyboard tapping, wind noise, and electromagnetic interference. It is particularly suitable for scenarios with complex environmental interference sources, such as financial customer service halls and hospital waiting areas.
[0148] Framing the denoised audio signal involves segmenting the continuous time-domain signal into a series of short-term analysis frames based on fixed time windows (e.g., 25ms) and overlapping intervals (e.g., 10ms), ensuring signal stability within each frame. This step uses a sliding window method to transform the denoised continuous audio signal into analyzable units. The framing results in a series of short-term subsequences with a temporal structure, providing the foundation for subsequent frequency-domain feature extraction and time-domain trend modeling.
[0149] Extracting the fundamental frequency envelope feature from an audio frame sequence involves calculating the fundamental frequency (F0) trend for each frame, generating a time series that reflects the fluctuations in the amplitude and periodic structure of the speech's fundamental frequency. This feature differs from conventional Mel-Frequency Cepstral Coefficients (MFCCs) in that it focuses more on changes at the sound source level, such as rising intonation, interrogative tone, and articulation clarity. It is a crucial indicator for identifying forged speech and lack of natural speech rhythm in cloning attacks. This can be achieved by generating a fundamental frequency estimate based on the YIN algorithm, CREPE neural network, or autoregressive linear prediction method, and then extracting the global contour using an envelope function.
[0150] The process of combining fundamental frequency envelope feature sequences into an acoustic feature vector involves matrix concatenation or statistical aggregation of the fundamental frequency features corresponding to all time frames to construct a temporally representative and representative vector representation. This vector can be input into downstream detection models as the primary acoustic indicator for anomaly detection. Combination methods can include average pooling, time-series flattening, or stacking for multi-dimensional expansion to enhance the representation of speech pitch and rhythm variations and adapt to the input requirements of different models.
[0151] This embodiment extracts acoustic features based on the fundamental frequency envelope from the original audio signal, and combines it with noise reduction and framing processing to effectively preserve the sound source level information in the speech that is difficult to forge against disturbances. In different environments, this feature has strong robustness and recognizability, which can improve the detection model's recognition accuracy for voice-changing attacks, cloned voices, and abnormal intonations. Compared with traditional methods that rely solely on spectral features, this acoustic expression method has stronger discrimination and generalization capabilities, especially in dialect systems and noisy environments. It performs stably, which helps to improve the overall effect of speech authenticity verification.
[0152] In one embodiment, the above step S50 includes:
[0153] S501, identifying text content corresponding to the audio signal to be detected and generating a text transcription;
[0154] S502, extracting semantic features of the text transcription to generate text semantic features;
[0155] S503, collecting information about a device generating the audio signal to be detected;
[0156] S504, analyzing the user interaction behavior pattern associated with the audio signal to be detected to generate a user operation behavior sequence;
[0157] S505, obtaining environmental sensor data when the audio signal to be detected is generated;
[0158] S506 , combining the text semantic features, generation device information, user operation behavior sequence, and environmental sensor data to generate a non-acoustic feature set.
[0159] In this embodiment, identifying the text content corresponding to the audio signal to be detected is to perform speech recognition processing on the voice information contained in the audio signal to convert it into a structured text expression. This process usually relies on an automatic speech recognition (ASR) model, such as an acoustic modeling structure based on CTC (Connectionist Temporal Classification) or an end-to-end speech recognition network based on the Transformer architecture. The recognition accuracy depends on the model's adaptability to voice accents, speaking speed, and noise. The recognized text information not only provides a basis for semantic understanding, but also provides contextual input for subsequent comparison and analysis.
[0160] Extracting text semantic features involves obtaining deep semantic representations from transcribed text. This process can utilize pre-trained language models (such as BERT and RoBERTa) to generate context-sensitive semantic embedding vectors, or it can combine sentiment analysis, keyword extraction, command recognition, and other subtasks to perform multi-layer expression extraction. This approach can capture the presence of unusual trading instructions, emotional fluctuations, sensitive content, and other semantic information in the text, which can be used to determine the effectiveness of cloning attacks.
[0161] The device information used to collect audio signals refers to the recording and analysis of the hardware and software configuration that generates the audio, including the device model, system version, codec parameters, microphone type, etc. This data can be obtained through application layer logs, communication protocol header information, browser fingerprints, or API call paths. This information can help determine whether the audio is coming from an unauthorized device or a suspicious environment, further assisting in verifying its authenticity.
[0162] Analyzing user interaction patterns associated with audio signals involves capturing user interface interactions triggered before, during, or after the audio is generated, such as click paths, scrolling paths, page dwell time, and button trigger frequency, to construct a user action sequence. This data can be collected and modeled through front-end tracking, log analysis, or behavioral modeling modules, generating a behavioral data structure that can be used to identify abnormal manipulation patterns or traces of script generation. This sequence feature is invaluable in identifying automated synthesis or fraudulent operations.
[0163] Acquiring environmental sensor data during audio signal generation involves extracting sensory information from the device's current environment when the audio is produced, such as temperature, light intensity, gravity, and network status around the microphone. This information is collected through the device's built-in sensor call interface. This data is used to determine whether the audio is generated in a real environment or is potentially fabricated. It can be cross-validated with acoustic signal timestamps and contextual behavior.
[0164] Combining text semantic features, generating device information, user action sequences, and environmental sensor data represents these multi-source heterogeneous data as a complete set of non-acoustic features using a unified structure. This combination can be achieved through multi-channel tensor stacking, multimodal graph structure fusion, or weight alignment based on an attention mechanism. Ultimately, this creates a multi-dimensional fusion feature representation of semantics, behavior, device, and environment that provides complementary verification capabilities for audio authenticity.
[0165] This embodiment constructs a set of non-acoustic features encompassing text semantics, device origin, user interaction patterns, and environmental conditions, effectively addressing the limited adaptability to adversarial samples found in traditional detection methods that rely solely on the speech itself. This multi-dimensional information collaboration approach cross-validates the likelihood of speech synthesis or forgery from multiple perspectives, including language content consistency, device credibility, behavioral rationality, and the authenticity of the generated environment. This helps improve the detection system's recognition accuracy against unstructured attacks, social engineering-style audio cloning attacks, and scripted batch attacks, while also enhancing the detection model's generalization robustness and contextual understanding capabilities.
[0166] In one embodiment, the above step S60 includes:
[0167] S601, performing vector normalization processing on the acoustic feature to generate a standardized acoustic feature vector;
[0168] S602, performing embedding mapping processing on the text semantic features in the non-acoustic features to generate a semantic vector;
[0169] S603, performing structured mapping processing on the generating device information in the non-acoustic features to generate a device feature vector;
[0170] S604, encoding the user operation behavior sequence in the non-acoustic feature to generate a behavior feature vector;
[0171] S605, encoding the environmental sensor data in the non-acoustic feature to generate an environmental feature vector;
[0172] S606 , dimensionally aligning and concatenating the standardized acoustic feature vector, the semantic vector, the device feature vector, the behavior feature vector, and the environment feature vector to generate a multi-dimensional feature vector.
[0173] In this embodiment, vector normalization is first required to convert acoustic information into a form suitable for model processing. This is to eliminate the impact of dimensional differences or numerical distribution deviations between different dimensions, so that subsequent model processing will not be distorted due to inconsistent feature numerical scales. The normalization method can use Z-score standard deviation normalization, minimum and maximum linear scaling, or uniform adjustment of mean and variance distribution through the BatchNorm module to generate a continuous acoustic vector representation that is friendly to model training. This processing has a positive effect on overcoming the offset caused by differences in audio device quality or environmental noise.
[0174] Embedding mapping for text semantic features is the process of converting discrete text information into a continuous spatial representation. The core goal is to maintain the consistency of the semantic structure of the language in the vector space. This mapping method can be based on pre-trained Transformer language models, such as BERT or ERNIE. By inputting a context sentence, it outputs a representation vector in a high-dimensional semantic space. These semantic vectors not only capture word meaning but also reflect contextual dependencies and relationships, making them valuable for identifying potential semantic anomalies and logical inconsistencies in synthesized speech.
[0175] The purpose of structured mapping of generated device information is to encode the raw, unstructured hardware or software configuration parameters into fixed-length feature representations. This can be achieved using one-hot encoding, hash encoding, or mapping to a sparse vector after looking up the corresponding index value in a device type dictionary. Alternatively, a concatenated vector representation can be constructed using a predefined field structure. The mapping results should ensure that speech generated by similar devices has measurable feature similarity, while illegitimate devices or impersonation sources will exhibit structural differences.
[0176] Sequence modeling methods such as RNN, Transformer, or one-dimensional convolution can be used to encode user action sequences. Pattern features, such as action rhythm, temporal distribution, and input density, can be extracted from these sequences. During the encoding process, the raw interaction data can be timestamp aligned, frequency counted, and windowed before being mapped into a vectorized sequence representation to characterize the rationality of the interaction process behind the generated audio. This feature can significantly enhance anomaly detection capabilities for audio data that simulates user click behavior or is automatically generated by scripts.
[0177] The encoding process of environmental sensor data aims to transform non-speech environmental information into supplementary features for determining audio authenticity. Common methods include normalizing various sensor values and concatenating them into multidimensional vectors, using convolutional networks to extract temporal features, or constructing graph structures to holistically model the environmental state. This process can capture whether authentic generation conditions were met at the moment of generation, thereby identifying suspicious background environments or data tampering.
[0178] After the feature expressions in these multiple dimensions are generated, dimensional alignment and concatenation are required. Dimension alignment can be achieved by using zero-padding, projection matrices, or attention fusion to ensure that features from different sources are uniformly encoded in the same dimensional space while maintaining the integrity of the feature information. The concatenation operation assembles all processed feature vectors into a continuous multidimensional vector structure for subsequent input into the model for classification or regression. This combined structure preserves the temporal information of the speech while integrating multimodal information from semantics, devices, behaviors, and the environment, resulting in stronger expressiveness and greater robustness against attacks.
[0179] This embodiment integrates standardized acoustic features with semantic, device, behavior, and environmental information vectors from non-acoustic dimensions, then performs dimensional fusion to construct a high-dimensional feature representation with multimodal expression capabilities, significantly enhancing the audio detection model's adaptability to complex attack scenarios. This feature fusion approach effectively mitigates the misjudgments and missed detections caused by traditional detection methods that rely solely on acoustic features. Furthermore, by integrating non-acoustic factors, it builds a complete understanding of the attacker's behavior chain and the generated environment, enabling three-dimensional recognition of cloned voices, counterfeit devices, and simulated interactions, effectively improving the comprehensiveness, stability, and generalization capabilities of detection.
[0180] In one embodiment, the above step S70 includes:
[0181] S701, inputting the multidimensional feature vector into the feature extraction layer of the audio detection model;
[0182] S702, performing multimodal feature fusion at the feature extraction layer to generate a fused feature representation;
[0183] S703, processing the fused feature representation through the time series analysis layer of the audio detection model to generate a time series feature code;
[0184] S704: determining, at the classification output layer of the audio detection model, an anomaly confidence score based on the temporal feature encoding;
[0185] S705: Convert the anomaly confidence score into an anomaly index.
[0186] In this embodiment, when the constructed multi-dimensional feature vector is input into the audio detection model, it first enters the feature extraction layer of the model, which is used to perform fusion perception and preliminary modeling of the multi-source heterogeneous features of the input. The feature extraction layer can be composed of a multi-channel convolutional neural network structure or a multimodal attention mechanism, which is used to process the different distribution characteristics of acoustic vectors, semantic vectors, device feature vectors, behavioral feature vectors and environmental feature vectors respectively. At this stage, the model automatically selects the corresponding sub-channel according to the dimension of the input data to perform independent transformation or extract common representations through a shared parameter model to ensure semantic consistency and local correlation between different modalities.
[0187] Multimodal feature fusion can be implemented using channel attention mechanisms (such as SE-block), cross-modal attention (such as BilinearAttentionPooling), or a unified transformation module (such as the Transformer Fusion Encoder). The fusion process not only preserves the discriminative aspects of the original features of each modality but also explores possible cross-dependencies between them, such as the synchronization of semantic and acoustic changes, and the structural association between device features and the speech spectrum. The fused feature representation reflects cross-modal collaborative information and is more sensitive to identifying potential fraud clues in synthesized speech.
[0188] Next, the fused feature representation is fed into the time series analysis layer, which can employ an LSTM, GRU, Temporal Convolution Network (TCN), or a positional encoding-based Transformer architecture to process the temporal relationships formed by multidimensional features as they change over time frames. This processing layer aims to capture the dynamic temporal patterns of audio signals and characteristics of attack behavior, such as sudden changes in pitch and inconsistent semantic rhythm. Within the multidimensional fused features, acoustic variations, behavioral rhythm, and environmental noise interference can manifest as typical temporal local anomalies, which can be extracted as discriminative temporal codes through time series modeling.
[0189] The temporal feature encoding is then passed to the classification output layer of the audio detection model. In this layer, one or more fully connected layers are usually configured to compress the high-dimensional representation into a specific confidence score through softmax, sigmoid, or normalized activation functions. This confidence value can be understood as the model's probability prediction result that the current input is abnormal audio. To enhance interpretability and threshold management, this score must also be mapped into a clear anomaly indicator through nonlinear transformation or normalization methods, such as a binary label based on a fixed threshold, a grade label based on interval division, or an attack risk value output in the form of a percentage.
[0190] Ultimately, the anomaly confidence score is converted into an anomaly indicator through a specific threshold mapping or scoring function, which can serve as a key basis for subsequent response strategy triggering, model performance evaluation, and sample review.
[0191] Furthermore, to ensure that the audio detection model can continuously adapt to evolving attack methods and changes in environmental variables after deployment, the dynamic update process incorporates online optimization and semi-supervised learning mechanisms to build a closed loop for adaptive detection model production. This process typically includes multiple collaborative processing units, including parameter update triggering logic, incremental sample acquisition mechanism, model weight reassessment strategy, and output verification logic.
[0192] During the dynamic update process, a dynamic feedback sample pool is first constructed by collecting difficult examples, misclassified examples, and examples with outliers in confidence levels encountered by the detection model during the deployment phase. These feedback samples are automatically classified using a posterior label correction mechanism and anomaly distribution clustering mechanism, allowing them to be used to construct a pseudo-label set without manual intervention. Combined with a gradient alignment module and robust regularization strategy, this allows the model to fine-tune the discrimination boundaries of new examples without interfering with existing model performance.
[0193] To avoid the risk of model migration caused by adversarial perturbation samples, the dynamic update process typically introduces a perturbation compression mapping mechanism in a specific feature space, confining the update training process to a low-risk subspace. Before updating the model structure parameters, the perturbation increment is calculated by a constrained optimization module. If the increment exceeds the parameter stability threshold, the offset weight is suppressed by a regularizer. During the update process, old and new samples in the joint training set are mixed according to an adaptive weighting function to generate a dynamically optimized data batch.
[0194] The updated model requires integration verification. This typically involves introducing a lightweight performance verification module into the online inference side. This module constructs a timing consistency check set using the input data stream within a sliding time window to ensure that the model meets established thresholds for both functional and timing stability. Finally, the verified model is deployed to the main online audio detection channel, with the model registrar completing version switching and archiving historical weights.
[0195] Through a dynamic update process, the audio detection model continuously absorbs new adversarial examples and real-world sample feedback from changing environments, improving its ability to identify variant attacks while ensuring model performance stability. This process achieves the goal of dynamically enhancing model robustness without requiring extensive manual annotation intervention by constructing pseudo-labeling, adaptive perturbation suppression, and pre- and post-update difference constraint mechanisms. The updated model maintains stable detection performance in diverse scenarios such as low signal-to-noise ratios, edge deployments, mixed dialect speech, and unknown device input, effectively shortening the defense capability iteration cycle and improving the system's security response efficiency in real-world applications.
[0196] Example: In a financial services voice system, the voice generated when a user submits a transaction instruction via telephone banking serves as the audio signal to be detected. The system first calls the voice acquisition interface to receive the raw audio and then generates a raw audio text dataset based on the recognized user pronunciation. Using a generative adversarial network, it simulates the voice cloning perturbations that a real attacker might employ. These perturbations include adding background white noise to the raw audio, adjusting the speech rate to simulate vocal rhythm differences, and applying frequency domain offset perturbations to simulate vocal feature drift, thereby generating a set of deceptive adversarial samples.
[0197] The original audio-text dataset and the aforementioned adversarial sample set are combined into a joint training dataset. The gradient alignment module dynamically adjusts the weight of the adversarial samples in training, allowing the audio detection model to complete targeted reinforcement learning. The model then accepts incoming audio signals to be detected and performs noise reduction, framing, and fundamental frequency envelope extraction to generate acoustic features. Simultaneously, it identifies the textual content corresponding to the speech, extracts the user's semantic intent, collects device information and historical user interaction sequences, and analyzes background sound characteristics in the current environment to form a set of non-acoustic features.
[0198] These acoustic and non-acoustic features are normalized, structured, and vector-mapped before being concatenated into a multidimensional feature vector. This vector is then fed into the audio detection model. Feature fusion, time series analysis, and classification layer processing generate anomaly confidence scores, which are ultimately converted into anomaly indicators. If the model determines that the audio presents a risk of counterfeiting, it implements a tiered response based on the anomaly indicator, including interception instructions, high-priority manual review, and account freezing.
[0199] The system also feeds the audio input, feature status, and response records from this processing process into the model's dynamic update mechanism. If the model experiences significant confidence fluctuations or high-frequency anomalies during processing, it automatically collects relevant audio and behavioral data to build a feedback sample pool, triggering the model's online incremental update process to ensure that model parameters can be adjusted promptly as attack patterns change.
[0200] On the hospital's intelligent voice consultation platform, patients submit a description of their health symptoms via a remote voice terminal, triggering an automated initial diagnosis service. While collecting the patient's voice input, the system performs speech recognition to generate a text description, building a dataset of raw audio and text. To account for potential deepfake voice deception attempts, such as attackers cloning patients' voices to induce false consultations or steal private consultation records, the platform automatically simulates typical adversarial attacks such as noise perturbations, speed-shifting interference, and frequency domain shifts using a generative adversarial network (GAN), generating a set of adversarial examples for voice cloning.
[0201] The model is jointly trained with real speech and adversarial samples, integrating real medical interaction data and typical attack scenarios during training to ensure that it can recognize potential synthesized speech in real diagnosis and treatment contexts. After a new voice is connected, the platform first performs noise reduction and frame processing on the voice content, and extracts acoustic signal patterns that are highly correlated with medical semantic features (such as fundamental frequency envelope features such as unstable breath and prolonged speech); at the same time, the system obtains the patient's voice terminal device model, historical conversation behavior patterns and semantic deviation trends, and reads the real-time monitoring of vital sign sensor data from the smart device (such as abnormal background heart rate, etc.) as non-acoustic features.
[0202] After the multi-dimensional features are concatenated, they are fed into the audio detection model for fusion processing, generating an anomaly confidence score. The final anomaly indicator is then formed based on whether the speech deviates from the true characteristics. For example, if the model detects high-frequency templated statements or inconsistent ambient noise, a high-confidence anomaly indicator is output, and different levels of response can be implemented: a mild response may prompt the user to repeat the statement or switch to manual consultation; a moderate response may block the instruction and log it; and a severe response may automatically mark the interaction as fraudulent and trigger the hospital's security audit process.
[0203] Likewise, all features and responses during model processing are automatically recorded and used as feedback data for subsequent dynamic updates. The system regularly collects samples of users' latest voices, emerging cloning technologies in medical scenarios, and speech from different dialects, automatically triggering incremental model updates and adaptations. This improves attack resistance in complex voice scenarios such as heterogeneous terminals, remote diagnosis and treatment, and chronic disease management.
[0204] This embodiment inputs multidimensional feature vectors into an audio detection model capable of multimodal fusion and time series modeling. The model then sequentially completes feature extraction, fusion modeling, dynamic change identification, and confidence quantification. This not only improves the model's ability to learn new synthetic attack features, but also enhances its ability to discriminate against nonlinear cross-modal interference and time series perturbations. The resulting anomaly indicators comprehensively reflect the degree of anomaly in audio across multiple dimensions, including content, source, behavior, and environment. This provides a stable and reliable basis for subsequent response decisions, effectively reducing missed detection rates and improving response efficiency and accuracy in actual deployment scenarios.
[0205] In one embodiment, an audio signal authenticity verification device is provided, which corresponds to the audio signal authenticity verification method in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the audio signal authenticity verification device of the present invention. These modules include a data acquisition module 10, an adversarial sample generation module 20, an adversarial training module 30, an acoustic feature extraction module 40, a non-acoustic feature extraction module 50, a feature fusion module 60, an anomaly identification module 70, and a response control module 80. Each functional module is described in detail below:
[0206] The data collection module 10 is used to collect original audio and corresponding text to generate an original audio text dataset;
[0207] An adversarial sample generation module 20 is configured to generate an adversarial sample set based on the original audio text dataset through a generative adversarial network;
[0208] An adversarial training module 30 is configured to input the original audio text dataset and the adversarial sample set into an audio detection model for joint training to obtain an adversarially trained audio detection model;
[0209] The acoustic feature extraction module 40 is used to obtain the audio signal to be detected and extract the acoustic features of the audio signal to be detected;
[0210] a non-acoustic feature extraction module 50, configured to obtain non-acoustic features associated with the audio signal to be detected;
[0211] A feature fusion module 60 is configured to construct a multi-dimensional feature vector based on the acoustic features and the non-acoustic features;
[0212] Anomaly identification module 70, configured to input the multi-dimensional feature vector into an audio detection model to generate an anomaly indicator;
[0213] The response control module 80 is configured to execute a hierarchical response operation based on the abnormality indicator.
[0214] In one embodiment, the adversarial sample generation module 20 is specifically configured to:
[0215] Build a generative adversarial network;
[0216] Simulating a noise injection attack in the adversarial generative network, adding noise perturbation to the audio signal of the original audio text dataset, and generating a noise-perturbed speech sample;
[0217] Simulating a variable speed perturbation attack in the generative adversarial network, performing variable speed perturbation processing on the audio signal in the original audio text dataset, and generating a variable speed perturbation speech sample;
[0218] Simulating a frequency domain offset attack in the adversarial generative network, performing frequency domain offset processing on the audio signal in the original audio text dataset, and generating a frequency domain perturbed speech sample;
[0219] The noise-perturbed speech samples, the speed-varied perturbed speech samples, and the frequency-domain perturbed speech samples are integrated to generate an adversarial sample set.
[0220] In one embodiment, the adversarial training module 30 is specifically configured to:
[0221] Set the initial network parameters of the audio detection model;
[0222] Merging the original audio text dataset and the adversarial sample set to generate a joint training dataset;
[0223] Applying a gradient alignment module to adjust the weights of adversarial examples in the joint training dataset;
[0224] Optimizing the network parameters of the audio detection model using the weighted joint training dataset to generate an optimized audio detection model;
[0225] Test the optimized audio detection model on the dialect dataset to generate performance verification results;
[0226] When the performance verification result reaches a preset performance threshold, the optimized audio detection model is output as the adversarially trained audio detection model.
[0227] In one embodiment, the acoustic feature extraction module 40 is specifically configured to:
[0228] receiving an audio input signal to be detected;
[0229] Performing noise reduction processing on the audio input signal to generate a noise-reduced audio signal;
[0230] Performing frame processing on the noise-reduced audio signal to generate an audio frame sequence;
[0231] Extracting fundamental frequency envelope features of the audio frame sequence to generate a fundamental frequency envelope feature sequence;
[0232] The fundamental frequency envelope feature sequences are combined into an acoustic feature vector.
[0233] In one embodiment, the non-acoustic feature extraction module 50 is specifically configured to:
[0234] Identifying text content corresponding to the audio signal to be detected and generating a text transcription;
[0235] Extracting semantic features of the text transcription to generate text semantic features;
[0236] Collecting information about a device generating the audio signal to be detected;
[0237] Analyzing a user interaction behavior pattern associated with the audio signal to be detected to generate a user operation behavior sequence;
[0238] Acquiring environmental sensor data when the audio signal to be detected is generated;
[0239] The text semantic features, generation device information, user operation behavior sequence and environmental sensor data are combined to generate a non-acoustic feature set.
[0240] In one embodiment, the feature fusion module 60 is specifically configured to:
[0241] performing vector normalization processing on the acoustic features to generate a standardized acoustic feature vector;
[0242] Performing embedding mapping processing on the text semantic features in the non-acoustic features to generate a semantic vector;
[0243] Performing structured mapping processing on the generating device information in the non-acoustic features to generate a device feature vector;
[0244] Encoding the user operation behavior sequence in the non-acoustic features to generate a behavior feature vector;
[0245] Encoding the environmental sensor data in the non-acoustic features to generate an environmental feature vector;
[0246] The standardized acoustic feature vector, the semantic vector, the device feature vector, the behavior feature vector, and the environment feature vector are dimensionally aligned and concatenated to generate a multidimensional feature vector.
[0247] In one embodiment, the abnormality determination module 70 is specifically configured to:
[0248] Inputting the multidimensional feature vector into a feature extraction layer of the audio detection model;
[0249] Performing multimodal feature fusion at the feature extraction layer to generate a fused feature representation;
[0250] Processing the fused feature representation through the time series analysis layer of the audio detection model to generate a time series feature code;
[0251] determining, at a classification output layer of the audio detection model, an anomaly confidence score based on the temporal feature encoding;
[0252] The anomaly confidence score is converted into an anomaly indicator.
[0253] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an audio signal authenticity verification method.
[0254] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side method for verifying the authenticity of an audio signal
[0255] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0256] Collect original audio and corresponding text to generate an original audio text dataset;
[0257] Generate a set of adversarial samples based on the original audio text dataset through a generative adversarial network;
[0258] Inputting the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model;
[0259] Acquiring an audio signal to be detected and extracting acoustic features of the audio signal to be detected;
[0260] Acquiring non-acoustic features associated with the audio signal to be detected;
[0261] Constructing a multidimensional feature vector according to the acoustic features and the non-acoustic features;
[0262] Inputting the multidimensional feature vector into an audio detection model to generate an anomaly indicator;
[0263] Based on the abnormal indicators, a graded response operation is performed.
[0264] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0265] Collect original audio and corresponding text to generate an original audio text dataset;
[0266] Generate a set of adversarial samples based on the original audio text dataset through a generative adversarial network;
[0267] Inputting the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model;
[0268] Acquiring an audio signal to be detected and extracting acoustic features of the audio signal to be detected;
[0269] Acquiring non-acoustic features associated with the audio signal to be detected;
[0270] Constructing a multidimensional feature vector according to the acoustic features and the non-acoustic features;
[0271] Inputting the multidimensional feature vector into an audio detection model to generate an anomaly indicator;
[0272] Based on the abnormal indicators, a graded response operation is performed.
[0273] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0274] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0275] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0276] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for verifying the authenticity of an audio signal, characterized in that: The following steps are involved: Collect original audio and corresponding text to generate an original audio text dataset; Generate a set of adversarial samples based on the original audio text dataset through a generative adversarial network; Inputting the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model; Acquiring an audio signal to be detected and extracting acoustic features of the audio signal to be detected; Acquiring non-acoustic features associated with the audio signal to be detected; Constructing a multidimensional feature vector according to the acoustic features and the non-acoustic features; Inputting the multidimensional feature vector into an audio detection model to generate an anomaly indicator; Based on the abnormal indicators, a graded response operation is performed.
2. The method for verifying the authenticity of an audio signal according to claim 1, wherein: Generate an adversarial sample set based on the original audio text dataset through a generative adversarial network, including: Build a generative adversarial network; Simulating a noise injection attack in the adversarial generative network, adding noise perturbation to the audio signal of the original audio text dataset, and generating a noise-perturbed speech sample; Simulating a variable speed perturbation attack in the generative adversarial network, performing variable speed perturbation processing on the audio signal in the original audio text dataset, and generating a variable speed perturbation speech sample; Simulating a frequency domain offset attack in the adversarial generative network, performing frequency domain offset processing on the audio signal in the original audio text dataset, and generating a frequency domain perturbed speech sample; The noise-perturbed speech samples, the speed-varied perturbed speech samples, and the frequency-domain perturbed speech samples are integrated to generate an adversarial sample set.
3. The method for verifying the authenticity of an audio signal according to claim 1, wherein: The original audio text dataset and the adversarial sample set are input into the audio detection model for joint training to obtain an adversarially trained audio detection model, including: Set the initial network parameters of the audio detection model; Merging the original audio text dataset and the adversarial sample set to generate a joint training dataset; Applying a gradient alignment module to adjust the weights of adversarial examples in the joint training dataset; Optimizing the network parameters of the audio detection model using the weighted joint training dataset to generate an optimized audio detection model; Test the optimized audio detection model on the dialect dataset to generate performance verification results; When the performance verification result reaches a preset performance threshold, the optimized audio detection model is output as the adversarially trained audio detection model.
4. The method for verifying the authenticity of an audio signal according to claim 1, wherein: Acquiring an audio signal to be detected and extracting acoustic features of the audio signal to be detected, including: receiving an audio input signal to be detected; Performing noise reduction processing on the audio input signal to generate a noise-reduced audio signal; Performing frame processing on the noise-reduced audio signal to generate an audio frame sequence; Extracting fundamental frequency envelope features of the audio frame sequence to generate a fundamental frequency envelope feature sequence; The fundamental frequency envelope feature sequences are combined into an acoustic feature vector.
5. The method for verifying the authenticity of an audio signal according to claim 1, wherein: Acquiring non-acoustic features associated with the audio signal to be detected, including: Identifying text content corresponding to the audio signal to be detected and generating a text transcription; Extracting semantic features of the text transcription to generate text semantic features; Collecting information about a device generating the audio signal to be detected; Analyzing a user interaction behavior pattern associated with the audio signal to be detected to generate a user operation behavior sequence; Acquiring environmental sensor data when the audio signal to be detected is generated; The text semantic features, generation device information, user operation behavior sequence and environmental sensor data are combined to generate a non-acoustic feature set.
6. The method for verifying the authenticity of an audio signal according to claim 1, wherein: Constructing a multidimensional feature vector according to the acoustic features and the non-acoustic features, including: performing vector normalization processing on the acoustic features to generate a standardized acoustic feature vector; Performing embedding mapping processing on the text semantic features in the non-acoustic features to generate a semantic vector; Performing structured mapping processing on the generating device information in the non-acoustic features to generate a device feature vector; Encoding the user operation behavior sequence in the non-acoustic features to generate a behavior feature vector; Encoding the environmental sensor data in the non-acoustic features to generate an environmental feature vector; The standardized acoustic feature vector, the semantic vector, the device feature vector, the behavior feature vector, and the environment feature vector are dimensionally aligned and concatenated to generate a multidimensional feature vector.
7. The method for verifying the authenticity of an audio signal according to claim 1, wherein: The multi-dimensional feature vector is input into the audio detection model to generate anomaly indicators, including: Inputting the multidimensional feature vector into a feature extraction layer of the audio detection model; Performing multimodal feature fusion at the feature extraction layer to generate a fused feature representation; Processing the fused feature representation through the time series analysis layer of the audio detection model to generate a time series feature code; determining, at a classification output layer of the audio detection model, an anomaly confidence score based on the temporal feature encoding; The anomaly confidence score is converted into an anomaly indicator.
8. An audio signal authenticity verification device, characterized in that: The audio signal authenticity verification device comprises: The data acquisition module is used to collect original audio and corresponding text to generate an original audio text dataset; An adversarial sample generation module, configured to generate an adversarial sample set based on the original audio text dataset through a generative adversarial network; An adversarial training module is used to input the original audio text dataset and the adversarial sample set into the audio detection model for joint training to obtain an adversarially trained audio detection model; An acoustic feature extraction module, configured to obtain an audio signal to be detected and extract acoustic features of the audio signal to be detected; a non-acoustic feature extraction module, configured to obtain non-acoustic features associated with the audio signal to be detected; A feature fusion module, configured to construct a multi-dimensional feature vector based on the acoustic features and the non-acoustic features; an abnormality identification module, configured to input the multidimensional feature vector into an audio detection model to generate an abnormality indicator; The response control module is used to perform a hierarchical response operation based on the abnormal indicators.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an audio signal authenticity verification program stored in the memory and executable on the processor. When the audio signal authenticity verification program is executed by the processor, the steps of the audio signal authenticity verification method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores an audio signal authenticity verification program, which, when executed by a processor, implements the steps of the audio signal authenticity verification method according to any one of claims 1 to 7.
Citation Information
Cited By
Block chain evidence storage method and system based on audio authenticity identification technology
CN121191538A
Wind power plant data source authentication method and device based on SVM-UGR and computer equipment
CN121210927A
Wind farm data source authentication method, device, and computer equipment based on SVM-UGR
CN121210927B
Neural network-based textile industry broken yarn identification method and system
CN121234141A
A neural network-based yarn breakage recognition method and system for the textile industry
CN121234141B