Speaker authentication and deep forgery detection cascade voice privacy protection device
By integrating the ASV and ADD module cascade structure into the voice front-end device, the device-level lack of voice privacy protection in the existing technology is solved, the privacy processing of voice data in the collection stage is realized, the practicality and deployment efficiency of the voice security system are improved, and counterfeiters are prevented from impersonating users. It is suitable for intelligent voice scenarios.
Patent Information
- Application Number
- CN202510936407.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
Existing voice privacy protection technologies lack standardized implementation at the physical device level, making them difficult to deploy in edge devices. Furthermore, the algorithm processing path can be easily bypassed, and they are unable to provide a reliable and tamper-proof protection process. They lack flexibility and adaptability, and are unable to distinguish the privacy protection needs of different users, resulting in insufficient protection or excessive processing.
It adopts the cascade structure of ASV and ADD modules and is integrated into the voice front-end device. Through the microphone, ASV module, ADD module, control unit and audio output module, it realizes identity authentication and voice authenticity judgment, and builds a physical voice privacy protection system, which is suitable for intelligent voice scenarios.
It realizes privacy processing in the voice data collection stage, reduces the risk of privacy leakage, improves the practicality and deployment efficiency of the voice security system, provides active authentication and passive detection, prevents counterfeiters from impersonating users, and realizes double-layer judgment of identity comparison. It is suitable for scenarios such as voice interaction and voice payment.
Smart Images

Figure CN120808793A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech signal processing and privacy protection, in particular to a speaker authentication and deep spoofing detection cascaded speech privacy protection device. BACKGROUND
[0002] With the rapid popularization of speech-related applications such as voice assistants, remote meetings, voice search, intelligent customer service, etc., the collection and transmission of speech data in daily life and work have become extremely frequent. At the same time, the speaker identity information, biological feature information (such as voiceprint features, gender, age, emotional state, etc.) carried in the speech have gradually become sensitive information that can be identified and misused, bringing serious challenges to user privacy protection.
[0003] Currently, most of the research on speech privacy protection is focused on the software algorithm level, such as using voice changing technology, anonymous recoding, speech disguise, etc. to disturb or replace the speech signal to cover the user's real identity. However, these methods mostly run in the form of software in the cloud or terminal operating system, which has the following significant shortcomings:
[0004] (1) Lack of standardized implementation at the physical device level, difficult to directly embed in intelligent hardware;
[0005] (2) Difficult to deploy in edge devices or resource-limited scenarios (such as smart microphones, vehicle-mounted systems);
[0006] (3) Algorithm processing path is easy to bypass, cannot provide reliable and tamper-proof protection process;
[0007] (4) Usually a "one-size-fits-all" protection method, lack of flexibility and adaptability.
[0008] Necessity and advantages of ASV+ADD cascaded structure:
[0009]
[0010] In addition, there is currently no commercial system that can integrate ASV (Automatic Speaker Verification) module and ADD (Audio Deep Spoofing Monitoring) module at the physical device level, thus realizing the dynamic cascaded privacy protection mechanism of "first identifying user identity, then processing speech anonymously as needed". This lack causes the existing system to be unable to distinguish the privacy protection needs of different users, resulting in insufficient protection or excessive processing, seriously affecting the use experience and system controllability.
[0011] Therefore, it is urgent to develop a modular hardware device with simple structure, complete function and embeddable in a voice front-end device, integrate an automatic speaker verification (ASV) module and an audio deepfake detection (ADD) module in an integrated package in a cascade structure, ensure privacy processing of voice data in the collection stage, greatly reduce the risk of voice privacy leakage, and improve the practicability and deployment efficiency of the voice security system. SUMMARY
[0012] The present application aims to provide a speaker authentication and deepfake detection cascade voice privacy protection device with clear structure, independent modules and perfect function. The device embeds an automatic speaker verification (ASV) module and an audio deepfake detection (ADD) module in a cascade structure in a voice processing path, completes user identity confirmation and voice authenticity judgment at the voice input end, thereby constructing a physical voice privacy protection system with "recognition + anti-fake" capability, and is suitable for intelligent voice scenarios such as voice interaction, voice payment, voice communication and other scenarios requiring identity verification.
[0013] The present application provides a speaker authentication and deepfake detection cascade voice privacy protection device, which comprises a microphone, an ASV module, an ADD module, a control unit and an audio output module.
[0014] The input end of the microphone receives a voice input signal, and the output end is connected to the input end of the ASV module. The output end of the ASV module is connected to the input end of the ADD module and the control unit. The output end of the ADD module is connected to the input end of the control unit. The output end of the control unit is connected to the input end of the audio output module. The voice input signal is first collected by the microphone and then input to the ASV module for identity verification, and the voice data is input to the ADD module for fake detection. The outputs of the two are transmitted to the control unit, which determines whether to allow the audio signal output according to a preset logic.
[0015] The ASV module comprises a voice input channel, a feature extraction module, an embedding model module, a similarity matching module and an ASV output module. The input end of the voice input channel is connected to the output end of the microphone. The output end of the voice input channel is connected to the input end of the feature extraction module. The output end of the feature extraction module is connected to the input end of the embedding model module. The output end of the embedding model module is connected to the input end of the similarity matching module. The output end of the similarity matching module is connected to the input end of the ASV output module.
[0016] The voice input channel receives I2S input from the microphone, collects and converts sound signals in the environment into processable digital signals, i.e. collects speaker voice using the microphone, captures analog signals in the real world, and converts analog voice signals into high-fidelity digital audio signals through an I2S (Inter-IC Sound) interface.
[0017] The feature extraction module converts the time domain signal into a Mel spectrum, that is, converts a high-fidelity digital audio signal from a speech input channel into a frequency domain feature suitable for speech recognition or speaker recognition; by using a short-time Fourier transform (STFT) to frame and analyze the frequency components of the time domain signal, a Mel filter bank is further used to generate a Mel spectrum diagram, so that it is closer to the human ear hearing perception characteristics;
[0018] The embedding model module is based on an ECAPA-TDNN neural network, which maps the input frequency domain features into low-dimensional speaker embedding vectors through a deep neural network. ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation Time Delay Neural Network) is an advanced speaker recognition network with good timing modeling and attention mechanism. It accepts Mel spectrum input, automatically learns deep voice features to distinguish different speakers, and finally outputs a fixed-length vector (usually 192 or 256 dimensions) representing the identity features of the current speaker, i.e. a voiceprint embedding;
[0019] The similarity matching module is used for comparison with the registered voiceprint template, compares the current speaker embedding with the registered multiple voiceprint templates in the database, calculates the similarity score, and can use comparison methods such as cosine similarity, Euclidean distance, etc. The distance between the current embedding vector and each registered embedding is compared; the smaller the distance or the higher the similarity, the more likely it is that the two voiceprints come from the same person; a matching threshold can be set to determine whether the recognition result is reliable, or output multiple possible candidates and their scores;
[0020] The ASV output module is used to output the identity ID and confidence information; according to the matching result, the final identity recognition result and the confidence are output. The output includes: the recognized identity ID (or user label), the similarity score (indicating the confidence); if the matching score is lower than the threshold, it can be determined as "unknown speaker" or rejection;
[0021] The whole process of the ASV module is as follows:
[0022] (1) Speech acquisition stage
[0023] The user speaks through the microphone, and the microphone module converts the speech signal into an analog electrical signal and digitizes it through an I2S interface to form a processable speech data stream;
[0024] (2) Feature extraction stage
[0025] The digital speech signal is input into the feature extraction module, which performs frame division, STFT, Mel filtering, etc. to extract acoustic features such as Mel spectrum diagrams representing speech;
[0026] (3) Embedding modeling stage
[0027] These spectral features are sent to the ECAPA-TDNN embedding model to generate a fixed-length, low-dimensional "speaker embedding vector" that highly condenses the speaker's voice features;
[0028] (4) Similarity matching stage
[0029] The embedding vector is matched with the registered voiceprint template library to calculate the similarity score between the current embedding and each template;
[0030] (5) Identity confirmation and output stage
[0031] The template with the highest matching score and exceeding the preset threshold is determined as the user, and the ASV output module returns the recognized identity ID and its matching confidence for subsequent system use;
[0032] The ADD module structure includes an acoustic feature extractor, a spectral and phase graph analysis module, a deep spoofing detection subnetwork module, and a true or false judgment output module. The input end of the acoustic feature extractor is connected to the ASV output module, the output end of the acoustic feature extractor is connected to the input end of the spectral and phase graph analysis module, the output end of the spectral and phase graph analysis module is connected to the input end of the deep spoofing detection subnetwork module, and the output end of the deep spoofing detection subnetwork module is connected to the input end of the true or false judgment output module.
[0033] The acoustic feature extractor is used for shared or independent extractors with the ASV module, responsible for converting the original speech signal into acoustic features suitable for downstream tasks (including ASV and deep spoofing detection);
[0034] The spectral and phase graph analysis module is used to identify the synthesis traces in the speech, i.e., to further analyze the amplitude spectrum and phase spectrum obtained from the acoustic feature extractor to find abnormal traces of deep synthesis in the speech. Multiple channels are constructed to process the amplitude graph and phase graph, and sometimes high-order statistical features are extracted as spoofing clues;
[0035] The deep spoofing detection subnetwork module adopts a convolutional neural network (CNN) structure, responsible for modeling the spectral and phase features to determine whether the speech is spoofed. It is composed of multiple convolutional layers, batch normalization, activation functions, and pooling operations, which can automatically extract complex spatial and frequency features to distinguish the subtle differences between real and synthesized speech, and output a high-dimensional embedding or directly a binary classification probability value (spoofed / real);
[0036] The true or false judgment output module is used to output the true or false judgment flag; it receives the output information from the deep spoofing detection subnetwork and finally outputs the true or false judgment result;
[0037] The whole process of the ADD module is shown as follows:
[0038] (1) Voice signal input
[0039] The user voice is first sent to the acoustic feature extractor to generate unified spectral and phase features;
[0040] (2) Spectrum and phase diagram analysis
[0041] The amplitude spectrum and phase spectrum generated by the extractor are sent to the spectrum and phase diagram analysis module; potential abnormal features for forgery detection are extracted through Fourier analysis, statistical modeling, etc., including spectral discontinuity, phase anomaly, etc.
[0042] (3) Forgery detection subnetwork processing
[0043] The analyzed features are input into the CNN structure of the forgery detection subnetwork; the network automatically learns how to distinguish between real and synthetic speech, and outputs an intermediate representation or probability indicating "whether it is forged";
[0044] (4) True or false judgment output
[0045] The final result is post-processed by the output module to output the true or false label and the confidence score; the system can set a threshold based on the confidence to determine: high confidence forgery is rejected, low confidence can be manually reviewed, etc.
[0046] The control unit receives the identity information of the ASV module output module; receives the forgery judgment result of the ADD output module; determines whether to output the audio signal according to the strategy; controls the audio output module to start or block.
[0047] In the present application, the microphone, the ASV module, the ADD module, the control unit and the audio output module are placed in the shell, the shell adopts a layered structure design, the modules adopt a horizontal series connection and vertical control layout, the microphone, the ASV module and the ADD module are placed at the top of the shell, and the control unit and the audio output module are placed in the middle part.
[0048] In the present application, the bottom of the shell is provided with a power supply and a heat dissipation module, and the ASV module, the ADD module, the control unit and the audio output module are respectively connected to the power supply.
[0049] In the present application, the control unit adopts an STM32F4 series microcontroller, which has strong I / O interface capability and low delay response.
[0050] In the present application, the audio output module supports a USB audio port, a 3.5mm analog audio port, or is directly connected to a terminal device (such as a sound box or a recording device).
[0051] In the application, the audio output module adopts one or several of the control LED, alarm or Bluetooth / Wi-Fi communication module.
[0052] In the application, the shell is made of aluminum alloy or plastic and copper heat dissipation fin material, and the shell has electromagnetic shielding on the outside, which is suitable for high interference industrial or terminal device environment.
[0053] In the application, the ASV module can be integrated in the ESP32-S3 or STM32F7 series chip, and has the characteristics of low power consumption and neural network acceleration.
[0054] The application has the following advantages:
[0055] The application adopts ASV and ADD module cascade, realizes "identity + authenticity" double judgment at the physical device level for the first time, improves the security through signal fusion strategy for joint decision-making, has clear structure and reasonable function partition, can realize modular deployment and replacement, and is suitable for various voice collection terminal devices, such as voice assistants, intelligent door locks, voice payment terminals and the like.
[0056] This combination can realize:
[0057] 1. Active authentication (ASV) + passive detection (ADD);
[0058] 2. Preventing counterfeiters from using deep synthesis voice to "pretend to be target users";
[0059] 3. Real-time filtering and identity comparison double-layer judgment on any voice request entering the system;
[0060] 4. Realize the complete closed loop of audio identity security protection of intelligent devices on the edge side. BRIEF DESCRIPTION OF DRAWINGS
[0061] Figure 1 It is a structure diagram of the voice privacy protection device of ASV and ADD cascade.
[0062] Figure 2 It is a structure flow chart of the ASV module, which shows the processing flow from voice input to identity judgment in the ASV module, including feature extraction, embedding model, similarity matching and identity output submodules.
[0063] Figure 3 It is a structure diagram of the ADD module, which shows the order of each substructure in the ADD module, including acoustic feature extraction, spectrum analysis, forgery detection subnetwork and authenticity judgment module.
[0064] Figure 4This is a connection diagram between the control unit and the audio output module, showing how the control unit receives the ASV output result (ID / confidence) and the forgery judgment result of the ADD module, and controls the on and off of the output channel.
[0065] Figure 5 This is a cross-sectional view of the structure of the present invention; it shows the physical layout structure and data flow of each core module (microphone, ASV module, ADD module, control unit, output module, power supply and heat dissipation channel) in the device inside the shell.
[0066] The numbers in the figure are: 1 is the microphone, 2 is the ASV module (automatic speaker verification module), 3 is the ADD module (audio forgery detection module), 4 is the control unit, 5 is the audio output module, 6 is the voice input channel, 7 is the feature extraction layer module, 8 is the embedding model module, 9 is the similarity matching module, 10 is the ASV output module, 11 is the acoustic feature extractor, 12 is the spectrum and phase analysis module, 13 is the deep forgery detection subnet module, 14 is the ADD output module, 15 is the power supply and heat dissipation channel, and 16 is the shell. DETAILED DESCRIPTION
[0067] In order to further understand the present invention, a preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings. However, those skilled in the art should understand that these drawings and descriptions are only for illustration and do not limit the scope of protection of the present invention.
[0068] Example 1:
[0069] like Figure 1 As shown, the present invention provides a device comprising: a microphone 1, an automatic speaker verification module (ASV module) 2, an automatic speaker verification module (ASV module) 3, a control unit 4, and an audio output module 5. The microphone 1 is connected to the automatic speaker verification module (ASV module) 2 via an I2S or USB audio interface, the output of the automatic speaker verification module (ASV module) 2 is connected to an audio forgery detection module (ADD module) 3, the judgment results of the automatic speaker verification module (ASV module) 2 and the audio forgery detection module (ADD module) 3 are respectively input to the control unit 4, and the control unit 4 outputs a control signal to the audio output module 5 to control whether it transmits an audio signal.
[0070] The specific implementation process is as follows:
[0071] (1) Microphone 1 collects voice signals and inputs them into ASV module 2 via the audio bus;
[0072] (2) Inside ASV module 2 Figure 2As shown, the feature extraction layer module 7 is used to convert the original speech signal into an acoustic representation representing the speaker characteristics, such as a mel-spectrogram, to provide input for subsequent modeling; the embedding model module 8 (such as ECAPA-TDNN) receives the feature map and extracts a fixed-length speaker embedding vector to express the identity information contained in the speech; the similarity matching module 9 compares the embedding vector with the voiceprint template in the registration database to calculate the similarity score; the ASV output module 10 outputs the recognized identity ID and its corresponding confidence according to the comparison result, which is used for identity verification or access control and other practical applications.
[0073] (3) At the same time, the speech signal is also input to the ADD module 3, as shown in Figure 3 The structure includes an acoustic feature extractor module 11 for extracting spectral and phase features for forgery detection from the speech signal; a spectral and phase analysis module 12 further processes these features to capture possible synthetic or forged traces in the speech; a deep forgery detection subnetwork module 13 models the analysis results based on a convolutional neural network to determine whether the speech is forged; and a true-false judgment output module 14 gives true or false labels and corresponding confidence according to the model output, which is used for authenticity verification of the speech content.
[0074] (4) The ASV module output result 10 and the ADD module output result 14 are transmitted to the control unit 4 as shown in Figure 2 The control unit is implemented using an STM32F4 series MCU;
[0075] (5) The control unit 4 determines whether the following conditions are met:
[0076] (5.1) The identity ID is an authorized user;
[0077] (5.2) The confidence is higher than a set threshold;
[0078] (5.3) The ADD determination is a real speech;
[0079] If the above three conditions are met, the audio output module 5 is turned on;
[0080] (6) The audio output module 5 includes a USB audio interface or a 3.5mm output port for transmitting audio to an external terminal;
[0081] (7) The entire device has a unified power supply and heat dissipation channel 15 at the bottom. The power supply module can output 3.3V / 5V voltage, and the air duct structure is located on both sides of the device, which has good heat dissipation performance.
[0082] In order to realize compact structure, PCB multilayer wiring board is used to connect between modules, the shell 16 is integrally formed by aluminum alloy, all signal transmission uses shielding wire or SPI / UART and the like, and the anti-interference performance is ensured. The inner bottom of the shell 16 is provided with a power supply and a heat dissipation module 15, the ASV module 2, the ADD module 3, the control unit 4 and the audio output module 5 are respectively connected with the power supply.
Claims
1. A cascaded voice privacy protection device for speaker authentication and deepfake detection, comprising a microphone, an ASV module, an ADD module, a control unit, and an audio output module, characterized in that: The microphone's input receives voice input signals, and its output is connected to the ASV module's input. The ASV module's output is connected to the inputs of the ADD module and the control unit, respectively. The ADD module's output is connected to the control unit's input, and the control unit's output is connected to the audio output module's input. The voice input signal is first collected by the microphone and then input to the ASV module for identity verification. At the same time, the voice data is input to the ADD module for forgery detection. The outputs of both are transmitted to the control unit, which decides whether to allow the audio signal output based on preset logic. The ASV module includes a speech input channel, a feature extraction module, an embedding model module, a similarity matching module and an ASV output module. The input end of the speech input channel is connected to the output end of the microphone, the output end of the speech input channel is connected to the input end of the feature extraction module, the output end of the feature extraction module is connected to the input end of the embedding model module, the output end of the embedding model module is connected to the input end of the similarity matching module, and the output end of the similarity matching module is connected to the input end of the ASV output module. The voice input channel receives I2S input from a microphone, collects sound signals from the environment, and converts them into processable digital signals. This means that the microphone collects the speaker's voice, captures real-world analog signals, and converts the analog voice signals into high-fidelity digital audio signals through the I2S (Inter-IC Sound) interface. The feature extraction module converts the time-domain signal into a Mel-spectrogram, essentially transforming the high-fidelity digital audio signal from the speech input channel into frequency-domain features suitable for speech recognition or speaker identification. It uses a short-time Fourier transform (STFT) to frame the time-domain signal and analyze its frequency components. It then uses a Mel filter bank to generate a Mel-spectrogram, which more closely matches the human auditory perception. The embedding model module is based on the ECAPA-TDNN neural network. It uses a deep neural network to map input frequency-domain features into low-dimensional speaker embedding vectors. ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation Time Delay Neural Network) is an advanced speaker recognition network with advanced temporal modeling and attention mechanisms. It accepts Mel-spectrum input, automatically learns the deep speech features that distinguish different speakers, and ultimately outputs a fixed-length vector (typically 192 or 256 dimensions) representing the current speaker's identity, known as the voiceprint embedding. The similarity matching module is used to compare with the registered voiceprint template, compare the current speaker embedding with multiple voiceprint templates registered in the database, and calculate the similarity score. It can use comparison methods (such as cosine similarity, Euclidean distance, etc.) to compare the distance between the current embedding vector and each registered embedding; the smaller the distance or the higher the similarity, the more likely the two voiceprints are from the same person; the matching threshold can be set to determine whether the recognition result is reliable, or multiple possible candidates and their scores can be output; The ASV output module is used to output identity ID and confidence information; The final identification result and confidence level are output based on the matching results. The output includes: the identified identity ID (or user tag), similarity score (indicating credibility); if the matching score is lower than the threshold, it can be judged as "unknown speaker" or rejected; The entire process of the ASV module is as follows: (1) Voice collection stage The user speaks through the microphone, and the microphone module converts the voice signal into an analog electrical signal and digitizes it through the I2S interface to form a processable voice data stream; (2) Feature extraction stage The digital speech signal is input into the feature extraction module, which performs frame segmentation, STFT, and Mel filtering to extract acoustic features such as the Mel-spectrogram representing the speech. (3) Embedded modeling stage These spectral features are fed into the ECAPA-TDNN embedding model to generate a fixed-length, low-dimensional "speaker embedding vector" that highly condenses the speaker's speech characteristics. (4) Similarity matching stage This embedding vector is matched with the registered voiceprint template library for similarity, and the similarity score between the current embedding and each template is calculated; (5) Identity confirmation and output stage The template with the highest matching score that exceeds the preset threshold is determined to be the user. The ASV output module returns the recognized identity ID and its matching confidence for subsequent system use. The ADD module structure includes an acoustic feature extractor, a spectrum and phase image analysis module, a deep fake detection subnet module, and an authenticity determination output module. The input end of the acoustic feature extractor is connected to the ASV output module, the output end of the acoustic feature extractor is connected to the input end of the spectrum and phase image analysis module, the output end of the spectrum and phase image analysis module is connected to the input end of the deep fake detection subnet module, and the output end of the deep fake detection subnet module is connected to the input end of the authenticity determination output module. The acoustic feature extractor is used to share with the ASV module or as an independent extractor, responsible for converting the original speech signal into acoustic features suitable for downstream tasks (including ASV and deep fake detection); The spectrum and phase map analysis module is used to identify synthesis traces in speech. This module further analyzes the amplitude and phase spectra obtained from the acoustic feature extractor to identify any abnormal signs of deep synthesis. Multiple channels are constructed to process amplitude and phase maps separately, and sometimes high-order statistical features are extracted as clues to forgery. The deepfake detection subnet module uses a convolutional neural network (CNN) structure to model spectral and phase features to determine whether the speech is forged. Composed of multiple convolutional layers, batch normalization, activation functions, and pooling operations, it can automatically extract complex spatial and frequency features, discern subtle differences between real and synthesized speech, and output a high-dimensional embedding or a binary probability value (fake / real). The authenticity determination output module is used to output the authenticity determination flag; it receives the output information from the deep fake detection subnet and ultimately outputs the authenticity determination result; The entire process of the ADD module is as follows: (1) Voice signal input The user's speech is first fed into the acoustic feature extractor to generate unified spectrum and phase features; (2) Spectrum and phase diagram analysis The amplitude spectrum and phase spectrum generated by the extractor are fed into the spectrum and phase diagram analysis module. This module uses Fourier analysis and statistical modeling to extract potential anomaly features for counterfeit detection, such as spectrum discontinuities and phase anomalies. (3) Forgery detection subnet processing The analyzed features are input into the forgery detection subnetwork of the CNN structure; the network automatically learns how to distinguish between real and synthesized speech and outputs an intermediate representation or probability indicating whether it is forged. (4) Authenticity determination output The final result is post-processed by the output module, which outputs a true or false label and a confidence score. The system can set a threshold based on the confidence level to make judgments: high-confidence forgeries are rejected, low-confidence ones can be manually reviewed, etc. The control unit receives the identity information of the ASV module output module; receives the forgery judgment result of the ADD output module; judges whether to output the audio signal according to the strategy; and controls the audio output module to be turned on or off.
2. A speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1, characterized in that The microphone, ASV module, ADD module, control unit and audio output module are all placed inside the shell. The shell adopts a layered structure design, and the modules are arranged in horizontal series and vertically controlled. The microphone, ASV module and ADD module are all placed at the top of the shell, and the control unit and audio output module are placed in the middle.
3. A speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1, characterized in that A power supply and heat dissipation module is provided at the bottom of the shell, and the ASV module, ADD module, control unit and audio output module are respectively connected to the power supply.
4. A speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1, characterized in that The control unit adopts the STM32F4 series microcontroller, which has strong I / O interface capabilities and low latency response.
5. A speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1, characterized in that The audio output module supports USB audio port, 3.5mm analog audio port, or direct connection with terminal devices.
6. A speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1, characterized in that The audio output module adopts one or more of the control LED, alarm or Bluetooth / Wi-Fi communication module.
7. The speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1 is characterized in that The housing is made of aluminum alloy or plastic with copper heat sink material. The exterior of the housing is electromagnetically shielded and is suitable for use in high-interference industrial or terminal equipment environments.
8. The speaker authentication and deep fake detection cascade voice privacy protection device according to claim 1 is characterized in that The ASV module can be integrated into ESP32-S3 or STM32F7 series chips, and has features such as low power consumption and neural network acceleration.
Citation Information
Cited By
Fake voice detection method based on time-frequency fusion network
CN121884859A
A method for detecting counterfeit speech based on a time-frequency fusion network
CN121884859B