Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Speaker recognition system" patented technology

A speaker verification method based on SASFV aggregation model

ActiveCN120766685BSpeech analysisSpeaker recognition systemNetwork generation
The application discloses a speaker verification method based on a SASFV aggregation model, and relates to the field of speech recognition.The method extracts a log Mel spectrogram through short-time Fourier transform and Mel filtering, generates frame-level features by using an ERes2Net network, introduces a SASFV aggregation model to generate fixed-length speaker-level features in combination with a Fisher Vector variable, a self-attention mechanism and a statistical method, and finally determines the identity of a speaker by using a cosine distance.The application solves the problem that the prior art cannot effectively represent and aggregate features in a short speech task, and significantly improves the accuracy, robustness and performance of a speaker recognition system.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

A robust speaker recognition method based on spectrogram denoising and adversarial learning

ActiveCN116469394BSpeech analysisNeural architecturesData setSpeaker recognition system
The present invention provides a robust speaker recognition method based on spectrogram denoising and adversarial learning. First, a spectrogram dataset of clean speech and a noisy spectrogram dataset after the clean speech is noisy are collected; a U-net with a multi-level encoding and decoding structure is trained using a mean square error loss function to remove noise interference from the mel-spectrogram of the noisy speech signal to obtain an enhanced mel-spectrogram; a conditional generative adversarial network based on a time-delay neural network (TDNN-CGAN) is trained using a least squares loss function, a time-delay neural network (TDNN) is used as a generator in the TDNN-CGAN to extract deep features of the enhanced mel-spectrogram, and a multi-layer perceptron (MLP) is used as a discriminator in the TDNN-CGAN; finally, a speaker classifier is trained using cross-entropy loss to identify the speaker's identity, thereby realizing speaker recognition in a noisy environment. The deep features extracted from the noisy speech by the present invention are close to the deep features extracted from the clean speech, thereby improving the performance of the speaker recognition system in a noisy environment.
Owner:NANCHANG UNIV

Speaker recognition model frequency modulation trigger injection method

InactiveCN121459823ASpeech analysisFrequency spectrumSpeaker recognition system
The invention discloses a speaker recognition model frequency modulation trigger injection method, particularly relates to the technical field of voice signal processing and artificial intelligence safety, and is used for solving the problems that an existing speaker recognition model back door injection method is insufficient in concealment, poor in black box environment adaptability, prone to distortion of trigger characteristics in the physical transmission process and poor in safety. And mainstream defense strategies such as model fine tuning, pruning and spectrum detection are difficult to resist. The method comprises the following steps: constructing a piecewise linear periodic low-frequency modulation curve for original voice, generating a smooth phase track according to a frequency and phase relationship, constructing a sinusoidal modulator to generate a hidden sample, mixing normal samples in proportion and uniformly marking, training a model by a black box, testing reproduction modulation and inputting the model; according to the method, the defects of insufficient concealment, black box adaptability, physical robustness and defensive resistance of an existing method are overcome, high-concealment, high-robustness and wide-application backdoor injection is realized, and reliable support is provided for security evaluation of a speaker recognition system.
Owner:SUZHOU VOCATIONAL INSTITUTE OF INDUSTRIAL TECHNOLOGY

A speech adversarial defense method and system for a speaker recognition system

ActiveCN119943057BInternal combustion piston enginesSpeech analysisSpeaker recognition systemEngineering
The application provides a speech confrontation defense method and system for a speaker recognition system, proposes a new type of confrontation purification framework SA-Net, and the key idea is to adopt a'subtraction first and then addition' strategy at the feature level. The subtraction step filters out non-robust features by analyzing the distribution of speaker features, thereby compressing the survival space of the confrontation noise. The addition step reconstructs the complete speech signal, so that the speaker recognition system can accurately identify without additional fine-tuning or retraining. The average defense accuracy of the application when resisting adaptive attacks on two open source SRSs reaches 87.8%, while maintaining a normal recognition accuracy of 98.5%, which is 29.3% and 2.8% higher than Parallel WaveGAN respectively. In addition, the application has strong defense ability and wide applicability, and can be used as a plug-and-play defense line for SRS in various deployments.
Owner:WUHAN UNIV

Speaker recognition method and system

The present invention relates to a speaker recognition method and system. The method comprises obtaining a trained recurrent generative adversarial network; obtaining real-time audio data; extracting frame-level Mel-spectrogram features of the real-time audio data, performing voice activity detection on the extracted frame-level Mel-spectrogram features of the real-time audio data, and determining frame-level Mel-spectrogram features containing speech in the real-time audio data; mapping the frame-level Mel-spectrogram features containing speech in the real-time audio data and the frame-level Mel-spectrogram features containing speech in the registered audio data using a common speech frame-level Mel-spectrogram feature generator in the trained recurrent generative adversarial network, and determining a first output result and a second output result; and determining a recognition result based on the first and second output results and a neural network model of the speaker recognition system. The present invention can improve the recognition effect when recognizing a speaker using common pronunciation.
Owner:NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF

Real-time speaker identification system utilizing meta learning to process short utterances in an open-set environment

ActiveUS12406673B2Speech analysisNatural language processingSpeaker recognition system
The invention is a speaker identification system, which is provided to train a speaker model based on a meta-learning approach. Through the training of a plurality of episodes, the speaker model is updated by backpropagating the gradients of a composite objective function comprised of two loss functions and each episode consists of a support set of long utterances and a query set of short utterances. With this single speaker model, the invention converts an input utterance into a speaker embedding vector. This enables the speaker identification system to identify different enrolled speakers solely through the comparison of speaker embedding vectors, effectively blocking spoofing attacks and impostor intrusion. Consequently, the invention is characterized by its lightweight nature, real-time response, suitability for short utterances and open set environments, and can be implemented using low-cost embedded hardware.
Owner:NATIONAL YUNLIN UNIVERSITY OF SCIENCE AND TECHNOLOGY

A text-independent speaker verification method based on identity information and semantic information disentanglement

ActiveCN116543775BDisentanglement implementationImprove generalization abilityInternal combustion piston enginesSpeech recognitionData setSpeaker recognition system
The present application relates to the technical field of voiceprint recognition, and particularly relates to a text-independent speaker verification method based on identity information and semantic information disentanglement. The technical scheme of the present application is completed by building a neural network, training a model, and testing a result. sch The semantic content is represented by a fixed-length feature f spk The correlation between the voiceprint feature f sch and the semantic feature f sch is calculated using a disentanglement method such as mutual information to achieve disentanglement between the voiceprint feature and the semantic feature, obtain more accurate and robust speaker voiceprint features, and solve the problem of poor system generalization performance caused by interference of semantic content factors in the data set extracted by the speaker recognition system.
Owner:SHANXI UNIV

Anonymous speaker identification method of double-branch attention mechanism feature fusion module

The invention discloses an anonymous speaker identification method based on a double-branch attention mechanism, and the method comprises the steps: firstly, extracting a 80-dimensional Mel feature vector of an audio; then, after data preprocessing, introducing a double-attention mechanism module to perform attention enhancement on the Mel feature vector, and highlighting the key frequency band or feature dimension for speaker identity recognition in the Mel feature vector; weighted features obtained by an SE branch and an ECA branch are fused to form a new feature vector, and fine-grained features, captured by the ECA branch, of adjacent channels are associated with global features of the SE branch, so that a speaker recognition system can decouple more speaker identity information. And finally, evaluating the performance of the speaker recognition system by using the anonymized speech, and verifying the effectiveness of the anonymized speaker recognition method based on the double-branch attention mechanism. According to the method, feature fusion is carried out by using the double attention mechanism modules, so that the performance of a speaker recognition system is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Speaker confirmation method based on SASFV aggregation model

ActiveCN120766685ASpeech analysisSpeaker recognition systemNetwork generation
The invention discloses a speaker confirmation method based on an SASFV aggregation model, and relates to the field of speech recognition, and the method comprises the steps: extracting a logarithmic Mel spectrogram through short-time Fourier transform and Mel filtering, generating a frame-level feature through an ERes2Net network, and carrying out the recognition of the frame-level feature; and an SASFV aggregation model is introduced, and a Fisher Vector variable, a self-attention mechanism and a statistical method are combined to generate speaker-level features with a fixed length, and finally, the identity of a speaker is judged through a cosine distance. According to the invention, the problem that features cannot be effectively represented and aggregated in a short voice task in the prior art is solved, and the accuracy, robustness and performance of a speaker recognition system are remarkably improved.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

A defense training method against adversarial samples of a speaker recognition system

PendingCN122337210ASpeaker recognition systemEngineering
This invention relates to the field of speech adversarial example defense, specifically a novel feature fusion defense training method based on speaker facial features, comprising: (1) constructing an audio-video dataset; (2) concatenating audio-video vectors to build a feature fusion module with a multi-head attention mechanism; (3) training the speaker recognition system by retaining only the audio feature portion of the cross-modal fusion features; and (4) generating adversarial examples using FGSM, PGD, CW, FakeBob, SirenAttack, and Kenansville for defense performance testing. This invention introduces speaker facial feature identity consistency information during speaker identity registration to train the speaker recognition system. In the case of unknown attack methods, cross-modal information enhances the robustness and accuracy of the speaker recognition system, achieving defense against adversarial examples of unknown attack methods.
Owner:BEIJING UNIV OF POSTS & TELECOMM +1