Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Speaker verification" patented technology

Ensemble learning (EL)-based speaker verification method

ActiveUS12633292B2Ensemble learningSpeech analysisSpeaker verificationData acquisition
Provided is an ensemble learning (EL)-based speaker verification method. The method includes: data acquisition and preprocessing; selecting and training a group of basic models, and optimizing model parameters; performing similarity scoring on an acquired pair of speaker feature embedding via the group of basic models; constructing a detection cost function (DCF); generating a weight, and performing weighted fusion on scoring results of the group of basic models based on the weight, to obtain a final ensemble model for speaker verification; based on a near-speaking or far-speaking test scenario of a voice sample, the scenario is distinguished and input into the ensemble model, to obtain a final similarity score after weighted fusion; and determining, based on a threshold, whether there is a same speaker, where if the similarity score is greater than the threshold, it is determined that there is a same speaker.
Owner:HANGZHOU DIANZI UNIV

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Quantization method of speaker verification model, electronic device and storage medium

The application discloses a speaker verification model quantization method, an electronic device and a storage medium, wherein the speaker verification model quantization method comprises the following steps: acquiring real-value weights of all layers of a speaker verification model; and mapping the real-value weights of all layers to a fixed integer set; or dynamically determining binary weights corresponding to the real-value weights of each layer to better match the real-value weight distribution. The method of the application embodiment proposes two new quantization strategies, namely static quantization and adaptive quantization. Furthermore, for static quantization, the application embodiment proposes a weight regularization technique to maintain maximum information entropy and reduce information loss. Furthermore, the application embodiment also proposes an adaptive quantization scheme, which can dynamically determine the best binary of each layer to achieve better alignment with the real-value weight distribution.
Owner:AISPEECH CO LTD

Speaker anonymization-oriented stylegan2-based F0 cloning method

PendingCN121506160ASpeech analysisAlgorithmSpeaker verification
The invention relates to the technical field of voice processing, in particular to a speaker anonymization-oriented stylegan2-based F0 cloning method, which mainly realizes source speaker anonymization protection and comprises the following steps of: preprocessing an original fundamental frequency sequence, and converting a random noise vector into a representation vector of a style space; and a StyleGAN2 architecture is introduced to carry out layer-by-layer feature processing and fusion. A probabilistic fusion strategy is adopted, content features and constant features are added, then layer-by-layer stylization processing is carried out, feature modulation is carried out on each layer by using a style vector, the time sequence resolution is gradually recovered, and the anonymization target of the identity features of the speaker is effectively changed. And finally, combining the generative fundamental frequency (F0), the text information and the anonymized identity features to synthesize anonymized voice so as to hide the identity information of the speaker. And the effectiveness of the speaker anonymization-oriented StyleGAN2-based F0 cloning method is verified through an automatic speaker verification model and an automatic speech recognition system.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Consistency check for large language model continuous conversations

In addition to speaker verification, an LLM continuous conversation check employs a spoken speed check using a phonemes-based spoken speed calculation, and acoustic energy check, and a signal-to-noise estimation to determine whether first and second audio inputs include utterances forming a continuous conversation by the user with the LLM. Results from the various components of the continuous conversation check are fused based on automatically assigned weights in making the determination. Continuous conversation detection for LLMs is therefore more robust, particularly for a very short second utterance. Optionally a distance to microphone check or a semantic consistency check may also be employed.
Owner:SAMSUNG ELECTRONICS CO LTD

Speaker feature vector extraction method, speaker feature vector verification method, speaker feature vector extraction device, speaker feature vector verification device and speaker feature vector product

The invention provides a speaker feature vector extraction method, a speaker feature vector verification method, a speaker feature vector extraction device, a speaker feature vector verification device and a product, and the speaker feature vector extraction method comprises the steps: processing an obtained input voice signal to obtain a basic acoustic feature; processing the basic acoustic features through at least three speaker attention weight generation modules in sequence; the speaker attention weight module comprises a compressed excitation residual block and a speaker conditional attention module, and the output end of the compressed excitation residual block is connected with the input end of the speaker conditional attention module; the multi-level feature aggregation module outputs a spliced statistical feature vector according to the output features of each speaker conditional attention module; and the full connection layer outputs a speaker feature vector according to the spliced statistical feature vector. The speaker condition attention mechanism is introduced to improve the accuracy of the feature vector of the speaker, and the problem that the verification is not accurate enough in a complex scene such as relatively high speaker similarity is solved when the method is used for speaker verification.
Owner:SHENZHEN JUSHENG TECHNOLOGY CO LTD +1

Robust intelligent synthesized speech speaker verification model training method and system

ActiveCN116597843BInternal combustion piston enginesSpeech analysisSpeaker verificationNoise
The application provides a kind of robust intelligent synthetic speech speaker confirmation model training method, system, storage medium and electronic equipment, it is related to speech processing technical field.The robust intelligent synthetic speech speaker confirmation model provided by the present application is used to improve the speaker confirmation accuracy of intelligent synthetic speech under strong background noise condition, comprising speech enhancement network, feature extraction network and feature enhancement module.In the training phase of the model, the training set of noisy intelligent synthetic speech data set is transported into speech enhancement network and feature extraction network after STFT feature preprocessing and Fbank feature preprocessing, and the two networks are connected through feature enhancement module for joint training to obtain speaker embedding features with noise-robustness.In the test phase, the model is tested based on the test set of noisy intelligent synthetic speech data set;The optimal model selection is carried out in the mode of repeating the foregoing training and testing alternately until the training iteration number reaches the set maximum value.
Owner:HEFEI UNIV OF TECH

Method and System for Personalising Speaker Verification Models

ActiveGB2630805BPersonalizationSpeaker verification
A computer-implemented server-performed method for personalising a trained speaker verification machine learning model for voice authentication of specific users obtains a “positive” audio sample of t
Owner:SAMSUNG ELECTRONICS CO LTD

Method for Obtaining Enrolment Data for a Speaker Verification Model

ActiveGB2636095BSpeech analysisSpeaker verificationAudio frequency
A method for obtaining enrolment data for a speaker verification mode, comprises capturing audio samples of utterances spoken by a speaker, generating an embedding vector for some samples, identifying
Owner:SAMSUNG ELECTRONICS CO LTD

Speaker verification method and system based on dual-stream low-rank adaptive and adversarial decoupling

The application discloses a speaker verification method and system based on double-flow low-rank adaptive and anti-decoupling, comprising: based on a pre-trained speech network, a double-flow low-rank adaptive anti-decoupling network is constructed, and the original weight parameters of the pre-trained speech network are frozen; based on a language feature extraction branch and a speaker feature extraction branch, original speech data is subjected to feature extraction respectively to obtain language features and speaker features; the language features are input into the shared discriminator to perform language classification prediction to complete language boundary anchoring; after gradient inversion processing of the speaker features, the speaker features are input into the shared discriminator which has completed language boundary anchoring to perform anti-decoupling; identity recognition is performed based on the speaker features subjected to anti-decoupling constraint; corresponding training losses are calculated respectively, and iterative updating is performed based on the training losses. The application can improve the acceptance rate of cross-language speech of the same person and the rejection rate of the same language speech of different persons, and is suitable for high-precision speaker verification in a multi-language environment.
Owner:NANJING UNIV

A time-varying speaker template updating method based on risk gating

PendingCN122337207ASpeaker verificationEngineering
This invention discloses a time-varying speaker template update method based on risk gating, belonging to the field of speaker verification and speech signal processing technology. The method first constructs an initial speaker template and anchor template from registered speech, then verifies the test speech. When the update conditions are met, an attractive update is performed first, then candidate negative samples are obtained. A risk gating quantity is constructed through uncertainty-driven or deterministic-driven gating branches to determine the repulsion strength. Projective repulsion update is then performed in the negative direction, followed by step size pruning and anchor point retraction. This method can reduce template drift caused by erroneous updates and improve the stability and discriminative power of time-varying speaker verification.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

A text-independent speaker verification method based on identity information and semantic information disentanglement

ActiveCN116543775BDisentanglement implementationImprove generalization abilityInternal combustion piston enginesSpeech recognitionData setSpeaker recognition system
The present application relates to the technical field of voiceprint recognition, and particularly relates to a text-independent speaker verification method based on identity information and semantic information disentanglement. The technical scheme of the present application is completed by building a neural network, training a model, and testing a result. sch The semantic content is represented by a fixed-length feature f spk The correlation between the voiceprint feature f sch and the semantic feature f sch is calculated using a disentanglement method such as mutual information to achieve disentanglement between the voiceprint feature and the semantic feature, obtain more accurate and robust speaker voiceprint features, and solve the problem of poor system generalization performance caused by interference of semantic content factors in the data set extracted by the speaker recognition system.
Owner:SHANXI UNIV

Far-field speaker verification method based on self-distillation pre-training and meta-learning fine-tuning

ActiveCN116863937BSpeech analysisPattern recognitionSpeaker verification
The application discloses a far-field speaker verification method based on self-distillation pre-training and meta-learning fine-tuning, and the process is as follows: log mel spectrum features are extracted from near-field training speech as the input of the network, and a self-distillation learning method is used to pre-train a deep neural network; then, log mel spectrum features of far-field training speech are used as the input of the network, and a meta-learning method is used to fine-tune the pre-trained network; then, log mel spectrum features of the registered speech are input into the pre-trained and fine-tuned network to obtain a transformed feature vector, and log mel spectrum features of the test speech are input into the fine-tuned and pre-trained network to obtain a transformed feature vector; finally, the distance between the transformed feature vector of the test speech and the transformed feature vector of the registered speech is calculated, and it is judged whether the two are from the same speaker. The method improves the network training efficiency and improves the speaker verification accuracy when the recording environment of the registered speech and the test speech is inconsistent.
Owner:SOUTH CHINA UNIV OF TECH

Voice analysis identity authentication method based on sound field reconstruction

ActiveCN116504251BEnsure consistencyNo additional costSpeech analysisHigh level techniquesSpeaker verificationVoice analysis
The application discloses a speech analysis identity authentication method based on sound field reconstruction, which comprises distance sensing, sound field reconstruction, sound field extraction, and model training and inference. The distance sensing is achieved by emitting a chirp signal through a loudspeaker and receiving the signal, and the distance from the user to the mobile phone is obtained by using a cross-correlation method. The sound field reconstruction is achieved by establishing a distance-dependent impulse response database, obtaining the corresponding transfer function according to the measured distance, and then reconstructing the sound field at the verification position to the sound field at the registration position. The sound field extraction is achieved by separately processing the signals of the two sound channels of the speech signal after the sound field reconstruction, and extracting the field texture. The model training and inference are achieved by using the registration field texture and the reconstructed field texture in the verification stage to construct a speech authentication model. The method can effectively solve the problem that the original sound field method is sensitive to distance when verifying the speaker, and the user does not need to maintain a fixed distance during verification as during registration.
Owner:ZHEJIANG UNIV

A method and device for verifying speaker identity in a far-field scene, and an electronic device

ActiveCN116343798BSpeech analysisManufacturing computing systemsSpeaker verificationEngineering
The present disclosure provides a method and device for verifying the identity of a speaker in a far-field scene, and an electronic device, wherein the method comprises: obtaining target identity information selected by a target user and a voice of the target user; dividing the voice of the target user into a plurality of target voice segments, and obtaining a plurality of target voiceprint feature vectors corresponding to the plurality of target voice segments respectively by using a trained speaker verification model, wherein the trained speaker verification model uses a multi-layer structure of an SE Res-D Block backbone network; and performing similarity comparison between the plurality of target voiceprint feature vectors and a target voiceprint feature space corresponding to the target identity information to verify the identity of the target user. Through the present disclosure, the problem that the related art directly mixes near-field voice data and far-field voice data into a model for learning and training, resulting in low accuracy of the model when performing far-field speaker verification, is solved, and the effect of improving the accuracy of far-field speaker verification is achieved.
Owner:KONKA GROUP

Attentive scoring function for speaker identification

ActiveUS12670912B2Speaker verificationSpeech sound
A speaker verification method includes receiving audio data corresponding to an utterance, processing the audio data to generate a reference attentive d-vector representing voice characteristics of the utterance, the evaluation ad-vector includes ne style classes each including a respective value vector concatenated with a corresponding routing vector. The method also includes generating using a self-attention mechanism, at least one multi-condition attention score that indicates a likelihood that the evaluation ad-vector matches a respective reference ad-vector associated with a respective user. The method also includes identifying the speaker of the utterance as the respective user associated with the respective reference ad-vector based on the multi-condition attention score.
Owner:GOOGLE LLC

Consistency check for large language model continuous conversations

PCT designated stageWO2026029296A1Machine learningSpeech recognitionAcoustic energySpeaker verification
In addition to speaker verification, an LLM continuous conversation check employs a spoken speed check using a phonemes-based spoken speed calculation, and acoustic energy check, and a signal-to-noise estimation to determine whether first and second audio inputs include utterances forming a continuous conversation by the user with the LLM. Results from the various components of the continuous conversation check are fused based on automatically assigned weights in making the determination. Continuous conversation detection for LLMs is therefore more robust, particularly for a very short second utterance. Optionally a distance to microphone check or a semantic consistency check may also be employed.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent synthesized speech speaker verification method based on group-res2block network

ActiveCN116092470BStrengthen contextual connectionsimprove accuracyFeature extractionSpeaker verification
The application discloses a kind of based on Group-Res2Block network's intelligent synthetic speech speaker confirmation method, comprising:1 obtains intelligent synthetic speech data set;2 establishes the feature extraction network model based on Group-Res2Block;3 trains the feature extraction network model based on Group-Res2Block established;4 realizes prediction using the model established, to reach the purpose of confirming the speaker subject of intelligent synthetic speech.This application can maximize to obtain the common effective voiceprint features of natural human voice and intelligent synthetic speech, so as to effectively realize intelligent synthetic speech speaker confirmation, and can improve confirmation accuracy.
Owner:HEFEI UNIV OF TECH

Method of verifying a speaker using a neural network, associated device and motor vehicle.

ActiveFR3158191B1Speech analysisSpeaker verificationEngineering
A speaker verification method, implemented in a motor vehicle, comprising the following steps: Receiving a verification input (110_EV80,160) representative of the speaker's voice; Producing a verification output (122_SV80,2) from the verification input (110_EV80,160) by a first electronic convolutional neural network (111_100), comprising the following steps: Determining a working matrix (110_T80,160,16) comprising a convolution operation Conv0 followed by an attention operation ATT0, where: TH,W,C = ATT0 (Conv0 (EH,W)), TH,W,C is the working matrix (110_T80,160,16), EVH,W is the verification input (0_EVH,W); Producing an intermediate output (121_SI80,160,16) from the working matrix (110_T80,160,16), Production of the verification output (122_SV80,2) from the intermediate output (121_SI80,160,16), Comparison of a distance between the verification output (122_SV80,2) and a reference output with a threshold. Figure for the abbreviation: Figure 2,
Owner:STELLANTIS AUTO SAS +1