Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Speaker verification" patented technology

A method for generating a speaker diary based on audio-visual fusion clustering

ActiveCN119964596BSpeech recognitionSpeech segmentationSpeaker verification
The application discloses a speaker diary generation method based on audio-visual fusion clustering, and aims to solve the problem of "who speaks at what time" in a multi-speaker scene. The method is realized through the following steps: first, an overlapping-aware speech segmentation model is used to segment the audio segment, solving the problem of overlapping speech; second, an advanced speaker verification model is used to extract the speaker voiceprint features of each audio segment and a speaker score matrix generated through face tracking and speaker detection; then, through an audio-video joint clustering method, the number of clusters is optimized according to the audio features and visual information, and K-means clustering is used to complete speaker clustering; the experimental results show that the system adopting the method achieves the lowest diary error rate (DER) on the Ego4D validation set.
Owner:HUNAN UNIV

Ensemble learning (EL)-based speaker verification method

ActiveUS12633292B2Ensemble learningSpeech analysisSpeaker verificationData acquisition
Provided is an ensemble learning (EL)-based speaker verification method. The method includes: data acquisition and preprocessing; selecting and training a group of basic models, and optimizing model parameters; performing similarity scoring on an acquired pair of speaker feature embedding via the group of basic models; constructing a detection cost function (DCF); generating a weight, and performing weighted fusion on scoring results of the group of basic models based on the weight, to obtain a final ensemble model for speaker verification; based on a near-speaking or far-speaking test scenario of a voice sample, the scenario is distinguished and input into the ensemble model, to obtain a final similarity score after weighted fusion; and determining, based on a threshold, whether there is a same speaker, where if the similarity score is greater than the threshold, it is determined that there is a same speaker.
Owner:HANGZHOU DIANZI UNIV

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Quantization method of speaker verification model, electronic device and storage medium

The application discloses a speaker verification model quantization method, an electronic device and a storage medium, wherein the speaker verification model quantization method comprises the following steps: acquiring real-value weights of all layers of a speaker verification model; and mapping the real-value weights of all layers to a fixed integer set; or dynamically determining binary weights corresponding to the real-value weights of each layer to better match the real-value weight distribution. The method of the application embodiment proposes two new quantization strategies, namely static quantization and adaptive quantization. Furthermore, for static quantization, the application embodiment proposes a weight regularization technique to maintain maximum information entropy and reduce information loss. Furthermore, the application embodiment also proposes an adaptive quantization scheme, which can dynamically determine the best binary of each layer to achieve better alignment with the real-value weight distribution.
Owner:AISPEECH CO LTD

System and method for speaker verification for voice assistant

ActiveUS12475886B2Speech recognitionSpeaker verificationSpeech sound
A method includes obtaining audio data and identifying an utterance of a wake word or phrase in the audio data. The method also includes generating an embedding vector based on the utterance from the audio data and accessing a set of previously-generated vectors representing previous utterances of the wake word or phrase. The method further includes performing clustering on the embedding vector and the set of previously-generated vectors to identify a cluster including the embedding vector, where the identified cluster is associated with a speaker. The method also includes updating a speaker vector associated with the speaker based on the embedding vector and determining, using a speaker verification model, a similarity score between the updated speaker vector and the embedding vector. In addition, the method includes determining, based on the similarity score, whether a speaker providing the utterance matches the speaker associated with the identified cluster.
Owner:SAMSUNG ELECTRONICS CO LTD

Adversarial network optimization method and system for short-utterance speaker verification

ActiveCN114530156BSpeech analysisPattern recognitionSpeaker verification
The embodiment of the specification provides a generative adversarial network optimization method and system for short speech speaker verification, wherein the method comprises the following steps: acquiring a plurality of pairs of long and short speech acoustic feature samples; inputting the short speech acoustic feature sample into a generator for splicing to obtain a generated pseudo long speech acoustic feature sample; inputting the pseudo long speech acoustic feature sample and the acquired long speech acoustic feature sample into a speaker verification model respectively, outputting a pseudo identity feature sample and a true identity feature sample through the speaker verification model; inputting the true identity feature sample and the pseudo identity feature sample into a discriminator and a classifier, calculating the loss of the discriminator and the classifier through a loss function, and updating the parameters of the discriminator, the classifier and the generator through back propagation optimization. The problem that the discrimination effect of a speaker verification system becomes poor as the speech duration becomes shorter is solved.
Owner:STATE GRID CORPORATION OF CHINA +1

Speaker verification model training methods, devices, media and equipment

This invention discloses a method, apparatus, medium, and device for training a speaker verification model. The method includes: extracting waveforms from acquired speech audio data to obtain audio waveform data corresponding to the speech audio data; inputting the audio waveform data into a preset classification model to output a predicted label corresponding to the speech audio data; determining the loss value of the preset classification model based on the anti-counterfeiting label and the predicted label corresponding to the speech audio data; and adjusting the parameters of the preset classification model using the loss value to obtain a speaker verification model. This invention adopts a supervised learning approach, using real anti-counterfeiting labels to assist in determining the predicted label of the model, and trains the entire model by minimizing the error between the output and the real label, thereby improving training efficiency and reducing costs.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL +1

A speaker verification method based on SASFV aggregation model

ActiveCN120766685BSpeech analysisSpeaker recognition systemNetwork generation
The application discloses a speaker verification method based on a SASFV aggregation model, and relates to the field of speech recognition.The method extracts a log Mel spectrogram through short-time Fourier transform and Mel filtering, generates frame-level features by using an ERes2Net network, introduces a SASFV aggregation model to generate fixed-length speaker-level features in combination with a Fisher Vector variable, a self-attention mechanism and a statistical method, and finally determines the identity of a speaker by using a cosine distance.The application solves the problem that the prior art cannot effectively represent and aggregate features in a short speech task, and significantly improves the accuracy, robustness and performance of a speaker recognition system.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Speaker anonymization-oriented stylegan2-based F0 cloning method

PendingCN121506160ASpeech analysisAlgorithmSpeaker verification
The invention relates to the technical field of voice processing, in particular to a speaker anonymization-oriented stylegan2-based F0 cloning method, which mainly realizes source speaker anonymization protection and comprises the following steps of: preprocessing an original fundamental frequency sequence, and converting a random noise vector into a representation vector of a style space; and a StyleGAN2 architecture is introduced to carry out layer-by-layer feature processing and fusion. A probabilistic fusion strategy is adopted, content features and constant features are added, then layer-by-layer stylization processing is carried out, feature modulation is carried out on each layer by using a style vector, the time sequence resolution is gradually recovered, and the anonymization target of the identity features of the speaker is effectively changed. And finally, combining the generative fundamental frequency (F0), the text information and the anonymized identity features to synthesize anonymized voice so as to hide the identity information of the speaker. And the effectiveness of the speaker anonymization-oriented StyleGAN2-based F0 cloning method is verified through an automatic speaker verification model and an automatic speech recognition system.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Method and apparatus for speaker verification based on next-tdnn

ActiveUS20250316273A1Speech analysisMachine learningSpeaker verificationEngineering
A method and an apparatus are disclosed for a speaker verification. The apparatus may comprise at least one processor configured to execute instructions to perform generating, based on an utterance of a speaker to a deep learning-based speaker verification model, a speaker embedding with a preset dimension, verifying, based on the speaker embedding, the speaker, and detecting, based on the verified speaker, an identity of a user associated with the apparatus.
Owner:HYUNDAI MOTOR CO LTD +2

Consistency check for large language model continuous conversations

In addition to speaker verification, an LLM continuous conversation check employs a spoken speed check using a phonemes-based spoken speed calculation, and acoustic energy check, and a signal-to-noise estimation to determine whether first and second audio inputs include utterances forming a continuous conversation by the user with the LLM. Results from the various components of the continuous conversation check are fused based on automatically assigned weights in making the determination. Continuous conversation detection for LLMs is therefore more robust, particularly for a very short second utterance. Optionally a distance to microphone check or a semantic consistency check may also be employed.
Owner:SAMSUNG ELECTRONICS CO LTD

Method for training multi-modal speaker verification model and electronic device

ActiveCN120148521BSpeech recognitionSpeaker verificationFeature coding
The application discloses a training method of a multi-modal speaker authentication model and an electronic device, relates to the field of speaker authentication, and comprises the following steps: obtaining an audio-video sequence sample corresponding to a given speaker; performing lip feature coding on each audio frame in the audio-video sequence sample in sequence to determine a corresponding audio sample lip latent vector sequence; performing feature coding on a speaker mouth region in each video frame in the audio-video sequence sample in sequence to determine a corresponding video sample lip latent vector sequence; calculating an audio-video sequence similarity corresponding to the audio sample lip latent vector sequence and the video sample lip latent vector sequence; and training a speaker authentication model according to the audio-video sequence similarity and an audio-video speaker alignment label. Thus, the similarity of the speaker audio and the lip shape graph of the person to be detected is compared by matching the speaker audio and the lip shape graph into the same dimension vector through the latent vector, so that the accuracy of the authentication result can be effectively improved.
Owner:AISPEECH CO LTD

Speaker feature vector extraction method, speaker feature vector verification method, speaker feature vector extraction device, speaker feature vector verification device and speaker feature vector product

The invention provides a speaker feature vector extraction method, a speaker feature vector verification method, a speaker feature vector extraction device, a speaker feature vector verification device and a product, and the speaker feature vector extraction method comprises the steps: processing an obtained input voice signal to obtain a basic acoustic feature; processing the basic acoustic features through at least three speaker attention weight generation modules in sequence; the speaker attention weight module comprises a compressed excitation residual block and a speaker conditional attention module, and the output end of the compressed excitation residual block is connected with the input end of the speaker conditional attention module; the multi-level feature aggregation module outputs a spliced statistical feature vector according to the output features of each speaker conditional attention module; and the full connection layer outputs a speaker feature vector according to the spliced statistical feature vector. The speaker condition attention mechanism is introduced to improve the accuracy of the feature vector of the speaker, and the problem that the verification is not accurate enough in a complex scene such as relatively high speaker similarity is solved when the method is used for speaker verification.
Owner:SHENZHEN JUSHENG TECHNOLOGY CO LTD +1

Robust intelligent synthesized speech speaker verification model training method and system

ActiveCN116597843BInternal combustion piston enginesSpeech analysisSpeaker verificationNoise
The application provides a kind of robust intelligent synthetic speech speaker confirmation model training method, system, storage medium and electronic equipment, it is related to speech processing technical field.The robust intelligent synthetic speech speaker confirmation model provided by the present application is used to improve the speaker confirmation accuracy of intelligent synthetic speech under strong background noise condition, comprising speech enhancement network, feature extraction network and feature enhancement module.In the training phase of the model, the training set of noisy intelligent synthetic speech data set is transported into speech enhancement network and feature extraction network after STFT feature preprocessing and Fbank feature preprocessing, and the two networks are connected through feature enhancement module for joint training to obtain speaker embedding features with noise-robustness.In the test phase, the model is tested based on the test set of noisy intelligent synthetic speech data set;The optimal model selection is carried out in the mode of repeating the foregoing training and testing alternately until the training iteration number reaches the set maximum value.
Owner:HEFEI UNIV OF TECH

Speaker recognition method based on speaker voice micro-movements

ActiveCN118918900BSpeech analysisDigital dataSpeaker verification
The present application relates to the technical field of electronic digital data processing, in particular to a speaker recognition method based on speaker voice micro-motion, comprising: after the voice stream is preprocessed, Fbank features are extracted and sent into a teacher network and a student network respectively to obtain respective corresponding feature embeddings; the feature embeddings obtained by the teacher network and the student network are sent into a loss function and back propagation is performed; the student network is normally iterated, and the teacher network is iterated through an EMA sliding average method; the voiceprint feature information obtained through an ECAPA-TDNN voiceprint model is aggregated and classified with the speaker voice micro-motion information obtained by training the accent data, and speaker recognition is performed; the present application improves the generalization performance of the model by using data enhancement and other methods, avoids fitting in the channel features, does not require manual labeling, and helps the speaker verification model to achieve the ability to identify speakers in a larger population by introducing new features.
Owner:NANJING LONGYUAN INFORMATION TECH CO LTD +1

Method and System for Personalising Speaker Verification Models

ActiveGB2630805BPersonalizationSpeaker verification
A computer-implemented server-performed method for personalising a trained speaker verification machine learning model for voice authentication of specific users obtains a “positive” audio sample of t
Owner:SAMSUNG ELECTRONICS CO LTD

Zero sample ASMR generation method and system based on large language model

PendingCN121260178ASpeech synthesisData setSpeaker verification
The invention discloses a zero sample ASMR generation method and system based on a large language model, and aims to solve the problems that in the prior art, personalized ASMR voice cannot be generated in a zero sample mode, and a high-quality ASMR special data set is lacked. The method comprises the following steps that speaker prompt and task signals of texts to be synthesized and normal or ASMR style voice are obtained; retrieving a matching task cue from the virtual pool of speakers based on the speaker verification system; pre-training the large language model to generate a voice token sequence containing a target style; the stream matching acoustic decoder fuses the voice token sequence and speaker acoustic information to generate a target Mel spectrum; the vocoder synthesizes a target audio; the system comprises an input module, a virtual speaker pool module, a large language model style coding module, a stream matching decoding module and a vocoder module, and further comprises a data set module for storing a DeepASMR-DB data set covering 9 types of themes, Chinese and English and over 670-hour voices. According to the method, zero-sample ASMR generation is realized, the tone and real breath sound of a speaker are reserved, and the audio fidelity is high.
Owner:SHANGHAI JIAOTONG UNIV

Method for Obtaining Enrolment Data for a Speaker Verification Model

ActiveGB2636095BSpeech analysisSpeaker verificationAudio frequency
A method for obtaining enrolment data for a speaker verification mode, comprises capturing audio samples of utterances spoken by a speaker, generating an embedding vector for some samples, identifying
Owner:SAMSUNG ELECTRONICS CO LTD

Automatic generation and / or use of text-dependent speaker verification features

ActiveUS12451140B2Digital data authenticationSound input/outputSpeaker verificationSpeech sound
Implementations relate to automatic generation of speaker features for each of one or more particular text-dependent speaker verifications (TD-SVs) for a user. Implementations can generate speaker features for a particular TD-SV using instances of audio data that each capture a corresponding spoken utterance of the user during normal non-enrollment interactions with an automated assistant via one or more respective assistant devices. For example, a portion of an instance of audio data can be used in response to: (a) determining that recognized term(s) for the spoken utterance captured by that the portion correspond to the particular TD-SV; and (b) determining that an authentication measure, for the user and for the spoken utterance, satisfies a threshold. Implementations additionally or alternatively relate to utilization of speaker features, for each of one or more particular TD-SVs for a user, in determining whether to authenticate a spoken utterance for the user.
Owner:GOOGLE LLC

Speaker verification model adversarial training method and system based on general adversarial disturbance

PendingCN121122324ASpeech analysisSpeaker verificationNoise
The invention relates to the technical field of speech recognition, in particular to a speaker verification model adversarial training method and system based on general adversarial disturbance, and the method comprises the steps: obtaining an indication label and random Gaussian noise with a fixed length; according to the indication label, converting the random Gaussian noise into general adversarial disturbance by using an adversarial disturbance generator, and adding the general adversarial disturbance to the clean audio sample in different modes to obtain an audio adversarial sample; using a discriminator to obtain a probability value of whether the audio confrontation sample is a target audio; and inputting the audio confrontation sample and the target audio into the speaker verification model to obtain an audio embedding vector, and obtaining similarity loss of the speaker verification model so as to perform confrontation training on the speaker verification model. Therefore, the problem that in confrontation training, confrontation disturbance is high in dependency on samples, confrontation disturbance numerical values are single, and the effect in a physical domain is poor, so that the robustness of a speaker verification model is poor is solved.
Owner:ARMY ENG UNIV OF PLA

Speaker verification method and system based on dual-stream low-rank adaptive and adversarial decoupling

The application discloses a speaker verification method and system based on double-flow low-rank adaptive and anti-decoupling, comprising: based on a pre-trained speech network, a double-flow low-rank adaptive anti-decoupling network is constructed, and the original weight parameters of the pre-trained speech network are frozen; based on a language feature extraction branch and a speaker feature extraction branch, original speech data is subjected to feature extraction respectively to obtain language features and speaker features; the language features are input into the shared discriminator to perform language classification prediction to complete language boundary anchoring; after gradient inversion processing of the speaker features, the speaker features are input into the shared discriminator which has completed language boundary anchoring to perform anti-decoupling; identity recognition is performed based on the speaker features subjected to anti-decoupling constraint; corresponding training losses are calculated respectively, and iterative updating is performed based on the training losses. The application can improve the acceptance rate of cross-language speech of the same person and the rejection rate of the same language speech of different persons, and is suitable for high-precision speaker verification in a multi-language environment.
Owner:NANJING UNIV

A time-varying speaker template updating method based on risk gating

PendingCN122337207ASpeaker verificationEngineering
This invention discloses a time-varying speaker template update method based on risk gating, belonging to the field of speaker verification and speech signal processing technology. The method first constructs an initial speaker template and anchor template from registered speech, then verifies the test speech. When the update conditions are met, an attractive update is performed first, then candidate negative samples are obtained. A risk gating quantity is constructed through uncertainty-driven or deterministic-driven gating branches to determine the repulsion strength. Projective repulsion update is then performed in the negative direction, followed by step size pruning and anchor point retraction. This method can reduce template drift caused by erroneous updates and improve the stability and discriminative power of time-varying speaker verification.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Speech synthesis method, speech synthesis device, electronic device, and storage medium

ActiveCN119811361BSpeech synthesisSynthesis methodsSpeaker verification
The voice synthesis method, voice synthesis device, electronic equipment and storage medium provided by the present application relate to the fields of artificial intelligence and financial technology. The method comprises the following steps: extracting speaker features from sample voice data sets through a feature embedder in an initial speaker verification model to obtain sample set speaker features; classifying the sample set speaker features through a feature classifier in the initial speaker verification model to obtain a sample verification speaker category, then adjusting parameters of the initial speaker verification model to obtain a target speaker verification model; extracting speaker features from target voice data through a feature embedder in the target speaker verification model to obtain target speaker features; and generating voice according to the target speaker features and target text features to obtain target synthesized voice data. The present application can alleviate the adverse effects of background noise in voice data and improve the accuracy of voice synthesis.
Owner:PING AN TECH (SHENZHEN) CO LTD

A text-independent speaker verification method based on identity information and semantic information disentanglement

ActiveCN116543775BDisentanglement implementationImprove generalization abilityInternal combustion piston enginesSpeech recognitionData setSpeaker recognition system
The present application relates to the technical field of voiceprint recognition, and particularly relates to a text-independent speaker verification method based on identity information and semantic information disentanglement. The technical scheme of the present application is completed by building a neural network, training a model, and testing a result. sch The semantic content is represented by a fixed-length feature f spk The correlation between the voiceprint feature f sch and the semantic feature f sch is calculated using a disentanglement method such as mutual information to achieve disentanglement between the voiceprint feature and the semantic feature, obtain more accurate and robust speaker voiceprint features, and solve the problem of poor system generalization performance caused by interference of semantic content factors in the data set extracted by the speaker recognition system.
Owner:SHANXI UNIV

Far-field speaker verification method based on self-distillation pre-training and meta-learning fine-tuning

ActiveCN116863937BSpeech analysisPattern recognitionSpeaker verification
The application discloses a far-field speaker verification method based on self-distillation pre-training and meta-learning fine-tuning, and the process is as follows: log mel spectrum features are extracted from near-field training speech as the input of the network, and a self-distillation learning method is used to pre-train a deep neural network; then, log mel spectrum features of far-field training speech are used as the input of the network, and a meta-learning method is used to fine-tune the pre-trained network; then, log mel spectrum features of the registered speech are input into the pre-trained and fine-tuned network to obtain a transformed feature vector, and log mel spectrum features of the test speech are input into the fine-tuned and pre-trained network to obtain a transformed feature vector; finally, the distance between the transformed feature vector of the test speech and the transformed feature vector of the registered speech is calculated, and it is judged whether the two are from the same speaker. The method improves the network training efficiency and improves the speaker verification accuracy when the recording environment of the registered speech and the test speech is inconsistent.
Owner:SOUTH CHINA UNIV OF TECH

Method and apparatus for speaker verification based on next-TDNN

ActiveUS12718821B2Speaker verificationEngineering
A method and an apparatus are disclosed for a speaker verification. The apparatus may comprise at least one processor configured to execute instructions to perform generating, based on an utterance of a speaker to a deep learning-based speaker verification model, a speaker embedding with a preset dimension, verifying, based on the speaker embedding, the speaker, and detecting, based on the verified speaker, an identity of a user associated with the apparatus.
Owner:HYUNDAI MOTOR CO LTD +2

Voice analysis identity authentication method based on sound field reconstruction

ActiveCN116504251BEnsure consistencyNo additional costSpeech analysisHigh level techniquesSpeaker verificationVoice analysis
The application discloses a speech analysis identity authentication method based on sound field reconstruction, which comprises distance sensing, sound field reconstruction, sound field extraction, and model training and inference. The distance sensing is achieved by emitting a chirp signal through a loudspeaker and receiving the signal, and the distance from the user to the mobile phone is obtained by using a cross-correlation method. The sound field reconstruction is achieved by establishing a distance-dependent impulse response database, obtaining the corresponding transfer function according to the measured distance, and then reconstructing the sound field at the verification position to the sound field at the registration position. The sound field extraction is achieved by separately processing the signals of the two sound channels of the speech signal after the sound field reconstruction, and extracting the field texture. The model training and inference are achieved by using the registration field texture and the reconstructed field texture in the verification stage to construct a speech authentication model. The method can effectively solve the problem that the original sound field method is sensitive to distance when verifying the speaker, and the user does not need to maintain a fixed distance during verification as during registration.
Owner:ZHEJIANG UNIV

Frame-level multi-channel speaker verification method under large-scale self-organizing microphone array

The application discloses a frame-level multi-channel speaker verification method under a large-scale self-organizing microphone array, a space-time processing block is added before a pooling layer of a single-channel speaker verification system, and context relationships in a channel, between channels and across time are modeled respectively, and the performance of a far-field ASV is further improved. The method comprises the following steps: 1) a space-time processing block composed of a cross-frame processing layer (CFL) and a cross-channel processing layer (CCL) is added before the pooling layer; and 2) in order to make the channel weight of a noise channel zero, a softmax operator of the cross-channel processing layer is improved into a sparsemax operator. Results on a Libri-adhoc-simu dataset show that a multi-channel ASV system of the STB realizes an equal error rate (EER) of less than 33% of an oracle one-best baseline; and results on a Libri-adhoc40 dataset show that the multi-channel ASV system of the STB realizes an equal error rate of less than 27% of the oracle one-best baseline, and also realizes an equal error rate of less than 9% of a speech-level cross-channel self-attention ASV system, and achieves superior performance.
Owner:NORTHWESTERN POLYTECHNICAL UNIV +1

Speaker verification with multitask speech models

ActiveUS12494206B2Speech analysisNerve networkSpeaker verification
A method includes obtaining a speaker identification (SID) model trained to predict speaker embeddings from utterances spoken by different speakers, the SID model includes a trained audio encoder and a trained SID head. The method also includes receiving a plurality of synthetic speech detection (SSD) training utterances that include a set of human-originated speech samples and a set of synthetic speech samples. The method also includes training, using the trained audio encoder, a SSD head on the SSD training utterances to learn to detect the presence of synthetic speech in audio encodings encoded by the trained audio encoder. The operations also include providing, for execution on a computing device, a multitask neural network model for performing both SID tasks and SSD tasks on input audio data in parallel.
Owner:GOOGLE LLC

A method and device for verifying speaker identity in a far-field scene, and an electronic device

ActiveCN116343798BSpeech analysisManufacturing computing systemsSpeaker verificationEngineering
The present disclosure provides a method and device for verifying the identity of a speaker in a far-field scene, and an electronic device, wherein the method comprises: obtaining target identity information selected by a target user and a voice of the target user; dividing the voice of the target user into a plurality of target voice segments, and obtaining a plurality of target voiceprint feature vectors corresponding to the plurality of target voice segments respectively by using a trained speaker verification model, wherein the trained speaker verification model uses a multi-layer structure of an SE Res-D Block backbone network; and performing similarity comparison between the plurality of target voiceprint feature vectors and a target voiceprint feature space corresponding to the target identity information to verify the identity of the target user. Through the present disclosure, the problem that the related art directly mixes near-field voice data and far-field voice data into a model for learning and training, resulting in low accuracy of the model when performing far-field speaker verification, is solved, and the effect of improving the accuracy of far-field speaker verification is achieved.
Owner:KONKA GROUP