Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

194 results about "Speaker identification" patented technology

LLM as a transcription filter

A user electronic device comprising: one or more microphones configured to capture raw audio data; and one or more processors and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising: receiving the raw audio data captured by the one or more microphones; processing the raw audio data using a speech transcriber to generate a live transcription of the raw audio data that comprises a plurality of text tokens; processing the raw audio data to generate a speaker identification output that identifies, for each of the text tokens, a respective speaker for each of the text tokens in the live transcription; and processing a first input comprising (i) a first input prompt and (ii) an input text generated from the live transcription using a language model neural network to generate a modified transcription.
Owner:GOOGLE LLC

Short voice-based voiceprint clustering method guided by speaker recognition pre-training model

PendingCN120375834ASpeech analysisSpeech segmentationFeature Dimension
The invention discloses a voiceprint clustering method guided by a speaker recognition pre-training model based on short voices, and the method comprises the following steps: obtaining an original voice signal, and randomly combining a plurality of enhancement strategies to achieve data enhancement; performing voice segmentation on the voice signal after data enhancement based on a uniform segmentation mode; extracting voiceprint features of the segmented voice based on an attention mechanism of global time-frequency domain context modeling; and obtaining a clustering result based on K-means clustering and spectral clustering, matching the clustering result with a real speaker tag, performing reverse transmission based on angle-dependent AAM-Softmax loss, and outputting a voiceprint clustering result. According to the method, the influence of environmental interference on feature extraction can be overcome, feature dimensions which are more effective for identity identification can be screened, a clustering output effect which is superior to that of a mainstream algorithm can be obtained with a relatively low parameter quantity, and the robustness under a noise interference condition is improved.
Owner:SOUTH CHINA UNIV OF TECH

Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

The invention relates to a simulation digital human real-time intelligent voice interaction system and a simulation digital human real-time intelligent voice interaction method based on vision and a large model, and aims to solve the problems of inaccurate target speaker recognition, high response delay and the like in digital human voice interaction in a complex scene. The system circles an effective recognition range through a camera, triggers audio collection in combination with face detection, locks a target speaker and reduces noise by using lip movement recognition and sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large language model (LLM) and knowledge retrieval enhancement (RAG) technologies, generates low-delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video stream and pushing the video stream to a front end for rendering in real time. Accurate pickup, low-delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.
Owner:UNICOM (HENAN) IND INTERNET CO LTD

Zero-configuration adaptive speaker recognition method and system

The invention discloses a zero-configuration adaptive speaker recognition method and system, and relates to the technical field of voice signal processing. The method comprises the following steps: receiving an audio stream, carrying out voice activity detection, obtaining a single-person voice segment, extracting a voiceprint embedding vector of the single-person voice segment, and carrying out online clustering to generate a speaker identity pool; and the identity pool is updated by calculating the multi-dimensional fusion similarity between the remaining single-person voice segments and the voice model, and temporary identity tags corresponding to the single-person voice segments are output. According to the speaker recognition method provided by the invention, a voiceprint template does not need to be registered in advance, and the flexibility and real-time performance of the system are improved.
Owner:北京文聿科技有限公司

Systems and Methods for Digital Transcript Creation Using Automated Speech Recognition

This disclosure relates generally to systems, methods, and computer readable media for providing improved insights and annotations to enhance recorded audio, video, and / or written transcriptions of testimony. For example, in some embodiments, a method is disclosed for correlating non-verbal cues recognized from an audio and / or video recording of testimony to the corresponding testimony transcript locations. In other embodiments, a method is disclosed for providing testimony-specific artificial intelligence-based insights and annotations to a testimony transcript, e.g., based on the use of machine learning, natural language processing, and / or other techniques. In still other embodiments, a method is disclosed for providing smart citations to a testimony transcript, e.g., which track the location of semantic constructs within the transcript over the course of various modifications being made to the transcript. In yet other embodiments, a method is disclosed for providing intelligent speaker identification-related insights and annotations to an audio recording of a testimony transcript.
Owner:MAGNA LEGAL SERVICES LLC

Land-air communication speaker recognition method facing short voice and complex noise

The invention relates to the technical field of voiceprint recognition, and discloses an air-ground conversation speaker recognition method for short voice and complex noise, which dynamically suppresses background noise through a self-adaptive filter in combination with a minimum mean square algorithm, and improves the voice signal quality. A multi-scale feature fusion technology is adopted, short-time Fourier transform, filter bank energy features and multi-scale convolution are combined, a random time and frequency shielding data enhancement strategy is introduced, and the short speech feature expression ability is enhanced; bi-GRU is integrated in the ECAPA-TDNN model, and context information of multiple rounds of conversations is captured through context time sequence modeling, so that the recognition stability in a dynamic scene is improved; and an AAM-Softmax loss function is adopted to strengthen the distinguishing capability of the model, and finally identity judgment is realized through cosine similarity matching. According to the method, the voiceprint recognition precision and robustness under complex noise and short voice conditions are remarkably improved, and the safety risk caused by identity misjudgment in aviation communication is effectively reduced.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

System and methods for audio data analysis and tagging

A system for automated processing and analysis of audio files for large data sets in a cloud environment. A unified analytic environment can integrate audio machine learning models for processing and analysis with a knowledge management system, including graph presentations of tracked entities, linked to audio files and / or associated translations and transcripts. Entities within such data can be searched or filtered and proposed for tracking, or identified as tracked objects. These features can allow triage and prioritization of audio files for analysis. User interfaces can facilitate feedback on transcription and translation outputs, thereby improving present outputs and future inputs and outputs. Entities speaking or referred to can be found, tagged, and distinguished in audio files (e.g., using speaker identification in audio files, text searching in transcripts, etc.) Users can provide feedback and input on various aspects of a system, to enhance or adjust initial automated or other machine learning outputs.
Owner:PALANTIR TECHNOLOGIES INC

Speaker recognition method based on bipartite graph matching and electronic equipment

The invention discloses a bipartite graph matching-based speaker recognition method and electronic equipment, and belongs to the technical field of audio recognition. The method comprises the following steps: acquiring an audio stream, and segmenting the audio stream into continuous target duration audio clips in real time; determining a voiceprint embedding vector corresponding to each audio clip of the target duration; determining an embedding matrix based on the voiceprint embedding vector; determining a target voiceprint feature corresponding to each target duration audio clip according to a bipartite graph method; calculating the similarity between the voiceprint embedding vectors; and grouping the voiceprint embedding vectors based on the similarity between the voiceprint embedding vectors, and clustering the voiceprint embedding vectors corresponding to the same speaker to obtain a clustering result and generate a target recognition report. According to the invention, the calculation complexity can be reduced and the identification precision can be improved.
Owner:BEIJING TONGXIANG QIANFANG TECHNOLOGY CO LTD

Speaker speech recognition method

The invention relates to a speaker voice recognition method, and relates to the technical field of telephone voice signal processing. The method comprises the following steps: acquiring a to-be-identified call audio in real time, and segmenting the to-be-identified call audio into a plurality of audio blocks; for each audio block, extracting voice features of the audio block, and inputting the voice features into a pre-trained speaker voice recognition model; outputting the affiliation probability that the audio block belongs to the target speaker through the speaker speech recognition model; and if the affiliation probability of the audio block belonging to the target speaker is greater than a preset speaker affiliation probability threshold, judging that the audio block belongs to the target speaker. According to the invention, target speaker identification in a complex scene can be realized.
Owner:SHENZHEN HUOLI TIAN HUI TECH CO LTD

Simultaneous interpretation device, simultaneous interpretation system, simultaneous interpretation processing method, and non-transitory computer readable storage medium

Provided is a simultaneous interpretation system that performs machine translation processing and speaker identification processing in real time. In the simultaneous interpretation system, the segment processing unit of the simultaneous interpretation device performs high-speed and highly accurate segment processing to obtain sentence data and also obtains a time range in which a word sequence included in the sentence data was uttered, thus, allowing for performing machine translation processing and speaker identification processing in real time. In other words, in the simultaneous interpretation system, the machine translation processing unit performs machine translation processing on sentence data obtained through high-speed and highly accurate segment processing, and performs, in parallel, processing for predicting a speaker who spoke during the period specified by the time range data based on the inputted video stream and the time range data, thus allowing for performing machine translation processing and speaker identification processing in real time.
Owner:NAT INST OF INFORMATION & COMM TECH

Multitask speech emotion recognition method based on parallel processing hybrid expert network

The invention belongs to the technical field of speech recognition, and particularly discloses a multi-task speech emotion recognition method based on a parallel processing hybrid expert network, and the method comprises the steps: carrying out the feature extraction of an input original speech, and obtaining a shared feature; constructing a parallel processing hybrid expert network, inputting the shared features into the parallel processing hybrid expert network, carrying out sentiment classification and speaker recognition on the shared features in parallel, and respectively outputting a predicted sentiment value and a predicted speaker; inputting the shared features into a full connection layer for voice recognition processing to obtain a text prediction value; and obtaining a loss function according to the predicted emotion value, the predicted speaker and the text predicted value in combination with the corresponding labels, and optimizing the model parameters according to the loss function to obtain the parallel processing hybrid expert network after optimization training. According to the invention, the speech emotion recognition performance and recognition accuracy can be improved.
Owner:HUAZHONG NORMAL UNIV

Conference summary generation method and device

The invention provides a conference summary generation method and device, and the method comprises the steps: obtaining an audio and video file of a target conference, and the audio and video file comprises an audio file and / or a video file; performing voice activity detection processing, text conversion processing and speaker recognition processing on the audio and video file to obtain a target dialogue text; constructing prompt words based on the application scene of the target conference and the target dialogue text; and inputting the cue word into a conference summary generation large model, and generating a conference summary corresponding to the target conference. According to the invention, the problem of low conference summary generation efficiency is solved, and more time is saved; the generation result of the conference summary generation large model is optimized by adopting a prompt word engineering mode, so that the accuracy of generating the conference summary can be further improved; the conference summary is generated by adopting the conference summary generation large model, compared with a traditional neural language programming technology, feature engineering construction is omitted, and the performance is better.
Owner:SANY HEAVY MACHINERY

Speaker recognition method in medical real-time speech recognition scene

ActiveCN121938378Aconsistent with auditory logicReduce the risk of number fluctuationsSpeech recognitionAutomatic speechSpeech sound
The invention discloses a speaker recognition method in a medical real-time speech recognition scene, and the method specifically comprises the steps: carrying out the parallel distribution of a received PCM audio stream through an audio diverter in a gateway layer, and transmitting the audio stream to an automatic speech recognition link and a speaker separation link; in the automatic speech recognition link, outputting a recognition text and Token-level timestamps corresponding to each minimum semantic unit in the recognition text; in the speaker separation link, extracting a corresponding speaker embedding vector based on a preloaded speaker embedding model, and performing online clustering processing on the embedding vector to generate a speaker identity identifier; token-level boundary adsorption labeling processing is executed based on the Token-level timestamp and an online clustering result, so that a speaker switching position is determined; and carrying out one-to-one mapping on the speaker tags and the recognition text according to the Token-level timestamps, merging Tokens with the same speaker tags continuously, and outputting a transliteration text with the speaker tags.
Owner:ZOE SOFT CORP LTD

Speaker recognition method and system, terminal and medium

The invention provides a speaker recognition method and system, a terminal and a medium, and the method comprises the steps: constructing and training a voice feature extraction depth model through employing a first time-frequency fusion residual network, and constructing a voice feature extraction lightweight model through employing a second time-frequency fusion residual network; and taking the voice feature extraction depth model as teacher network training to obtain a convergent voice feature extraction lightweight model, so as to further perform feature extraction on the to-be-recognized voice data subjected to logarithm Mel spectrum feature extraction to generate a corresponding to-be-recognized voice feature vector, comparing the target speech feature vector with each template speech feature vector in a speech library, and determining a target speech feature vector and a target speaker; according to the speech feature extraction lightweight model, the calculation cost of the model and the occupied equipment memory are greatly reduced, the frequency domain information and the time domain information in the speech data to be recognized can be comprehensively and accurately captured, and the speech feature extraction lightweight model has high recognition precision and strong anti-noise capability.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +4

Speaker identification method and system based on time-frequency domain dynamic characteristic matrix

The invention provides a speaker recognition method and system based on a time-frequency domain dynamic feature matrix, and the method comprises the steps: mapping a time dynamic feature sequence of an original voice into a two-dimensional image, calculating the similarity through a similarity matrix, and enhancing the time domain dynamic features in the time dynamic feature sequence through a self-adaptive weighting method; performing short-time Fourier transform on the original voice to obtain a frequency spectrum value, calculating a frequency domain dynamic characteristic, and dynamically adjusting a similarity threshold value; training in a convolutional neural network model CNN, and extracting acoustic features in a traditional mode to obtain initial speaker features of the original speech; and the feature fusion classifier calculates the category probability distribution of the speaker according to the weighted adaptive fusion feature vector of the full connection layer, and takes the category with the maximum probability as a final result. According to the invention, the accuracy and robustness of speaker identification in a complex scene are improved, the adaptability to different voice scenes is enhanced, and the information for identifying the identity of the speaker in the voice signal is better reserved.
Owner:XIAMEN UNIV

Audio speaker recognition method and system, storage medium and electronic equipment

The invention discloses an audio speaker recognition method and system, a storage medium and electronic equipment, and belongs to the field of audio processing and artificial intelligence, and the method comprises the steps: carrying out the preprocessing of an input audio, extracting the sound channel information of the input audio, and judging whether the audio type is a single sound channel or a double sound channel; for the dual-channel audio, judging whether the audio type is a pseudo dual-channel or a true dual-channel by comparing at least two acoustic characteristic parameters of a left channel and a right channel; selecting a processing strategy according to the audio type; sequentially performing noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition and punctuation recovery on the selected sound channel; performing speaker role marking on an audio content recognition result by adopting a large language model; and combining the content fragments of the same speaker role to generate structured output. Through multidimensional acoustic feature analysis and cooperation of the deep learning model and the large language model, the recognition accuracy and the automation degree are improved, and the method can adapt to various complex audio scenes.
Owner:SHANGHAI-CHONGQING ARTIFICIAL INTELLIGENCE RES INST

Selective disablement of noise cancelation for conversations

A noise canceling disablement system is provided to enables a user to hear select conversational speech directed at the user while wearing a noise cancelling hearable device. The disablement system automatically at least partial disables the noise canceling feature of the hearable device in response to recognizing conversational speech of a speaking person within a detected conversation zone of the user. In some cases, triggering of the noise canceling disablement further requires the speaking person to be identified by the system as significant person of the user.
Owner:SONY GROUP CORP

Automatic speech recognition result optimization system and method, medium and equipment

The invention provides an automatic speech recognition result optimization system and method, a medium and equipment, and the system comprises an ASR preprocessing module which carries out the initial ASR recognition of an original audio, obtains a preliminary recognition text, carries out the named entity recognition of the preliminary recognition text, enables a recognition result to be matched with a hot word with similar pronunciation in a hot word library, and carries out the recognition of the hot word; inputting the hot words into an ASR model for secondary identification; the ASR text post-processing module is used for carrying out semantic error correction on the text output by the ASR model by utilizing a large language model and verifying the reasonability of the text; and the ASR speaker recognition module is used for segmenting and numbering the audio by using a speaker separation technology, and mapping the number with a specific name. According to the invention, by adopting a mode of entity word extraction and hot word matching, the problem that a speech recognition system cannot screen related hot words in advance is solved, and the recognition effect of an ASR system is improved.
Owner:SHANGHAI SHENGHEKUN INFORMATION TECH CO LTD

Intelligent conference summary generation method and device, equipment and storage medium

The invention provides an intelligent conference summary generation method and device, equipment and a storage medium, and relates to the technical field of natural language processing. According to the method provided by the invention, clear capture and accurate identification of audios are realized by using the multi-channel microphone array; and inputting the audio with the speaker identifier into the speech recognition model, and converting the audio into a structured initial summary text with a timestamp. Performing semantic analysis on the initial summary text through a large language model, and extracting key text information; a rule engine and a large language model are combined for combined judgment, so that the information accuracy is ensured; the initial summary text is converted into a first summary text in a preset text style, so that the text specialty is improved; humanized feature injection enhances the language naturalness of the second summary text; and the second summary text is converted into the structured target summary text, so that the conference summary better meets the requirements of professional scenes such as financial risk control conferences and medical consultation conferences, and the accuracy and availability of intelligent generation of the conference summary are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speaker recognition feature extraction method and system based on self-supervised multi-task learning

The invention discloses a speaker recognition feature extraction method and system based on self-supervised multi-task learning, and belongs to the technical field of biological recognition. In order to solve the problems that the cost is high when a data set with a label is manufactured, and the specific speaker information is difficult to judge under the condition of no supervision information prompt because a voice signal contains various kinds of complex information and also has information of a special hierarchical structure. According to the method, a multi-task learning encoder is adopted to extract a feature sequence, and the feature sequence is used for speaker identification; in the encoder training process, feature sequences corresponding to audios in a training set are simultaneously sent to a plurality of processing tasks for processing, and the processing tasks comprise an original waveform task, a logarithmic power spectrum task, a filter bank feature task and a Mel-frequency cepstrum coefficient task. And a local information LIM identity prediction task and a global information GIM identity prediction task. And performing back propagation by using the loss average value of all the tasks to complete training.
Owner:HARBIN UNIV OF SCI & TECH

Unlearnable speech data generation method against model fine-tuning attacks

This invention discloses a method for generating non-learnable speech data against model fine-tuning attacks, comprising: constructing a proxy speaker recognition model based on a pre-trained speech representation network and a differentiable classifier head; initializing perturbation variables and calculating a temporal saliency mask; determining target identity labels for source samples using a candidate buffer pool and cosine similarity rules; simulating an attacker in the inner layer optimization, and fine-tuning the model using protected speech and real speaker labels; in the outer layer optimization, using the model parameters obtained from the inner layer optimization as a temporary model, calculating the outer layer loss for the target identity labels, and solving the gradient of the perturbation according to a first-order approximation strategy; updating the perturbation and performing projection and spectral difference constraints; and outputting protected speech through repeated iterations. This invention improves the effectiveness against target identity attacks, enhances identity obfuscation capabilities, reduces optimization waste caused by invalid targets, and improves the stability and practicality of the speech identity protection process.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Audio processing method and device, electronic equipment and storage medium

The invention relates to the technical field of audio processing, in particular to an audio processing method and device, electronic equipment and a storage medium. The method comprises the following steps: a main audio acquisition device sets any device as a main speaking device in response to a main speaking mode starting operation triggered for any device, and sets a main speaking identifier of the main speaking device; if the audio data collected by at least one device is obtained, screening out a main speaking device from the devices based on the main speaking identifier and the device identifier carried by each audio data; if the main speaking equipment is screened out, preset operation is executed on the at least one piece of audio data to obtain audio data to be transmitted, and the preset operation comprises at least one of the following operations: enhancement processing is performed on the audio data collected by the main speaking equipment, and suppression processing is performed on other audio data; and transmitting the to-be-transmitted audio data to the terminal equipment for playing. According to the embodiment of the invention, in the speaker mode, the interference of other sounds on the speaker sound is reduced, and the sound collection effect of the cascade equipment is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Selective disablement of noise cancelation for conversations

A noise canceling disablement system is provided to enable a user to hear select conversational speech directed at the user while wearing a noise cancelling hearable device. The disablement system automatically at least partially disables the noise canceling feature of the hearable device in response to recognizing conversational speech of a speaking person within a detected conversation zone of the user. In some cases, triggering of the noise canceling disablement further requires the speaking person to be identified by the system as a significant person of the user.
Owner:SONY GROUP CORP

System and method for associated narrative based transcription speaker identification

Techniques for associated narrative based transcription speaker identification are provided. A narrative of an incident is received at a computing device. The narrative describes an incident. An identification of at least one person involved in the incident is extracted from the narrative. The identification includes a specific identifier for the at least one person. Semantic information is extracted from the narrative. A transcript of media capturing the incident is received at the computing device. The transcript includes a generic identifier for at least one speaker whose speech was transcribed. The generic identifier for the at least one speaker whose speech was transcribed is correlated with the identification based on the semantic information. The generic identifier for the at least one speaker in the transcript is replaced with the specific identifier included in the identification.
Owner:MOTOROLA SOLUTIONS INC

Systems and methods for authenticating and calibrating passive speakers with a graphical user interface

Systems and methods for detecting and configuring passive speakers within a playback system using a graphical user interface are disclosed. In one embodiment, a method of for detecting and configuration passive speakers in a playback system using a mobile device includes deriving speaker identification data concerning one or more passive speakers connected to an audio device in a playback system based upon at least an electrical signal sent to and returned from the one or more passive speakers, where the electrical signal is sent by the audio device including an audio stage comprising one or more amplifiers, and where the speaker identification data comprises information identifying a type of speaker, and displaying a graphical user interface screen on a mobile device based upon the identified type of speaker, where the displayed information and selectable options are dependent upon the identified type of speaker.
Owner:SONOS INC

Electronic device and method for providing translation function by using same

This electronic device may display, on a display, a user interface for providing a translation function of a plurality of speakers, and when a first voice signal is received through a microphone and a first user interaction is detected, train, on the basis of the first voice signal, a speaker recognition model for a characteristic and language of a speaker corresponding to the first voice signal. The electronic device may determine whether a second voice signal is received through the microphone and whether a second user interaction is detected. When the second user interaction is not detected, the electronic device may determine whether there is a training-completed speaker recognition model corresponding to to the second voice signal. When there is the training-completed speaker recognition model corresponding to the second voice signal, the electronic device may translate the second voice signal into a target language on the basis of the training-completed speaker recognition model, and display same on the display.
Owner:SAMSUNG ELECTRONICS CO LTD

A Speaker Recognition Method Based on Channel Attention in a Communication Scenario

The present application relates to a speaker recognition method based on channel attention in a communication scenario. The method includes: constructing a speaker recognition model, which includes a feature extraction backbone network and a speaker classification network connected in sequence; embedding a channel attention mechanism based on cross-network layer feature aggregation into the feature extraction backbone network in the form of multiple channel attention network modules, and the channel attention network module includes a learnable dictionary encoding unit and an information aggregation unit; optimizing and training the speaker recognition model embedded with the channel attention network module, and using the trained speaker recognition model to perform the speaker recognition task in the communication scenario. This method can accurately perceive the importance of each channel feature in the network by representing hidden layer feature information at multiple levels, so as to perform feature selection and differential modeling more efficiently, which has important value for improving the discriminability of speaker representations and the accuracy of speaker recognition.
Owner:NAT UNIV OF DEFENSE TECH

A Speaker Identification Method Based on the Large Model Cam++ of Speech Recognition and Voiceprint Recognition

The present application relates to the technical field of speech and voiceprint recognition, and provides a method for distinguishing speakers based on speech recognition and voiceprint recognition big model Cam++, the method comprising: obtaining the start timestamp and end timestamp corresponding to each sentence in the input audio through the speech recognition big model Cam++, and dividing the audio segment corresponding to each sentence according to the start timestamp and the end timestamp; filtering the audio segments whose duration is less than the preset duration threshold; inputting the filtered audio segments into the voiceprint recognition big model Cam++ to obtain the voiceprint feature vector of each audio segment; obtaining the initial K‑means clustering number; initial K‑means clustering, extracting the deviated abnormal feature vector; performing secondary processing on the extracted deviated abnormal feature vector; clustering and merging groups; processing all ungrouped segments. The present application uses an advanced big model to improve the accuracy of the big model's own voiceprint recognition, while improving the granularity of distinguishing speakers, performing multi-layer K‑means distinction, and improving the accuracy of the number of speaker groupings.
Owner:SSE INFORMATION NETWORK LTD

End-to-end speaker recognition using deep neural network

The present invention is directed to a deep neural network (DNN) having a triplet network architecture, which is suitable to perform speaker recognition. In particular, the DNN includes three feed-forward neural networks, which are trained according to a batch process utilizing a cohort set of negative training samples. After each batch of training samples is processed, the DNN may be trained according to a loss function, e.g., utilizing a cosine measure of similarity between respective samples, along with positive and negative margins, to provide a robust representation of voiceprints.
Owner:PINDROP SECURITY INC