Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

84 results about "Speaker identification" patented technology

Speaker recognition method in medical real-time speech recognition scene

ActiveCN121938378Aconsistent with auditory logicReduce the risk of number fluctuationsSpeech recognitionAutomatic speechSpeech sound
The invention discloses a speaker recognition method in a medical real-time speech recognition scene, and the method specifically comprises the steps: carrying out the parallel distribution of a received PCM audio stream through an audio diverter in a gateway layer, and transmitting the audio stream to an automatic speech recognition link and a speaker separation link; in the automatic speech recognition link, outputting a recognition text and Token-level timestamps corresponding to each minimum semantic unit in the recognition text; in the speaker separation link, extracting a corresponding speaker embedding vector based on a preloaded speaker embedding model, and performing online clustering processing on the embedding vector to generate a speaker identity identifier; token-level boundary adsorption labeling processing is executed based on the Token-level timestamp and an online clustering result, so that a speaker switching position is determined; and carrying out one-to-one mapping on the speaker tags and the recognition text according to the Token-level timestamps, merging Tokens with the same speaker tags continuously, and outputting a transliteration text with the speaker tags.
Owner:ZOE SOFT CORP LTD

Intelligent conference summary generation method and device, equipment and storage medium

The invention provides an intelligent conference summary generation method and device, equipment and a storage medium, and relates to the technical field of natural language processing. According to the method provided by the invention, clear capture and accurate identification of audios are realized by using the multi-channel microphone array; and inputting the audio with the speaker identifier into the speech recognition model, and converting the audio into a structured initial summary text with a timestamp. Performing semantic analysis on the initial summary text through a large language model, and extracting key text information; a rule engine and a large language model are combined for combined judgment, so that the information accuracy is ensured; the initial summary text is converted into a first summary text in a preset text style, so that the text specialty is improved; humanized feature injection enhances the language naturalness of the second summary text; and the second summary text is converted into the structured target summary text, so that the conference summary better meets the requirements of professional scenes such as financial risk control conferences and medical consultation conferences, and the accuracy and availability of intelligent generation of the conference summary are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Unlearnable speech data generation method against model fine-tuning attacks

This invention discloses a method for generating non-learnable speech data against model fine-tuning attacks, comprising: constructing a proxy speaker recognition model based on a pre-trained speech representation network and a differentiable classifier head; initializing perturbation variables and calculating a temporal saliency mask; determining target identity labels for source samples using a candidate buffer pool and cosine similarity rules; simulating an attacker in the inner layer optimization, and fine-tuning the model using protected speech and real speaker labels; in the outer layer optimization, using the model parameters obtained from the inner layer optimization as a temporary model, calculating the outer layer loss for the target identity labels, and solving the gradient of the perturbation according to a first-order approximation strategy; updating the perturbation and performing projection and spectral difference constraints; and outputting protected speech through repeated iterations. This invention improves the effectiveness against target identity attacks, enhances identity obfuscation capabilities, reduces optimization waste caused by invalid targets, and improves the stability and practicality of the speech identity protection process.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

System and method for associated narrative based transcription speaker identification

Techniques for associated narrative based transcription speaker identification are provided. A narrative of an incident is received at a computing device. The narrative describes an incident. An identification of at least one person involved in the incident is extracted from the narrative. The identification includes a specific identifier for the at least one person. Semantic information is extracted from the narrative. A transcript of media capturing the incident is received at the computing device. The transcript includes a generic identifier for at least one speaker whose speech was transcribed. The generic identifier for the at least one speaker whose speech was transcribed is correlated with the identification based on the semantic information. The generic identifier for the at least one speaker in the transcript is replaced with the specific identifier included in the identification.
Owner:MOTOROLA SOLUTIONS INC

Electronic device and method for providing translation function by using same

This electronic device may display, on a display, a user interface for providing a translation function of a plurality of speakers, and when a first voice signal is received through a microphone and a first user interaction is detected, train, on the basis of the first voice signal, a speaker recognition model for a characteristic and language of a speaker corresponding to the first voice signal. The electronic device may determine whether a second voice signal is received through the microphone and whether a second user interaction is detected. When the second user interaction is not detected, the electronic device may determine whether there is a training-completed speaker recognition model corresponding to to the second voice signal. When there is the training-completed speaker recognition model corresponding to the second voice signal, the electronic device may translate the second voice signal into a target language on the basis of the training-completed speaker recognition model, and display same on the display.
Owner:SAMSUNG ELECTRONICS CO LTD

End-to-end speaker recognition using deep neural network

The present invention is directed to a deep neural network (DNN) having a triplet network architecture, which is suitable to perform speaker recognition. In particular, the DNN includes three feed-forward neural networks, which are trained according to a batch process utilizing a cohort set of negative training samples. After each batch of training samples is processed, the DNN may be trained according to a loss function, e.g., utilizing a cosine measure of similarity between respective samples, along with positive and negative margins, to provide a robust representation of voiceprints.
Owner:PINDROP SECURITY INC

Method and system for identifying a speaker of interest in an audio

The present method (300) identifies a speaker of interest in an audio file through a systematic approach. The process begins by receiving an input audio file via a processor (201). The audio file is then split into one or more chunks, followed by the extraction of relevant features from each chunk. Using a transformer encoder model, embeddings of the speaker of interest are generated based on these extracted features. The method identifies one or more nearest neighbours from various data structures corresponding to potential speakers, utilizing a classification model based on the generated embeddings. A set of nearest neighbours is then identified, ensuring that the count exceeds a predefined threshold and that the distance of each neighbour remains below a specified nearest-neighbour distance threshold. Finally, the method provides an identification of the speaker of interest as one of the recognized persons, enhancing speaker recognition capabilities in audio analysis.
Owner:ONIBER SOFTWARE PTE LTD

Training methods, systems, media, and cross-channel and dialect voiceprint recognition models

The present application relates to the technical field of speaker recognition, in particular to a training method, system, medium and voiceprint recognition model across channels and dialects. The voiceprint recognition model has a first input layer, a first convolutional layer, a voiceprint feature extraction network layer, a feedforward network layer and a first output layer connected in turn. When training the voiceprint recognition model, the following steps are included: step S1, constructing a text feature extraction network, a second output layer and a third output layer; step S2, constructing a training set X; step S3, initializing the parameters of the voiceprint recognition model; step S4, training the voiceprint recognition model and constructing a loss function L; step S5, updating the parameters of the voiceprint recognition model; step S6, repeating steps S4 and S5 until the loss function is optimal, and completing the training of the voiceprint recognition model. The present application can preferably improve the generalization and robustness of the trained model.
Owner:NANJING PUBLIC SECURITY BUREAU

Systems and methods for generating video based on informational audio data

In one embodiment, a computer-implemented method may include receiving a media file, extracting, using an artificial intelligence engine including one or more trained machine learning models, one or more audio features from the media file. The one or more audio features include at least one of a time-synchronized transcript, speaker recognition data, mood data, index data, visual asset speaker data, color palette data, and written description speaker data. The method may include generating, based on the one or more audio features, a video, wherein the video is presented via a media player on a user interface of a computing device.
Owner:MUSIXMATCH SPA

Device and method with target speaker identification

A processor-implemented method includes: extracting a target speaker voice feature based on an input voice of a target speaker; determining an utterance scenario of the input voice based on the target speaker voice feature; generating a final target speaker voice feature based on the determined utterance scenario; and determining whether the target speaker corresponds to a user based on the final target speaker voice feature and a final user voice feature, wherein the determined utterance scenario comprises either one of a single-speaker scenario and a multiple-speaker scenario.
Owner:SAMSUNG ELECTRONICS CO LTD

Speaker recognition model frequency modulation trigger injection method

InactiveCN121459823ASpeech analysisFrequency spectrumSpeaker recognition system
The invention discloses a speaker recognition model frequency modulation trigger injection method, particularly relates to the technical field of voice signal processing and artificial intelligence safety, and is used for solving the problems that an existing speaker recognition model back door injection method is insufficient in concealment, poor in black box environment adaptability, prone to distortion of trigger characteristics in the physical transmission process and poor in safety. And mainstream defense strategies such as model fine tuning, pruning and spectrum detection are difficult to resist. The method comprises the following steps: constructing a piecewise linear periodic low-frequency modulation curve for original voice, generating a smooth phase track according to a frequency and phase relationship, constructing a sinusoidal modulator to generate a hidden sample, mixing normal samples in proportion and uniformly marking, training a model by a black box, testing reproduction modulation and inputting the model; according to the method, the defects of insufficient concealment, black box adaptability, physical robustness and defensive resistance of an existing method are overcome, high-concealment, high-robustness and wide-application backdoor injection is realized, and reliable support is provided for security evaluation of a speaker recognition system.
Owner:SUZHOU VOCATIONAL INSTITUTE OF INDUSTRIAL TECHNOLOGY

Unsupervised construction method and device for speaker recognition model, equipment and medium

The invention discloses an unsupervised construction method and device for a speaker recognition model, equipment and a medium, and relates to the technical field of information processing, and the method comprises the steps: extracting the features of a speaker of an original audio data set through a voice pre-training model; performing clustering analysis based on the speaker feature set through a target clustering algorithm, and screening clustering results to obtain a first pseudo-tag set and a first pseudo-tag data set composed of the first pseudo-tag set and the original audio data set; taking the first pseudo tag set as a supervision signal, training the initial speaker recognition model by using the first pseudo tag data set to obtain a first trained speaker recognition model, predicting the original audio data set, and correcting a first prediction result; and training the first trained speaker recognition model according to the first corrected pseudo tag set to obtain a first optimized speaker recognition model, and performing iterative optimization based on the original audio data set until a preset stop condition is met to obtain a target speaker recognition model.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

End-to-end voiceprint recognition method and voiceprint recognition device

Disclosed are an end-to-end voiceprint recognition method and a voiceprint recognition device. The voiceprint recognition method comprises: based on received input speech, using a speaker voice extraction module of an end-to-end deep learning network to perform a speaker voice extraction task to extract speech features of a target speaker; and based on the speech features of the target speaker, using a speaker recognition module of the end-to-end deep learning network to perform a speaker recognition task to recognize the target speaker in the received input speech.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Speaker recognition methods, devices, equipment, and media based on pre-trained speech models

This application discloses a speaker recognition method, apparatus, device, and medium based on a pre-trained speech model, relating to the field of audio recognition technology. The method includes: inserting a target adapter after the target layer of a pre-trained speech model, and processing the target hidden representation output by the target layer using the target adapter; fusing the adapted representations to obtain a fused feature representation, and obtaining a main task loss based on the fused feature representation; acquiring several speech segments, acquiring target speech features corresponding to each speech segment, and adjusting the distance between the target speech features in the target vector space to obtain a content-identity decoupling contrast loss; adjusting the parameters of the target adapter according to the main task loss and the content-identity decoupling contrast loss, and using the adjusted adapter for speaker recognition. By introducing the content-identity decoupling contrast loss, interference from changes in speech content on speaker embedding is suppressed.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Line speaker recognition method and related device

The invention discloses a line speaker recognition method and a related device, and the method comprises the steps: firstly obtaining a video file and a line list, then inputting a target text fragment in the line list and a context text fragment into a large language model, determining N target texts with complete semantics, and carrying out the recognition of a line speaker through the N target texts; and finally, P speakers used for expressing the N target texts are determined in the video text, identities of the P speakers are determined, speaker identities corresponding to the N target texts are obtained, the target texts with complete semantics serve as a whole, speaker identity recognition corresponding to each target text is achieved in the video file, and the recognition efficiency is improved. Compared with a mode of recognizing the line speaker by taking each text segment as an independent recognition unit in the related technology, the method and the device have the advantages that too short duration or wrong recognition of the line corresponding to the same speaker can be avoided, and the accuracy of line speaker recognition is effectively improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Speaker recognition method and related apparatus, device and storage medium

The application discloses a speaker recognition method and related device, equipment and storage medium, wherein the speaker recognition method comprises: obtaining a first voice of a first speaker and a second voice of a second speaker; determining whether the first speaker and the second speaker are the same speaker based on a feature distance between a first voice feature of the first voice and a second voice feature of the second voice; wherein the voice feature is extracted by a feature extraction model, the feature extraction model is trained based on a sample voice set to minimize the training loss, the sample voice set contains sample voices labeled with the sample speaker, the training loss is positively correlated with a first feature distance between the sample voice feature of the reference voice and the positive example voice feature, and negatively correlated with a second feature distance between the sample voice feature of the reference voice and the negative example voice feature. The above scheme can improve the accuracy of the feature extraction model in extracting voice features, thereby improving the speaker recognition accuracy.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Speaker recognition method based on speaker voice micro-movements

ActiveCN118918900BSpeech analysisDigital dataSpeaker verification
The present application relates to the technical field of electronic digital data processing, in particular to a speaker recognition method based on speaker voice micro-motion, comprising: after the voice stream is preprocessed, Fbank features are extracted and sent into a teacher network and a student network respectively to obtain respective corresponding feature embeddings; the feature embeddings obtained by the teacher network and the student network are sent into a loss function and back propagation is performed; the student network is normally iterated, and the teacher network is iterated through an EMA sliding average method; the voiceprint feature information obtained through an ECAPA-TDNN voiceprint model is aggregated and classified with the speaker voice micro-motion information obtained by training the accent data, and speaker recognition is performed; the present application improves the generalization performance of the model by using data enhancement and other methods, avoids fitting in the channel features, does not require manual labeling, and helps the speaker verification model to achieve the ability to identify speakers in a larger population by introducing new features.
Owner:NANJING LONGYUAN INFORMATION TECH CO LTD +1

Artificial intelligence-based automatic marine rescue signal identification system

PCT designated stageWO2026010013A1Semantic analysisAlarmsEngineeringData rescue
An automatic marine rescue signal identification system according to an embodiment of the present disclosure is connected to a ship terminal through a network and capable of identifying a speaker of a rescue signal. The automatic marine rescue signal identification system comprises: a sound receiving unit for receiving sound information including voice and sounds generated from a ship from the ship terminal; a sound transmission unit for digitizing the sound information received from the sound receiving unit to generate and transmit sound data; a rescue signal identification unit for extracting voice data corresponding to a human voice from the sound data received from the sound transmission unit, and for identifying a rescue signal from the extracted voice data; and a speaker identification unit for identifying a speaker of the rescue signal if the rescue signal identification unit identifies the rescue signal from the voice data.
Owner:REPUBLIC OF KOREA (KOREA COAST GUARD COMISSIONER)

Speaker recognition model training method and device based on dynamic weighted mixed loss

The invention provides a speaker recognition model training method and device based on dynamic weighted mixed loss. The method comprises the following steps: training a voiceprint feature extraction network by using a first loss function; freezing network parameters of the voiceprint feature extraction network, extracting an embedded vector of the training set sample based on the voiceprint feature extraction network, and training a generative PLDA network based on the embedded vector; the network parameters of the differentiable PLDA network are initialized based on the network parameters of the generative PLDA network, and the network parameters of the voiceprint feature extraction network are unfrozen; constructing a mixed loss function based on the dynamic smooth weight; and performing joint back propagation updating on the voiceprint feature extraction network and the differentiable PLDA network based on the mixed loss function to obtain a speaker recognition model. And an information feedback closed loop from the rear end to the front end is constructed, so that the final model can achieve better comprehensive performance on a plurality of key indexes, and the accuracy of speaker recognition is further improved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Speaker-identification model for controlling operation of a media player

In one aspect, an example method includes (i) obtaining, by a media player of a media presentation system, an audio signal, where the audio signal includes a voice command and is obtained using a microphone of the media presentation system; (ii) identifying, by the media player, which of multiple speakers of a household uttered the voice command using the audio signal and a speaker-identification model; (iii) performing, by the media player, an action corresponding to the voice command; and (iv) based on the identifying of the speaker using the audio signal and the speaker-identification model, selecting, by the media player, a user profile associated with the identified speaker within a streaming channel so as to bypass a profile selection screen of the streaming channel.
Owner:ROKU INC

Intermediate data for inter-device speech processing

ActiveUS12531070B2Speech recognitionLimited speechData stream
Some speech processing systems may handle some commands on-device rather than sending the audio data to a second device or system for processing. The first device may have limited speech processing capabilities sufficient for handling common language and / or commands, while the second device (e.g., an edge device and / or a remote system) may call on additional language models, entity libraries, skill components, etc. to perform additional tasks. An intermediate data generator may facilitate dividing speech processing operations between devices by generating a stream of data that includes a first-pass ASR output (e.g., a word or sub-word lattice) and other characteristics of the audio data such as whisper detection, speaker identification, media signatures, etc. The second device can perform the additional processing using the data stream; e.g., without using the audio data. Thus, privacy may be enhanced by processing the audio data locally without sending it to other devices / systems.
Owner:AMAZON TECH INC

A target speaker recognition and language learning method and system

The present application relates to the technical field of speech signal processing, in particular to a target speaker recognition and language learning method and system. The method comprises: constructing a unified time axis of multi-modal input data; extracting voiceprint, semantic and context features; fusing the above features to construct a target speaker probability field; performing cross-modal time alignment based on the probability field; extracting and enhancing the target speech accordingly; constructing a target speaker personalized speech model using the extracted speech and generating a synthesized speech; and finally realizing closed-loop optimization based on language learning feedback. Through multi-modal information fusion and probability field modeling, the present application significantly improves the accuracy and robustness of target speaker recognition in complex scenarios, and realizes a complete technical closed loop from accurate extraction, personalized speech modeling to adaptive learning optimization, which is especially suitable for high-quality language teaching applications.
Owner:PEKING UNIV

Speaker identification method and device

The embodiment of the invention provides a spokesman recognition method and device, computer equipment, a computer readable storage medium and a computer program product, and belongs to the technical field of voice recognition. The method comprises the following steps: acquiring an original audio slice sequence generated by a voice recognition system, wherein each original audio slice in the original audio slice sequence comprises identification information of an initial spokesman and voice feature data of the initial spokesman; clustering and merging all the initial spokesmen based on the identification information of the initial spokesmen and the voice feature data to generate a candidate spokesman set; determining the speaking contribution degree of each candidate spokesman; and taking the candidate spokesmen whose speech contribution degrees are not lower than a preset contribution degree threshold value in the candidate spokesman set as target spokesmen to obtain a target spokesman set. According to the technical scheme, the accuracy of the spokesman classification result can be improved.
Owner:SHANGHAI BILIBILI TECH CO LTD

A global context-based speaker analysis correction method

The application provides a speaker analysis error correction method combining global context, and belongs to the technical field of speech signal processing. In view of the problems that in the prior art, speech recognition and speaker analysis are independent of each other, and context utilization is insufficient, leading to low recognition accuracy and semantic consistency, first, the input speech is framed, windowed, denoised and speech activity detected, and multi-dimensional acoustic features are extracted to construct a feature sequence; then, a speech recognition model is used to generate candidate texts and screening is performed to obtain recognition results and time alignment information; further, historical speech features and text information are combined to construct a context representation and fusion; on this basis, speaker embedding is extracted, an initial division result is obtained through similarity calculation and clustering, and the context is used for correction; finally, a structured result containing a speech segment, recognized text and speaker identification is output, so that the recognition accuracy and consistency are improved.
Owner:HUNAN UNIV

A speaker recognition method based on reparameterization cross parallel convolutional neural network

ActiveCN116030802Breduce dependenceImprove computing speedSpeech recognitionNetwork modelSpeech sound
The application relates to a speaker recognition method based on a reparameterization cross-parallel convolutional neural network, and belongs to the field of speech recognition, and comprises the following steps: S1, constructing a structure reparameterization convolutional network RepCNN; S2, building a cross-parallel convolutional neural network based on the RepCNN; S3, extracting two features with different frequency domain resolutions from speech frames of a speaker by adopting two mel filter banks with different numbers of triangular filters; S4, taking the two features as inputs of the cross-parallel convolutional neural network, outputting enhanced features and S5, performing feature fusion on the two features, and outputting a speaker deep feature vector. The application reduces the dependence of the model on computing resources, improves the operation rate and accuracy of the model, and realizes the lightweight of the network model.
Owner:CHONGQING UNIV OF POSTS & TELECOMM